The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 7.4

In this chapter · 5 sections
Term help

Hyperscaler XPUs: TPU, Trainium/Inferentia, Maia, MTIA

A hyperscaler XPU leaves compiler releases and capacity allocation under the cloud’s control; choose the offered service when its workload economics pay for an exit that can require both a port and replacement capacity.

POWER-BOUNDGOODPUTDENSITY-RAMP

What you'll decide here

  1. Whether XLA/JAX on TPU or Neuron on Trainium supports the model and its custom kernels, and whether the measured benefit pays for a tested migration and exit plan; CUDA portability also requires a working port.
  2. Whether you can access an XPU at all: TPU, Trainium, Maia, and MTIA are captive — rentable through one cloud or not rentable at all — so the procurement question is which hyperscaler you are willing to anchor to, not which chip you prefer.
  3. Whether the perf/watt and tokens/$/W advantage of an inference-optimized XPU is large and durable enough to justify re-tooling your serving stack, against a roadmap whose cadence and software maturity you cannot audit from outside.
  4. Whether to rent provider silicon, as Anthropic does with AWS Trainium, or fund a design program such as OpenAI/Broadcom: the latter buys design control by taking on NRE, compiler staffing and tape-out risk.
  5. Which topology the named generation supplies: Ironwood’s optically switched 3D torus, Trainium3’s NeuronLink fabric, or merchant NVLink/UALink. Domain size and fault isolation set what can run and recover together; TPU 8i’s Boardfly is a separate topology.

The previous two chapters covered silicon you can buy: NVIDIA and AMD sell merchant accelerators to anyone with a purchase order. The hyperscaler XPUs you can only rent, or only build. Google's TPU, AWS's Trainium and Inferentia, Microsoft's Maia, and Meta's MTIA are not products on a price list; they are vertically-integrated cost structures — a chip co-designed with a compiler, wired into a proprietary fabric, deployed in racks the same company owns, serving models the same company (or its anchor tenant) runs. TrendForce's January 2026 forecast puts ASIC-based servers at roughly 28% of 2026 AI server shipments — the highest share since 2023 — with custom-ASIC unit growth running near 45% year-over-year, nearly triple the rate of merchant GPUs (TrendForce, January 2026; Tom's Hardware, May 2026). The merchant-GPU monopoly is not collapsing, but the inference half of the market is quietly defecting.

For most operators you do not select an XPU — you inherit one when you pick a cloud. You cannot buy a TPU; you rent a Cloud TPU slice and accept XLA. You cannot buy a Trainium; you rent a Trn instance and accept the Neuron SDK. The decisions here are therefore second-order: stable workload vs portability, captive vs merchant, renting the anchor's economics vs building your own. Underneath all of them runs one structural trend — the model owner concluding that the cheapest way to serve its own tokens is to design the chip that serves them.

Why a hyperscaler builds its own chip

Raw FLOPS alone do not decide: NVIDIA spreads R&D across merchant volume, while Google and AWS can optimize an XPU around their own workloads. What moves a model-owner is three pressures felt at full force. First, cost-per-token at volume. When you serve hundreds of billions of tokens a day against a model architecture you control and rarely change, a workload-specialized XPU that remains compiler-programmable through a stack such as XLA or Neuron can beat a general-purpose GPU on tokens/$/W when that workload lowers cleanly — and at hyperscale, single-digit-percent efficiency gains are nine-figure annual line items. Second, supply independence. NVIDIA allocation can be a binding constraint for a named delivery window (see Chapter 7.6 on HBM and Chapter 2.3 on long-lead procurement); a hyperscaler with its own tape-out and its own TSMC wafer-and-CoWoS allocation is not standing in NVIDIA's queue. Third, the NVIDIA margin. Avoiding NVIDIA's purchase price adds design, software, yield and sustaining costs; gross margin is not recoverable program savings.

All three pressures require scale to pay off, which is why most operators should not build. Custom silicon carries a program-specific NRE and development schedule, a hard minimum-volume threshold below which per-unit cost is worse than just renting GPUs, and a permanent software-engineering tax to keep the compiler competitive (see Chapter 7.5 for the ASIC break-even math). An entity with a durable workload, committed volume and a funded design/software organization can test that bar. An external renter instead qualifies the offered instance, region and reservation; internal-only silicon is not rentable capacity.

Google TPU: the most mature XPU, and the deepest lock-in

The TPU is the only hyperscaler XPU with a decade of generations behind it, and it is the existence proof that the model-owner-builds-silicon thesis works at frontier scale — Gemini is trained and served on TPUs end to end. The 2026 flagship is Ironwood (TPU v7), previewed in November 2025, made generally available March 31, 2026, and explicitly positioned as 'the first Google TPU for the age of inference.' Each chip delivers 4,614 FP8 TFLOPS and 192 GB of HBM3E at 7.37 TB/s — a 6x memory jump over Trillium (TPU v6), the largest single-generation memory expansion in TPU history — at roughly 2x the perf/watt of the prior generation (Google Cloud; SemiAnalysis, Nov 2025). Ironwood is what you can rent; the eighth generation is announced but not the same thing. On April 22, 2026 Google detailed TPU 8t — a pre-training part with a 9,600-chip 3D-torus superpod, SparseCore embedding offload, native FP4, and the two-layer Virgo scale-out fabric — and TPU 8i, a sampling-and-serving part with 3x the on-chip SRAM, a Collectives Acceleration Engine replacing Ironwood's SparseCores, and a high-radix Boardfly ICI topology that holds a chip-to-chip diameter of seven hops where a 1,024-chip torus needs sixteen. Read those as announced architecture, not schedulable capacity: the training/serving split changes the shape of the pod you will eventually buy, not the one you can buy now.

What separates TPU from every GPU fabric is the interconnect. TPUs are wired as a 3D torus over the Inter-Chip Interconnect (ICI, 9.6 Tb/s per chip on Ironwood), and the torus is stitched together by optical circuit switches (OCS) — Google's own MEMS-mirror optical switches that physically reconfigure which chips are neighbors. An Ironwood pod scales to 9,216 chips delivering 42.5 FP8 ExaFLOPS with ~1.77 PB of pooled HBM. The OCS decouples the logical topology a job sees from the physical wiring, which buys two things a switched fat-tree cannot: per-job topology-on-demand (a job requests a torus shape, the OCS provisions it), and fast routing-around-failures (a dead chip is optically bypassed rather than failing the slice). This is why Google can run ~9,200-chip coherent domains without the all-to-all switch silicon that dominates a comparable GPU cluster — the optics are the scale-up fabric. → Chapter 8.2 treats OCS-vs-switched scale-up in full.

The lock-in runs correspondingly deep. There is no CUDA on a TPU; the path is JAX or TensorFlow lowered through XLA, the compiler that turns a high-level graph into TPU machine code. XLA is genuinely excellent for the dense and MoE transformer shapes Google cares about — but it is a compiler-first model, not a kernel-first one. You express computation and trust XLA to schedule it, and where the compiler's choice is not good enough you drop into Pallas, JAX's kernel language for hand-written TPU kernels with explicit control over blocking and memory movement. What you do not get is a PTX-equivalent assembly layer or CUDA's two decades of third-party tuned kernels, so the honest comparison is supported ops, compiler constraints, profiling and performance portability — not custom kernels versus none. For a stable architecture this is a feature (the compiler optimizes across the whole graph); for a research workload chasing a new operator every week it is friction. And TPUs are captive to Google Cloud — you cannot buy one, cannot colocate one, and your exit option is a full re-port to another stack. → Chapter 7.9 quantifies the XLA-vs-CUDA switching cost.

AWS Trainium / Inferentia and the anchor-tenant economics

AWS runs a two-chip split: Inferentia for inference and Trainium for training, both programmed through the Neuron SDK, which sits under PyTorch and JAX. The 2026 story is Trainium3, announced at re:Invent 2025. Each Trainium3 chip delivers 2.52 FP8 PFLOPS with 144 GB of HBM3E at 4.9 TB/s — the first 3nm AI accelerator AWS has shipped. The system unit is the Trn3 UltraServer: up to 144 chips for 362 FP8 PFLOPS, 20.7 TB of HBM, and 706 TB/s of aggregate bandwidth, delivering ~4.4x the performance and 4x the perf/watt of the Trn2 UltraServer (AWS, Dec 2025). The scale-up fabric is NeuronLink-v4 with a NeuronSwitch-v1 all-to-all topology at 2 TB/s per chip — AWS taking a vertical scale-up like NVIDIA's, against an open-Ethernet scale-out (EFA).

What makes Trainium strategically different from TPU is the anchor-tenant model. Google builds TPUs primarily for itself; AWS builds Trainium substantially for Anthropic. Project Rainier — the Trainium2 cluster in St. Joseph County, Indiana, an $11B AWS site — went from announcement (re:Invent, Dec 2024) to a live cluster of ~500,000 Trainium2 chips in under twelve months, and by 2026 Anthropic and AWS report running over one million Trainium2 chips to train and serve Claude, with the partnership committed toward up to 5 GW of compute (AWS; Anthropic, 2026). External adopters such as Uber have cited cost savings on the order of 50% versus NVIDIA on suitable workloads. The anchor tenant de-risks the silicon: a guaranteed million-chip buyer amortizes the NRE before the first external customer ever rents a Trn instance. By Q2 2026 the chip business had become a headline in its own right — ~$25B annualized revenue across Trainium, Graviton, and Nitro (blended, not Trainium-only), ~$225B of multi-year commitments, and Trainium4 pointed at 2027 (Amazon, Aug 2026). The anchor is no longer exclusive, though: Anthropic signed for up to 2 GW of AMD MI450/Helios in July 2026 (Chapter 7.3), so the Rainier tenant now dual-sources. The consequence for everyone else: you are renting into economics that were sized for someone else's model, on a Neuron stack whose roadmap follows the anchor's needs, not yours.

Hyperscaler XPU comparison (2026-current)
XPUFlagship (2026)Per-chip peakHBM / BWScale-up fabricSoftwareAccess modelPrimary workload
Google TPUIronwood (v7)4,614 FP8 TFLOPS192 GB HBM3E / 7.37 TB/s3D torus over ICI + optical circuit switch (OCS); 9,216-chip pod, 42.5 ExaFLOPSJAX / TF via XLACaptive — Google Cloud onlyInference-first (also trains Gemini)
AWS TrainiumTrainium3 (Trn3)2.52 FP8 PFLOPS144 GB HBM3E / 4.9 TB/sNeuronLink-v4 all-to-all (NeuronSwitch-v1, 2 TB/s/chip); 144-chip UltraServerNeuron SDK (PyTorch/JAX)Captive — AWS only; anchor: AnthropicTraining + inference (Claude)
AWS InferentiaInferentia2Inference-tuned (lower TFLOPS, high efficiency)32 GB HBM / moderateNeuronLink (smaller domains)Neuron SDKCaptive — AWS onlyCost-optimized online inference
Microsoft MaiaMaia 200>10 PFLOPS FP4 / 5 PFLOPS FP8216 GB HBM3E / 7 TB/sEthernet-based; wider rack + Sidekick closed-loop liquidMaia SDK / Triton pathCaptive — Azure first-partyServing GPT-class + Copilot
Meta MTIAMTIA 300-seriesRecsys/inference-tunedHBM (3nm, CoWoS from 300-gen)Internal Meta fabricPyTorch-native internal stackNot rentable — Meta internal onlyRecommendation + GenAI inference
Per-chip and per-pod/system figures from vendor primaries and independent analysis; see keynumbers for sources and vintages. 'Captive' = rentable through one cloud only, not purchasable. Maia/MTIA figures are first-party-deployment specs, not rentable instances.

The 'Access model' column is what governs strategy: every XPU here is captive, and two of them (MTIA, and effectively the first-party tiers of Maia) are not rentable at all. Choosing an XPU is therefore downstream of choosing a cloud, and choosing a cloud is downstream of choosing a software stack you are willing to live inside. The perf/watt numbers matter, but what you are really picking is an ecosystem.

Maia, MTIA, and the model-owner-builds-silicon wave

The third cohort is the newest, and the clearest signal of where the industry is heading. Microsoft Maia 200 deployed in early 2026: TSMC 3nm, 140B+ transistors, >10 PFLOPS FP4 / 5 PFLOPS FP8, 216 GB HBM3E at 7 TB/s in a 750 W envelope, which Microsoft claims delivers ~30% better performance-per-dollar than the best hardware in its existing fleet. Maia ships with a co-designed system: a wider rack and the Sidekick closed-loop liquid-cooling sidecar that lets Azure retrofit Maia into existing halls without re-plumbing the building (a deliberate density-ramp hedge — see Chapter 5.10 on liquid retrofits). In 2026 Maia 200 serves GPT-class models for OpenAI and powers Microsoft 365 Copilot. Meta MTIA took the opposite, most-aggressive roadmap stance: in early 2026 Meta disclosed four new generations (MTIA 300 through 500) for deployment through 2027, moving to 3nm with CoWoS packaging — silicon built for Meta's recommendation and GenAI inference and never sold to anyone.

The most consequential entrant is not a cloud at all. OpenAI, partnered with Broadcom in a ~$10B program, taped out its first custom inference ASIC ('Jalapeño') in roughly nine months, targeting deployment from late 2026 (VentureBeat / Tom's Hardware, 2026). This is the model-owner-builds-silicon trend in its purest form: the entity that owns the most-served model in the world deciding that the cheapest tokens are the ones running on a chip it designed. And it reveals the quiet kingmaker — Broadcom is the design-and-SerDes partner behind OpenAI's Jalapeño (with Marvell playing a similar role elsewhere). The merchant-silicon disruption is not a swarm of independent chip startups; it is a handful of model owners renting Broadcom's and Marvell's packaging, SerDes, and physical-design expertise to convert their workload knowledge into captive silicon. → Chapter 7.5 on the merchant-silicon ASIC economics; Chapter 8.3 on the SerDes and switch-ASIC supply chain underneath all of it.

192 GB
HBM3E per Ironwood (TPU v7) chip @ 7.37 TB/s; 4,614 FP8 TFLOPS; 6x memory vs Trillium
Scope & caveats

Ironwood HBM capacity

9,216
chips per Ironwood pod = 42.5 FP8 ExaFLOPS, ~1.77 PB pooled HBM, OCS 3D-torus
2.52 PFLOPS
FP8 per Trainium3 chip; 144 GB HBM3E @ 4.9 TB/s; first 3nm AWS accelerator
144 chips
per Trn3 UltraServer = 362 FP8 PFLOPS, 20.7 TB HBM, 4x perf/watt vs Trn2
>1M
Trainium2 chips on Project Rainier serving and training Claude (from ~500k at activation, <12 mo)
216 GB
HBM3E per Microsoft Maia 200 @ 7 TB/s; >10 PFLOPS FP4 / 5 PFLOPS FP8, 750W, +30% perf/$
27.8% forecast for 2026forecast
ASIC share of AI server shipments in 2026 (highest since 2023); custom-ASIC units +~45% YoY
Scope & caveats

Forecast of server-unit shipments, not an observed accelerator, revenue, energy or workload share; retain the historical forecast date.

Rent an XPU, build one, or stay on merchant GPUs

For an operator who is not a frontier model owner, the practical question collapses to three options, and the right one is a function of workload stability and your willingness to enter a single vendor's gravity well.

Stay on merchant GPUs (NVIDIA/AMD). A candidate for workloads that churn architecture, need supported portable paths, or depend on its kernel ecosystem. You pay the NVIDIA margin and stand in the allocation queue; in exchange you keep CUDA velocity, and a multi-vendor exit still requires a tested ROCm or other backend. → Chapter 7.2, Chapter 7.3.

Rent an XPU through its captive cloud. The right call when you have a stable high-volume inference workload and the XPU's tokens/$/W beats GPUs by enough to repay the port — a TPU slice for a JAX-native serving stack, a Trn/Inf instance for a PyTorch model that lowers cleanly through Neuron. The consequence is cloud lock-in: your serving stack now assumes that vendor's compiler, fabric, and roadmap, and your exit is a re-port. → Chapter 7.9, Chapter 7.11.

Build your own. Reserved for entities that own the model, ship at hyperscale, and can clear the program-specific NRE and development schedule, and the permanent compiler-team headcount — almost always via a Broadcom/Marvell partnership rather than from scratch. The potential payoff is a separately contracted allocation and a lower cost-per-token after program costs; the risk is a workload-specialized design's compiler lock-in against a model architecture that may move under you. → Chapter 7.5.

Rent-XPU vs build-silicon vs merchant-GPU — the fork and its downstream cost
PathUp-front costSoftware burdenSupply postureExit costBest-fit
Merchant GPU (buy/rent)Capex or opex; pays NVIDIA marginLowest — CUDA/ROCm ecosystemStands in allocation queueLowest — multi-vendor, portableArchitecture churn, portability, frontier research
Rent XPU (TPU/Trn/Inf)Opex only; below-GPU $/token at volumeModerate — XLA or Neuron portInherits the cloud's allocationHigh — re-port to exit the cloudStable high-volume workload on one cloud
Build own siliconFunded NRE, tape-out schedule and compiler teamHighest — own the full stackOwn allocation risk; shared foundry/CoWoS suppliers can remainSunk — committed to the architectureModel owner funding design (Google; OpenAI/Broadcom)
Heuristic decision frame for the model-owner-builds-silicon question. 'Break-even volume' is directional; the real threshold is workload-specific (see Chapter 7.5).
Deep dive: why TPU's OCS torus and Trainium's all-to-all are opposite bets — and what each costs you

The two most mature XPUs made opposite scale-up bets, and the difference is instructive because it is the same fork merchant-GPU buyers face between NVLink-style all-to-all and a torus. Google’s Ironwood TPU uses a 3D torus stitched by optical circuit switches. Google’s TPU 8t/8i disclosure separates 8t’s torus from 8i’s Boardfly topology. Each chip talks directly only to its torus neighbors over ICI; the OCS layer physically reconfigures which chips are neighbors, so a job requests a logical topology and the optics provision it. The win: 9,216-chip Ironwood pods with distributed memory and no central all-to-all switch silicon, per-job topology-on-demand, and supported optical reconfiguration around unavailable resources; this does not promise transparent continuation of a running job. The cost: a torus has higher hop-count and more constrained collective patterns than a fully-connected switch, so the compiler (XLA) must be exceptionally good at mapping collectives onto the torus — which it is, for the shapes Google runs, and which is exactly why TPU is hard to use for shapes it does not run.

AWS's Trainium3 UltraServer took a switched-fabric approach: a NeuronSwitch all-to-all fabric (NeuronLink-v4, 2 TB/s/chip) inside the UltraServer. The win: low-hop, uniform any-to-any bandwidth across the 144-chip domain, which makes tensor- and expert-parallel mapping straightforward and forgiving — closer to how a GPU cluster behaves. The cost: switch silicon and the copper/optics reach that bounds how large the all-to-all domain can grow before you fall back to scale-out (EFA) for the rest. Neither bet is wrong; they encode different convictions about whether the compiler or the fabric should absorb the complexity. For the operator, the practical read-through is the same as the merchant-GPU case in Chapter 8.2: the scale-up domain size — not the per-chip FLOPS — sets the supported parallelization plan and the blast radius when a chip dies.

Deep dive: the Neuron / XLA software maturity tax (the gap behind the paper FLOPS)

The published peak FLOPS of an XPU is the easy number. The hard number is realized MFU — the fraction of peak you actually achieve on your model — and it depends on the compiler, kernels, memory and fabric as well as silicon. Compiler age helps explain the ecosystem; it does not measure this workload. XLA is a decade mature: for dense and MoE transformers in JAX/TF, realized efficiency on TPU is competitive with a well-tuned CUDA stack, because Google has had years to co-evolve the compiler with its own models. Neuron is younger and narrower: it lowers PyTorch and JAX well for the model families AWS and Anthropic care about, but a novel operator, an unusual attention pattern, or a custom kernel can hit a wall — either failing to lower efficiently or requiring AWS to add support on a roadmap you do not control. Independent analysis (SemiAnalysis, Trainium3 deep dive) repeatedly frames Neuron's software maturity, not the chip, as the gating factor versus NVIDIA.

The consequence for the decision: the perf/watt and tokens/$/W advantages quoted in a vendor keynote assume the workload already lowers well on that XPU. Budget the port as a real engineering project, validate quality-qualified goodput, latency and recovery on your model before committing volume, and treat any custom-kernel dependency as a funded maintenance obligation on the XPU’s supported custom-kernel interface. A 4x-perf/watt chip running your model at half its achievable MFU is not a 4x chip. → Chapter 7.9 qualifies the execution tuple across CUDA / ROCm / XLA / Neuron.

This chapter sits inside the accelerator selection arc. The merchant accelerators these XPUs compete with are in Chapter 7.2 (NVIDIA) and Chapter 7.3 (AMD); the taxonomy that frames GPU-vs-XPU is Chapter 7.1. The ASIC economics — NRE, lead time, minimum volume, and the Broadcom/Marvell merchant-silicon model — are quantified in Chapter 7.5. For a design using HBM and CoWoS, qualify the memory lot and package slot separately using Chapter 7.6 and Chapter 7.7; the software lock-in (CUDA / ROCm / XLA / Neuron) and the realized-MFU gap is Chapter 7.9; the qualified configuration from Chapter 7.11 enters the lifecycle cost model in Chapter 1.8; procurement qualification stays in Chapter 7.11. The scale-up fabrics — TPU's OCS torus vs NeuronLink/NVLink/UALink all-to-all — are engineered in Chapter 8.2, and the SerDes/switch-ASIC supply chain underneath them is Chapter 8.3. The inference-share and power-bound forces driving the XPU wave trace back to Chapter 1.3 and Chapter 3.2; long-lead procurement and anchor-tenant supply strategy is Chapter 2.3.
Cite this chapter
Fehn, J. (2026). Hyperscaler XPUs: TPU, Trainium/Inferentia, Maia, MTIA (Chapter 7.4). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-4-hyperscaler-xpus-tpu-trainium-inferentia-maia-mtia (accessed 2026-09-29).
@misc{aidc-7-4,
  author       = {Fehn, Jacob},
  title        = {Hyperscaler XPUs: TPU, Trainium/Inferentia, Maia, MTIA (Chapter 7.4)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-4-hyperscaler-xpus-tpu-trainium-inferentia-maia-mtia},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit