The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 7.1

In this chapter · 5 sections
Term help

Accelerator Landscape & Taxonomy

Choosing an accelerator commits you on four independent axes — ownership and access, execution architecture, workload specialization, and memory semantics — and the four reference groupings (merchant GPU, systolic TPU, hyperscaler XPU, inference ASIC) are points in that space, each carrying its own software stack, scale-up fabric, and lock-in horizon.

GOODPUTPOWER-BOUND

What you'll decide here

  1. Which accelerator family you are designing around — merchant GPU, systolic-array TPU, hyperscaler custom XPU, or inference-specialized ASIC — because that single choice sets your software stack, your scale-up fabric, and your switching cost for the life of the asset.
  2. Whether to buy or rent a NVIDIA/CUDA or AMD/ROCm system, rent a provider XPU, or fund a custom program: design control costs NRE, a compiler team and a multi-year tape-out schedule, while portability still needs testing.
  3. Which datasheet number actually governs your workload — dense vs sparse, peak vs sustained, and at which precision — so you size against a goodput-realistic figure and not a marketing headline that you will never reach.
  4. How much of the value sits above the die: the Broadcom/Marvell design-partner relationship and TSMC fab/CoWoS allocation that any custom program is hostage to, regardless of whose logo is on the chip.
  5. Whether a single-family fleet (operational simplicity, deepest software) or a heterogeneous fleet (price leverage, supply resilience) matches your scale, your team, and your exposure to one vendor's roadmap and allocation.

Part 7 begins with a multi-axis classification: what kind of accelerator is this? Not the part number — the architectural family. Merchant/captive availability, GPU/systolic/dataflow architecture and training/inference specialization are independent, overlapping axes rather than four mutually exclusive families. The familiar labels summarize different bets about how flexible the silicon should be, who writes the compiler, who owns the scale-up fabric, and how long you are locked in. The family choice propagates the way the workload fork in Chapter 1.1 does: it sets the software stack you train your team on, the interconnect you plumb your racks for, the memory supply chain you are exposed to, and the depreciation argument your CFO will fight over.

Compare products on the properties that drive consequences — programmability, who controls the toolchain, and the scale-up domain — and on the two business models layered on top: merchant (you buy a chip and pay the designer's margin for a portable stack) versus captive (you co-design a chip and eat the NRE to replace a merchant purchase price with NRE, software, yield, packaging, capacity and lifecycle obligations). Two firms — Broadcom and Marvell — carry nearly every captive program, and one fab, TSMC, holds every family hostage. The chapter closes on a costly literacy gap in accelerator procurement: reading a datasheet without being lied to by it. Dense vs sparse, peak vs sustained, and the precision games inflate a headline FLOPS number by 4x before you have run a single token.

The four families, by what actually differs

The labels below are not four mutually exclusive accelerator families. They are reference groupings across four independent axes: ownership and access, execution architecture, workload specialization, and memory semantics / scale-up domain. At a given TSMC node and HBM generation, peak FLOPS alone does not settle those choices:

  • Execution architecture and workload specialization. SIMT, systolic, and reconfigurable-dataflow designs are execution choices; training-general, inference-specialized, and operator-specific targets are a separate workload axis. A compiler-programmable TPU or XPU can be workload-specialized without being fixed-function.
  • Ownership, access, and toolchain control. Merchant purchase, cloud rental, and captive internal deployment determine how you obtain the silicon; CUDA, XLA, and Neuron determine the software lock-in. A merchant GPU ships with a mature stack you can hire for, while a captive XPU's compiler exists to serve one owner's models and cloud instances.
  • Memory semantics and scale-up domain. Load/store addressability, compiler-managed collectives, and message passing are different semantics; domain size is how many accelerators retain the selected semantics before traffic falls onto a slower scale-out network. NVLink, ICI, NeuronLink, and UALink therefore compare on both semantics and reach. The fabric engineering itself lives in Part 8; here those are independent taxonomy axes.

Those independent axes order the table below. Read each axis independently before accepting the bundle of consequences its leftmost column implies.

Representative accelerator groupings across independent axes → what you inherit
Reference groupingCanonical 2026 partsExecution architecture / programmabilityOwnership, access / toolchain controlMemory semantics / scale-up domainBest-fit workload / consequence
Merchant GPUNVIDIA Blackwell Ultra GB300/B300 (2026 volume), GB200 (installed base), Vera Rubin VR200 (ramping); AMD MI355X/MI400Fully general SIMT; runs any kernelVendor stack, but portable + hireable (CUDA / ROCm)NVLink 72 (NVL72) → 576; UALink/Ethernet (AMD)Highest unit margin paid for optionality, ecosystem, resale liquidity
Systolic TPUGoogle TPU v7 Ironwood (eighth-gen TPU 8t / 8i announced Apr 2026)Compiler-programmable systolic dataflow; strongest on GEMM, narrower off supported opsXLA / JAX, captive to one operator's cloudICI 3D-torus; 9,216-chip OCS podBest perf/watt on supported ops; rent-only, no merchant market
Hyperscaler XPUAWS Trainium3, Microsoft Maia, Meta MTIA, OpenAI/BroadcomWorkload-specialized and compiler-programmable through the vendor stackCaptive SDK (Neuron, etc.); thin third-party supportNeuronLink / vendor scale-up (e.g. 144-chip UltraServer)Anchor-tenant economics; escapes merchant margin, eats SDK-maturity tax
Inference ASICGroq LPU, SambaNova RDU, AWS Inferentia, d-Matrix, TenstorrentServing-specialized; compiler-programmable or reconfigurable dataflow by designBespoke / emerging; limited framework reachVendor-specific; often small or noneLowest cost-per-token in its niche; fixed-function obsolescence risk
Representative 2026-current groupings, not mutually exclusive silicon families. Read ownership/access, execution architecture, workload specialization, and memory semantics/scale-up domain as independent axes. Part 8 engineers the fabrics; Chapter 7.9 deepens software control.

Each row of the table is a set of inheritances. Choose the systolic TPU and you have also chosen XLA/JAX and a Google Cloud rental relationship — there is no TPU you can buy, rack, and resell. Choose a hyperscaler XPU and you have chosen anchor-tenant economics: the part exists because one operator's internal demand justified the tape-out, and your access to it is a function of their spare capacity and their SDK's maturity, not a merchant price list. Choose an inference ASIC and you may win dramatically on cost-per-token in a narrow serving regime — and lose the moment the dominant model architecture shifts under a fixed-function design. The families do not trade off on one axis; they trade flexibility for efficiency for control, and your scale and software org set the answer, not a FLOPS chart. → the per-family deep dives: NVIDIA in Chapter 7.2, AMD in Chapter 7.3, hyperscaler XPUs in Chapter 7.4, custom ASICs in Chapter 7.5.

Why a GPU and a TPU are not the same animal

The deepest split in the taxonomy is architectural, and it is worth making concrete because it explains the perf/watt gap and the portability gap at once. A GPU is a SIMT (single-instruction, multiple-thread) machine: thousands of cores execute the same instruction across different data, with a large register file, programmable caches, and integrated Tensor Cores for matrix math. Its virtue is generality — kernels supported by its instruction set and software can run — and its vice is that generality costs silicon area and power on control logic, scheduling, and data movement that a more specialized datapath can spend on math.

A TPU is built around a systolic array: a 2D grid of multiply-accumulate units through which data flows rhythmically, each cell passing partial sums to its neighbor so that operands are reused across the array without re-fetching from memory. For dense matrix multiplication — the dominant op in transformer training and serving — this is extraordinarily efficient: it minimizes the data movement that dominates the energy budget, which is why Google's TPU v7 Ironwood reaches roughly 4,614 FP8 TFLOPS per chip at a perf/watt the company markets aggressively against Blackwell (Google Cloud / TrendForce, Nov 2025). The cost is rigidity: ops without an efficient implementation for the array's dataflow — irregular sparsity, dynamic control flow, exotic attention variants — can run poorly or fall back to slower paths; XLA compilation boundaries and Pallas custom kernels determine which paths are available. GPU control and scheduling consume area and energy, while TPU dataflow specialization can make an irregular workload harder to map. Neither tax is a vendor-wide percentage: compile the actual graph, inspect fallbacks, and time its dynamic shapes. → numerics and precision in Chapter 7.10.

~1.44 EFLOPS
GB200 NVL72 rack FP4 with sparsity (dense and other-precision figures are lower — the datasheet-reading trap)
4,614 FP8 TFLOPS
per-chip TPU v7 Ironwood; 9,216-chip OCS pod = 42.5 FP8 ExaFLOPS
2.52 PFLOPS
per-chip Trainium3 MXFP8, 144 GB HBM3E; ~4x perf/watt vs Trn2 UltraServer
~10.1 PFLOPS / 288 GB
AMD MI355X FP4 (no sparsity) and HBM3E capacity; 1,400 W peak board power
5–6 yr book / 4–6 yr estimate / 2–3 yr bear caseforecast
GPU book life vs frontier-economic life — the depreciation fight the family choice feeds (CONTESTED)
Scope & caveats

Accounting policy, published useful-life estimates and the guide's obsolescence stress case are separate quantities.

$270k → $256.4k per nodeestimate
Historical H100 compute-node BOM before/after host and DPU optimization (SemiAnalysis, October 2024)
Scope & caveats

Historical 1,024-H100 cluster with 128 eight-GPU nodes. The compute-node subtotal excludes rack/cluster fabric, networked storage, software and facilities. This is an analyst estimate for that configuration, not a 2026 quote or a complete installed-cost range.

Merchant vs captive — and the firms in the middle

The business-model layer sits on top of the architectural one and is just as consequential. Merchant silicon is sold to anyone: NVIDIA and AMD design a chip, TSMC fabs it, and the buyer pays a gross margin — NVIDIA's consolidated gross margin has run above 70% — in exchange for a portable software stack, a hireable skills market, a deep secondary market, and someone else's roadmap risk. Captive silicon is designed by the operator who will run it. Google, Amazon, Microsoft, Meta, and now OpenAI build accelerators tuned to their own models and clouds, evaluated on total program economics rather than a merchant purchase price. The prize is enormous at hyperscale: at a million-accelerator fleet, the relevant comparison is total delivered cost per useful unit of work, including NRE, software, yield, packaging, capacity, operations and lifecycle risk; a vendor gross-margin percentage is not buyer savings.

But almost no operator designs the whole chip alone. The physical-design, SerDes, packaging, and tape-out expertise sits with two merchant-silicon houses, and nearly every captive program flows through one of them. Broadcom (≈55–60% of the custom-ASIC market) is the design partner behind Google's TPU line, Meta's MTIA, and OpenAI's first in-house accelerator; Marvell (≈15%) serves AWS Trainium/Inferentia. Together they hold roughly 70–75% of the custom-ASIC market (J.P. Morgan via TrendForce; Tom's Hardware, 2025–2026). The strategic reading: 'building your own silicon' rarely means vertical independence — it means substituting a merchant purchase with a design-partner, foundry, packaging and software program and, underneath both, the same single fab.

Merchant vs captive → the business-model fork
DimensionMerchant (NVIDIA, AMD)Captive XPU (TPU, Trainium, Maia, MTIA)
Who pays the marginYou — the vendor's margin (NVIDIA's consolidated gross margin has run above 70%; not a segment figure)You escape the merchant markup and pay NRE + design-partner fee instead; net savings only if total program economics beat the purchase price
Up-front costPurchase price onlyHundreds of $M NRE; mask sets; multi-program commitment
Lead time to volumeOrder against an allocation queueProgram-specific design, qualification and allocation milestones
SoftwarePortable, mature, hireable (CUDA / ROCm)Captive compiler (XLA / Neuron); thin third-party support
Resale / secondary marketDeep — underwrites residual value and GPU-backed debtEffectively none; asset is captive and non-fungible
Who it makes sense forAlmost everyone below frontier-self-build scaleOperators with own-model volume + a silicon/compiler org
The business-model layer beneath the architectural one. Design and delivery budgets belong to the named program; Chapter 7.5 screens its committed volume.

Reading a datasheet without being lied to

The costly literacy gap in accelerator procurement is taking the headline FLOPS number at face value. Vendors quote the largest defensible figure, and the gap between that figure and what your workload sustains can be a factor of four or more before you have run a single token. Three traps recur, and each one inflates the number in a different way.

Trap 1 — dense vs sparse. The biggest headline numbers usually assume structured sparsity (commonly 2:4 — two of every four weights zeroed), which doubles the quoted matmul throughput. NVIDIA's prior-generation GB200 NVL72 rack is marketed at ~1.44 ExaFLOPS FP4 with sparsity; the dense figure is half that, and most production workloads do not realize the full sparsity speed-up. The GB300 NVL72 datasheet a 2026 buyer is actually quoting is written the same way. AMD, by contrast, quotes MI355X FP4 at 10.1 PFLOPS without sparsity — so a naive 'their number is smaller' comparison is comparing a dense figure against a sparse one. Always normalize to dense, at the same precision, before comparing two vendors.

Trap 2 — precision inflation. A chip's biggest FLOPS number is at its lowest-precision format. Drop from FP16 to FP8 and the number doubles; drop to FP4/FP6 and it doubles again. MI355X is a clean illustration: ~2.5 PFLOPS FP16, ~5 PFLOPS FP8, ~10.1 PFLOPS FP4/FP6 — same silicon, a 4x spread purely from the precision the marketing slide chose. If your training run needs BF16/FP8 for stability, the FP4 headline is a number you will never see. Match the quoted precision to the precision your workload actually runs at. → the precision ladder in Chapter 7.10.

Trap 3 — peak vs sustained. Peak FLOPS assumes every multiply-accumulate unit is fed every cycle. Real workloads are throttled by memory bandwidth, collective-communication stalls, kernel launch overhead, and thermal limits. The honest metric is Model FLOPS Utilization (MFU) — sustained useful FLOPS over peak — which Meta’s July 2024 Llama 3 Table 4 measures at 41% BF16 for 405B pre-training on 16,384 H100 GPUs at 8,192-token sequence length, and the goodput that survives failures and restarts is lower still; 90% versus 96% is an illustrative sensitivity, not an industry-average versus best-in-class baseline. A part with a higher peak but a worse compiler and a thinner memory pipe can lose on sustained throughput to a part with a lower headline. This is why measured benchmarks diverge from spec sheets, and why a procurement RFP must demand sustained-MFU numbers on your models, not peak FLOPS on the vendor's. → realized-MFU gap and switching cost in Chapter 7.9; the qualified configuration from Chapter 7.11 enters the lifecycle cost model in Chapter 1.8.

Deep dive: why an inference ASIC can win on cost-per-token and still be the wrong buy

Inference-specialized silicon — Groq's LPU, SambaNova's RDU, AWS Inferentia, d-Matrix, Tenstorrent and others — narrows the dataflow further than even a TPU, optimizing for one serving regime (often low-latency single-stream decode, or high-throughput batched prefill). Within that regime the results can be striking: deterministic latency, very high tokens-per-second, and a cost-per-token well below a general GPU because none of the silicon is spent on training flexibility, large register files, or speculative generality. For a stable, high-volume serving workload on a fixed model architecture, this is a genuine win.

The catch is fixed-function obsolescence risk, and it is the inference-ASIC version of the cooling-cliff one-way door. A part designed around today's dominant attention pattern, today's KV-cache layout, and today's quantization scheme is exposed when the model architecture shifts — and in this field it shifts yearly. A GPU absorbs an architectural change by recompiling a kernel; a fixed-function ASIC may need a new tape-out, which is another design/verification/tape-out cycle and a new NRE bill. The reconfigurable-dataflow ASICs (SambaNova, Tenstorrent) hedge this by keeping the dataflow programmable, trading some peak efficiency for the ability to track model evolution. The decision therefore mirrors the merchant/captive fork: buy the inference ASIC only for a workload whose shape you are confident will outlive the silicon's design cycle, and keep a GPU pool for everything still moving. This is the hybrid fleet, justified. → custom-ASIC economics and the fixed-function-vs-reprogrammable trade in Chapter 7.5.

Single-family vs heterogeneous fleet

The last decision the taxonomy forces is fleet composition, and it is a goodput-vs-resilience trade. A single-family fleet — an all-NVIDIA deployment, for example — buys operational simplicity: one software stack, one scheduler integration, one set of failure modes, one driver matrix, a shared pool of engineers, and one resale channel. The price is total exposure to one vendor's roadmap, one vendor's allocation queue, and one vendor's pricing power. A heterogeneous fleet — GPUs for flexible training, a TPU or Trainium pool for steady high-volume work, an inference ASIC for a stable serving tier — buys price leverage (a credible second source disciplines the incumbent's quote), supply resilience (independently qualified allocations can soften a family-specific shock), and workload-fit efficiency. The price is multiplied operational complexity: every additional family is another compiler to maintain, another set of kernels to port, another realized-MFU gap to measure, another on-call runbook.

The rule of thumb that survives contact with operators: match fleet diversity to scale and to the size of your software org. Below a few thousand accelerators, a second stack still has to earn its recurring engineering and on-call cost; accelerator count alone does not set the threshold. At hyperscale, heterogeneity can pay because the allocation and pricing exposure of a single-vendor fleet at a million-chip scale is an existential risk, and a funded compiler team can pay the integration cost. The middle is genuinely hard, and it is where most of the bad decisions get made — a mid-size operator adds a second family for price leverage it is too small to realize, and drowns in the operational complexity it was too small to absorb. → switching-cost quantification in Chapter 7.9; technical eligibility and RFP construction in Chapter 7.11; buy/rent economics in Chapter 1.8.

Deep dive: the taxonomy as a supply-chain map, not just an architecture map

It is tempting to read the four families purely as engineering categories. The more useful reading in 2026 is as a supply-chain dependency map, because that is what actually gates delivery. Trace a named 2026 accelerator through its stack; many families share chokepoints, but the product determines which ones. The logic die: record the foundry and node, rather than assigning TSMC N3/N3P to every generation. The advanced packaging that stitches logic to memory: TSMC CoWoS for qualified products using that flow, or the named alternative. The memory: a three-supplier HBM oligopoly (SK hynix, Samsung, Micron), with capacity commitments that differ by supplier, generation and delivery window. The captive-program design IP: Broadcom or Marvell in partner programs, or an in-house design team. The merchant alternative: NVIDIA or AMD, themselves at the front of the same TSMC queue.

The consequence for a strategist is that family choice does not diversify your upstream risk as much as it appears to. Switching from NVIDIA GPUs to a Broadcom-designed custom ASIC changes your margin structure and your software stack, but it may leave you on the same TSMC wafers, CoWoS assembly flow and HBM suppliers — it may even put you deeper into the same queue behind the merchant vendors who pre-booked capacity. For an HBM/CoWoS product, additional supply diversification requires a qualified alternative at the packaging and memory layer as well as an architecture choice, which is why those layers — not the choice of GPU vs XPU — are treated as the real allocation gate in Chapter 7.6, Chapter 7.7, and the procurement strategy in Chapter 2.3.

Each family gets a full treatment downstream: NVIDIA's generational cadence in Chapter 7.2, AMD's open-challenger position and the ROCm-maturity tax in Chapter 7.3, the hyperscaler XPUs (TPU, Trainium, Maia, MTIA) and anchor-tenant economics in Chapter 7.4, and the custom-ASIC build threshold (NRE, lead time, minimum-volume) in Chapter 7.5. The binding-constraint layers under all four — HBM in Chapter 7.6, advanced packaging in Chapter 7.7. The software lock-in that the family choice commits you to, and the realized-MFU gap, in Chapter 7.9; the precision ladder behind the datasheet traps in Chapter 7.10. Selection and the technical RFP use this taxonomy; the qualified configuration from Chapter 7.11 enters the lifecycle cost model in Chapter 1.8, including depreciation and buy/rent appraisal; and the consolidated per-generation perf/watt and cost-per-token roadmap in Chapter 16.2.
Cite this chapter
Fehn, J. (2026). Accelerator Landscape & Taxonomy (Chapter 7.1). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-1-accelerator-landscape-and-taxonomy (accessed 2026-09-29).
@misc{aidc-7-1,
  author       = {Fehn, Jacob},
  title        = {Accelerator Landscape & Taxonomy (Chapter 7.1)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-1-accelerator-landscape-and-taxonomy},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit