The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 1.3

In this chapter · 10 sections
Term help

Inference Data Centers: Bursty, Distributed, Always-On

An inference data center serves many independent requests against a latency SLO — always-on, close to users, and sized to goodput-per-dollar and tokens-per-watt at real batch sizes.

GOODPUTPOWER-BOUNDDENSITY-RAMP

What you'll decide here

  1. Which inference sub-mode dominates your facility — interactive/online, batch/offline, or agentic/long-context — because each sets a different SLO, a different prefill:decode mix, and a different fleet-sizing math.
  2. Whether to disaggregate prefill from decode (and pay the KV-cache transfer tax) or co-locate them — the single architectural fork that most shapes accelerator mix, fabric, and goodput in a 2026 serving stack.
  3. How far to oversubscribe power and fabric: inference's uncorrelated per-request peaks can open power headroom (~21% in one scoped example) and fabric oversubscription that a synchronized training trace might not pass — sized from the measured traffic matrix and tail-latency SLO, not the workload label; money you either capture or strand.
  4. Where to site the fleet against a latency budget, and how many regions: proximity-to-users binds only where the complete serving path fails from a cheaper eligible site, and it trades directly against energy cost.
  5. GPU vs inference-ASIC selection scored on tokens-per-watt and tokens-per-dollar at your real batch sizes and context lengths — not peak FLOPS — because the deflation of token prices punishes a wrong silicon bet fast.

For most operators, inference is the business. Training builds the asset; inference is what earns against it, request by request, token by token, every hour of every day. In its November 2025 outlook, Deloitte forecasts inference at roughly two-thirds of all AI compute in 2026 — up from about half in 2025 and a third in 2023 (Deloitte TMT Predictions 2026). That is a share of compute, not of installed capacity, energy, or instantaneous power. Yet the reflex in the industry is still to design the inference fleet as if it were a training cluster with the dial turned down. That reflex is expensive. An inference data center optimizes a fundamentally different objective function, and nearly every subsystem decision flips sign because of it.

Chapter 1.1 split the archetypes on one axis: training-shaped facilities optimize a single tightly-coupled job; inference-shaped facilities optimize many independent requests against a latency SLO. This chapter works the consequences of that difference all the way down — the inference taxonomy and the economics shift, the reasoning-driven demand multiplier reshaping the decode-heavy future, latency-driven siting, the prefill/decode disaggregation that has become the defining 2026 serving pattern, the reliability posture, and a GPU-vs-ASIC selection that is genuinely contestable for the first time. The serving engineering — batching, scheduling, the disaggregation tax in detail — is owned by Chapter 10.11; here we cover what those choices mean for the building.

The inference taxonomy: three sub-modes, three design bases

"Inference" covers at least three distinct workloads, and they pull the facility in different directions. The mistake that recurs is treating a fleet as homogeneous when its dominant sub-mode silently dictates the SLO, the prefill:decode ratio, the memory hierarchy, and the fleet size.

Interactive / online inference is a human (or an interactive agent) waiting on the output in real time. It is governed by two latency metrics — time-to-first-token (TTFT), set by the prefill of the prompt, and time-per-output-token (TPOT), set by decode — and it is bursty and always-on: traffic can swing from 30% to 90% of capacity in minutes on a diurnal-plus-spike pattern. The design basis is high availability, proximity to users, and enough headroom to absorb the peak without breaching the SLO.

Batch / offline inference — embeddings generation, document and corpus processing, synthetic-data creation, evaluation sweeps, nightly re-scoring — has no user waiting. It is throughput-bound, not latency-bound, so it tolerates queuing, interruption, and aggressive oversubscription, and it is the natural consumer of spot capacity, off-peak power, and curtailable interconnections. It is the cheapest inference to host and the most flexible to schedule, which makes it the load you shift to soak up the headroom the interactive fleet leaves on the table.

Agentic / long-context inference is the fast-growing third mode and the one that breaks naive sizing. An agent issues many model calls per user action, carries long and growing context (tool outputs, retrieved documents, prior turns reaching toward 1M+ tokens), and interleaves reasoning with tool use. It inflates the prefill share (huge prompts), explodes the KV-cache footprint (long context held live across many concurrent sessions), and shifts the GPU:CPU ratio toward more host work for orchestration and tool calls. A fleet sized on single-turn chat assumptions is undersized for agents on both memory and prefill compute.

Inference revenue: price accepted serving demand

That Deloitte share is a forecast of compute, not an observed fleet allocation. Capacity forecasts measure another denominator. McKinsey’s 2025–2030 base case has inference capacity rising from ~20.9 GW to ~93.3 GW by 2030 — a ~35% CAGR — against training's ~23.1 GW to ~62.2 GW at ~22%. The crossover is in motion: training is still the larger installed pool in the base year, but inference compounds ~13 points faster and overtakes it well before 2030 on the same forecast. The consequence for siting and procurement is direct. Training capacity concentrates in a handful of gigawatt campuses chasing cheap firm power; latency-constrained inference capacity distributes toward demand, into more, smaller, latency-sited halls.

The dated price evidence also supplies a downside stress. Ramp's customer cohort paid about $2.50 per million tokens in March 2025, down from about $10 per million roughly one year earlier — a Jevons-paradox dynamic where unit cost collapses while aggregate spend rises because demand more than compensates. That cohort price is not a universal market average. For an inference operator this is a business risk to test against the demand and commitment terms: a fleet scoped to today's $/Mtoken can be underwater in a year if its tokens-per-dollar does not improve on the same curve. So the inference design basis fixates on efficiency per token — tokens-per-watt and tokens-per-dollar at real batch sizes — rather than peak throughput. The full unit-economics build-up and the deflation risk are scored in Chapter 1.8.

The serving ledger: a profitable hour can still fail acceptance

$100/fleet-hourmodeled
Full paid fleet-hour cost
Idle and failover reserve remain in the numerator.
Scope & caveats

1.3 ledger only, 64 installed GPUs including idle and reserve. Allocated IT/facility capital recovery, energy, fabric, support and serving software; common downstream application cost excluded from both output ratios. Assumed sensitivity: 80–150 $/fleet-hour; bounds explained in the case ledger.

$0.040/requestmodeled
Net receipt per billable accepted request
An extra attempt never creates an extra customer receipt.
Scope & caveats

1.3 ledger only. Net of discounts and credits, one billable accepted response per original request; failed attempts, retries and free-tier responses earn no receipt. Assumed sensitivity: 0.03–0.05 $/request; bounds explained in the case ledger.

Reconcile attempts to originals: 3,600 + 200 = 3,800 attempts; 3,420 + 180 = 3,600 original outcomes. Accepted-original share = 3,420/3,600 = 95%; active/paid GPU-hours = 48/64 = 75%. Neither activity nor an attempt counter proves service attainment. Accepted output = 3,420 × 500 = 1,710,000 tokens; discarded or otherwise unaccepted generation = 2,000,000 − 1,710,000 = 290,000 tokens. Billable output = 3,000 × 500 = 1,500,000 tokens. Accepted output, billable output and generated output are three different denominators.

Let C be the displayed fleet-hour cost and P the displayed receipt per billable accepted request. Cost per million generated tokens is C/2.00, about $50; cost per million accepted output tokens is C/1.71, about $58; cost per million billable output tokens is C/1.50, about $67. Receipts = 3,000P = about $120 for the hour, leaving about $20 over the stated capacity cost before excluded common application costs. Dividing by generated tokens would conceal the retries and rejected output without reducing the bill.

Select HOLD: 95% misses the 97% service requirement even though the hour has positive capacity contribution. The acceptance crossover is ceiling(0.97 × 3,600) = 3,492 distinct accepted originals, 72 more than this trial. Reaching that count at the same quality, latency and paid capacity makes the trial eligible; adding 72 retries without additional accepted originals changes nothing. Independently, contribution changes sign when P = C/3,000, about $0.033 per billable request. Carry the cost boundary, original-request acceptance, useful output and billing rules together to Chapter 1.8; do not extrapolate this hour into annual demand without a traffic ramp. MLCommons’ inference benchmark rules motivate explicit quality, scenario and latency boundaries; production event accounting is canonical in Chapter 14.1.

Inference sub-mode → requirements cascade
Sub-modeWhat waitsPrefill:decode tiltKV-cache pressureFabric / power oversubscriptionResilience objectiveSiting driver
Interactive / onlineA user, on TTFT + TPOT (sub-second to seconds)Balanced, decode-heavy for reasoningHigh — many live sessions held concurrentlyFabric 2:1-3:1 examples; power headroom measured, not assumed (~21% in the 2024 study)No load loss for maintenance or a single component fault when either event breaches the serving SLO; geo-redundancy contains site lossSub-50 ms proximity to users; latency-first
Batch / offlineNothing — throughput-boundPrefill-heavy (large corpora, short outputs)Low-moderate; reuse/prefix caching helpsHeavily oversubscribed; cost-optimizedRestart- and queue-tolerant; a single distribution path is acceptable when retry latency stays inside the service objectiveCheapest curtailable power; off-peak
Agentic / long-contextA user or pipeline, across many chained callsPrefill-heavy + long decode; tool-call CPU workVery high — 1M+ token context held liveFabric moderate; power lumpy, harder to capPreserve live session state through maintenance and a single component fault when chained-call loss breaches the service objectiveProximity + KV-storage tier near compute
How the dominant inference sub-mode propagates into the rest of the fleet. Latency thresholds and oversubscription figures carry their own vintages and scopes — the ~21% power headroom is a 2024 measurement of specific production fleets, not an allowance a new fleet inherits; see keynumbers for sources.

The table is a cascade, like the one in Chapter 1.1: the left two columns are what you choose; the right-hand columns are hypotheses to qualify. An interactive-dominant fleet must test geo-distribution, a high KV-cache budget, and a willingness to oversubscribe power but not availability. Its revenue-critical service objective often justifies stronger facility topology, but the required maintenance and fault outcomes, not the workload label, select that topology. A batch-dominant fleet can consolidate and accept curtailment when restart and delayed completion still meet its deadline. The agentic column is the one most operators under-provision today, because it looks like interactive chat until the context lengths and call counts reveal a memory-and-prefill problem the single-turn sizing never anticipated.

Reasoning and test-time compute: the demand multiplier

The largest force reshaping inference demand since 2025 is the rise of reasoning models that spend compute at inference time to improve answers — the post-o1, post-R1 paradigm. A reasoning model emits a long internal chain of thought before its visible answer, turning a request that once decoded a few hundred tokens into one that may decode tens of thousands. This is test-time compute, and it has three structural consequences for the fleet that compound:

  • The prefill:decode ratio shifts toward decode. Long chains of thought are autoregressive generation, which is memory-bandwidth-bound, not compute-bound. Decode now dominates the token budget for reasoning traffic, which changes which silicon and which memory hierarchy you want — and tilts the disaggregation math below.
  • KV-cache pressure inflates. Every token of context and every token generated must keep its keys and values resident for attention. Long-decode reasoning, multiplied across many concurrent sessions, makes the KV-cache a first-class capacity constraint — often the binding one — not an afterthought. This is the demand driver behind the new inference-memory tier in Chapter 9.7.
  • Effective demand per request multiplies. If the average request decodes 10-50x more tokens, a fleet sized on the old token-per-request assumption is undersized by roughly the same factor at constant request volume. Reasoning is, in effect, a demand multiplier hiding inside a flat user count.

Treat the decode-heavy future as the base case rather than a tail risk. It argues for memory-bandwidth-rich silicon, a deep KV-cache hierarchy, and fleet headroom for a per-request token budget that keeps climbing. The serving-engineering levers that absorb this — continuous batching, chunked prefill, speculative decoding — live in Chapter 10.11.

~2/3forecast
Deloitte's 2026 forecast for inference share of AI compute (½ in 2025, ⅓ in 2023)
price the serving opportunity, then plan capacity from accepted customer work rather than a market fraction
Scope & caveats

Deloitte’s November 2025 prediction for inference as a share of AI compute in 2026. Not an observed fleet share, installed capacity, energy or instantaneous electrical draw. The forecast does not allocate an individual fleet.

A forecast for calendar 2026, not an observed 2026 outcome.

20.9 → 93.3 GWforecast
AI inference capacity to 2030 (~35% CAGR) vs training 23.1 → 62.2 GW (~22%)
inference is the faster-growing market — bet your build on serving capacity, not training
Scope & caveats

McKinsey base-case scenario for AI workload capacity growth to 2030, not observed capacity. Compare it only against other capacity forecasts on the same denominator.

>$50Bforecast
market for inference-optimized chips in 2026; most inference stays in data centers, not at the edge
the chip market you're buying into — and a sign the edge won't absorb serving load
Scope & caveats

Forecast 2026 market for inference-optimized chips; not realized revenue or an edge-compute market.

~21% vs ~3%
power-oversubscription headroom: inference (uncorrelated per-request peaks) vs training (synchronous peaks)
inference safely sells more compute per megawatt — capacity training can't unlock
Scope & caveats

Measured on the production LLM fleets and power-management platform characterized by POLCA (Patel et al., ASPLOS 2024). It is not a deployable allowance for a 2026 reasoning, MoE, or disaggregated fleet: derive that from the proposed fleet's measured coincident demand, its tested protection and capping response, and the serving degradation you will allow.

2:1–3:1 examples; ~31% modeled 2:1 cost deltaderived
reported 2:1–3:1 inference examples; derive each tier from measured traffic, placement, failure headroom, and tail-latency SLO
right-sizing the inference fabric strips a third off network capex you'd overbuild
Scope & caveats

Examples, not workload defaults. Distributed MoE inference, KV movement, or prefill/decode disaggregation can require high bisection; derive each tier from measured traffic, placement, failure headroom, and the tail-latency SLO.

192 GB / 7.4 TB/s
HBM3E per Ironwood TPU v7 (inference-era ASIC); 9,216-chip pods, 42.5 FP8 ExaFLOPS, 4,614 FP8 TFLOPS/chip
purpose-built inference ASICs are a real alternative to GPUs — and a vendor-lock-in fork
Scope & caveats

Decimal GB, as Google's Ironwood announcement states; the same figure is carried at its canonical home in Chapter 7.4 (claim hbm3e-per-ironwood-tpu-v7-chip-7-37-tb-s-4-614, 192 GB at 7.37 TB/s). 192 GiB would be ~206 GB and is not what the cited source says.

~$10/M -> ~$2.50/M in Ramp cohort
Ramp customer-cohort price per million tokens: ~4× decline over the cohort year; not a market average
your price can fall several-fold a year — today's margin evaporates unless cost falls faster
Scope & caveats

Ramp customer cohort; keep the separate guide-derived self-hosted scenario in its own claim.

$42.3B; inference $23.3B > training $19Bforecast
Gartner forecast: AI-optimized IaaS spending in 2026, with inference dollars above training
the inference thesis now has a dated spend print, not just a compute-share estimate — different denominator, same direction

Latency-driven siting and regional distribution

Training siting chases the cheapest firm megawatt and the coldest free-cooling climate, indifferent to where users are. Inference inverts the hierarchy: the binding constraint is the latency budget back to the user, and that turns siting into a coverage problem. Interactive inference must allocate the illustrative 30/50/100 ms service budgets (treated in full for the edge in Chapter 1.5); a centralized fleet whose actual routed propagation and processing exceed 50 ms cannot meet a sub-50 ms latency budget no matter how fast the silicon, because the speed of light in fiber is not negotiable — roughly 0.5 ms per 100 km one-way (about 1 ms round-trip), before any switching or queuing.

The consequence is a different geographic footprint. Where training consolidates into a few gigawatt campuses, inference distributes into more, smaller, latency-sited regions — typically Tier-1 and Tier-2 metros within fiber reach of population centers. Siting closer to users buys proximity and pays for it in energy cost (test a 2-4x metro-versus-rural price scenario against actual delivered tariffs) and in the operational overhead of running many sites instead of one. Site too centrally and you breach the SLO in distant regions; site too widely and you fragment your fleet below the scale where batching efficiency and utilization hold up. The fiber-and-latency screen that governs this is Chapter 3.6; the market-cluster playbook is Chapter 3.13; the reordered siting hierarchy is Chapter 3.1.

Disaggregated inference: prefill vs decode, and KV-cache as a resource

The defining serving architecture of 2026 is the recognition that the two phases of a request want different hardware. Prefill — processing the entire prompt to produce the first token and the initial KV-cache — is compute-bound: it saturates the accelerator's matrix engines and scales with prompt length. Decode — generating each subsequent token autoregressively — is memory-bandwidth-bound: each step reads the whole model and the growing KV-cache to emit one token, leaving the matrix engines mostly idle. Run both on the same GPU and they interfere: a long prefill stalls the decode stream of every other request sharing the device, and you cannot tune the hardware for either because it is doing both.

Disaggregated serving splits them onto separate pools — a prefill pool sized for compute, a decode pool sized for memory bandwidth and capacity — connected by a fast fabric that ships the KV-cache from prefill to decode. By early 2026 this is supported across every major open-source engine (vLLM, SGLang, TensorRT-LLM) and orchestrated by frameworks like NVIDIA Dynamo and llm-d, with NIXL the standard KV-transfer transport over RDMA/NVLink. The payoff is higher goodput per dollar and the ability to mix silicon — expensive compute-dense parts for prefill, cheaper memory-rich parts for decode. The cost is the disaggregation tax: every request now pays a KV-cache transfer across the fabric, which only nets out when the pools are large enough and the fabric fast enough that the transfer is cheaper than the interference it eliminates.

The facility consequence is that the KV-cache becomes a first-class, tiered resource — held in HBM where it is hottest, spilled to local NVMe and Ethernet-attached flash as it cools, reused across requests that share a prefix. That memory hierarchy is its own subsystem now, engineered in Chapter 9.7. For the building, disaggregation means the inference hall is no longer a uniform sea of identical nodes; it is a heterogeneous fleet whose pool ratios you tune to your prefill:decode mix — which, per the reasoning shift above, is moving.

Deep dive: when disaggregation pays — and when co-location wins

Disaggregation wins when its reduced phase interference exceeds the capacity, transfer and operational cost of separate pools. The KV-cache that prefill produces can be large — gigabytes for a long-context request — and shipping it to a decode worker consumes fabric bandwidth and adds latency to TTFT. The trade pays when three conditions hold together: (1) request volume is high enough that prefill and decode pools each stay busy independently (a small fleet cannot keep two specialized pools utilized); (2) the prefill:decode mix is imbalanced enough that co-location wastes one resource (uniform short chat is the weakest case; long-prompt or long-decode reasoning is the strongest); and (3) the interconnect is fast enough — NVLink within a rack, high-bandwidth RDMA across racks — that the transfer is cheaper than the head-of-line blocking it removes.

The consequence of getting it wrong cuts both ways. Disaggregate too small and you strand capacity in under-utilized pools and pay the transfer tax for no benefit — co-location would have been simpler and cheaper. Co-locate at scale and you cap your goodput on phase interference and forfeit the ability to run cheaper memory-rich silicon for decode. A large interactive fleet must demonstrate the win at equal quality, offered load and tail-latency limits; for small or batch-dominant fleets, co-location with continuous batching is often the better-scoped choice. The engine-level mechanics — chunked prefill as a middle path, KV-aware routing, the transfer-tax accounting — are in Chapter 10.11. Include rejected and incomplete streams, cache transfer, warm reserve and paid idle time in the comparison, then hand the accepted-output ledger to Chapter 1.8.

Oversubscription: qualify the workload and failed states

Two of inference's defining properties — loose coupling and uncorrelated per-request load — open headroom that a synchronous training job forbids, and capturing it is one of the clearest efficiency levers in the fleet.

Fabric. Many local inference requests fit inside one node or a small scale-up domain, so some designs validate upper-tier 2:1–3:1 oversubscription; a named 2:1-versus-1:1 model estimates roughly 31% lower back-end cost. Those are scoped examples, not an inference default. Distributed MoE, KV movement, or prefill/decode disaggregation can require high bisection. Derive each tier from measured request, KV and expert-parallel traffic, placement, failure headroom, and the tail-latency SLO, then validate it on the target fabric. Topology and oversubscription are engineered in Chapter 8.5.

Power. Training's synchronous steps make thousands of GPUs draw their peak in lockstep, leaving only ~3% power-oversubscription headroom — the facility must provision near the synchronized peak. Inference's per-request peaks are uncorrelated, so the aggregate load is far smoother — ~21% of unused headroom in the production fleets the 2024 study characterized. That is evidence the headroom exists in inference, not an allowance your fleet inherits: derive deployable oversubscription from your own measured coincident demand, the capping and protection response you have actually tested, and the serving degradation you will allow (Uptime Institute Journal; arXiv power-profile studies). The result is density-per-megawatt that training cannot match — but it must be engineered with capping and ride-through, because reasoning and agentic traffic make the per-request peaks lumpier than single-turn chat. The grid-facing transient behavior this implies is in Chapter 12.2's reliability frame.

Reliability and uptime: define the surviving service

Training and inference expose different outage consequences, but neither workload label selects a facility topology. Checkpoint-and-resume can reduce a training job’s lost work and recovery cost; request retry, replica capacity, zone routing, and regional failover can mask an inference-site interruption. In both cases, define the maintenance and fault states, allowable transfer interruption, post-event capacity, path and control independence, common modes, recovery SLO, and contractual consequence.

An always-on inference business may place high value on uninterrupted single-site service, but that does not automatically require 2N or Tier-IV-class power: sufficient independent replica, zone, or region capacity can satisfy the end-to-end serving objective with a different site topology. Conversely, a training contract or recovery limit can require path continuity despite checkpointing. Select the site and fleet designs together, then prove the named states. The quantitative comparison is in Chapter 12.5.

GPU versus inference ASIC: accepted output per watt and dollar

Merchant training is dominated by one answer, the rack-scale NVLink domain, though it is not the only one — AMD Instinct systems have carried published end-to-end large-model training programs (AMD/Zyphra, 2025–2026), so the axis to score is software migration cost, demonstrated scale, and time-to-target-quality rather than the absence of an alternative. On the NVIDIA side, GB300 NVL72 shipped in 2025 and is deploying through 2026 alongside GB200; Vera Rubin NVL72 entered full production in August 2026 (NVIDIA Q2 FY27). The workload rewards raw FLOPS, the largest scale-up domain, and the most mature collective-comms software. Inference is the first archetype where the silicon choice is a genuine, contestable fork, because the scoring metric changes from peak FLOPS to tokens-per-watt and tokens-per-dollar at your real batch sizes, context lengths, and SLOs. On that scoreboard, purpose-built inference ASICs become competitive, and several are now in volume. The GPU they are scored against is that same rack: the 72-GPU NVLink domain is what makes prefill/decode disaggregation economic, which is why Azure built its first GB300 cluster for OpenAI reasoning-model inference (October 2025) and AWS's P6e-GB300 instances reached GA in December 2025.

Google's Ironwood (TPU v7) is explicitly "the first TPU for the age of inference": 192 GB of HBM3E at 7.4 TB/s per chip, 4,614 FP8 TFLOPS, scaling to 9,216-chip pods delivering 42.5 FP8 ExaFLOPS — a memory-bandwidth-rich design aimed squarely at the decode-heavy future. AWS's Inferentia/Trainium line and the hyperscaler XPUs (Maia, MTIA) target the same tokens-per-dollar objective for captive fleets. At Cloud Next 2026 (April) Google previewed the eighth-generation split behind it — TPU 8t (training: 3D torus, 9,600-chip superpods, Google-claimed up to 2.7x training $/perf vs Ironwood) and TPU 8i (inference: a new all-to-all 'Boardfly' interconnect, up to 80% better inference $/perf at low-latency MoE serving; both up to 2x perf/W — vendor figures, Apr 2026). Ironwood remains the GA part in 2026. These are surveyed in Chapter 7.4; the qualified configuration from Chapter 7.11 enters the lifecycle cost model in Chapter 1.8.

An ASIC can win decisively on tokens-per-watt for a stable, high-volume model family — but it carries software-ecosystem and lock-in risk (the CUDA moat is real, and a custom part with a thin software stack can leave throughput on the table), and it is a bet that the model architecture you optimize for stays put. Against the Ramp customer cohort's roughly $10/M → $2.50/M decline through March 2025 (not a universal market average) and reasoning reshaping the prefill:decode mix, a part optimized for last year's workload can be the wrong part this year. The rational posture for most operators is to score the fork at their own batch sizes and context lengths — peak-FLOPS datasheets are nearly useless here — and to keep the silicon decision as reversible as the procurement decision, because both move faster than a facility's depreciation clock.

Deep dive: why peak FLOPS is the wrong number for inference silicon

A training buyer can almost get away with comparing peak FLOPS, because training keeps the matrix engines busy. An inference buyer cannot, because the workload's two phases stress different parts of the chip and the datasheet headline reflects neither at the operating point that matters. Decode — the dominant phase for reasoning and chat — is memory-bandwidth-bound: it emits one token per pass over the model weights plus the KV-cache, so it is gated by HBM bandwidth and capacity, and a chip with twice the FLOPS but the same bandwidth decodes no faster while that critical path stays bandwidth-bound. Batch hard enough and the matrix-multiply share can turn compute-bound again while attention stays bandwidth-sensitive, so rank silicon on a measured throughput/latency curve across the concurrency range you actually serve, not on the rule. Ironwood's datasheet reads accordingly — 192 GB of HBM at 7.4 TB/s rather than a FLOPS headline — and memory, not compute, is the figure of merit for the decode pool.

Prefill is compute-bound and does reward FLOPS — but only at the prompt lengths and batch sizes you actually run, and the achievable fraction of peak (the realized MFU) varies enormously with the software stack and the kernel quality for your model shape. The consequence: the only defensible silicon comparison is a measured tokens-per-watt and tokens-per-dollar at your SLO, prefill:decode mix, and context distribution. Two parts with identical datasheets can differ by 2x in realized goodput once batching, quantization (FP8/FP4), and KV-cache behavior are accounted for. The benchmarking discipline and the cross-vendor MFU gaps are in Chapter 7.11; precision and quantization in Chapter 7.10.

Anti-patterns

The recurring inference mis-scopes all trace to one root cause: designing the inference fleet as a de-rated training cluster instead of from its own objective function. Four are worth naming:

  • Sizing fabric from the word inference. A named 1:1-versus-2:1 model estimates a ~31% back-end cost difference, but the workload label does not set the ratio. Node-local serving may validate upper-tier oversubscription; distributed MoE, KV movement, or prefill/decode disaggregation may require more bisection. Derive each tier from measured traffic, placement, failure headroom, and the tail-latency SLO.
  • Sizing the whole fleet to single-turn chat. Reasoning and agentic traffic multiply the per-request token budget and explode KV-cache pressure; a fleet sized on old token-per-request assumptions can exhaust memory or decode capacity, with the factor set by the model, cache policy and admitted concurrency. Replay the declared mix, including long prompts, decode tails and tool waits.
  • Centralizing inference for power economics. Consolidating an interactive fleet onto one cheap-power gigawatt campus breaches the latency SLO for users outside its measured path budget. Distribute toward demand and pay the energy premium knowingly.
  • Buying peak FLOPS instead of tokens-per-watt. Selecting inference silicon on datasheet FLOPS ignores the memory-bandwidth limit when decode is in that regime; the chip that wins on paper can lose 2x on realized goodput at your batch sizes and context lengths.
Inference sits inside the archetype framework of Chapter 1.1 and opposite the training treatment of Chapter 1.2; the hybrid middle (RL as inference-heavy training) is Chapter 1.4 and the latency-bound extreme is Chapter 1.5. The serving engineering this chapter defers — batching, chunked prefill, disaggregation tax, goodput-optimal scheduling — is owned by Chapter 10.11; the KV-cache memory hierarchy by Chapter 9.7. Latency-driven siting connects to Chapter 3.1, Chapter 3.6, and Chapter 3.13; fabric oversubscription to Chapter 8.5; silicon selection to Chapter 7.4 and Chapter 7.11; the reliability flip to Chapter 12.2; and the unit economics and deflation risk to Chapter 1.8.
Cite this chapter
Fehn, J. (2026). Inference Data Centers: Bursty, Distributed, Always-On (Chapter 1.3). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-1-strategy-workload-archetypes-and-economics/1-3-inference-data-centers-bursty-distributed-always-on (accessed 2026-09-29).
@misc{aidc-1-3,
  author       = {Fehn, Jacob},
  title        = {Inference Data Centers: Bursty, Distributed, Always-On (Chapter 1.3)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-1-strategy-workload-archetypes-and-economics/1-3-inference-data-centers-bursty-distributed-always-on},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit