Chapter 7.8
In this chapter · 5 sections
Host CPUs, GPU:CPU Ratios & System Composition
Host throughput, memory and attach architecture decide whether accelerators stay fed; name the host/socket/package denominator in every CPU-to-accelerator ratio, then size the host and independent CPU tier from measured work and latency.
What you'll decide here
- Whether to use Grace/Vera over NVLink-C2C or EPYC/Xeon over PCIe, comparing supported coherence, migration, latency, service boundaries and supply coupling; CPU ISA alone does not determine coherence.
- The GPU:CPU ratio your workload actually needs — the profiled host throughput per workload; an eight-accelerator dual-socket host is 8 per host but 4 per CPU package, so vendor projections require a named denominator — because it sets host core count, host memory, and per-rack CPU spend.
- Which host-memory bargain to buy: Grace’s soldered LPDDR5X, Vera’s detachable SOCAMM, or socketed DDR5/MRDIMM you can size and service. Compare usable capacity, bandwidth, RAS and the exact OEM repair unit.
- Whether your fleet is one composition or several — a dense coherent trainer SKU and a CPU-heavy agentic-inference SKU are different machines, and pretending one part serves both strands GPUs or strands cores.
- Whether the system-composition decision is bundled (you take the vendor's rack-scale CPU:GPU ratio as shipped) or disaggregated (you compose host and accelerator yourself) — and what that costs you in flexibility versus integration risk.
An accelerator does not run a workload by itself. It runs as the muscle of a system — a host CPU that loads the model, drives the data pipeline, schedules kernels, services interrupts, runs the OS and the container runtime, terminates the storage and management network, and increasingly does real application work: tokenization, retrieval, sandboxed tool execution, agent orchestration, and the reward-evaluation loop of reinforcement learning. The question this chapter answers is not "which GPU" — that is Chapter 7.2 through Chapter 7.5. It is the question that sits one level up: what machine do you wrap around the accelerator, and in what ratio. Get it wrong in one direction and the host starves the GPUs — they idle waiting on data that the CPU pipeline cannot deliver, and you have paid for the most expensive silicon in the building to wait. Get it wrong in the other direction and you have bought host cores and host memory that no kernel ever touches — dead capex riding on a power budget you are otherwise fighting to conserve.
This was a quiet, settled question for the training era. The answer was "as little CPU as you can get away with": eight GPUs per host, the cheapest x86 that could keep the pipeline fed, and as much of the silicon budget as possible spent on accelerators. Two forces broke that settlement. First, coherent host-attach — NVIDIA's NVLink-C2C welding a Grace (and now Vera) Arm CPU to the GPU as a single memory-coherent superchip — turned the host from a discrete component into part of the accelerator's memory hierarchy. Second, and with wider consequences, agentic inference and RL shifted real, sustained compute back onto the host. The orchestration layer that schedules sub-agents, routes tool calls, runs retrieval, and decides whether a task is done is CPU work, and it does not fit in the slivers of host left over from an 8:1 training box. The industry is re-deciding system composition in real time, and the fork has consequences that propagate into power, density, supply, and serviceability for the life of the fleet.
The two host-attach models
There are two physically and architecturally distinct ways to attach a host CPU to accelerators, and the choice is the first fork in system composition. They differ in the interconnect, the memory model the programmer sees, the failure blast radius, and — critically in 2026 — in how tightly they couple you to a single vendor's supply chain.
Coherent host-attach (the superchip model). NVIDIA's Grace-Blackwell and Vera-Rubin packages connect an Arm host CPU to the GPU(s) over NVLink-C2C, a chip-to-chip link delivering ~900 GB/s of coherent bandwidth on Grace-class parts (rising on Vera). The word coherent is doing real work here: the CPU's LPDDR5X and the GPU's HBM live in one cache-coherent address space, so the GPU can reach host memory and the CPU can reach HBM with hardware coherence, no explicit copy, no staging through a bounce buffer. The host memory becomes a fast, large NUMA tier behind the GPU — the natural home for KV-cache spillover, large embedding tables, MoE expert weights that do not fit in HBM, and the working set of long-context and agentic inference. The cost: the CPU is welded to the GPU. You take the ratio the package ships (2 GPUs : 1 Grace in GB200; the Vera-Rubin pairing in 2026), the LPDDR5X capacity is what the package ships (soldered on Grace, field-replaceable SOCAMM modules on Vera), and you are buying one vendor's CPU, GPU, and interconnect as an indivisible unit.
Discrete host-attach (the PCIe model). The classic HGX/8-GPU baseboard hangs accelerators off a pair of x86 sockets (AMD EPYC or Intel Xeon) over PCIe — Gen5 at ~64 GB/s per direction per x16 (~128 GB/s bidirectional), Gen6 at ~128 GB/s per direction. The CPU and GPU keep separate memory spaces; moving data between them uses DMA or a supported unified-memory/page-migration path over PCIe or, for GPU-to-GPU, a sideband over NVLink that bypasses the host entirely. You give up coherent host memory — host RAM is not a transparent extension of HBM, and PCIe bandwidth is an order of magnitude below NVLink-C2C — but you gain composability. You choose the CPU vendor, the core count, the socketed DDR5 capacity (terabytes if you want them), and the GPU:CPU ratio independently, and you can service or upgrade the host without touching the accelerator. For workloads that do not need coherent host memory — most discrete-GPU inference, and any fleet that values dual-source supply over the last increment of host-GPU bandwidth — this is the pragmatic default.
| Axis | Coherent Arm (Grace / Vera + NVLink-C2C) | Discrete x86 (EPYC / Xeon + PCIe) | What the choice costs you |
|---|---|---|---|
| Host-to-GPU link | NVLink-C2C ~900 GB/s bidirectional and coherent (rising on Vera) | PCIe Gen5 ~128 GB/s / Gen6 ~256 GB/s bidirectional per x16 | ~7x vs Gen5 / ~3.5x vs Gen6 host-GPU bandwidth, and coherence vs explicit copies |
| Memory model | One cache-coherent address space; HBM + host LPDDR5X unified | Separate CPU and GPU spaces; DMA across PCIe | Coherent host RAM as an HBM spill tier vs manual staging |
| Host memory | LPDDR5X — Grace 480 GB soldered; Vera up to 1.5 TB in detachable SOCAMM modules | Socketed DDR5 / MRDIMM — multiple TB, field-serviceable | Bandwidth-per-watt vs capacity, RAS, and serviceability |
| GPU:CPU ratio | State denominator: GB200 pairs 2 GPUs per Grace CPU package; other platforms differ | State denominator and socket/package count; an 8-accelerator dual-socket host is 4 accelerators per CPU package | Take the vendor's ratio vs size it to your workload |
| CPU ISA / vendor | Arm (NVIDIA Olympus/Neoverse cores), single-source | x86, dual-source (AMD + Intel) | CUDA-tight integration vs portability and supply hedging |
| Blast radius / service | CPU + GPU fail and refresh as one welded unit | Host serviceable independently of accelerators | Tighter integration vs independent repair and upgrade |
The GPU:CPU ratio — and why agentic AI is rebalancing it
The GPU:CPU ratio is the number that most concisely captures system composition. A common training host carried eight accelerators and two CPU sockets/packages: eight accelerators per host, or 4:1 when the denominator is CPU packages. Ratios must name their denominator. In this example, the CPU's job was to keep the data pipeline full and stay out of the way. The host did almost no application compute — pre-training is GPU-bound on collectives and matmul, and the CPU's contribution is loading shards, augmenting data, and launching kernels. The rational move was to minimize host spend: the dollars and the watts belonged to the GPUs. Even host memory got cut — neocloud reference builds trimmed training-server RAM from 2 TB to 1 TB precisely because the trainer never used it — a host-BOM option alongside a smaller CPU and removing two BlueField-3 DPUs. SemiAnalysis’s October 3, 2024 example priced those changes at $270k → $256.4k per H100 server, a $13.6k saving in its 1,024-GPU scenario. A new quote must preserve any required DPU isolation or networking function.
Agentic inference and RL break the constant. The work that defines an AI agent — decomposing a request into sub-tasks, scheduling them, routing tool calls, passing state between sub-agents, running retrieval, executing code in sandboxes, and evaluating whether the original goal is met — is overwhelmingly host-side, control-heavy compute. It is branchy, latency-sensitive, and does not vectorize onto a GPU. RL training compounds this: every action an agent takes during a rollout has to be executed and scored, and that evaluation loop runs on the CPU. As agents proliferate, the host stops being a feeder and becomes a co-processor doing sustained work, and the ratio that kept GPUs the scarce resource per node no longer holds. TrendForce projects the ratio shifting from the training-era 4:1–8:1 toward lower accelerator-to-CPU-package ratios in some projected agentic deployments; Intel describes a similar directional shift, but vendor projections require a named denominator and workload profile. Arm goes further on the aggregate: as agents proliferate, orchestration, tool execution and state handling consume host cores while the GPU waits on them. Arm’s addressable-market projection does not specify this rack’s CPU utilization, node power or admission policy.
That rebalance has a concrete consequence. A fleet scoped at 8:1 for training and then repurposed for agentic serving will be CPU-starved: the orchestration and tool-execution layer saturates the thin host, queues form ahead of the GPUs, tail latency blows past the SLO, and accelerators sit underutilized behind a host bottleneck. The fix is not more GPUs — it is more CPU per GPU, which means a different node, often a different rack SKU, and a host-side power and memory budget the training composition never planned for. This is why NVIDIA designed Vera explicitly "for agents": 88 custom Olympus cores, 176 threads via spatial multithreading, and up to 1.5 TB of LPDDR5X at ~1.2 TB/s — a host built to do real work, not just feed.
Host memory: Grace LPDDR5X, Vera SOCAMM, and socketed DDR5/MRDIMM
The memory the host carries is the third axis of composition, and it tracks the host-attach fork tightly. The coherent superchips use LPDDR5X — the low-power DRAM from the mobile world — soldered on Grace, but in detachable SOCAMM modules on Vera. The reasons are bandwidth-per-watt and density: NVIDIA’s GB200 profile pairs Grace with up to 480 GB of LPDDR5X at up to 512 GB/s; the host power setting belongs to the selected OEM profile, and NVIDIA's framing is roughly 2x the bandwidth at half the power of conventional socketed server memory. Vera pushes this to up to 1.5 TB at ~1.2 TB/s. In a power-bound rack where every watt spent on host memory is a watt not spent on accelerators, LPDDR5X's efficiency is the whole point. The price you pay depends on the generation, and the two are not the same bargain. Grace's LPDDR5X is soldered: capacity is fixed at manufacture and a memory fault is a package fault, with no module to pull and replace. Vera changes that — NVIDIA packages its LPDDR5X in SOCAMM modules specified as detachable, field-replaceable and capacity-upgradable, which is the point of the design. Serviceability and error protection are separate axes in any case: Grace documents ECC and channel sparing, so compare the RAS features and the repair unit each SKU actually supports rather than assuming LPDDR means no server-grade protection.
The discrete x86 hosts use socketed DDR5, and increasingly MRDIMM (multiplexed-rank DIMMs) on the Intel side to claw back the bandwidth gap — Granite Rapids reaches MRDIMM-8800 against AMD's DDR5-6000/6400 on EPYC Turin. Socketed memory offers composability (compare Vera’s detachable SOCAMM on its own terms): you size capacity to the workload (multiple terabytes for embedding tables, large KV-cache, or in-memory retrieval indices), you get full server-grade RAS, and you field-service or upgrade memory without touching the accelerators. The cost is power and density — registered DDR5 burns more per gigabyte than LPDDR5X, and the DIMM slots consume board area and a slice of the rack's thermal and electrical budget. The decision mirrors the host-attach fork: take a Grace superchip and you take its soldered LPDDR5X bargain; Vera’s SOCAMM changes the repair unit; take the discrete host and you can spend power to buy capacity, bandwidth, RAS, and serviceability on your own terms.
Scope & caveats
900 GB/s is the bidirectional aggregate. Compare only against bidirectional PCIe figures (~128 GB/s Gen5 x16, ~256 GB/s Gen6 x16), giving ~7x and ~3.5x — not 7–14x.
Scope & caveats
GB200 NVL72 product profile; do not import a different Grace system’s 546 GB/s into this one.
System composition as a fleet decision, not a node decision
Composition is not a single answer. A real AI operation runs more than one workload shape, and the host-attach model, the GPU:CPU ratio, and the host memory that suit a pre-training trainer may differ from those of an agentic-inference fleet. So system composition is plural: compose distinct node and rack SKUs for the workloads that actually dominate your draw.
A training SKU may accept the coherent superchip ratio as shipped (2:1 on the GB200/GB300 NVL72 generation), profiled host application compute, host memory sized to the data pipeline not to the model, and every spare watt routed to accelerators. An agentic-inference SKU may need a larger host pool — many cores for orchestration and tool execution, large host memory for KV-cache and retrieval state, and a GPU:CPU ratio or separate CPU pool sized from CPU-seconds, memory and queueing evidence. A batch-inference SKU sits between them, throughput-bound and tolerant of a thinner host. Forcing one composition across mismatched profiles is how you end up simultaneously paying for unused host memory on the trainers and starving the agentic fleet of the CPU it needs. This is the same plurality the archetype framework demands at facility scale (Chapter 1.1) — here it lands inside the node.
Scope & caveats
CPU work excludes waiting. Unsupported host workloads, allocations and ranges are explained in the opening callout; Chapter 14.1 owns accounting. Neither CUDA profiling nor occupancy proves these inputs or p99.
CPU demand = 100 requests/s × 0.20 CPU-s/request = 20 busy cores. Provision ceil(20/0.70) = 29 cores, so the 32-core pool passes and 16 cores fail. Memory = 30 × 2.0 GB + 40 GB = 100 GB, within 128 GB. The 16-core service ceiling is 16 × 0.70/0.20 = 56 requests/s; the 32-core ceiling is 32 × 0.70/0.20, about 110 requests/s. At 120 requests/s, ceil(120 × 0.20/0.70) = 35 cores: the pool must grow or the CPU work must fall.
Scope & caveats
CPU demand = 100 requests/s × 0.20 CPU-s/request = 20 busy cores. Provision ceil(20/0.70) = 29 cores, so the 32-core pool passes and 16 cores fail. Memory = 30 × 2.0 GB + 40 GB = 100 GB, within 128 GB. The 16-core service ceiling is 16 × 0.70/0.20 = 56 requests/s; the 32-core ceiling is 32 × 0.70/0.20, about 110 requests/s. At 120 requests/s, ceil(120 × 0.20/0.70) = 35 cores: the pool must grow or the CPU work must fall.
Add the CPU pool in the assumed base case, then trace queues, scheduler locality and request latency before adding GPUs. At or below 56 requests/s the thin host clears this CPU screen and the added pool may buy no service benefit. Grace’s fixed memory and Vera’s replaceable SOCAMM have different repair consequences; choose their supported FRUs independently of the CPU arithmetic.
Method: NVIDIA CUDA best-practices profiling guidance. Chapter 7.11 owns the next handoff.
Deep dive: why the host starves the GPU (and how to see it before you buy)
The failure mode this chapter exists to prevent is GPU starvation by the host, and it is subtle because the symptom — low GPU utilization — looks like a workload problem, not a composition problem. The mechanism: an accelerator can only do work that has been staged for it. If the host cannot decode the next batch, run the tokenizer, service the retrieval call, execute the agent's tool, and launch the kernel fast enough, the GPU drains its queue and idles. On a training box, decoding, augmentation and input staging can bind too; a profiled host may instead have cycles to spare. On an agentic-inference box it can bind, because every request fans out into host-side orchestration and tool calls that contend for the same thin host the training composition specced.
You can see this before you buy by profiling the host-side critical path, not just the GPU. Measure the CPU time per request spent in tokenization, retrieval, sandbox execution, and orchestration, and the host memory footprint of the KV-cache and retrieval state. Separate CPU busy time from network, storage and tool wait. A CPU queue or saturated CPU service demand indicates a compute bottleneck; memory pressure and eviction indicate a capacity bottleneck. Either way, an inherited host configuration can strand accelerators. For a demonstrated CPU bottleneck, the fix is to move down the ratio (more CPU per GPU), widen the host (more cores), or fatten host memory — and to recognize that adding GPUs to a CPU-bound fleet buys nothing but idle silicon. The goodput discipline of Chapter 14.1 has a direct analogue here: delivered goodput under the service contract is the number that pays the bill.
Deep dive: coherent host memory as an HBM extension — what it actually buys
The headline case for the coherent superchip is that host LPDDR5X becomes a usable extension of HBM. The distinction between real value and marketing here turns on when host memory is actually reached. HBM is the scarce, expensive, capacity-constrained tier (Chapter 7.6) — a Blackwell Ultra (B300/GB300) GPU carries 288 GB, while an HGX B200 SXM GPU carries 180 GB, and a Grace host carries 480 GB of LPDDR5X. Across NVLink-C2C's ~900 GB/s coherent link, that host memory is reachable from the GPU as a second NUMA tier: slower than HBM, with capacity shared among the attached GPUs and, crucially, addressable without an explicit copy. For workloads whose working set exceeds HBM — long-context inference with multi-gigabyte KV-cache, wide-MoE serving where the full expert set dwarfs HBM, embedding-heavy recommendation, or in-memory retrieval over large indices — this lets you spill gracefully into a coherent tier instead of falling off a cliff to PCIe-attached host memory or to storage.
The discrete PCIe host reaches host memory across a far narrower path: a separate space at ~64–128 GB/s per x16, addressed by explicit DMA or, under supported CUDA Unified Memory and Linux HMM configurations, by automatic page migration. Automatic migration removes the hand-staging burden; it does not remove the bandwidth gap or the fault-and-migration overhead that hardware coherence avoids. Compare the two on addressability, coherence, migration cost, bandwidth and software prerequisites — not on whether a programmer has to write the copy. So the coherent model's payoff is real precisely for the memory-hungry inference workloads that are growing fastest — and that is the same pressure (CXL pooling and memory-semantic fabrics, Chapter 8.2) reshaping the memory hierarchy above the host. The spill benefit is zero when state fits in HBM; shared structures, synchronization and data staging can still exercise coherence. Buying coherence you do not exercise is the inverse error of starving a host you over-loaded — both come from specifying composition without profiling the workload.
Take the ratio, or compose it yourself
Underneath all of it sits one more decision: how much of system composition do you let the vendor make for you. The rack-scale coherent platforms (GB200/GB300 NVL72 today, Vera Rubin next) ship as integrated units with the CPU:GPU ratio, the host CPU, the host memory, and the interconnect all fixed at the factory. You buy a rack, not a node, and the composition is a given. That is genuinely valuable — the integration risk of welding 72 GPUs, 36 CPUs, NVSwitch trays, and a liquid-cooling loop into one coherent domain is enormous, and the vendor has absorbed it. But it means the GPU:CPU ratio, the host ISA, and the host memory are not yours to tune. If your agentic workload wants a fatter host than the package ships, the lever inside the rack is a different SKU, not a different host. The lever outside it is to stop treating the accelerator-attached host as the whole CPU tier: tool execution, RL environments, preprocessing and orchestration can run on a network-connected CPU pool sized independently — NVIDIA positions Vera as standalone CPU infrastructure as well as a GPU host. A fixed CPU:GPU ratio inside a superchip does not fix the CPU:GPU ratio of the service; what decides is latency, data locality, isolation and transfer volume.
The discrete path keeps composition in your hands. You build the node — choosing AMD or Intel, the core count, the socketed memory, the GPU count per host, the ratio — and you own both the flexibility and the integration risk. For a fleet with heterogeneous workloads and a procurement strategy that refuses to single-source, this is the path that lets composition track the workload. The 2026 reality is that most large operators run both: bundled coherent racks where the workload genuinely needs the coherent domain (training, memory-hungry inference), and composed discrete nodes where composability and dual-source supply matter more. The decision is not which path is correct in the abstract — it is which workloads belong on which path, and that lands you right back at the fleet-plurality discipline above. → procurement and fleet composition in Chapter 7.11.
Cite this chapter
Fehn, J. (2026). Host CPUs, GPU:CPU Ratios & System Composition (Chapter 7.8). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-8-host-cpus-gpu-cpu-ratios-and-system-composition (accessed 2026-09-29).
@misc{aidc-7-8,
author = {Fehn, Jacob},
title = {Host CPUs, GPU:CPU Ratios & System Composition (Chapter 7.8)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-8-host-cpus-gpu-cpu-ratios-and-system-composition},
note = {Accessed 2026-09-29}
}