The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Part 7

Part 7

Compute, Silicon & System Integration

15 chapters

7.1
Accelerator Landscape & Taxonomy
Choosing an accelerator commits you on four independent axes — ownership and access, execution architecture, workload specialization, and memory semantics — and the four reference groupings (merchant GPU, systolic TPU, hyperscaler XPU, inference ASIC) are points in that space, each carrying its own software stack, scale-up fabric, and lock-in horizon.
7.2
NVIDIA Accelerators: Hopper → Blackwell → Vera Rubin → Rubin Ultra → Feynman
NVIDIA publishes an annual platform roadmap; the purchase can include the chip, rack or pod, and each boundary changes the power, cooling, fabric and refresh work you inherit.
7.3
AMD Instinct & the Open Challenger
AMD’s Instinct is a credible second source once the quoted system clears memory, software, fabric and delivery gates; its useful-output benefit has to pay for the ROCm porting and recurring release support that the datasheet never prices.
7.4
Hyperscaler XPUs: TPU, Trainium/Inferentia, Maia, MTIA
A hyperscaler XPU leaves compiler releases and capacity allocation under the cloud’s control; choose the offered service when its workload economics pay for an exit that can require both a port and replacement capacity.
7.5
Custom ASICs & the Merchant-Silicon Disruption
Custom silicon pays only past a sustained-demand threshold where qualified lifetime savings repay design, software and delay costs; below it, the NRE and lead time buy a chip your roadmap has already obsoleted, and a workload change during design can strand the tape-out and compiler investment before volume repays them.
7.6
HBM: The Binding Constraint on AI Compute
In 2026 HBM, not the GPU die, binds AI compute — allocated years ahead and dominating the accelerator bill of materials — so size it for the model, KV state and runtime reserve, test its sustained bandwidth against the workload, and qualify memory and packaging allocations separately before relying on delivery.
7.7
Advanced Packaging & the Integration Substrate
Packaging yield on the largest interposers still caps how many accelerators ship, so choose a package whose memory layout, die crossings, thermal path and yield can meet the workload and delivery plan; an interposer roadmap is not a qualified accelerator.
7.8
Host CPUs, GPU:CPU Ratios & System Composition
Host throughput, memory and attach architecture decide whether accelerators stay fed; name the host/socket/package denominator in every CPU-to-accelerator ratio, then size the host and independent CPU tier from measured work and latency.
7.9
Software Ecosystems & Lock-In
Buying an accelerator commits you to CUDA, ROCm, XLA or Neuron; kernel ports, collective tuning and continuing releases are switching costs the datasheet never priced, so measure them with the useful output of the executable path.
7.10
Precision, Quantization & the Compute-Memory Tradeoff
Fewer bits can spend quality headroom before a service reaches its quality floor; count scale metadata and runtime state in memory savings, and credit throughput only to the hardware’s supported native path.
7.11
Accelerator Selection, TCO & Procurement Strategy
Select the complete accelerator configuration that clears memory, useful-output, power and delivery gates; the RFP fixes those obligations and hands qualified costs and service operands to Chapter 1.8.
7.12
On-Package Power Delivery & Power Integrity
At a sub-1 V core rail carrying roughly 2,000 A, a short resistive path still wastes I²R power and excess droop throttles the accelerator; budget the upstream 48 V shelf and the propagated transient separately.
7.13
The Rack as Integration Unit
The rack is now the unit you buy, ship, power, cool, cable and certify as one object, so buying one commits the floor, busbar, manifold and cabling as well as the equipment; every populated and reserved interface must fit, or a refresh inherits building work that replacing the rack alone cannot supply.
7.14
Server & System Integration
Where you enter the integration pipeline — DGX, HGX-OEM, ODM-direct, or OCP self-design — sets who owns burn-in, the acceptance gate, and the RMA, and the days between a powered shell and a producing cluster.
7.15
Deployment Velocity & Cabling at Scale
Time-to-goodput is set on the floor by how fast racks land and links light; pre-terminated cabling and off-line optics screening are the two levers that keep mis-cabling off the acceptance critical path.