Guide › Part 1
Part 1
Strategy, Workload Archetypes & Economics
8 chapters
1.11.21.31.41.51.61.71.8
The Archetype Decision Framework: Workload Is the Master Variable
The workload you intend to run deterministically sets density, cooling, fabric, redundancy, and siting; fix the archetype before any subsystem decision, because most mis-scopes cannot be undone without re-pouring concrete.
Training Data Centers: Synchronous, Dense, Checkpointable
A training cluster is one synchronous supercomputer moving at its slowest GPU's pace; design for goodput per megawatt — dense, liquid-cooled, non-blocking, checkpointable — before you cut steel.
Inference Data Centers: Bursty, Distributed, Always-On
An inference data center serves many independent requests against a latency SLO — always-on, close to users, and sized to goodput-per-dollar and tokens-per-watt at real batch sizes.
Post-Training, Fine-Tuning & RL: The Hybrid Middle
Rollout generation — autoregressive inference — dominates post-training cost, so the cluster is a fleet of inference engines feeding a small trainer; scope each half to its own profile.
Edge Inference & Distributed Micro-Datacenters
Edge inference trades cheap power for proximity to the user; it pays only when the latency budget is physical — one a centralized region cannot meet.
Procurement Archetypes: Build vs Buy vs Rent
Procure capacity to the durability of your demand: own the durable, contracted load; rent the uncertain load, paying rental's premium as the price of the option to walk away.
The Requirements-and-Consequences Matrix
Name the workload archetype and the subsystem commitments follow in sequence; this matrix turns that one input into a numbered, signed design basis for cooling, fabric, storage, redundancy, and siting.
Business Models, Economics & ROI
Four numbers decide an AI data center's return — capex per watt, assumed depreciation life, achieved utilization, and post-deflation price — and any one, wrong, strands the asset.