Guide › Part 13
Part 13
Commissioning & Go-Live
10 chapters
13.113.213.313.413.513.613.713.813.913.10
Commissioning Fundamentals, Levels & Program Governance
Commissioning converts design intent into evidence; an AI factory needs two interlocked acceptance tracks — facility and cluster — whose gates you sequence deliberately or the schedule sequences for you.
Documentation, Scripts & Acceptance Test Plans
Every commissioning test needs a pre-agreed observable gate and a witnessed result; an AI building earns acceptance when facility and cluster evidence close the same power, cooling and workload boundary.
Electrical Power Acceptance (L3/L4)
Electrical acceptance forces the paper power chain to prove, with instrumented evidence, that it holds a GPU cluster's synchronized megawatt-scale swings — a load no static load bank reproduces.
Commissioning On-Site Generation & Microgrid Controls
Behind the meter you are the grid, and commissioning proves that generation, controls and storage can sustain the declared AI load, including the campus-wide duty for a gigawatt-class site, through the required source states and the instant the utility lets go.
Cooling Acceptance: Air, Liquid-to-Chip & CDU Commissioning
Commission the cooling protection chain in layers: surrogate load and injection/HIL prove controls and fail-safe actions, liquid-cooled load banks on the manifolds prove the liquid path, and staged normal workloads close operating evidence without making live GPUs the fault target.
Level 5 Integrated Systems Testing (IST) & Failure-Mode Demonstration
IST proves design-basis faults at a boundary whose applicable permissions and methods are agreed with the utility, AHJ, insurer and Cx authority; controlled trips, injection/HIL and load banks prove the utility-loss chain, and staged workloads close the fixtures' remaining operating evidence.
Network Fabric Commissioning & Validation
Fabric defects surface quietly, one marginal optic and one mis-cabled rail at a time; every defect you fail to screen at layer 1 returns as a straggler or a stalled all-reduce in production.
GPU Node Burn-In, Diagnostics & Stress Validation
Burn-in forces a GPU fleet's infant-mortality defects into a bounded acceptance window under the selected synthetic stress before a bad node can restart a synchronous job; release follows executed diagnostic coverage, correctness and thermal evidence for that hardware/software cohort.
Cluster-Scale Benchmarking, Reference Training & Storage/Scheduler Validation
A cluster's performance is accepted through the production scheduler and storage: training needs ML Productivity Goodput measured over a sustained reference training run, inference needs correct responses within its latency contract, and neither can average away a failed safety, correctness, isolation or recovery gate.
Staged Power/Load Ramp, Go-Live & Handover to Operations
Go-live ramps megawatts and synchronized GPU load in stages through an operational-readiness gate; the failure modes are outpacing what grid and cooling can absorb, and handing over before operations is ready.