Chapter 13.1
In this chapter · 5 sections
Commissioning Fundamentals, Levels & Program Governance
Commissioning converts design intent into evidence; an AI factory needs two interlocked acceptance tracks — facility and cluster — whose gates you sequence deliberately or the schedule sequences for you.
What you'll decide here
- Whether to fund a true Level 5 integrated systems test at scoping — staged load banks at scale, schedule float, and CxA scope paid for now — or accept that the integration risk moves onto a live, revenue-bearing cluster; Chapter 13.6 then supplies the capability ladder and the approved method for each required outcome.
- Which resilience model the IST must prove — service each Tier III capacity component/path without disrupting the accepted IT load, or prove Tier IV single-fault and compartment boundaries, including specified fire/flood isolation — because Chapter 12.1's contracted states write the failure-mode matrix.
- How the two parallel tracks — facility Cx and IT/cluster validation — interlock, and where the explicit overlapping gates sit (mechanical-Cx ↔ GPU burn-in; electrical acceptance ↔ load-bank IST ↔ first real workload).
- Who owns the program: an independent Commissioning Authority (CxA) engaged at design, not a general-contractor afterthought — and which governing documents (OPR, BOD, SOO) the whole acceptance chain traces back to.
- Which adopted standards spine you commission against — ASHRAE Guideline 0 and Guideline 1.1 for their process/HVAC&R scopes, Uptime Tier or BICSI 002 for resilience — and which owner-specified cluster criteria fill the gap; proposed Guideline 1.6P supplies no published requirement.
Commissioning (Cx) is the discipline that refuses to take the design's word for it. Every prior part of this guide makes claims — that the power chain will ride through a utility dip, that the cold plates will hold the GPUs under sustained load, that the fabric will carry a non-blocking all-reduce, that the standby plant will pick up the building before the UPS runs flat. Commissioning is the structured process of generating the evidence that those claims are true under the conditions that matter, before the facility carries revenue load. It is the bridge between Part 2's design and Part 14's day-2 operations, and it is the last point in the project where a latent integration defect costs hours to fix rather than a multi-day outage on a live cluster.
For a conventional enterprise data center this is a mature, well-codified ritual. For an AI factory it is not, and the gap is the subject of this entire Part. The reason is structural: an AI facility is two machines pretending to be one. There is the facility — substation, switchgear, generators, UPS/BESS, CDUs, chillers, piping, BMS — which the commissioning industry knows how to accept. And there is the cluster — tens of thousands of GPUs, the scale-up and scale-out fabric, storage, and the scheduler — which the facility commissioning standards barely acknowledge exists. The two tracks run in parallel, share critical resources (power and cooling), and must be interlocked at specific gates, or each track signs off in isolation against a load the other half cannot actually deliver. What follows sets out the levels, the two tracks, the roles, and the governing documents that the rest of Part 13 builds on.
Sequence the L1–L5 program
Mission-critical commissioning is organized as a ladder of increasing integration, conventionally Level 1 through Level 5 (with some programs adding a Level 0 design-review and a Level 6 post-occupancy / seasonal stage). Each rung tests a wider boundary than the last; you do not climb to the next rung until the current one is signed and its deficiencies closed. The ladder's rule is that integration faults are found at the lowest level at which they can possibly appear — a mis-wired CT caught at L3 standalone test is a morning's rework; the same fault discovered at L5 during a black-building test invalidates the run and resets days of scripted work.
Freeze the owner’s level names and integration boundaries in the contract; Chapter 13.6 supplies the sole L1–L5 capability ladder. The two parallel tracks (next section) each have their own analogue of this ladder, and Chapter 13.2 turns each rung into a quantitative script with explicit pass/fail gates.
Chapter 13.6 assigns the factory, installation, standalone, discipline and integrated evidence to the L1–L5 ladder. Budget that sequence on both tracks. Uptime Intelligence’s AI in Practice Paper 4: Level 4 & 5 Commissioning (May 2026) describes AI-aware banks on power and liquid loops, scriptable GPU power-versus-time curves, third-party-witnessed L5 and decommissioning procedures written during Cx. Those mechanisms belong in the fixture and witness scope before the schedule is frozen.
Two governance failures are expensive to recover from: deferring a discipline’s own failure tests into the integrated campaign, and truncating that campaign when schedule slips. Require the UPS transfer and battery ride-through evidence before the generator/cooling/BMS handoff test, and name the owner of each dependency in Chapter 13.6’s ladder. A generator can start and a cooling plant can ride through while the BMS still hands the load off in the wrong order. Calling the individual passes “good enough” pushes that integration risk onto day two.
Concurrent maintainability vs fault tolerance
The L5 test matrix is dictated by the resilience claim the facility was designed to meet, and the two claims that matter are the Uptime Institute's concurrent maintainability (Tier III) and fault tolerance (Tier IV). They sound similar and they are not, and commissioning is where the difference becomes physical.
Concurrent maintainability (Tier III) means any single capacity component or distribution path can be removed from service — for planned maintenance — with no impact on the IT load. The IST must therefore prove that you can take down each path in turn, one at a time, deliberately, and the load never notices. It does not require surviving an unplanned simultaneous fault. Fault tolerance (Tier IV) is the stronger claim: the facility survives any single unplanned worst-case failure — including the loss of an entire distribution path to fire or flood — with no load impact, while still being concurrently maintainable. Tier IV adds compartmentalization and continuous cooling requirements that Tier III does not, and the IST must demonstrate the auto-response to a fault, not just an orderly manual switchover.
The consequence for the commissioning program is direct: the resilience target you committed to in design writes your failure-mode demonstration script. A Tier III program scripts orderly, one-path-at-a-time maintenance scenarios. A fault-tolerant-topology program scripts approved, risk-assessed scenarios representing the required fault set, and continuous-cooling ride-through — a materially larger, riskier, longer IST. Choosing the topology is a design and capital decision (Chapter 12.1); commissioning proves each claim with the safest valid method: inspection, secondary injection, hardware-in-the-loop, simulation, staged load banks, or a specifically approved live fault test.
The two parallel tracks and how they interlock
What makes AI-factory commissioning different from every commissioning program before it is structural: there are two acceptance ladders running at once, on different schedules, owned by different organizations, governed by different standards — and they are coupled by shared physical resources.
The facility track (Chapters 13.3–13.6) is the classical mission-critical Cx program: utility energization, switchgear, generators, UPS/BESS, CDUs, chillers, piping, controls, culminating in the L5 IST. It is owned by the owner's CxA and the MEP trades, governed by ASHRAE Guideline 0 and Guideline 1.1 alongside Standard 202's process requirements (the data-center-specific Guideline 1.6P is still a proposal), Uptime, and BICSI 002. The cluster track (Chapters 13.7–13.9) is the IT validation program: fabric commissioning (BER, link-flap, topology, bandwidth/latency), GPU node burn-in and diagnostics, cluster-scale NCCL/collective acceptance, and a reference/proxy training run. It is owned by the platform/SRE/ML-infra team, governed by vendor deployment guides and de-facto standards like SemiAnalysis ClusterMAX — none of which the facility CxA typically touches.
If those two tracks each sign off in isolation, you have certified nothing useful. The facility track accepts a power-and-cooling envelope using load banks — usually air-cooled resistive banks, whose step granularity, control bandwidth and air-side heat rejection set the load they can actually present. The cluster track accepts compute using real GPUs that draw a violent, synchronized, switching load and reject heat into liquid. Neither track, alone, exercises the seam between them. The interlock gates are where that seam is tested.
| Interlock gate | Facility track delivers | Cluster track delivers | What the seam actually tests | Chapter |
|---|---|---|---|---|
| Mechanical-Cx ↔ GPU burn-in | Flushed, filled, leak-tested liquid loop; CDU at flow/temperature setpoint | Nodes drawing real heat flux into the cold plates | Whether the loop, CDU controls and worst-case branch hold at coincident duty; liquid banks test the loop, real nodes confirm die/contact behavior | 13.5 |
| Electrical acceptance ↔ load-bank IST | Energized power chain proven concurrently maintainable / fault tolerant | (Not yet present — emulated by load banks) | Redundancy topology under the declared steady and switched-load profiles; the fixture's step size and bandwidth bound transient coverage | 13.3 |
| Load-bank IST ↔ first real workload | Building proven against resistive/reactive/AI-emulating load banks | Proxy training run imposing real synchronized power and thermal dynamics | Residual power and thermal dynamics under a staged production-representative run; protective functions remain proven by injection/HIL or isolated controlled tests | 13.6 / 13.9 |
The middle row is the defining limitation of facility commissioning for AI. A conventional air-cooled resistive bank is a deliberately boring load: elements switched in steps, drawing a steady current at unity power factor and rejecting its heat straight to the room air. What it can and cannot reproduce is set by its step granularity, control bandwidth and heat-rejection path, not by the word 'resistive' — liquid-loop rejection and fast, repetitive waveforms need a bank specified for them. A frontier GPU cluster can exceed that conventional bank's profile — tens of thousands of accelerators that ramp from idle to full and back in milliseconds in lockstep across a synchronous step, imposing power transients and harmonic content that a step-switched bank does not reproduce, and dumping that heat into a liquid loop the load bank never touches. The facility track can prove the power chain is wired correctly and redundant; it cannot prove the chain absorbs the dynamic load swing of a real all-reduce, because the selected air-rejecting, step-switched fixture cannot generate that full waveform; another fixture earns only the coverage its measurements establish. This is the dynamic-load realism gap; its canonical treatment is in Chapter 13.6, the transient physics behind it lives in Chapter 4.5. The mitigation is sequencing: liquid-loop emulation followed by the staged proxy training run of Chapter 13.9 closes the gap, which is why it must come before — not after — go-live.
Deep dive: why facility Cx and cluster Cx cannot simply run independently
The intuitive program-management instinct is to treat the two tracks as independent workstreams — let the MEP CxA finish the building, hand over a 'powered shell with cooling,' and then let the ML-infra team bring up the cluster on top. For a low-density, air-cooled enterprise hall that instinct is mostly fine. For a liquid-cooled AI factory it produces three specific failures.
One: the load realism handoff is silent. The facility team signs off cooling against load banks that reject to air. The first time the liquid loop sees a real cold-plate heat-flux transient is when an expensive cluster is already racked and running — which means the worst-case-branch thermal-hydraulic behavior, the CDU control-loop stability under a real step change, and the leak-detection response under real pressure cycling are all being discovered in production. The fix is the explicit mechanical-Cx ↔ burn-in overlap (Chapter 13.5): the loop is accepted first with liquid banks, then with real nodes drawing heat to confirm the die-to-coolant path, by sequencing GPU burn-in to begin while mechanical Cx is still open.
Two: the power-transient handoff is silent. Electrical acceptance proves the redundancy topology under a smooth load. The dynamic-load-swing tolerance — the UPS/BESS and any rack-level energy storage absorbing a synchronized GPU power step — is never exercised until the proxy run. A cluster that passed every facility gate can still trip protection or sag a bus the first time 10,000 GPUs enter a collective in lockstep. The mitigation stack (BBU → BESS → in-shelf capacitance under realistic dynamics) is validated against declared fixture waveforms first, then the proxy run, per Chapter 13.6; software timing leaves a workload-specific residual.
Three: ownership gaps become finger-pointing. When a node throttles, is it a GPU fault (cluster team), a coolant-temperature excursion (facility team), or a CDU control-tuning issue (the seam)? Without an interlocked program and a shared deficiency log (Chapter 13.2), each track's commissioning record shows 'pass' and the defect lives in the gap between them. The governance answer is a single integrated commissioning schedule with named interlock gates and a CxA whose scope explicitly spans the seam.
Roles and the governing documents (OPR / BOD / SOO)
Commissioning is only as good as the requirements it tests against, and those requirements live in a short chain of governing documents that every acceptance script must trace back to. The chain is deliberately a chain — each document derives from the one before it, so that a pass/fail gate in an L5 script can be followed all the way back to a stated owner intent.
- Owner's Project Requirements (OPR). The owner's intent in measurable terms: the availability target (and therefore the Tier or BICSI class), the density and ramp the building must accommodate, the environmental envelopes, the maintainability expectations, and the acceptance criteria the owner will hold the project to. Everything downstream is an answer to the OPR. For an AI factory the OPR must state the workload intent (training-shaped vs inference-shaped, per Chapter 1.1), because that is what makes the difference between commissioning for goodput and commissioning for availability.
- Basis of Design (BOD). The design engineer's documented explanation of how the proposed systems satisfy each OPR requirement — the topology, the redundancy scheme, the cooling architecture, the setpoints, and the assumptions. ASHRAE Guideline 1.1-2025 covers applying the commissioning process to new HVAC&R systems, and is where the BOD's development and verification is set out; it is not a BOD-only standard. The CxA verifies the BOD answers the OPR before construction, not after.
- Sequence of Operations (SOO). The control-logic specification — exactly how the BMS/EPMS/DCIM is supposed to behave in every normal, maintenance, and failure mode. The SOO is the script-writer's bible: an L4/L5 functional test is, in essence, a line-by-line proof that the real controls match the SOO under real conditions. A vague or incomplete SOO is the most common root cause of a failed IST, because there is nothing precise to test against.
The role that owns this chain is the Commissioning Authority (CxA) — and the governance decision that matters most is to engage the CxA at the design phase, as an independent party, not as a general-contractor self-check bolted on at the end. ASHRAE Guideline 0 (the process), ASHRAE Standard 202 (the formalized commissioning process / Cx-Process), and the proposed data-center-specific Guideline 1.6P — on ASHRAE's register as a proposed guideline, not a published one — provide process guidance with different scopes; this guide recommends an independent CxA verifying the OPR→BOD→SOO chain across the whole project lifecycle, from design review (L0) through post-occupancy. An owner who engages the CxA only to 'witness the IST' has already lost the design-phase reviews where the cheapest defects are caught.
Give each interlock an evidence owner and a release owner. The electrical lead supplies bus waveforms and relay events; the mechanical lead supplies branch flow, temperature, pressure and actual protective response; the platform lead supplies device health, workload correctness and recovery. The CxA joins those records against OPR/BOD/SOO requirements; operations accepts the restored state and the owner names the person authorized to release load. These are example responsibilities, not signatures. Record the installed rack interfaces, utility boundary and OT access/restore obligations from Chapter 11.11 in the same applicability register. A missing vendor result stays open under its named owner instead of disappearing between two green dashboards.
Scope & caveats
Secondary-source whole-facility planning estimate quoted by Savills in May 2024; no disclosed estimating population or method. No universal multiplier follows from Uptime Tier criteria.
Single-source planning heuristic: Savills (May 2024) attributes it to Dgtl Infra, which publishes no sample, geography, density or estimating method. The older ~10–25% inverts a McKinsey 2011 statement (10–20% saving moving Tier IV→III). Uptime Tiers are outcome-based, so no universal cost multiplier follows from the standard; re-estimate against the actual design.
Scope & caveats
Register observation, not a new publication date. Adopted editions determine project applicability; Guideline 1.1 covers new HVAC&R commissioning.
Scope & caveats
Whole-job interruption events for one named Meta run (Llama 3 405B, 16,384 H100s, 54 days). Not a per-GPU MTBF and not a facility-availability figure — the paper does not report facility availability.
The paper's attribution percentages do not reconcile against its own printed counts: Table 5 lists 148 faulty-GPU and 72 HBM3-attributed events among the 419 unplanned interruptions (35.3% and 17.2% of that base), so the quoted ~78% hardware and 58.7% GPU shares are not shares of the same 419 denominator this tile values. Use the counts, not the percentages, and state your denominator.
Why the AI factory breaks the conventional program
A conventional data-center commissioning program is a single, well-standardized ladder culminating in an IST against load banks, governed end-to-end by ASHRAE/Uptime/BICSI, signed by an independent CxA. That program, applied unchanged to an AI factory, certifies a building that has never seen its actual load and a cluster the standards never mention.
The three structural breaks are: (1) two tracks, not one — facility and cluster, on different schedules and standards, requiring explicit interlock gates; (2) load realism — the air-rejecting, steady load banks in a conventional scope cannot reproduce the dynamic power and liquid-side thermal behavior of real GPUs; switched and liquid fixtures extend that coverage, so a staged, product-representative run closes workload-specific normal-operation evidence that facility emulators do not cover; its scope and acceptance criteria follow the approved project risk and test plan; and (3) density ramp — GB300 NVL72 is deploying through 2026 alongside GB200, Vera Rubin entered full production in August 2026, and the announced ~600 kW Rubin Ultra/Kyber planning point follows in H2 2027, so the IST must validate not only the steady state but the headroom the ramp will consume (floor, water, electrical, cooling-plant turndown). A program that handles these three is a Part 13 program; one that ignores them hands a long deficiency list to day-2 operations — exactly the failure environment the 419-interruptions-in-54-days reality of a real training run cannot afford.
Deep dive: the commissioning schedule is a redundancy problem, not just a calendar
One subtlety that separates AI-factory commissioning from a fresh greenfield enterprise build: much AI capacity is energized in stages, with live blocks already carrying load while later blocks are still being commissioned. That turns the Cx schedule into a redundancy-engineering problem. An L5 IST that pulls utility or fails a generator on a shared distribution path can threaten the live blocks unless the test boundary is drawn to preserve their redundancy throughout. This is why staged energization (Chapter 13.10) and the IST sequence (Chapter 13.6) are co-designed: you cannot run a destructive integrated test on a path that a revenue-bearing block depends on without first proving you can isolate it.
The governance consequence is that the CxA and operations must agree, in writing, on which faults the IST is permitted to inject at which stage, and what the live-block protection is during each. A program that treats commissioning as a pre-occupancy gate that finishes before any load arrives does not match how AI capacity is actually brought online — it arrives block by block, against a power-bound interconnection clock that rewards energizing early and commissioning around live load. → staged ramp and the Operational Readiness gate in Chapter 13.10.
Choose the interlocked program whose required states have a valid test method, an evidence owner and a funded slot in the schedule. Finishing facility and cluster checklists separately is cheaper to organize, but the first synchronized workload then becomes an unscheduled test of their shared power, cooling and controls.
Cite this chapter
Fehn, J. (2026). Commissioning Fundamentals, Levels & Program Governance (Chapter 13.1). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-13-commissioning-and-go-live/13-1-commissioning-fundamentals-levels-and-program-governance (accessed 2026-09-29).
@misc{aidc-13-1,
author = {Fehn, Jacob},
title = {Commissioning Fundamentals, Levels & Program Governance (Chapter 13.1)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-13-commissioning-and-go-live/13-1-commissioning-fundamentals-levels-and-program-governance},
note = {Accessed 2026-09-29}
}