The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Commissioning & Go-Live › 13.2

Chapter 13.2

In this chapter · 6 sections
Term help

Documentation, Scripts & Acceptance Test Plans

Every commissioning test needs a pre-agreed observable gate and a witnessed result; an AI building earns acceptance when facility and cluster evidence close the same power, cooling and workload boundary.

GOODPUTDENSITY-RAMP

What you'll decide here

  1. Whether each script carries a hard, pre-agreed numerical or categorical acceptance gate, uncertainty rule, named witness and redline-on-fail rule — or a soft 'engineer's judgement' clause that turns a disputed result into a change order.
  2. Where the facility acceptance boundary stops and cluster acceptance starts — and who owns the seam between air- or liquid-rejecting fixtures and real GPU transients, using Chapter 13.6's capability ladder.
  3. What instrumentation and data-acquisition basis the pass/fail gates are read against — calibrated reference instruments and a captured time-series, or the building's own BMS sensors marking their own homework.
  4. How deficiencies are classified (A/B/C severity) and which classes are go-live blockers versus warranty-list items — because the punch list, not the script binder, is what actually gates handover.
  5. Whether you capture a signed baseline 'fingerprint' of every subsystem at acceptance, so day-2 drift has a reference — or accept that the first time you characterise the plant is the day it misbehaves.

Chapter 13.1 set the governance frame — the L1–L5 ladder, the two parallel facility and IT tracks, and the governing documents (OPR/BOD/SOO) that everything traces back to. Here governance becomes execution — the actual documents the commissioning agent writes, the scripts the field runs, and the gates that decide whether a system passes. This is the least glamorous material in Part 13, and it most reliably separates a building that goes live on schedule from one that slips a quarter while two firms argue about whether a 14-second generator transfer was a pass.

The question that recurs at every level is the same: did you write an observable, pre-agreed acceptance gate, or did you leave it soft? A script that says 'UPS shall transfer to battery without dropping the load' is unfalsifiable — what is 'the load' under no real load, and what voltage sag counts as a 'drop'? A script that says 'on loss of utility, bus voltage shall not sag below 90% nominal for more than 10 ms, verified against a calibrated power-quality analyser logging at ≥ 10 kHz, witnessed by Owner and CxA' makes the decision explicit, but those illustrative thresholds become a contract only after the owner freezes the terminals, waveform/RMS convention and uncertainty rule against the equipment limits. Chapter 13.3 supplies a fully disclosed source-transfer record. That precision decides whether a failing result is a deficiency the contractor fixes on their dime or a 'disagreement' that becomes a change order on yours. Price commissioning from the witness hours, temporary load, instrumentation and retest scope; the rework and downtime it prevents justify that expense only if the gates are hard enough to enforce. → Chapter 13.1.

Anatomy of a commissioning script

A commissioning script (variously a test procedure, ATP step, or functional-performance test) is not a checklist. A checklist asks 'did you do X?'; a script asks 'when you do X, does the measured result fall inside the gate?' Every well-formed script has the same eight fields, and the absence of any one of them is where disputes are born:

  • Unique ID and traceability — back to a specific OPR/BOD requirement and a SOO step, so a passing script proves a design intent, not just an action.
  • Pre-conditions — the exact system state, isolations, and safety lockouts that must hold before the step runs (the field's single most common shortcut, and the one that injures people).
  • Procedure — the numbered actions, written so a competent technician who has never seen the plant can execute them identically.
  • Expected result with an observable gate — a number, tolerance and unit, or a categorical result such as the correct breaker position or denied access. For a numerical gate, name the terminals, time window and uncertainty rule; 'within spec' leaves all three undecided.
  • Instrumentation and DAQ basis — which calibrated instrument reads the result, its accuracy class, and its in-date calibration certificate.
  • Actual result — the recorded measurement, with a timestamp and the captured waveform/trend reference, not a tick.
  • Pass / Fail / Deferred determination — decided against the gate, not against an opinion in the room.
  • Witness signatures — CxA, Owner's representative, and the responsible contractor, each signing that they observed the result, not that they trust it.

Whatever else varies, the gate is agreed and frozen before the test is run. The instant you negotiate the acceptance threshold while looking at a failing result, you have lost the leverage commissioning exists to give you. Freeze the gates at script-review (a formal owner sign-off of the procedures, weeks before energization); after that they are a contract, not a draft.

Facility ATP vs cluster ATP: two acceptance boundaries, one seam

An AI factory has two acceptance regimes that meet at a seam, and most program risk lives in that seam. The facility ATP (the L3/L4 mechanical/electrical track) proves the building: switchgear, generators, UPS/BESS, chillers, CDUs, pumps, BMS, under load-bank load, culminating in L5 integrated systems testing. The cluster ATP / SAT (the IT track) proves the machine: node burn-in, fabric BER and bandwidth, NCCL collectives, storage throughput, scheduler behaviour, and a reference workload — accepted against goodput, not against load-bank kilowatts. They run on different schedules, against different standards, witnessed by different parties, and they are not interchangeable. → facility electrical in Chapter 13.3, cooling in Chapter 13.5, IST in Chapter 13.6; cluster burn-in in Chapter 13.8, fabric in Chapter 13.7, benchmarking in Chapter 13.9.

The seam exists because of a physics gap that is canonical to AI commissioning: the facility ATP exercises the building with a load the real workload does not resemble. An air-rejecting resistive load bank held flat draws a unity-power-factor, thermally-steady load; switched banks add declared steps and liquid banks put heat into the loop. A synchronous training job draws a spiky, microsecond-scale, multi-megawatt-swinging load that rejects its heat into cold plates and a liquid loop. A flat, air-rejecting fixture can prove a declared steady load — 100 MW is an illustrative scope — but leaves UPS/BESS transient response and CDU branch heat untested. A switched fixture adds only its measured waveform coverage; a liquid-connected fixture adds only its measured branch duty and control bandwidth. Real nodes still confirm the die/contact path and software timing. That gap makes two tests essential: liquid-loop emulation first, then a staged SAT workload — acceptance work in their own right, rather than 'IT validation' tacked on at the end. The dynamic-load realism problem is the canonical subject of Chapter 13.6; the cooling load-realism limit is engineered in Chapter 13.5.

Facility ATP vs cluster ATP/SAT — what each can and cannot prove
DimensionFacility ATP (L3–L5)Cluster ATP / SAT (IT track)
Object under testPower, cooling, BMS — the buildingNodes, fabric, storage, scheduler — the machine
Applied loadResistive/reactive/AI-emulating load banksReal GPUs running burn-in + reference workload
Heat path exercisedAir or liquid rejection as specified; measure the actual fixture heat splitReal heat into cold plates and the full liquid loop
Primary acceptance metrickW held, °C delta-T, transfer ms, leak-free holdBER, busbw (GB/s), goodput %, SDC count, FIO IOPS
Governing standardsASHRAE Gd 0 / DC Cx guideline, Uptime, BICSI 002Vendor RA, ClusterMAX-class criteria, NCCL/MLPerf
Can provePlant holds declared steady/transient profiles; required redundancy statesHardware health, fabric integrity, real-workload goodput
Cannot proveExact GPU software timing and die/contact behavior outside the fixture's qualified envelopeFacility ride-through under utility loss (needs the plant)
The two acceptance regimes of an AI factory. The right-hand 'cannot prove' column is why the seam between them is where program risk concentrates.

Instrumentation and data acquisition: who reads the gate

A pass/fail gate is only as trustworthy as the instrument that reads it, and the recurring sin is letting the building's own BMS grade its own homework. The facility's permanent sensors are installed for control and trending, not for metrology: a BMS temperature point may carry ±1–2 °C uncertainty and a multi-second poll interval, which is useless for accepting a delta-T gate of ±1 °C or a transfer gate of ±10 ms. Acceptance reads against calibrated reference instruments — power-quality analysers, thermal imagers, ultrasonic and Coriolis flow meters, calibrated PT/RTD references, micro-ohmmeters — each with an in-date NIST-traceable (or national-lab-traceable) calibration certificate attached to the script. A result without a certificate behind the instrument is an anecdote.

Two DAQ decisions distinguish a serious program. First, sample rate must out-resolve the phenomenon: a UPS transfer or a generator pickup is a sub-100 ms event, so logging at hundreds of Hz to tens of kHz is required to even see the sag you are accepting against — a 1 Hz BMS trend will report 'no anomaly' through a transient that breached spec. Second, capture the full time-series, not the summary statistic: store the waveform and the trend, not just 'min 89.2%'. The captured series is what lets you adjudicate a disputed result after the fact, and it doubles as the baseline fingerprint discussed below. On AI factories this matters more than on legacy IT halls precisely because the loads are transient: the interesting failures live in the milliseconds, and a DAQ basis that cannot see milliseconds cannot accept against them. → fabric timing acceptance (PTP/IEEE-1588) as its own metrology problem in Chapter 8.7.

Freeze the evidence rule with the gate. For a lower limit, subtract the declared measurement uncertainty; for an upper limit, add it. A result that overlaps the limit is HOLD for better measurement or correction, not a rounded pass. Specify waveform channels, RMS window, event-trigger timing, clock alignment, missing-sample handling and artifact hashes. The source-transfer record in Chapter 13.3 shows this rule; the installed-fabric record in Chapter 13.7 uses the same evidence ownership and restoration discipline.

1e-12
Legacy ibdiagnet v2.13.0 --ber_test default; later devices use the supported PHY path
Scope & caveats

The 10^-12 default belongs to --ber_test for SwitchX, ConnectX-4 and ConnectX-3. The manual directs later devices to --get_phy_info; qualify BER/FEC fields and limits under the installed interface profile in Chapter 13.7.

72–168 hrguidance
2025 practitioner example: 72–168 hr GPU node soak; not a universal acceptance bound
Scope & caveats

Together AI and ClusterMAX described a 72–168 hr range in 2025 practitioner guidance. The project/OEM/contract test plan sets duration, intensity, statistical stopping rule, pass/re-soak gate, and vendor disposition.

90% vs 96% scenariomodeled
training-goodput sensitivity scenario: 90% vs 96% (illustrative — replace with the named fleet's measured goodput)
Sep 2026Guide analysis — stipulated sensitivity scenario; no claim of an industry measurement.register ↗
Scope & caveats

Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.

45 °C maximum liquid inlet; 65 °C maximum liquid return (separate limits)
QCT GB200 NVL72 QoolRack reference maxima: 45 °C liquid inlet and 65 °C liquid return; separate limits, not a selected operating pair
Scope & caveats

Exact QCT reference. The 45 °C liquid-inlet maximum and 65 °C liquid-return maximum are separate limits, not a prescribed 20 K operating rise. Select a supported operating point, approved fluid, liquid heat load, and design ΔT; ASHRAE W45 describes FWS supply capability, not this product's setpoint.

Separate acceptance maxima, not a prescribed 20 K operating rise; do not attribute these limits to HPE without an HPE document that states them.

~115 / ~17 kW
HPE GB200 NVL72 heat split (~115 kW liquid / ~17 kW air at 132 kW nominal) — the liquid-side load an air-rejecting facility load bank leaves untested
Scope & caveats

HPE GB200-specific. Keep separate from the GB300 NVL72 split: Lenovo Press LP2357 puts GB300 on ~90% liquid / ~10% air — roughly 13.5 kW on air at 135 kW rack TDP and ~15.5 kW at the ~155 kW peak, with the NVLink switch trays moved fully to liquid. Size a residual-air path from the rack you actually name.

Deficiency and punch-list management: the document that actually gates go-live

The binder of passed scripts is the visible deliverable; the deficiency log is the one that decides whether you go live. A serious program treats every failed or partially-passed step as a tracked deficiency with an owner, a root cause, a corrective action, a re-test reference, and a severity classification — and the severity classification is the lever. A flat punch list where a mislabelled valve sits at the same priority as a failed UPS transfer guarantees that go-live becomes a negotiation about which items 'really' matter, conducted under schedule pressure. Pre-agree the severity tiers and which tiers block.

  • Class A (blocker) — a life-safety defect or a failure of a core redundancy/ride-through claim. Go-live cannot proceed until closed and re-tested. Example: failed automatic transfer to generator; an EPO that does not trip; a leak-detection interlock that does not isolate.
  • Class B (conditional) — a real deficiency that does not defeat the design basis. Go-live may proceed only on a written, time-limited concession: a dated corrective-action plan, an owner, and a named person who accepts the residual risk until it closes. Example: a single redundant pump trending warm; a BMS alarm mis-mapped but functional.
  • Class C (warranty/punch) — items with no safety or operational impact, which roll to the warranty list. Classify by impact, not by category: an equipment or isolation label a technician needs to switch safely, and an as-built or single-line a responder needs at 3 a.m., are Class A or B however much they look like documentation. Example: cosmetic paint damage; a formatting correction on a non-operational document.

The consequence of getting the tiering wrong cuts both ways. Tier too loosely and you carry a Class-A ride-through gap into live operation, where the first real utility loss finds it. Tier too strictly and you hold a live-block go-date hostage to a paint scratch. Fix the tiering rules and the blocker list in writing when you freeze the gates — before anyone has a result to argue about. Zero open Class-A items, every open Class-B under a signed time-limited concession, and a complete safety-critical label and as-built set are the real go-live gate, and they feed directly into the Operational Readiness review and the handover package. → handover and the Operational Readiness gate in Chapter 13.10.

Baseline 'fingerprint' capture: acceptance as the birth of day-2

The most valuable artifact commissioning produces is one that has no pass/fail gate at all: the baseline fingerprint. At the moment a system is accepted, it is in its known-good state — clean filters, balanced flows, calibrated sensors, fresh firmware, characterised transients. Capture that state quantitatively and you have given day-2 operations a reference against which all future drift is measured. Skip it and the first time anyone characterises the plant is the day it misbehaves, with nothing to compare against.

A useful fingerprint is multi-domain and time-stamped: the captured transient waveforms from every transfer test; the as-accepted pump/fan curves and flow balance; per-rack and per-branch coolant flow and delta-T at known load; thermal images of every switchgear connection and busbar joint; PUE/WUE at the commissioned load point; per-node power-draw signatures and HBM/ECC baselines from burn-in; per-port BER and per-link bandwidth from the fabric; and NCCL busbw and goodput from the reference run. Together they seed the operational digital twin and the day-2 reliability program. Anomaly detection, predictive maintenance, and lemon-node ejection all need a 'normal' to deviate from, and acceptance is the first controlled observation of the system inside its tested envelope. → the operational twin and telemetry handoff in Chapter 14.2; the goodput baseline carried into operations in Chapter 14.1. Note the fingerprint is distinct from the design-validation digital twin of Chapter 2.7 — that one predicts behaviour pre-build; this one records measured reality at acceptance.

Deep dive: writing a UPS-transfer gate that survives the room

Consider the most frequently disputed facility script: loss-of-utility ride-through. The soft version — 'on utility loss, UPS shall support the load without interruption' — fails the moment the result is anything but obviously clean, because every term is undefined. Here is the same step as a hard gate, with the field limits still to be frozen:

Pre-conditions: use an approved isolated test boundary or test source with the specified load-bank profile; verify protection, backfeed prevention, communications, authority, hold points, and abort/restore criteria; place calibrated PQ instrumentation at the named buses; require the CxA, owner, electrical safety authority, and system operator roles identified by the project plan. Procedure: initiate the approved loss-of-source simulation or controlled transfer at the defined boundary; observe pickup and re-transfer without exposing an unapproved live load. Gate: apply the project SOO and equipment ride-through limits, including declared sampling/aggregation rules, to the captured waveform and event logs. Determination: pass only against the approved criteria; abort safely on any protection, stability, personnel, or boundary violation.

For AI factories specifically: holding the load bank flat makes this a steady, well-behaved load, so passing it proves that operating point. The hard case — a real multi-MW GPU power swing during the same transfer — is outside a flat bank profile; a switched fixture can present the specified step within its measured capability, which is why this script's gate must be read alongside the dynamic-load realism analysis of Chapter 13.6 and the electrical transient physics of Chapter 13.3. A green checkmark here is necessary, not sufficient; the script binder must say so explicitly so no one reads facility acceptance as workload acceptance.

Deep dive: digital Cx platforms — what they fix and what they cannot

Paper-and-PDF commissioning is collapsing under the document volume of a multi-hundred-MW AI campus, and digital Cx platforms (CxPlanner, Bluerithm, ProjectSight and peers) are now standard on hyperscale builds. What they genuinely fix: a single source of truth for thousands of scripts; templated, reusable test procedures that enforce consistency across identical blocks; real-time deficiency tracking with severity, owner, and re-test linkage; automated rollup of pass/fail status to a live program dashboard; mobile field execution with photo/waveform attachment at the point of test; and auto-generated turnover packages. On a campus where the same NVL72-block script runs hundreds of times, templating alone removes a class of transcription error that paper guarantees.

What they cannot fix, and must not be mistaken for: a digital platform makes a soft gate just as fast to sign off as a hard one. The tool enforces process completeness, not measurement rigour — it will happily collect a thousand signed scripts whose gates are 'satisfactory'. The platform is a force multiplier on whatever discipline you bring to the script content; bring soft gates and you have merely digitised the dispute. The rigour has to exist before the tool: freeze hard gates and a severity taxonomy first, then let the platform scale their execution. AI-assisted script generation (now appearing in these platforms) sharpens the warning — a generated procedure can read fluently and still ship an unfalsifiable gate, so the human review that converts every gate to an observable measurement or state remains the load-bearing step.

Sequencing: how scripts interlock across the program

Scripts form a dependency graph: a passing downstream script is only valid if its upstream prerequisites passed first. Electrical acceptance (L3/L4) must clear before integrated systems testing can apply real building load; cooling acceptance and the secondary-loop flush must clear before any GPU draws power into a cold plate; fabric BER and bandwidth must clear before NCCL collectives mean anything; node burn-in must clear before a reference training run is interpretable. The program is therefore a sequenced set of gates, each unlocking the next, with two deliberately overlapping seams that 13.1 flagged: mechanical-Cx ↔ GPU burn-in (liquid banks load the loop before real nodes confirm cold-plate/contact behavior) and facility-IST ↔ first-real-workload (a staged proxy/reference run is one product-representative source of workload evidence; dynamic/AI-emulating and liquid-cooled test loads close other portions without risking production hardware). Treat those overlaps as a single coordinated gate with shared acceptance criteria, not as a clean hand-off, or each side will accept to its own boundary and the seam will go untested. → the staged power/load ramp that walks these gates live in Chapter 13.10; the design-basis redundancy definitions the topology scripts validate against in Chapter 0.5.

Release against the frozen requirement, captured result and restored configuration together. A signed waveform without its terminals and uncertainty can settle nothing; a signed checklist without a retest dependency can describe equipment that no longer exists. Spend the witness time on that join before the revenue date makes every missing field a dispute.

This chapter is the documentation-and-gates layer beneath the governance frame of Chapter 13.1. The scripts it describes are executed by domain in the chapters that follow: electrical power acceptance in Chapter 13.3, cooling/CDU acceptance in Chapter 13.5, Level-5 IST and the dynamic-load realism gap in Chapter 13.6, network fabric in Chapter 13.7, GPU burn-in in Chapter 13.8, and cluster-scale benchmarking/goodput acceptance in Chapter 13.9. The fingerprint feeds the day-2 telemetry and twin of Chapter 14.2 and the goodput program of Chapter 14.1; the redundancy-topology gates trace to the design basis in Chapter 0.5 and the availability model in Chapter 12.5; go-live and handover land in Chapter 13.10.
Cite this chapter
Fehn, J. (2026). Documentation, Scripts & Acceptance Test Plans (Chapter 13.2). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-13-commissioning-and-go-live/13-2-documentation-scripts-and-acceptance-test-plans (accessed 2026-09-29).
@misc{aidc-13-2,
  author       = {Fehn, Jacob},
  title        = {Documentation, Scripts & Acceptance Test Plans (Chapter 13.2)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-13-commissioning-and-go-live/13-2-documentation-scripts-and-acceptance-test-plans},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit