The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Commissioning & Go-Live › 13.6

Chapter 13.6

In this chapter · 7 sections
Term help

Level 5 Integrated Systems Testing (IST) & Failure-Mode Demonstration

IST proves design-basis faults at a boundary whose applicable permissions and methods are agreed with the utility, AHJ, insurer and Cx authority; controlled trips, injection/HIL and load banks prove the utility-loss chain, and staged workloads close the fixtures' remaining operating evidence.

POWER-BOUNDGOODPUTDENSITY-RAMP

What you'll decide here

  1. What the IST master sequence proves and in what order — steady soak, single faults, utility-loss transfer and required combined faults — assigning each link to a controlled installed trip, injection/HIL or qualified analysis, with explicit coverage, abort limits and restoration records.
  2. Which resistive, reactive or programmable fixture delivers the required step, slew, repetition and air/liquid heat path — and which workload operating points remain untested after it passes.
  3. How far down the mitigation stack (BBU → facility BESS → in-rack capacitance → workload power-smoothing) IST exercises at the declared initial reserve and source states, given that a static load bank cannot trigger the transients that stack exists to absorb.
  4. Which Appendix-F failure modes need installed controlled tests, injection/HIL, factory witness or analysis — and who owns every remaining evidence gap before the affected load is released.
  5. The acceptance criteria that bridge load-bank IST to first-real-workload: the conditions for energizing GPUs, the instrumented ramp thresholds, and which open deficiencies block the proxy run or restrict its approved boundary.

Every level of commissioning before this one tested a subsystem in isolation against its own design intent: the switchgear in Chapter 13.3, the generators and microgrid in Chapter 13.4, the cooling plant and CDUs in Chapter 13.5. Level 5 Integrated Systems Testing (IST) verifies the integrated facility against the defined design-basis scenarios. Each scenario runs as a script — preconditions, witnesses, hold points, abort limits, expected result, restoration — and each design-basis fault has a valid evidence method: installed controlled tests on load banks, injection/HIL and verified analysis close different links, and none silently substitutes for an untested installed interface. IST is not authority to improvise concurrent faults beyond the design basis, or to put revenue-bearing GPUs behind the test.

And it is built on a lie of convenience. IST is run against load banks — heaters that consume the design power so the building has something to cool and power while you trip its redundancy. Load banks are how you can pull the plug on 100 MW without owning 100 MW of irreplaceable accelerators. But a load bank is a fundamentally different electrical and thermal object than a GPU cluster. A bank held flat draws a smooth, steady, controllable load; switched and programmable banks add bounded transients, while a GPU fleet supplies its software-synchronized, millisecond-scale waveform. That gap gets its canonical treatment here. After the IST master sequence and the failure-mode demonstration, the sections below set out what each load-bank technology can and cannot reproduce, across both the electrical and the thermal dimensions, and how a staged, production-representative run in Chapter 13.9 closes workload-specific normal-operation evidence that facility emulators cannot fully reproduce.

IST scope and the master sequence

IST does not begin until every feeding subsystem has its own L4 acceptance signed and its baseline fingerprint captured (Chapter 13.2). The reason is diagnostic, not bureaucratic: if a cooling pump fails during defined-boundary loss-of-source or transfer evidence, you must already know it passed its standalone test, or you cannot tell whether IST found an integration fault or merely an un-commissioned component. IST tests seams, not parts. Sequencing it after the parts pass is what makes a failure during IST interpretable.

The master sequence is a planned escalation. It typically runs steady-state proving first (the building holds design load at design conditions for a sustained soak), then single-fault scenarios (lose one utility feed, one generator, one UPS module, one CDU), then the utility-loss transfer chain — the feed opened at the approved boundary with the building on load banks, and the ride-through, generator start and mechanical-plant transfer all timed — then the combined-fault matrix the design basis specifies, which is what proves the redundancy topology actually delivers the tier it was sold as. Build the IST planning and execution window from fixture mobilization, required states, witness availability, thermal stabilization, battery recharge and retest dependencies; observe these durations on the accepted equipment rather than importing a generic weeks-to-months allowance. The deliverable is a witnessed, time-synchronized data record of how the integrated building behaved at every transition, not a pass/fail stamp.

Common L1–L5 test-capability ladder
Gate / boundaryEvidence that releases itCoverage carried forward
L1 / factory equipmentQualified fixture, delivered hardware/control version, waveform and calibration recordTransport/installation changes and site interfaces
L2 / installed equipmentAs-built identity, wiring, settings, cleanliness and pre-energization checksEnergized performance and automatic actions
L3 / standalone systemLocal functions and protection; injection closes the actual points exercisedCoupled source, load and control response
L4 / discipline performanceDeclared steady/step/slew/repetition, initial reserve and operating conditionsCross-discipline failures and common resources
L5 / integrated boundaryTime-aligned power, thermal, control and restoration evidence for each required stateNamed operating points for staged workload confirmation
Labels follow the owner contract. Example responsibilities: discipline lead supplies evidence, CxA reconciles requirements, operations confirms restoration, owner's authority releases the boundary. These are roles, not approvals.

Attach a capability receipt to each rung: fixture serial/configuration; kW and kVA/PF; step size, slew, repetition and harmonic spectrum; liquid/air heat split; branch placement; control bandwidth; instrument uncertainty; and the exact failure interface injected. A load-step claim without a measured waveform is as weak as a power rating without a terminal. Avtron's generator load-testing method includes switched-load testing; Chapter 13.5 applies the independent liquid-path requirements. A missing capability means another fixture or an explicit HOLD, not an unseen test accepted by its label.

Cascading and concurrent faults; thermal ride-through

A resilience claim is a claim about its named design-basis states, and IST is where it is cashed. A concurrently maintainable design is proved by taking each path out of service in turn with the load held; a fault-tolerant design by tripping each single fault with the load held; and the combined faults the design basis names — a utility loss during a planned maintenance isolation is the classic — are tripped in the order most likely to expose a shared dependency. 2N promises survival of the design-basis events, not of an arbitrary second fault improvised on the surviving path, so the hidden single points of failure — two nominally independent systems sharing a breaker, a controller, a cable tray, or a cooling loop — are hunted by tracing the as-built topology and by relay secondary injection, then confirmed by the approved scripted trip where authorized, or by explicitly scoped injection/HIL and installed-interface checks, with any untested response left open. The cascading-fault matrix is built directly from the facility FMEA (the consolidated catalog lives in Appendix F): each high-severity failure mode becomes an IST scenario, executed in the order most likely to surface a shared dependency.

For AI factories, the dominant integration risk is no longer electrical — it is thermal ride-through, the seam that did not meaningfully exist in legacy IT. In an air-cooled hall, a brief cooling lapse during power transfer is absorbed by the thermal mass of the room and the air; you have minutes. In a direct-to-chip liquid hall, the thermal mass at the die is almost nothing. A GB200-class rack accepts coolant inside the envelope its own vendor documentation names — QCT's QoolRack reference gives separate maxima of ~45 °C liquid inlet and ~65 °C liquid return, and ~25 °C was NVIDIA's launch operating example, not a floor (Chapter 13.5); ASHRAE W45 describes facility-water supply capability, not a product setpoint. Past that envelope the GPUs throttle, and a sustained loss of flow reaches the throttle in seconds. How deep the throttle goes and how long an excursion may last are properties of the named server, coolant profile and workload — read them from the product documentation and measure them in the approved test rather than carrying a generic percentage into the script. So the IST question is specific: during the worst-case power transition, does coolant flow and temperature stay inside the envelope, or does the cluster throttle or trip? Critical cooling therefore rides on the protected bus, and its re-energization time is measured against that window, not assumed. IST is where you measure that gap against a stopwatch, with a load bank standing in for the heat the GPUs would have produced. This is the explicit interlock between facility Cx and cluster burn-in flagged in Chapter 13.5.

BMS / DCIM / SCADA integration as a first-class test object

The control and monitoring stack is one of the things under test, not a witness to IST. Three layers must be proven to agree: the BMS (mechanical/electrical building automation), the SCADA / power-management system (switchgear, generators, the microgrid controller of Chapter 13.4), and the DCIM that operations will actually watch on day 2 (Chapter 14.2). The classic IST finding is not a hardware failure at all — it is that the building did the right thing while the DCIM displayed the wrong thing, or raised forty alarms for one event, or missed the event entirely because a Modbus/BACnet mapping was transposed during integration. An alarm flood is itself a failure: an operator who cannot find the root-cause alarm under a cascade of consequential ones will mis-diagnose the next real incident. IST validates the alarm hierarchy, the automatic control sequences (not just manual operation), and the point-to-point mapping from physical sensor to operator screen — the data path that the entire day-2 reliability program (Chapter 14.1) is built on top of.

The dynamic-load realism gap (the canonical treatment)

The central limitation of fixture-based IST: the load you test against is not the load you will run. A load bank exists to consume power and reject heat on command. A GPU cluster running a synchronous training step does something a load bank was never built to do — it swings its entire draw, in lockstep, across thousands of accelerators, on the cadence of the collective-communication pattern. When the all-reduce stalls compute, tens of megawatts can fall in milliseconds; when compute resumes, it returns just as fast. This is the transient physics made canonical in Chapter 4.5, and it is the reason the GB300 NVL72 adds ~65 J/GPU of in-shelf energy storage and facility BESS designs add their own power-smoothing role. In July 2025, NVIDIA reported up to ~30% lower peak power on its instrumented GB200/Megatron demonstration of these features. IST poses it directly: which load-bank technology do you commission against, and therefore which of these dynamics do you leave un-demonstrated?

Load-bank technology vs. what it can and cannot reproduce
Load-bank typeWhat it emulatesElectrical realismThermal realismWhat it CANNOT demonstrate
Resistive (air-rejecting)Real power (kW) at unity power factor; steady-state heat into airkW at unity PF; switched steps within declared step size, slew and repetitionHeats the room/air, not cold plates; no secondary-loop heat fluxTransients outside the switching envelope; reactive/harmonic profiles not reproduced; no liquid-loop heat
Reactive (R + L/C)Real + reactive power; lagging/leading PF; some inrush/rippleAdds PF/kVA stress on UPS, generator and AVR plus qualified steps; verify GPU harmonic and di/dt coverageAir or liquid sink as specified; impedance does not choose the heat pathUnrepresented waveforms and die/contact behavior; CDU/branch coverage depends on liquid connections and control bandwidth
Dynamic / AI-emulating (transistor or DC programmable)Programmed step-load and ramp profiles approximating workload swingsBest available proxy for di/dt and step loads; still a scripted approximationSpecify air/liquid split, branch connections and thermal-control bandwidthThe exact, cross-rack-synchronized power pattern of a real model on a real fabric; true cold-plate transient heat flux
Real GPUs (proxy training run)The actual workload on the actual fabricActual NCCL or platform-equivalent waveform for the tested software and operating pointReal die heat through real cold plates for the tested population and operating pointNo fault you do not trip — it runs under the staged ramp, never as a fault target
The decision fork at the heart of IST. 'Electrical realism' = ability to reproduce real-power magnitude and the millisecond di/dt swings of synchronized GPU collectives. 'Thermal realism' = ability to impose realistic worst-case heat flux into cold plates and the CDU/secondary loop. Capabilities are 2026-current practitioner ranges; see keynumbers for sources.

The table compares independent capabilities, not a ladder from resistors to perfect GPUs; price the missing waveform or heat path instead of buying a technology label. Resistive load banks are cheap, ubiquitous, and prove the steady-state power and cooling capacity — they are the right tool for the bulk of L4 and the soak portion of IST. Reactive banks add power-factor and ripple stress that resistive banks miss, exercising the UPS, the generator AVR, and protection devices closer to real conditions. Dynamic / AI-emulating banks — transistor-based or, for 800 VDC architectures, programmable DC loads — are a programmable option for the required step-load profile; their measured bandwidth, like that of switched resistive and reactive banks, decides which edges they can reproduce; they let you fire a scripted swing at the BBU and BESS mitigation stack and watch it absorb. Air-cooled electrical load banks leave the technology-cooling system unexercised. Liquid-cooled load banks and product-representative emulator racks validate the thermal-hydraulic loop before production GPUs arrive; a staged workload then supplies hardware- and workload-specific normal-operation evidence.

The mitigation stack under realistic dynamics

Modern AI facilities defend against GPU transients with a layered stack, and IST is where you decide how much of it you exercise on the real building versus accept on vendor witness test. The layers, from grid inward: the facility-level BESS (and any synchronous condenser) absorbing multi-megawatt swings and supporting ride-through (Chapter 4.5, Chapter 13.4); the BBU / UPS layer bridging power transfers; the in-rack / in-shelf capacitance (GB300's ~65 J/GPU, Vera Rubin's larger reservoir) catching the fastest edges; and the workload-side power smoothing — ramp-rate limiting and power capping via SMI/Redfish — that shaves the peak before it reaches the wire. The problem is recursive: a static load bank cannot create the transient the stack exists to absorb, so an IST run holding resistive or reactive banks flat proves only that operating point. Switched or programmable fixtures prove the stack's response inside their measured envelope; real workloads add the tested software-synchronized profile, never every possible future swing. That leaves an explicit risk allocation: which mitigation layers you require to be demonstrated absorbing a real swing during IST, and which you sign off by analysis and recharge-test, knowing the first true exercise comes during the proxy run.

65 J/GPU
GB300 NVL72 in-shelf energy storage for power smoothing; the separate ~30% peak-power result was measured on an instrumented GB200 rack running Megatron
~400 J/GPU
Vera Rubin power-smoothing reservoir target; facility BESS roles for transient/ride-through/DR
Scope & caveats

NVIDIA's stated design figure for the Vera Rubin power-shelf PSU capacitor system (NVIDIA Vera Rubin POD, 16 March 2026); the platform entered full production in August 2026. Vendor design statement, not an independent field measurement.

~1,500 MW
large-load loss over a six-fault, 82 s sequence on a 230 kV line (NoVA, Jul 2024; NERC Level 3 alert 2026) — the ride-through problem IST must prove against
Scope & caveats

Load loss as seen by the grid. NERC's incident review ('Load Details') found the affected data centers transferred their loads to backup power — static UPS, decentralized rack UPS, or DRUPS — in response to the disturbance. The figure is a loss of demand at the interconnection, not evidence that IT power was interrupted or that training jobs restarted.

The approximately 1,500 MW is the total customer-side load reduction coincident with the six-fault sequence; NERC reports approximately 1,260 MW as the sustained drop at the third voltage depression. The NERC-investigated canonical case. A second, larger occurrence followed on 2026-07-22: ~3.8 GW dropped on a single normally-cleared Ashburn 230 kV fault (see companion key number). Two vintages of the same failure mode, not a replacement figure.

45 °C maximum liquid inlet; 65 °C maximum liquid return (separate limits)
QCT GB200 NVL72 QoolRack reference maxima: 45 °C liquid inlet and 65 °C liquid return; separate limits, not a selected operating pair
Scope & caveats

Exact QCT reference. The 45 °C liquid-inlet maximum and 65 °C liquid-return maximum are separate limits, not a prescribed 20 K operating rise. Select a supported operating point, approved fluid, liquid heat load, and design ΔT; ASHRAE W45 describes FWS supply capability, not this product's setpoint.

Separate acceptance maxima, not a prescribed 20 K operating rise; do not attribute these limits to HPE without an HPE document that states them.

~3% training / ~21% inference
power-oversubscription headroom training vs inference — why transient behavior differs by workload IST cannot run
Scope & caveats

Observed headroom in the historical fleets POLCA studied (Patel et al., ASPLOS 2024); POLCA separately simulated 30% additional provisioned servers under specific controls. Neither figure is a deployable allowance for a 2026 reasoning, MoE, or disaggregated fleet: derive that from the proposed fleet's measured coincident demand, its tested protection and capping response, and the serving degradation you will allow.

~55%forecast
reported 2026 forecast for single-phase cold-plate/direct-to-chip share; PMR's published cold-plate category is broader, and market share does not select a project architecture
Scope & caveats

Reported forecast estimate, not a measured deployment census or a project cooling-selection rule. The cited PMR cold-plate category is broader than single-phase DTC.

Reported forecast estimate, not measured fleet share; PMR's published cold-plate category is broader than single-phase DTC and is not a project-selection rule.

Deep dive: close the remaining realism gap with staged, bounded evidence

Each evidence method has a boundary. A product-representative emulator or staged workload on the actual fabric and cooling path, treated in Chapter 13.9, can close electrical and thermal gaps that static load banks leave, while limiting exposure through staged scope and declared abort criteria. Electrically, it produces the exact cross-rack-synchronized sawtooth that the collective-communication pattern dictates, including the di/dt edges no scripted load profile fully captures, because the timing is set by NCCL and the network, not by a test engineer. Thermally, it dumps real heat into real cold plates, driving the CDU control loop, the secondary-loop pumps, and the worst-case branch through the transient regime that air-rejecting load banks structurally cannot reach.

The consequence for sequencing: IST and the proxy run are not redundant, and you cannot substitute one for the other. IST proves the building survives its design-basis faults, on safe load; the proxy run proves the building can carry the real workload's dynamics. Neither is run against the other's question. A facility that passed IST brilliantly can still fail its first proxy run because the CDU was never asked to track a real heat pulse, or because the power-smoothing config was tuned against a load profile that did not match the real model. Treat the proxy run as the final commissioning gate rather than a day-2 benchmark — and accept GPUs into the building only under the staged, instrumented ramp of Chapter 13.10 so that the first real dynamics arrive against a known-good electrical and thermal baseline.

Failure-mode demonstration: live trip vs. witnessed vs. analyzed

Not every failure mode in Appendix F can or should be physically tripped on the real building. Some are too destructive (you do not deliberately rupture a coolant line on a live hall to watch the leak-detection-and-isolation sequence), some are impractical (you cannot induce a real utility-side 230 kV fault on demand), and some are covered acceptably by a witnessed factory or vendor test plus an installed-and-healthy check. IST forces an explicit triage of the FMEA catalog into four buckets, and the allocation is a risk decision the owner signs, not the Cx agent.

FMEA demonstration triage at IST
Demonstration methodTypical failure modesWhat it provesResidual risk carried
Live trip, electrical chain (on load banks)Utility loss at the approved boundary; generator start and transfer; UPS/BESS ride-throughThe integrated response fires correctly, at full magnitude, in real timeUnrepresented waveforms, initial storage states and installed workload behavior
Live trip, thermal chain (on heaters or liquid banks)Thermal ride-through; cooling failover; critical-bus continuity; alarm hierarchy under cascadePower/cooling chains survive the fault; controls narrate itUnrepresented branch duty, die/contact behavior and fault timing
Witnessed factory / vendor testBESS cell-level behavior; breaker interruption ratings; generator load-acceptance curvesComponent meets spec under controlled conditionsInstalled configuration and integration seam require separate closure
Analysis + inspection + scoped injection/HILCoolant-line rupture cascade; real utility fault waveform; multi-MW grid-side ride-throughDesign intent is sound; protection is present and configuredUnvalidated model assumptions and excluded installed responses stay open
How each Appendix-F failure mode is proven. The owner signs the allocation and carries the residual risk of every mode not tripped.

The table's second column needs an actual coverage record. Thermal ride-through, cooling failover and the alarm hierarchy can be tested against commanded disturbances on liquid banks and injected signals; residual risk depends on the missing branch, waveform or control interface, not a universal 'medium' rating. That is the structural consequence of the realism gap: the failure modes most specific to AI factories are precisely the ones a load bank can demonstrate against a fault but not against the workload. Each required criterion still receives PASS, FAIL or HOLD; the release bridge identifies which subsequent test closes each HOLD.

Acceptance criteria: bridging load-bank IST to first-real-workload

A defensible IST acceptance package does three things the legacy IT version never had to. First, it states what was proven and against what load — every passed scenario annotated with the load-bank class used, so the residual-risk register is honest about which results carry the realism caveat. Second, it defines the instrumented thresholds that gate GPU energization: coolant temperature and flow stability under the worst observed transition, critical-bus voltage and timing within the named terminal limits, CDU re-energization time inside the throttle window, and a clean, de-duplicated alarm record. Third, it names the deficiencies knowingly carried into the proxy run — the specific thermal-hydraulic dynamics and workload-synchronized electrical behavior outside the fixtures' measured coverage — so the proxy run is explicitly chartered to close them rather than being treated as a victory lap.

  • Gate to energize GPUs: the utility-loss chain, transfers and design-basis fault set have passing evidence from approved methods at the stated boundary; thermal ride-through measured inside the coolant envelope; critical cooling confirmed on protected bus with re-energization time bounded; BMS/DCIM/SCADA agreement and a clean alarm hierarchy demonstrated.
  • Carried to the proxy run (Chapter 13.9): CDU and secondary-loop control response to real transient heat flux; the cross-rack-synchronized power swing against the full mitigation stack; worst-case-branch thermal-hydraulics under realistic, not air-rejected, load.
  • Carried to staged ramp (Chapter 13.10): validation of power-smoothing and ramp-rate-limit configuration against the actual model, preserving live-block redundancy as load builds.

The acceptance gate is therefore not 'IST passed' but 'IST passed, residual risk named, and the conditions for safely meeting the real workload are written down.' That handoff — from a building proven against faults to a building about to meet its first real dynamics — is the seam IST exists to make safe.

Buy the tests that close the required source, branch and control states, then release only their measured envelope. A bigger load-bank invoice does not close an unconnected liquid branch, and a successful GPU run does not clear an untested protective action. Choosing by that evidence lets the staged workload begin with a finite residual list instead of the building's first experiment.

IST sits inside the commissioning program governed by Chapter 13.1 and scripted per Chapter 13.2, downstream of electrical acceptance (Chapter 13.3), microgrid commissioning (Chapter 13.4), and cooling acceptance (Chapter 13.5). The transient physics it cannot fully reproduce is canonical in Chapter 4.5; the mitigation stack it exercises is engineered in Chapter 4.5 and Chapter 13.4. The FMEA catalog it demonstrates against is consolidated in Appendix F. The remaining installed workload evidence is supplied by the staged proxy training or inference run in Chapter 13.9 and the staged go-live ramp in Chapter 13.10; the simulation-and-digital-twin discipline that should make this test a confirmation rather than a discovery is Chapter 2.7; the goodput framing that justifies the whole exercise lives in Chapter 12.2, with the DCIM handoff in Chapter 14.2.
Cite this chapter
Fehn, J. (2026). Level 5 Integrated Systems Testing (IST) & Failure-Mode Demonstration (Chapter 13.6). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-13-commissioning-and-go-live/13-6-level-5-integrated-systems-testing-ist-and-failure-mode-demonstration (accessed 2026-09-29).
@misc{aidc-13-6,
  author       = {Fehn, Jacob},
  title        = {Level 5 Integrated Systems Testing (IST) & Failure-Mode Demonstration (Chapter 13.6)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-13-commissioning-and-go-live/13-6-level-5-integrated-systems-testing-ist-and-failure-mode-demonstration},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit