Chapter 14.14
In this chapter · 5 sections
Continuous & Re-Commissioning on a Live Campus
On a live campus, each change ticket and density step can invalidate part of the commissioning evidence; a standing re-commissioning program retains unaffected proof and retests changed behavior while protecting the revenue-bearing load.
What you'll decide here
- Which of the three integrated outcomes — loss-of-source and transfer, redundancy failover, thermal ride-through — you re-prove on the maintenance calendar, which on a drift trigger, and which rung of the evidence ladder (inspection, injection, simulation, test load, empty block, bounded live test) each one earns; a live source pull is the top rung, never the default.
- Which density, topology, equipment and firmware changes require full affected-system re-validation versus a bounded retest, and who signs the impact matrix, retained evidence and new service limits.
- How you isolate the test scope so a re-commissioning event stresses one block or fault domain at a time, never the whole factory — and whether your topology even permits that isolation.
- How the drift-detection loop from your DCIM and telemetry (Chapters 14.2 / 14.8) closes into a re-validation trigger automatically, rather than waiting for an outage to reveal the gap.
- Which Appendix F failure modes you actively demonstrate on the live campus on a recurring schedule, versus the ones you only model — and the cost of being wrong about that line.
The commissioning program in Part 13 ended at handover: the L5 integrated systems test passed, the approved utility-loss and system-transfer evidence accepted, the proxy training run hit its goodput gate, and operations took the keys. That is the moment the building was proven. It is also the moment the proof started to expire. Unlike a legacy enterprise hall, which drifts slowly, an AI campus mutates continuously. Racks get denser, CDUs get re-piped, firmware gets pushed fleet-wide, a UPS module gets RMA'd, a fault domain gets re-cabled, a setpoint gets nudged to chase PUE. Each of those changes invalidates some slice of the commissioning evidence, and none of them announce that they have done so. Re-commissioning is the discipline of re-proving the building against its own design intent, on a cadence and on triggers, while protecting operating load and people.
Re-commissioning runs under a constraint the original Cx team never faced: you cannot take the factory down to test it. The original IST happened in an empty building with load banks and a sacrificial proxy run; the re-test happens on an operating campus, so the evidence comes from test sources, injection, simulation and empty blocks first. That difference reorders everything — what you test, how you isolate it, when you accept the risk, and what you simply decline to demonstrate because the demonstration itself is more dangerous than the failure mode it would reveal.
Why the proof decays: the half-life of a commissioning result
A commissioning result is a statement of the form "under these conditions, this system behaves this way." It stays true exactly as long as the conditions hold. On a frontier campus the conditions never hold for long. Three forces erode the proof, and they erode it at very different rates.
Physical drift is the slow one. Coolant chemistry degrades and biofouls the cold plates; quick-disconnects seep; UPS and BESS cells age and lose ride-through margin; generator fuel polymerizes; CRAH coils foul; breaker contacts and bus connections develop resistance. None of this is visible on a single-line diagram, and all of it moves the real failover and ride-through behavior away from the commissioned number. This is the regime that classic retro-commissioning was invented for, and the DOE/LBNL evidence base is unambiguous that the savings and reliability left on the table are large — but that program was built for buildings that barely change, not for ones that double in density every other year.
Configuration drift is the fast one and the dangerous one. Every change ticket — a firmware bump, a setpoint edit, a re-cabled rail, a swapped PDU, a BMS logic tweak to silence a nuisance alarm — moves the system off the configuration that was commissioned. Individually each is trivial. In aggregate, after eighteen months of a live campus, the running configuration and the as-commissioned configuration have diverged so far that the original test evidence describes a building that no longer exists. The Uptime Institute's outage analysis keeps finding the same thing: power is the leading reported primary-cause category for impactful outages, while human error is a major controllable contributor and ignored or inadequate procedures dominate the human-error subset. A change made outside its procedure is also a change that never reached the commissioning baseline.
Step changes are the violent ones. A density step-up from one accelerator generation to the next does not drift the building off its proof — it invalidates the evidence whose electrical, thermal, structural or control boundary changed; unchanged evidence transfers only after an explicit applicability check. A hall commissioned for 40 kW air racks that now hosts 132 kW HPE GB200 NVL72 racks, GB300 NVL72 positions designed for up to 142 kW, or Vera Rubin cabinets provisioned at 330 kW — or a Vera Rubin hall being prepared for the ~600 kW Rubin Ultra / Kyber H2 2027 roadmap point — is a different thermal and electrical machine. The original IST’s loss-of-source and transfer evidence stops at the tested load envelope; unchanged protection, component and route evidence can still apply after the change-impact review. → Chapter 5.1 on the density wall the step crosses.
Periodic re-testing on a live campus
Three integrated capabilities decay after handover, and each is re-proved on its own clock. Loss-of-source and transfer — the utility-to-generator handoff and the BBU/BESS ride-through that bridges it — decays with cell aging, fuel degradation and setpoint drift, and its calendar is the one most plants already run: the monthly under-load exercise NFPA 110 requires of diesel sets (2019 edition §8.4.2: 30 minutes at not less than 30% of standby nameplate kW, or at the manufacturer's recommended minimum exhaust-gas temperature) and, where the monthly run cannot reach that load, the annual supplemental load-bank test on the sequence the adopted edition prescribes (2019 §8.4.2.3: not less than 50% for 30 continuous minutes then not less than 75% for 1 continuous hour, 1.5 hours total); IEEE 450/1188 capacity tests for lead-acid strings; and the OEM state-of-health protocol for lithium BBU/BESS, which those standards do not cover. Confirm which edition your AHJ has adopted before you write the calendar. Redundancy failover — that the N+1 and 2N paths still engage on demand — decays through stuck static-transfer switches, isolation valves left shut after maintenance and control cards that drifted; it is re-proved per fault domain on a rolling calendar and after every change that touched the path. Thermal ride-through — how long the liquid loop holds temperature when a CDU trips or a pump fails at real heat flux — is the AI-specific one and the one go-live could never fully prove, because facility load banks reject heat to air, not into cold plates (→ Chapter 13.5; Chapter 13.6). It is re-proved after any material change to the loop and at every density step.
The hard part is running any of them without taking the factory down, and one rule governs all three: never re-test the whole campus; re-test one block, one fault domain, one concurrent-maintainability boundary at a time, while the rest of the campus carries load and stands ready to absorb the block under test if it fails for real. Inside that boundary, choose the least hazardous method that still produces valid evidence, and climb the ladder only when the rung below cannot answer the question: inspection and functional checks; relay and secondary injection; simulation or hardware-in-the-loop against the commissioning fingerprint; staged transfer on a test source or load bank; an empty or emulator block for thermal work; and, last, a bounded live-load test under a specific hazard and business-risk approval with named hold points and abort limits. Interrupting a live utility source to prove a handoff is the top of that ladder, not the default, because the stakes have inverted since go-live: if the ride-through or the generator start now falls short, you do not lose a test, you drop a live training run and breach an SLA.
The thermal re-test deserves the extra sentence because it is the one where the ladder's bottom rungs are weakest. A pump-drop simulation against the commissioning fingerprint tells you what the controls will do; only product-representative heat — an emulator rack, or a block between tenants — tells you what the silicon will do, and a density step changes the heat duty, while a pump, topology or control change can be equally consequential. Plan the thermal re-test into the refresh, on the empty block, before the new racks go live.
| Re-test | What it re-proves | What degrades it | Live-campus risk | Isolation strategy | Typical cadence |
|---|---|---|---|---|---|
| Utility-loss / transfer | Ride-through and generator handoff at the approved boundary and load class | BBU/BESS cell aging; fuel degradation; setpoint drift | Dropped run + SLA breach if the handoff falls short | One fault domain at a time; injection, HIL or a test source first; a live source pull only by specific approval | Monthly under-load exercise, load-bank test where the monthly run falls short (NFPA 110); IEEE 450/1188 capacity tests for lead-acid, OEM SoH tests for lithium; after any change to the chain |
| Redundancy failover | N+1 / 2N paths still engage on demand | Stuck STS, shut valves, control drift, latent standby faults | Failed transfer escalates a maintainable event to an outage | Per fault domain; inspection and injection first, a bounded live transfer only where the path cannot be proved otherwise | Rolling per-block calendar plus every change that touched the path |
| Thermal ride-through | Liquid-loop hold time on CDU/pump loss at real heat flux | Coolant fouling, flow drift, control-loop retuning, density step | Thermal throttle or trip cascades across racks | Emulator rack or an empty block; bound any live-workload exposure | After any material loop change and at every density step; controls-only simulation between |
Re-commissioning triggers, ranked by risk
Re-commissioning fires on triggers, and the triggers are not equal: their rank sets the scope of re-validation, the sign-off authority required, and the amount of live load you are willing to put at risk to get the proof. Three dominate, in descending order of how much of the original commissioning evidence they invalidate; re-rank by the change-impact matrix when a topology or control change reaches further than the density step.
1 — Density step-up (highest risk). A new accelerator generation that raises rack power can invalidate evidence across several systems, because it changes the electrical and thermal machine simultaneously and at the rack scale where the original proof was most specific. The power chain that was commissioned for 132 kW racks now sees a different load profile, different transient (NVL72-class power swings and the BBU→BESS→GPU-capacitance smoothing stack), different heat flux, different CDU branch loading, different floor mass. Retain original IST evidence for unchanged, bounded conditions and list precisely which tests the new load profile invalidates. Where the changed load crosses the proved envelope, the impact matrix can require broad affected-block re-IST: re-validate the power chain transient behavior (→ Chapter 13.3 electrical acceptance), re-accept liquid duty and flush only where the changed wetted boundary, chemistry or procedure requires it (→ Chapter 13.5), re-run the thermal ride-through, re-baseline goodput. This is why Chapter 1.1 calls the density ramp the most expensive irreversible mistake of the era — and why the substrate (floor, water, electrical headroom) had to be provisioned for it at scoping time, because re-commissioning cannot create headroom that was never built.
2 — Topology change (high risk). Re-cabling a fault domain, splitting or merging blocks, re-piping a cooling distribution, changing a redundancy boundary, or repurposing a hall changes which failure is contained where — and containment is the whole point of the redundancy topology. The original commissioning proved a specific set of fault domains held; a topology change can silently merge two domains that were supposed to be independent, so a single fault now takes both. Topology changes demand a re-run of the failover tests across every boundary the change touched, plus a documentation re-baseline so the as-built single-lines match reality. Skip it and the redundancy claim is paper-true but physically false — the diagram says independent, the copper says shared.
3 — Equipment replacement (managed risk). Swapping a UPS module, a CDU, a generator, a PDU, or a pump under RMA is the most frequent trigger and the most procedural. The replaced unit was never commissioned in this building; the manufacturer's factory test is not your site acceptance. Equipment replacement demands a unit-level re-acceptance (the L3/L4 script for that component) plus a re-test of the redundancy path it sits in, because the act of swapping it is itself a configuration change that can leave an isolation valve shut or a control setpoint at default. A bounded equivalent replacement can justify a component/path retest; changed ratings or control dependencies can demand wider tests and higher authority. Release it only if the change-management discipline (→ Chapter 14.12) actually enforces the re-test rather than treating the swap as done when the unit powers on.
| Trigger | Risk rank | Evidence invalidated | Required re-validation | Sign-off authority |
|---|---|---|---|---|
| Density step-up (new generation) | 1 — highest | Only the electrical, thermal, structural, fabric or workload evidence whose boundary changed | Impact-selected electrical/thermal/structural tests and measured goodput; retain unaffected evidence | Cx authority + ops leadership; design-basis change |
| Topology change (re-cable / re-pipe / re-boundary) | 2 — high | Fault-domain containment and redundancy claims | Failover re-test on every boundary touched + as-built single-line re-baseline | Cx authority + reliability owner |
| Equipment replacement (RMA / swap) | 3 — managed | Unit-level acceptance + the redundancy path it sits in | Component L3/L4 re-acceptance + path failover re-test | Operations change-board |
| Firmware / software push (fleet-wide) | Cross-cutting | Control behavior, power/thermal response, fabric timing | Canary block soak + behavioral diff vs baseline fingerprint before fleet rollout | Operations + fleet-software owner |
Scope & caveats
Respondent-based survey result (Uptime Institute, Annual outage analysis 2026, Figure 4, n=96, 'primary cause of your most recent impactful incident'), not an event-by-event census of the industry. The 2026 report does republish 45% for 2025 outages — down from 54% in 2024 — and reports a fifth consecutive year of declining per-site outage frequency with the pace of improvement slowing. Uptime attributes part of the fall to electrical upgrades and part to other causes rising (fire-suppression share up six points year-over-year).
Scope & caveats
SemiAnalysis reported this for one 512-H100 cluster at a top-tier operator in October 2024. It is not a per-GPU rate, a scaling law, or a portable fleet baseline; do not extrapolate it by accelerator count.
Scope & caveats
Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.
The drift-detection → re-validation loop
Triggering re-commissioning on the calendar alone is necessary but not sufficient — it catches the slow physical drift but misses the fast configuration drift that does the real damage between scheduled tests. The mature program closes a loop: the same DCIM, telemetry, and observability stack that runs the campus (→ Chapter 14.2) continuously compares the live configuration and live behavior against the commissioning fingerprint — the baseline captured at go-live (→ Chapter 13.2) — and fires a re-validation trigger when the gap crosses a threshold.
The loop has four stages. Baseline: the as-commissioned fingerprint — power-chain transient signatures, CDU flow/thermal curves, fabric BER and timing, redundancy-path states, firmware inventory. Observe: the live telemetry stream, including the firmware/software lifecycle state from fleet management (→ Chapter 14.8). Diff: the automated comparison that flags when the running configuration has drifted off the fingerprint — a setpoint that no longer matches the SOO, a firmware version that diverged from the approved baseline, a failover path that has not been exercised within its window, a thermal margin that has eroded. Trigger: the diff escalates into a re-test scope sized to what drifted — a firmware divergence triggers a canary soak; an eroded thermal margin triggers a ride-through re-test; an unexercised path triggers a failover re-test.
Where you set the threshold is the design choice. Too tight and you re-commission constantly, putting live load at risk for noise and burning the operations team out. Too loose and you re-discover the gap as an outage. The right setting is workload-aware: checkpointable training and latency-bound inference can set different service-performance thresholds and outage budgets, but neither may loosen protection, OEM, chemistry, dew-point or personnel-safety limits. Price replayed GPU-hours for checkpointable training and missed latency-SLA responses, service credits and lost contribution for serving. → the goodput-vs-availability framing in Chapter 12.2.
Deep dive: a firmware push retests the behaviors and fleet boundaries it changes
A fleet firmware update looks like a software operation and is treated like one — schedule a maintenance window, push the image, confirm version. But a firmware change to a GPU, a BMC, a PDU controller, a CDU controller, or a UPS module changes the behavior the campus was commissioned around: power-capping response, thermal throttle curves, transient draw under NVL72-class swings, fabric timing, alarm logic. A push that subtly alters how thousands of accelerators respond to a power excursion has, in effect, re-configured the load the electrical system was commissioned against — within each cohort receiving it. Original transient evidence needs revalidation where the changed behavior leaves its proved envelope; unchanged cohorts retain their applicable evidence.
This is why a disciplined fleet-software program treats every behavior-affecting firmware push as a partial re-commissioning event with a canary: roll to one isolated block, soak it, and diff its power/thermal/fabric behavior against the commissioning fingerprint before the fleet rollout proceeds. The blast radius is what makes this trigger uniquely dangerous — an unbounded simultaneous firmware deployment can reach the whole campus and remove isolation options. Bounded waves limit that reach; a shared controller or a replacement changing redundancy can have equally broad consequences. The canary is an early gate; staged waves, dependency checks, automatic aborts and post-rollout monitoring must also catch effects that need scale or a rare operating condition. The OCP GPU firmware update specification (Redfish, PLDM-over-MCTP, secure out-of-band) gives the mechanism; the canary-and-diff discipline gives the safety. → firmware lifecycle in Chapter 14.8; change control in Chapter 14.12.
Tie-in: DR drills and the FMEA catalog
Re-commissioning does not live alone — it is one of three recurring proof activities that share the same machinery and should be planned together. The other two are disaster-recovery drills and FMEA-driven failure-mode demonstration.
DR drills (→ Chapter 12.3) are re-commissioning at the campus-and-region scale. A failover drill that fails workload from one site to another, or exercises geographic redundancy, is testing the same kind of claim that defined-boundary loss-of-source and transfer evidence does — "the redundant path engages on demand" — just at a larger fault domain. Co-scheduling them is efficient, and it keeps the drill from proving the wrong thing: co-schedule them only when the receiving site's own power evidence is current, because a DR failover that lands on a site whose generator handoff has drifted proves nothing.
The FMEA catalog (Appendix F) is the master register that decides what gets demonstrated, how, and how often. Every failure mode in the catalog needs a re-commissioning treatment, and assigning one is the operator's first pass over the register — Appendix F supplies the mode, trigger, propagation path, detection, blast radius, mitigation, recovery and owner; the treatment class, retest interval, EOP identifier and evidence reference are the columns you add for your own plant: live-demonstrable on a single block (re-test it on a schedule), sub-scale or next-empty-block demonstrable (prove it before that block goes live), or model-only (too dangerous to demonstrate on live load — prove it analytically and via the quantitative reliability model in Chapter 12.5). Where a random failure and an adversarial abuse share a physical consequence the catalog links them, but the evidence differs: a fault test proves recovery, an adversarial test must also prove authorization, detection and containment (→ Chapter 11.10). The demonstration line drawn earlier governs here too: demonstrate what you safely can on live load, model what you cannot, and never let a checklist push you across that line on a running factory.
For the valve-identity HOLD in Chapter 14.5, the first release is documentary: reconcile the OEM procedure, site P&ID and physical tags, then approve the work package in Chapter 14.12. After authorized isolated maintenance, retain the fluid record and inspect the repaired boundary for leakage; prove flow, pressure, temperature sensing, alarms and the affected failover path against the approved profile in Chapter 13.5. Do not infer hydraulic capacity from agreeing temperature sensors. The June 2026 ASHRAE/PNNL/NEMA framework is guidance for coordinating such evidence, not an adopted test mandate. Record the passing evidence and unchanged evidence retained before returning the affected service to normal duty.
Cite this chapter
Fehn, J. (2026). Continuous & Re-Commissioning on a Live Campus (Chapter 14.14). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-14-day-2-operations-upgrades-and-lifecycle/14-14-continuous-and-re-commissioning-on-a-live-campus (accessed 2026-09-30).
@misc{aidc-14-14,
author = {Fehn, Jacob},
title = {Continuous & Re-Commissioning on a Live Campus (Chapter 14.14)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-14-day-2-operations-upgrades-and-lifecycle/14-14-continuous-and-re-commissioning-on-a-live-campus},
note = {Accessed 2026-09-30}
}