Chapter 13.4
In this chapter · 7 sections
Commissioning On-Site Generation & Microgrid Controls
Behind the meter you are the grid, and commissioning proves that generation, controls and storage can sustain the declared AI load, including the campus-wide duty for a gigawatt-class site, through the required source states and the instant the utility lets go.
What you'll decide here
- Whether the site runs prime/island, grid-parallel, or grid-as-backup — because that one choice decides whether your microgrid controller must hold frequency and voltage on its own or merely follow the utility's.
- No-break island service versus a dead-bus break-before-make transfer — and whether the IT terminals stay inside their ride-through envelope on generation, inverter/BESS support or UPS/BBU during planned and unplanned separation.
- Who owns the seam: where facility commissioning (13.3) ends, where microgrid Cx begins, and where it must hand a stable, characterized bus to Integrated Systems Testing (13.6).
- How much of the transient-support stack — grid-forming BESS, synchronous condensers, GPU-side power smoothing — you validate at Cx versus discover under a real workload.
- Which installed tests, injection/HIL or qualified analysis prove loss of utility, loss of the largest generator and black start — budget witness time and controlled-trip exposure, qualify substitutes, and require restoration evidence for each excluded interface before its risk reaches the first real event.
For most of the data-center industry's history, on-site generation meant a yard of standby diesels that started, took the block load, ran for an hour during the monthly test, and went back to sleep. Commissioning was a checklist: crank it, transfer to it, load-bank it, prove the ATS. That world is gone for the AI build. When you site behind the meter to escape an interconnection queue that can run three to seven-plus years — and roughly 90 GW of behind-the-meter gas had been announced cumulatively by mid-2026 (~82 GW of it since January 2025) against a grid that simply cannot energize it in time (Cleanview / SemiAnalysis, 2026) — your generators are prime rather than standby: they run the building, and the utility, if present at all, is the backup. That inversion changes the entire character of acceptance: you are no longer commissioning a transfer switch, you are commissioning a power system — generation, synchronization, protection, storage, and a controller that must do what a utility control room does, autonomously, in milliseconds.
The job here is proving that power system before a single GPU depends on it. It sits deliberately in the sequence: after electrical power acceptance (Chapter 13.3) has energized and characterized the switchgear, UPS, and BESS as components, and before Integrated Systems Testing (Chapter 13.6) pulls the plug on the whole building. The engineering it validates — paralleling and load-sharing, fuel-gas conditioning and dual-fuel changeover, islanding transitions, microgrid-controller tuning, black start, and inertia/transient support — is designed in Chapter 4.8 (electrical integration), Chapter 4.9 (fuel and gas process), and Chapter 4.10 (grid-interactive behavior toward the point of interconnection). Here we prove it works, and mark where a wrong acceptance decision strands the most capital.
Prime/island acceptance: paralleling and load-sharing
The first acceptance gate is mechanical and electrical at once: getting multiple prime movers to share a load without fighting each other. A single 35 MW aeroderivative turbine (GE Vernova's LM2500XPRESS, the unit Crusoe ordered 29 of for its Abilene campus, is the 2025 reference point — 35 MW, five-minute start, black-start-capable and grid-independent by design) does not run a 200 MW hall. Eight or ten of them do, in parallel, on a common bus, and they must synchronize (match voltage, frequency, and phase angle within tight windows before the breaker closes) and then load-share (split real and reactive power proportionally, with no unit hogging kW or VARs).
Which load-sharing control philosophy you commission determines how the plant behaves the instant a GPU cluster steps load. Isochronous control holds frequency dead-flat at 60 Hz and is the natural mode for a true island, but it requires a master controller or a robust load-sharing line and is intolerant of communication faults. Droop control lets frequency sag with load (on a 60 Hz island, 4% droop = 2.4 Hz across 0–100%) and provides local sharing without a communication link. It is the usual grid-parallel choice, though not the only one; stability depends on gains, impedances and operating point. Base-load / import-export control keeps an isochronous governor with a load-control bias while paralleled to the utility, and hands over to island duty more gracefully. Against a stiff utility the droop curve sets the unit's power dispatch, not the bus frequency; declare the nominal frequency before quoting any droop band, and commission the mode you will actually run. Most AI microgrids commission both — isochronous-load-sharing when islanded, droop when grid-parallel — and the acceptance test must prove the controller hands off cleanly between them. Skip the dual-mode validation and you discover, the first time the utility drops, that the plant either hunts (units oscillating against each other) or dumps load because the mode transition was never exercised.
| Mode | Frequency behavior | Reference source | Primary use | What Cx must prove | Failure if skipped |
|---|---|---|---|---|---|
| Isochronous (single master) | Flat 60 Hz | Grid-forming master | Steady island, one prime mover leading | Master holds f under largest load step; backup-master failover | Loss of master collapses the island |
| Isochronous load-sharing | Flat 60 Hz | Distributed / load-share line | Multi-generator island | Proportional kW/VAR split; no hunting; comms-fault behavior | Units fight; oscillation; uneven wear |
| Droop | Stiff grid fixes frequency; island frequency follows the chosen droop and secondary control | Utility when connected; participating voltage/frequency-forming sources in an island | Grid-parallel operation (alternative: base-load / import-export with load-control bias) | Correct power dispatch vs grid; reverse-power protection | Reverse power; export trip; instability |
| Isochronous ↔ droop handoff | Prove reference continuity and recovery; no prescribed flat-to-sag sequence | Named source retains or takes the reference under the approved SOO | Seamless grid/island duty | Clean mode transition at island detect and reconnect | Load dump or hunting at the transition |
Fuel-supply and gas-process acceptance
An electron-side island is only as firm as its molecule-side supply, and this is the acceptance gate that the electrical team is most tempted to wave through because it lives in a different P&ID. Don't. The prime movers will not start, hold load, or pass an emissions stack test if the fuel gas is out of spec, and every aeroderivative turbine is fussy about its inlet. Commissioning the fuel chain (the engineering of which is laid out in Chapter 4.9) means accepting the conditioning train end to end: filtration and coalescing (no liquids, no particulates), dew-point control and inlet heating (superheat above the hydrocarbon dew point so nothing condenses in the fuel skid), Wobbe-index verification (the interchangeability number the turbine's combustion map is tuned to — a Wobbe excursion forces a derate or a trip), and on-site boost compression for aeroderivatives that demand higher fuel-gas pressure than the pipeline tap provides.
The decision with the longest tail is firmness: firm vs interruptible pipeline transport, and the dual-fuel hedge. Interruptible gas is cheaper and faster to contract, but a winter curtailment is exactly when the grid is also stressed and your island is doing real work. The hedge is dual-fuel changeover — gas to distillate (diesel) with on-site liquid storage — and the changeover is a commissioning test in its own right: prove the turbine transfers fuels under load, without a trip, within the changeover window, and that the on-site distillate inventory matches the firmness commitment (operators chasing 'five-nines' reliability size days, not hours, of on-site liquid backup). Accept the gas path on paper and you will find the gap at the worst possible time — a Wobbe trip or a failed changeover during the one curtailment event the whole island exists to survive.
Islanding transitions: seamless vs break-before-make
The islanding-transition choice maps directly onto the IT load's tolerance for a power interruption. Break-before-make islanding opens the utility breaker, lets the bus de-energize, then closes onto generation — a clean, simple, well-understood sequence with one fatal property for AI: there is a dead bus interval. Every millisecond of that interval the GPU racks ride on UPS or rack-level BBU, and a synchronous training run that loses power mid-step restarts from its last checkpoint. Seamless ('no-break') islanding keeps the declared bus energized; make-before-break describes a breaker sequence and is not a synonym: a grid-forming source — a BESS inverter, an already-paralleled synchronous machine, or both — is already holding the bus in parallel, the controller detects the grid anomaly, and it opens the utility breaker while the SOO-designated synchronous machine retains the reference or the qualified grid-forming inverter takes it. Done right, the transition keeps the measured voltage, frequency and transient at the declared bus and IT terminals inside the named equipment ride-through envelope; downstream UPS, BBU and capacitance may still absorb a measurable event.
The consequences cascade. Choose break-before-make and you have implicitly sized your UPS/BBU autonomy to cover the dead-bus interval plus generator pickup — and a tightly-coupled job restarts only if the downstream ride-through fails to cover the disturbance. Choose seamless and you have committed to a grid-forming source able to hold the bus through separation — a grid-forming BESS, an online synchronous generator carrying the reference, or a hybrid — sized against the residual mismatch at the instant of separation (the import or export you lose, the generation already synchronised, the reactive demand, the load's own response), not against the whole campus load; a planned separation at near-zero power exchange and an unplanned separation at heavy import are different sizing cases. It also commits you to a faster and more sophisticated controller, and a far more demanding commissioning test: you must prove the island forms with the load online, not on a dead bus you re-energize at leisure. The acceptance criterion is exact: measure voltage and frequency at the named acceptance points through the transition, and show each trace stays inside the ride-through limits of the equipment at those terminals. The ITI (CBEMA) curve is an AC voltage-envelope application note written for 120 V, 60 Hz single-phase equipment, and it states of itself that it is not a design specification — it supplies no frequency-versus-time envelope and does not transfer to an arbitrary facility bus or a DC rack. Take frequency and RoCoF against the protection settings instead, and declare the sampling, aggregation and measurement uncertainty before the test rather than after reading the traces. This test, more than any other in microgrid Cx, determines whether the building is goodput-grade or merely available.
Microgrid controller tuning: droop, regulation, dispatch
The microgrid controller is the brain, and IEEE 2030.7-2017 is the standard that defines what it must do — the control functions above component level: dispatch, unplanned and planned islanding, reconnection, and black start. Its companion, IEEE 2030.8-2018, defines how you test those functions, and a disciplined microgrid Cx plan cites both as the acceptance basis. Tuning the controller is genuine engineering, because the controller's loops interact: droop coefficients set steady-state power sharing; the secondary voltage/frequency regulation restores nominal after a droop excursion; and the dispatch logic decides, second by second, which generators run, which idle, and how the BESS is charged and discharged against the campus load and the GPU power profile.
The binding trade is response speed vs stability. Tune the controller too aggressive and it overshoots and oscillates against the turbine governors and the inverter control — a multi-loop instability whose power swings, if their frequency spectrum harmonizes with a utility's critical frequencies, can physically stress grid infrastructure — a documented risk for synchronous AI training loads (see the power-stabilization literature, arXiv 2508.14318). Tune it too soft and it cannot arrest the frequency excursion from a GPU load step before protection trips. There is no generic setting; the loops must be tuned against the measured dynamics of this plant, which is why controller tuning is a commissioning activity and not a factory pre-set. The deep-dive below walks the tuning sequence.
Deep dive: tuning the controller against measured plant dynamics
Controller tuning at Cx is a staged escalation: characterize each loop in isolation, then close the outer loops. Stage 1 — primary (droop/governor): with the island formed at light load, inject controlled load steps and measure the frequency and voltage excursion and recovery for each prime mover and inverter individually. Confirm the governor and AVR responses match the model from the Chapter 4.8 stability study; mismatch here means the model is wrong and everything downstream is suspect. Stage 2 — secondary (restoration): close the controller's voltage/frequency restoration loop and prove it returns the bus to nominal after a droop excursion without overshoot, with restoration time inside the dispatch-cycle budget. Stage 3 — load-sharing: with multiple sources online, prove proportional kW/VAR split holds through load steps and that no unit hunts. Stage 4 — dispatch: exercise the economic/availability dispatch logic across a synthetic day, including a forced loss-of-a-generator, and confirm the controller re-dispatches and the BESS bridges the gap without a frequency violation.
The reason this sequence matters: each stage's gains are inputs to the next, and a hidden instability in an inner loop is invisible until an outer loop excites it under a real transient. Tuning the whole thing at once — the temptation when schedule is tight — is how you ship a controller that passes every static test and then oscillates the first time a 200 MW training cluster steps its load. The transient physics you are tuning against are canonical in Chapter 4.5.
Black-start capability and restoration sequencing
Black start is the demonstration that the island can come back from nothing — no utility, no running generation, a dead campus — under its own power. It is the most demanding test in the program and the one most likely to be deferred to 'analysis' under schedule pressure, which is exactly why the required restoration states need valid evidence under the approved test boundary. The black-start source (an aeroderivative turbine with on-board black-start capability, or a grid-forming BESS energizing a starting bus) must cold-start, energize a dead bus, then sequence the restoration: bring up auxiliaries and fuel-gas compression, parallel additional generation, and pick up load in blocks small enough that no single step exceeds the running generation's transient capability.
Block-loading granularity trades restoration speed against transient risk. Large blocks restore the campus faster but risk a frequency excursion that trips the nascent island and forces a restart of the whole black-start sequence. Small blocks reduce each transient but restore load more slowly, and a slow restoration is lost goodput. The acceptance test must prove the sequence, not just the start: cold-start the source, energize the bus, and walk the documented block-loading steps with the BESS bridging each step's transient, demonstrating the frequency stays inside the protection window at every block. Critically, the cooling plant must restore in lockstep — there is no point energizing GPU racks faster than the CDUs and pumps can take heat, a coupling that becomes the IST problem in Chapter 13.6 and is why this test is sequenced before, not during, integrated testing.
BESS / synchronous-condenser inertia and transient support
An island built on inverter-based generation and aeroderivative turbines is low-inertia — it lacks the large spinning masses of a utility grid that resist frequency change. Low inertia means frequency moves fast under a load step, and an AI load steps hard: a synchronous cluster can swing tens of megawatts in milliseconds as a training step begins or ends. NERC's rare Level 3 alert came after roughly 1,500 MW of data-center load disconnected over a six-fault, 82-second sequence in Virginia in July 2024 — but that was a voltage-ride-through and protection-coordination event triggered by repeated grid faults, not a workload-synchronous swing. Inside the island, the same physics means the transient-support stack is not optional, and commissioning must validate every layer of it.
The stack is layered by timescale, and the acceptance plan must test each. Synchronous condensers (free-spinning synchronous machines) and flywheels add real rotating inertia and short-circuit strength — validate by measuring the rate-of-change-of-frequency (RoCoF) the island survives. Grid-forming BESS provides synthetic inertia and sub-second power injection — validate against measured load steps. And increasingly the GPU rack power shelves carry the fastest storage layer: NVIDIA's Vera Rubin power-smoothing system holds ~400 J of rack-level energy storage per GPU — roughly 6× the GB300's 65 J — with a closed-loop controller that tracks capacitor state of charge and reduces peak current demand by up to ~25% against sub-second transients (NVIDIA, March 2026). The prior generation's measured result is a different metric on a different platform: up to ~30% lower peak power on an instrumented GB200 rack running Megatron with the GB300 power-shelf features (NVIDIA, 2025). Carry the two separately. For commissioning, the more the transient is smoothed by workload controls, power-shelf capacitance and BESS, the less inertia the island must carry — but you can only bank that relief if you measure it at acceptance rather than assume it. Validate the layers bottom-up, and the load-bank-vs-real-workload realism gap (the canonical home of which is Chapter 13.6) is where you confront what the test load could and could not reproduce.
Scope & caveats
Tracker estimate of announced US generation capacity across 59 projects; not contracted output, operating supply or a forecast that all announcements will commission. Public summary reports about 2 GW operating and 1.2% under construction.
Announcement-stage stock, not built plant: Cleanview (mid-2026) counts ~2 GW operating across four projects (xAI Colossus 1+2 = 1,498 MW of it), ~1.2% under construction, 36% permitted, 60% announcement-only; ~2.8–3.2 GW operating expected by end-2026.
Scope & caveats
NVIDIA's stated design figure for the Vera Rubin power-shelf PSU capacitor system (NVIDIA Vera Rubin POD, 16 March 2026); the platform entered full production in August 2026. Vendor design statement, not an independent field measurement.
Scope & caveats
Load loss as seen by the grid. NERC's incident review ('Load Details') found the affected data centers transferred their loads to backup power — static UPS, decentralized rack UPS, or DRUPS — in response to the disturbance. The figure is a loss of demand at the interconnection, not evidence that IT power was interrupted or that training jobs restarted.
The approximately 1,500 MW is the total customer-side load reduction coincident with the six-fault sequence; NERC reports approximately 1,260 MW as the sustained drop at the third voltage depression. The NERC-investigated canonical case. A second, larger occurrence followed on 2026-07-22: ~3.8 GW dropped on a single normally-cleared Ashburn 230 kV fault (see companion key number). Two vintages of the same failure mode, not a replacement figure.
Scope & caveats
SemiAnalysis reported this for one 512-H100 cluster at a top-tier operator in October 2024. It is not a per-GPU rate, a scaling law, or a portable fleet baseline; do not extrapolate it by accelerator count.
Failure-mode demonstration at the seam
Microgrid Cx earns its place in the program by demonstrating failure, not just function. The three required outcome families for a credible island acceptance are: loss of utility (the planned and unplanned islanding test — prove the seamless or break-before-make transition holds the IT bus), loss of the largest generator (the N-1 contingency — prove the BESS bridges and the controller re-dispatches without a frequency violation), and black start (prove the campus comes back from dead). Each needs a method selected from installed controlled tests, injection/HIL and qualified analysis, with residual coverage stated; each installed demonstration costs schedule and carries the risk of a real trip — but each buys a tested restoration sequence and a characterized transient response that the IST in Chapter 13.6 can build on rather than re-discover.
For a first-of-its-kind low-inertia island feeding a novel AI load, use model uncertainty to target staged test sources, injection/HIL and approved isolated demonstrations; an untrusted model is a reason to acquire better evidence, not authority to improvise a live failure. The transition where the utility lets go, the contingency where a turbine trips off the bus, and the cold black start are precisely the events where a wrong assumption strands a gigawatt of capital. Witness them here, with a load bank standing in for the GPUs, and hand a stable, characterized bus to integrated testing.
E-02 separation trace. Lose 4.0 MW import at t =0: BESS supplies 4.0 MW for 0.50 s, then declines at 1.0 MW/s as generation increases from 6.0 to 10.0 MW at t =4.50 s. At every instant generator + BESS =10.0 MW. Output energy =4.0×0.50 +½×4.0×4.0 =10 MJ <20 MJ. Peak converter duty √(4.0²+1.0²) ≈4.1 MVA <5.5 MVA and real demand 4.0 MW <5.0 MW. The separate simulated bus bounds are 0.95–1.04 pu and 59.4–60.5 Hz, with no trip or reference loss.
Recharge before reconnection. At t =4.50–5.50 s, ramp generation 10→11 MW while charging rises 0→1 MW; generation minus charge remains 10 MW. Hold generation 11 MW and charge 1 MW for 11.50 s, then ramp both down together for 1.00 s to 10 MW generation and zero charge at t =18.00 s. Charge input =½×1×1 +1×11.50 +½×1×1 =12.50 MJ; 0.80×12.50 =10 MJ restores the spent reserve to 20 MJ. Charging with 1 Mvar needs about 1.4 MVA, below converter capacity. The declared bus extrema cover this entire sequence; energy balance alone does not prove stable control.
Keep reconnection inhibited until reserve is restored, the utility grants permission, and all synchronism/interlock checks pass. The assumed 0.02 pu/0.08 Hz/8° results fit the stated gates; on reconnection ramp generation back to 6 MW while import rises to 4 MW, keeping total supply at 10 MW. Planned zero-interchange separation, largest-generator loss and black start have separate records, including auxiliary, fuel and cooling restoration. A failed reference, channel or permissive stops the sequence at its pre-approved isolated/source state.
Result and flip. The simulation is a candidate; installed release stays HOLD for source/PQ, breaker, mode, dispatch, reserve and restoration records owned by the controls lead, protection engineer, CxA and operations. Hold load at 10 MW for the flip: 5.0 MW import means 5.0 MW initial generation, 15 MJ bridge energy and the real-power ceiling reached. At 5.1 MW import, initial generation is 4.9 MW; the 5.1 s ramp ends at 10 MW after 5.6 s, energy =5.1×0.50 +½×5.1×5.1 =15.555 MJ (about 16 MJ), converter duty about 5.2 MVA. Energy and MVA fit, but 5.1 MW exceeds the 5.0 MW real limit: FAIL. Reduce import or qualify more power support; buying energy alone cannot clear it. IEEE 2030.8 supplies the controller-test method; 4.8 owns integration and 13.6 owns coverage.
Cite this chapter
Fehn, J. (2026). Commissioning On-Site Generation & Microgrid Controls (Chapter 13.4). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-13-commissioning-and-go-live/13-4-commissioning-on-site-generation-and-microgrid-controls (accessed 2026-09-29).
@misc{aidc-13-4,
author = {Fehn, Jacob},
title = {Commissioning On-Site Generation & Microgrid Controls (Chapter 13.4)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-13-commissioning-and-go-live/13-4-commissioning-on-site-generation-and-microgrid-controls},
note = {Accessed 2026-09-29}
}