Chapter 2.1
In this chapter · 6 sections
Program & Project Management: The Integrated Master Schedule & Critical Path
The critical path of an AI build is the zero- or least-float dependency path to a common first-productive-run milestone; transformers, interconnection, permitting, building, or cluster integration may govern by configuration; a single slipped slot strands depreciating silicon, so the schedule itself is the asset under management. Utility service, facility acceptance and cluster acceptance retain separate contractual cash-flow milestones.
What you'll decide here
- Which single milestone you are managing the whole program toward — time-to-first-train (or first-token) — and therefore which of the parallel tracks (power, building, IT) you treat as the governing critical path versus the ones you keep off it with float. Keep utility billing, rent commencement and compute billing tied to their own signed acceptance conditions.
- For each applicable long-lead item, which need-by milestone and lead-duration percentile back-calculate its release date — with P50 used as an internal target forecast rather than the contractual need date — because the gap between those two dates is measured in quarters of revenue, and the deposit goes out before the design is frozen. Define whether the need-by event is factory release, dock delivery or accepted service before subtracting any duration.
- How you run the facility track and the cluster track as two schedules that must be bound by explicit integration milestones — the powered-shell handoff, energization, water-on, and the burn-in gate — rather than one monolithic Gantt that hides the seam where most slip happens. Give each handoff an owner, acceptance evidence and a protected test window.
- Which project-controls discipline (earned value, milestone-deposit cash curve, change-order and claims process) you stand up on day one, because owner controls retrofitted onto a hot project become a forensic exercise, not a steering tool. Reconcile physical progress, cash committed and forecast cost-to-complete before the next draw.
- What your stage-gate governance actually gates — which irreversible commitments (the interconnection deposit, the transformer PO, the GPU slot reservation) are released at which board approval, and where the assumptions-and-decisions register records what you bet and who owns the bet. Name the evidence that stops each release.
Part 1 decided what to build and whether the economics close. This chapter is where the abstraction ends and the calendar begins. An AI data center is a program with a deadline that is set not by the owner's ambition but by physics and supply chains: the day the cluster can take its first synchronous training step, or serve its first revenue token, under the workload’s acceptance conditions. Everything upstream of that day is a race, and everything about how you run the race is a sequence of decisions whose consequences are denominated in time. Because the asset may face the guide's contested 2–3-year bear-case economic clock, time converts directly into money. → the depreciation clock that prices every lost month is in Chapter 1.8.
This chapter applies that frame to schedule: the phase-gate lifecycle and the build reframed as a time-to-first-train race; the Integrated Master Schedule (IMS) and the critical path across three tracks that move at different speeds; schedule risk quantified with Monte Carlo and the P50/P90 dates the long poles force on you; the owner's project controls — earned value, milestone deposits, change orders and claims; the facility and cluster schedules bound by integration milestones; and the stage-gate governance and assumptions/decisions register that record what the program is actually betting. In a power-bound, allocation-constrained market, the schedule is the project, and the long poles are not the ones a traditional general contractor watches. Keep the cash-flow milestones distinct: utility energization can start utility billing, accepted facility service can start a landlord’s rent, and accepted compute service can precede billable use. Bind each to its own contract and signatories in the IMS.
The lifecycle and the phase-gate model
A data-center program moves through a recognizable sequence — scope and design basis → site control and entitlement → interconnection and power → procurement → construction → commissioning → go-live → operations — and the mature way to govern it is a phase-gate (stage-gate) model: each phase ends in a gate where capital is released, assumptions are tested, and the program either advances, holds, or kills. The gate exists to make the irreversible commitments explicit and to put a named owner and a dated decision on each one before the money leaves. → the reversible-vs-irreversible discipline this inherits is set in Chapter 1.1.
What makes the AI build different from a 2018 enterprise data center is that the gates are no longer evenly spaced. In a power-bound market the early gates — interconnection and long-lead procurement — release the commitments that set the finish date, while the late gates (fit-out, commissioning) govern execution against the governing dependency path established by the current IMS. The construction industry's instinct is to gate on design maturity; the AI program's reality is that you must gate the time-critical packages on power-service evidence and allocation evidence before the whole design is mature, once each released package’s interfaces are stable, or you arrive at a finished building with no megawatts and no GPUs. The phase-gate model has to be re-weighted accordingly: front-load the gates that release time-critical deposits, and accept that you are committing capital against assumptions you have not yet fully retired. On a permit-controlled path, an earlier transformer deposit cannot recover the permit date. Require the package’s approved interfaces and the remaining risk owner alongside the schedule benefit before releasing it.
Building the Integrated Master Schedule across three tracks
The Integrated Master Schedule is the single time-logic network that ties every deliverable, dependency, and milestone into one critical-path-method (CPM) model. The mistake that defines failed AI programs is running it as one undifferentiated Gantt. An AI data center is really three schedules braided together, each governed by a different physics and a different supplier ecosystem, each with its own critical path:
- The power track — interconnection studies and agreements, the utility's grid upgrades, the substation, load-serving HV/substation and medium-voltage transformers (plus GSUs only where generation/export is in scope), switchgear, and (increasingly) on-site or behind-the-meter generation as a bridge. This track is dominated by lead times the owner cannot compress: large power transformers at roughly 128 weeks and generator step-up units at ~144 weeks (Wood Mackenzie Q2 2025 survey), and the host utility's project-specific service studies, upgrades, agreements and energization milestones. It often governs, but the configured dependency network — not the track label — identifies the critical path to a common end milestone.
- The building track — entitlement and permits (the air permit is a recurring long pole where on-site gas is involved), earthworks, shell, mechanical/electrical/plumbing, and the cooling plant. A shell-and-core AI hall can be built in 18–24 months — with start/end events that differ from interconnection estimates; compare both to the same required milestone before assigning float.
- The IT / cluster track — the GPU allocation (a slot, not a purchase, negotiated quarters ahead), CoWoS/HBM-gated accelerator delivery, network fabric, storage, structured cabling, then rack-and-stack, fabric validation, burn-in, and the reference run. This track is gated by allocation, not by the owner's cash. → the allocation game lives in Chapter 2.3; the HBM constraint behind it in Chapter 7.6.
The IMS exists to expose the float between these tracks and the integration milestones where they must meet. Float is the schedule's shock absorber: each track's float is calculated from the project dependency network and selected supply/permitting path, and the discipline is to spend that float deliberately — sequencing the fit-out to land just-in-time against energization — rather than letting it evaporate into early-but-idle completion. What kills programs is the current zero- or least-float path consuming all the float silently while the team celebrates the building track finishing early — on a slab that has no power.
| Track | Governs | Typical long pole(s) | Indicative duration | Float vs the program critical path |
|---|---|---|---|---|
| Power | Megawatts at the rack, on a firm date | Host-utility service studies, agreements and upgrades; load-serving HV/substation transformers (~128–208 wk in cited surveys); HV switchgear (equipment-specific published range); GSU only for generation/export | Project-specific energization network; any bridge-generation duration starts and ends at separately declared milestones | Configuration-specific — calculate to the common first-productive-run milestone |
| Building | A weather-tight, plumbed, code-compliant hall | Air permit (where on-site gas); cooling plant; long-span steel | 18–24 months shell-to-MEP-complete | Configuration-specific — building duration and interconnection duration use different start events |
| IT / cluster | A validated cluster doing useful work | GPU allocation slot; CoWoS/HBM-gated delivery; the fabric | Allocation negotiated 2–4 quarters ahead; 6–10 wk bring-up after install | Calculate from allocation release through accepted reference run, linked to powered-shell and energization milestones |
Read the table as a sequencing problem. The power track sets the date; the building track must finish into that date with just enough float to absorb a slipped transformer; the cluster track cannot start meaningful integration until the powered-shell handoff, and then carries a bring-up tail that the inexperienced owner forgets to schedule. Making those three truths visible at once is what the IMS is for — so that effort and capital flow to whichever track is currently binding, which in 2026 is almost always power.
Schedule risk analysis: Monte Carlo, P50/P90, and the long poles
A deterministic CPM schedule produces a single finish date, and that date is a fiction — it is the result you get only if every activity lands on its point estimate, which collectively never happens. The mature program runs a quantitative schedule risk analysis (QSRA): assign a duration distribution (typically three-point — optimistic/most-likely/pessimistic) to each activity, model the correlations (a transformer delay and a switchgear delay are not independent — they share a strained supply chain), and run a Monte Carlo over the network a few thousand times. The output is not a date but a distribution, and the two numbers that matter are the P50 (the date you have a coin-flip chance of beating) and the P90 (the date you are 90% confident of beating).
The gap between P50 and P90 is dominated by a handful of long poles with long right-tails: the load-serving HV/substation transformer (and a GSU only where generation/export is in scope), the grid interconnection energization date, the air permit where on-site generation is in scope, and the GPU/HBM allocation. These are not normally distributed — they are long-right-tailed, because the failure modes (a transformer factory slot slips a quarter, a required utility study or upgrade milestone slips, an air-permit challenge adds eighteen months) move the date a lot, not a little. A schedule whose P50–P90 spread is six months is telling you that one of these poles can eat two quarters of revenue, and the deposit on that pole goes out the door before the design is frozen.
Scope & caveats
For Cedar, decide whether to promise first productive training by week 70 and when to release the transformer package. Assume all starts at week 0 on one calendar, no resource leveling or hidden float, and these durations in weeks: utility readiness 44; manufacture to factory release 40; shipping to site 4; building readiness 42; installation after both dock delivery and building readiness 6; integrated facility acceptance after utility and installation readiness 4; IT delivery 46; racking after IT delivery and building readiness 2; cluster validation after facility acceptance and racking 6. An eight-week common supply disruption affects utility equipment and transformer manufacture together with probability 20%. An independent permit delay adds 16 weeks to building readiness with probability 10%; no other risks are modeled. The required confidence is 90%. These condensed teaching durations expose the electrical, building and IT interfaces described here; they are not supplier quotes or Cedar project observations. The calendar origin, predecessor links, common supply event and independent permit branch are exact enumeration conventions, not frequency estimates. Every duration is a selected teaching constant: factory/shipping end just after building readiness; utility precedes installation completion; IT/racking can run alongside the electrical path. The supply and permit increments deliberately bracket the path crossover. The event probabilities are chosen weights inside [0,1], not measured base rates; their independence and the shared utility/factory event are explicit dependence assumptions. The promise and confidence are sponsor preferences. Chapter 2.1 owns this network; a real Cedar release uses quoted durations and a project QSRA.
Cedar’s critical path is a dependency calculation. In weeks, dock = release + manufacture + common delay + shipping; building = baseline readiness + permit delay; installation finish = max(dock, building) + install duration. Facility acceptance = max(utility readiness + common delay, installation finish) + acceptance duration. Racking = max(IT delivery, building) + rack duration; productive date = max(facility acceptance, racking) + validation duration. Multiply independent event weights, then accumulate the completion distribution. The first date that reaches the selected cumulative probability is its percentile.
Scope & caveats
Exact week-0 enumeration: d/p=0/0, 8/0, 0/16, 8/16 weeks with probabilities .72/.18/.08/.02. Dock=40+d+4; building=42+p; installation=max(dock,building)+6; facility=max(44+d,installation)+4; racking=max(46,building)+2; productive=max(facility,racking)+6. Facility dates 54/62/68/68 and productive dates 60/68/74/74 give cumulative .72/.90/1.00. P50=60, P90=68, Pr(finish≤70)=.90. Base utility float=50−44=6 weeks. Dock need=70−6−4−6=54; selected factory/shipping P90=40+8+4=52; latest release=54−52=week 2. At that release the productive outcomes are 62/70/74/74, still .90 by week 70; week-3 release gives 63/71/74/74 and only .72 by week 70. Crossover at the original week-0 release: 42+p>44+d, hence p>2+d; at the selected week-2 release: p>4+d. The 16-week permit event controls either release, ending at week 74.
Release the package by the date in the result and retain the permit branch in the promise. The no-disruption path runs through manufacture, shipping, installation, facility acceptance and cluster validation; utility float is measured to installation finish, so adding another contingency for the same supply event double-counts it. The crossover shown for the original calendar release differs from the selected later release. If building readiness becomes later than dock delivery, accelerating the transformer cannot recover the building path.
The GAO Schedule Assessment Guide supplies the dependency and risk method. This chapter owns the IMS; 2.3 receives the package release, 2.4 assigns custody/remedies, and 13.10 accepts the cluster. Rent commencement retains its own contract gate.
Scope & caveats
Load-serving substation/power transformers only. Generator step-up (GSU) transformers are carried as a separate register entry (~144–208 wk). The upper bound comes from large-unit and constrained-market quotes, not from GSU indices.
Indices diverge in mid-2026 for large power transformers generally; the GSU-specific divergence (VAWN 144 wk vs SemiAnalysis 3–4 yr) is recorded on the GSU claim.
Scope & caveats
2026 US build share under construction
Early-2026 tracking snapshot. JLL's midyear print (2026-08-11) reads much stronger on a different metric: 25 GW of North American net absorption in H1 2026 (a leasing metric, not energized MW), 66 GW under construction with 95% pre-committed, vacancy at 1% for a third year, 77% of construction in frontier markets. Do not conflate absorption/construction-pipeline metrics with calendar-year US energization targets.
Scope & caveats
Editorial planning range for a liquid-cooled AI hall against a 4-6-week air-cooled comparator. The cited Level-5 guide describes the workflow, not these durations; derive the project duration from the approved test scripts, phasing, witness plan, defect-correction allowance and retest scope.
Not an OEM, standards-body or measured-cohort benchmark and not an invariant minimum. Do not copy it into an IMS in place of a scripted test plan.
Scope & caveats
SemiAnalysis reported this for one 512-H100 cluster at a top-tier operator in October 2024. It is not a per-GPU rate, a scaling law, or a portable fleet baseline; do not extrapolate it by accelerator count.
Scope & caveats
This is the rental/IaaS denominator (SemiAnalysis, contested). Distinct and much larger is the lab token-revenue side: SemiAnalysis's Tokenomics model (Aug 2026) puts OpenAI/Anthropic API inference at >$100B/GW/year on a GB300 cluster against ~$12B/GW/year of rental cost — a model-derived figure sensitive to utilization and price mix, not an audited disclosure. Do not conflate lab API revenue with IaaS rental in one number.
Scope & caveats
Eligible new or uprated generation offering at least 250 MW accredited UCAP; study deposit $500,000, readiness deposit $15,000/MW. This is not a campus load-service application schedule.
Temporary generation track; host-utility load service has separate study, service and security obligations.
Owner's project controls: earned value, deposits, change and claims
A schedule you cannot measure against is a wish. Project controls is the owner-side discipline that turns the IMS into a steering instrument: a cost-and-schedule baseline, periodic measurement of progress against it, and a forecast that updates honestly. The backbone is earned value management (EVM) — comparing the budgeted cost of work performed (BCWP/EV) against the budgeted cost of work scheduled (BCWS/PV) and the actual cost (ACWP/AC), to derive a schedule performance index (SPI) and cost performance index (CPI). The value of EVM on an AI build is that it forces physical-percent-complete discipline and produces an estimate-at-completion early enough to act on, instead of a surprise at the end.
But EVM was built for labor-and-materials projects, and an AI data center's cost is dominated by a few enormous milestone-deposit equipment orders — the transformer, the switchgear, the turbines, the GPU allocation — paid against vendor manufacturing milestones, not against installed progress. This breaks naive EVM: booking the full PO value as "earned" on deposit overstates progress; booking nothing until delivery understates it for two years. The owner's controls function has to track a commitment/cash curve alongside the EVM curve — when each deposit is contractually due, what it secures (a factory slot, a queue position), and what its forfeiture costs if the program pivots. On AI builds the deposit schedule, not the construction draw, is the dominant near-term cash event. → deposit and slot-reservation instruments in Chapter 2.3; the contract that governs them in Chapter 2.4.
Change-order and claims management is the other half. AI programs change scope mid-flight more than any other large construction class — a GPU-generation jump (NVL72 to a denser successor) mid-design re-rates the cooling plant, the floor loading, and the busway; an interconnection re-study moves the energization date and cascades into the fit-out sequence. Each change is a fork with a schedule and cost consequence, and the owner who has not stood up a disciplined change-control board on day one ends up litigating those consequences as claims at the end. The cheap move is a tight baseline plus a fast, well-documented change process; the expensive move is a loose baseline that turns every density surprise into a dispute.
| Instrument | What it measures | What it catches early | AI-specific twist |
|---|---|---|---|
| Earned value (SPI/CPI) | Performed vs scheduled vs actual cost | Slip and overrun, via a real estimate-at-completion | Distorted by milestone-deposit equipment — needs physical-% rigor |
| Commitment / cash curve | When each deposit is due and what it secures | Forfeiture exposure if the program pivots | Deposits (transformer, GPU slot) dwarf the construction draw early |
| Critical-path & float report | Which track is binding; float remaining | Float being silently consumed by a long pole | Three braided tracks — must report per-track, not one number |
| Change-control board | Scope deltas, priced with schedule impact | Density/generation pivots before they become claims | GPU-gen jumps re-rate cooling/floor/power mid-design |
| Risk register & QSRA refresh | P50/P90 movement as risks retire or fire | A long pole's tail materializing | Long poles are correlated — model them jointly |
The facility-vs-cluster two-track schedule and its integration milestones
The chronically under-managed seam in an AI build is the boundary between the facility (the powered, cooled shell, delivered by the construction and MEP world) and the cluster (the GPUs, fabric, and software, delivered by the IT and platform world). These are two organizations, two cultures, two schedules, and two definitions of "done" — and the project lives or dies in how cleanly they are bound. The right structure is an explicit two-track schedule with a small set of named integration milestones where the tracks hand off, each with an unambiguous entry/exit gate and an owner. → the powered-shell delivery model that creates this seam is in Chapter 2.2.
The integration milestones that bind the two tracks, in order:
- Powered-shell handoff. The facility delivers a hall with conditioned space, structural floor capacity, and the power and cooling distribution stubbed to the white space — but not yet energized to the rack. This is the contractual seam between base-building and IT fit-out, and the cleanest place to split scope and risk.
- Energization (power-on). Medium-voltage power live to the in-row PDUs/busway, UPS and any on-site generation commissioned (L3/L4). Until this gate the cluster track cannot draw load; it is the most common place for the power track's slip to surface as a cluster-track delay. → electrical acceptance in Chapter 13.3.
- Water-on / cooling-ready. The facility cooling loop and CDUs flushed, leak-checked, balanced, and proven to spec — non-negotiable before energizing liquid-cooled racks, because coolant temperature and flow must remain inside each installed rack's OEM operating envelope. → CDU commissioning in Chapter 13.5.
- Integrated systems test (L5 IST). The facility proves it holds load and rides through faults under simulated full IT load. For a liquid-cooled AI hall, set the Level-5 duration from the approved test scripts, phasing, witness plan, defect-correction allowance, and retest scope. → IST in Chapter 13.6.
- Cluster burn-in and the reference run. Now the IT track owns the clock: node diagnostics, fabric BER validation, a burn-in campaign selected by the signed test plan and measured defect-discovery rule, and a reference training/inference run at goodput. This is first-train. → burn-in in Chapter 13.8; cluster-scale validation in Chapter 13.9.
The reason to make these milestones explicit rather than implicit is that the seam is where finger-pointing lives. When the building is "done" but the cluster is not earning, the question is always whose milestone slipped — and a program with named integration gates and per-gate owners answers it in a stand-up, while a program with one Gantt answers it in a claim.
Deep dive: why the cluster bring-up tail is the schedule everyone forgets
Construction-world schedules end at ready-for-service. A landlord’s rent can start at the contracted facility acceptance; compute-service receipts start when the customer’s capacity and useful-work acceptance conditions are met, and the gap between the two is a cluster bring-up tail that is routinely missing from the owner's IMS. The tail has hard, un-compressible content. After racks are powered and water flows, the fabric must be validated (an InfiniBand bit-error-rate sweep against a ~1e-12 threshold, per-port, across tens of thousands of links), nodes must be diagnosed and the inevitable dead-on-arrival GPUs and HBM swapped, and the cluster must burn in under a signed campaign whose intensity, exposure, pass/re-soak rules, and statistical stop condition fit the fleet. SemiAnalysis's October 2024 playbook recommends at least 3–4 weeks of factory high-temperature burn-in before deployment and separately reports ~7 days MTBF for one 512-H100 cluster at a top-tier operator; neither figure is a universal settling time or per-GPU rate. A reference run then demonstrates goodput.
The consequence of omitting this tail is a 6–10-week phantom delay between "building done" and "cluster earning" that the owner did not budget — six to ten weeks during which the GPU fleet depreciates and earns nothing. On a 200 MW hall at ~$12–13B/GW/yr, that tail is on the order of $277–500M of foregone revenue if it is a surprise instead of a plan. The fix is structural: put burn-in and the reference run on the IMS as critical-path activities, staff them, and manage time-to-first-train as the finish line — not ready-for-service. → the goodput target that defines a successful bring-up is in Chapter 13.9; the checkpoint math behind training's restart cost in Chapter 9.4.
Stage-gate governance, board approvals, and the assumptions register
The phase-gate model only protects the program if the gates actually gate something irreversible. The governance question is therefore concrete: at which board approval is each one-way-door commitment released? The interconnection-study deposit is committed before any building exists (under PJM's Expedited Interconnection Track, accepted by FERC on June 9, 2026 and effective July 31, 2026, a $500k study deposit that turns non-refundable once the generation request is valid, plus a $15k/MW readiness deposit; for the campus load itself, the host utility's own study fees and the security behind its service agreement); the HV transformer PO commits a factory slot 128–208 weeks out; the GPU allocation reservation commits a slot quarters ahead of silicon that is itself CoWoS/HBM-gated. Each of these is capital released against assumptions that have not been fully retired — which is exactly why the gate exists: to force the board to look at the assumption, name its owner, and accept the bet on the record.
The artifact that makes this auditable is the assumptions-and-decisions register — the schedule-and-commercial analogue of the design-basis document from scoping. It records, for every load-bearing assumption (the energization date, the transformer delivery date, the GPU-generation the cooling plant is sized for, the contracted-vs-merchant power split the financing assumes), what was assumed, who owns it, when it must be confirmed or it becomes a risk, and which downstream commitments depend on it. When a long pole's tail fires — a transformer slips a quarter — the register is what tells you, in minutes, which downstream dates and deposits move and who has to be told. Without it, the same event becomes a forensic reconstruction conducted under deposition.
Cite this chapter
Fehn, J. (2026). Program & Project Management: The Integrated Master Schedule & Critical Path (Chapter 2.1). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-2-project-delivery-schedule-procurement-contracts-and-risk/2-1-program-and-project-management-the-integrated-master-schedule-and-critical-p (accessed 2026-09-29).
@misc{aidc-2-1,
author = {Fehn, Jacob},
title = {Program & Project Management: The Integrated Master Schedule & Critical Path (Chapter 2.1)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-2-project-delivery-schedule-procurement-contracts-and-risk/2-1-program-and-project-management-the-integrated-master-schedule-and-critical-p},
note = {Accessed 2026-09-29}
}