Chapter 7.13
In this chapter · 7 sections
The Rack as Integration Unit
The rack is now the unit you buy, ship, power, cool, cable and certify as one object, so buying one commits the floor, busbar, manifold and cabling as well as the equipment; every populated and reserved interface must fit, or a refresh inherits building work that replacing the rack alone cannot supply.
What you'll decide here
- Whether the system implements 19-inch EIA-310, 21-inch OCP ORV3 or Open Rack Wide: busbar pitch, manifold geometry, cable trays and BBU shelves must mate, while mounting width remains distinct from the external and service envelope.
- Where the rack is integrated (factory L11 vs field) and therefore whether it ships wet or dry, at 1.5–3 t, with the cabling and coolant already plumbed — the call that sets your install velocity and your serviceability ceiling.
- How the four spines — power (busbar), cooling (manifold + QDs), network (copper spine vs structured fiber), and data (DCIM/digital twin) — share the same rear plane of a frame roughly 600 mm wide and more than 1,200 mm deep without colliding, because rack real estate is the constraint nobody budgets until it is gone.
- Whether the in-rack network is a copper NVLink spine (cheap, low-power, reach-limited) or structured fiber to a row manifold — the copper-vs-optics decision made canonical elsewhere but physically resolved here, at the rack.
- Which future power and fluid interfaces justify a priced reservation today, and when replacement is cheaper: you cannot retrofit missing width into a frame already cut for the wrong rack.
For thirty years the rack was the least interesting object in the data center: a 19-inch steel cabinet, 42 rack-units tall, that you bolted to the floor and slid servers into. It had no opinion about power, cooling, or networking — those arrived from elsewhere, and the rack merely held the boxes that consumed them. AI ended that. A GB200 NVL72 is best read not as 72 GPUs in a rack but as a rack that happens to contain 72 GPUs — factory-integrated compute trays, power shelves, DC busbar, liquid manifolds, NVLink spine, top-of-rack switches and service cabling, built and validated as one system. The HPE GB200 QuickSpecs gives 3,245 lb ≈ 1.47 t fully loaded with PGW; plan lifting against the selected shipping inventory. Separately, NVIDIA’s October 15, 2024 OCP design describes a 1,400 A busbar circuit: nominal 48 V × 1,400 A ≈ 67 kW, which cannot supply the complete HPE rack’s 132 kW nominal duty. The HPE profile specifies 50 V DC distribution and multiple power shelves; its mass and supply boundary stay together. For a rack-scale scale-up platform, the smallest thing you can meaningfully buy, ship, power, cool, cable, and accept is the whole frame — the integration unit. Air-cooled PCIe and HGX-class servers still sell by the box, and a workload that does not need a 72-GPU coherent domain should still be bought that way (Chapter 7.8).
This is the canonical home for the rack as an object — its anatomy, its form factors, and the way every other physical subsystem now lands on it — and it deliberately points outward. The power chain inside the rack is engineered in Part 4; the liquid loop in Part 5; the civil mass and anchoring in Chapter 6.7; the fabrics in Part 8; the operational telemetry in Part 14. The integration discipline is what stays here: which rack standard to build a facility around, and the fact that every interior subsystem — busbar pitch, manifold geometry, cable management, BBU shelf, DCIM model — is dimensioned to that frame and inherits its constraints for a decade. The frame outlives three GPU generations, so cutting it to the wrong width strands the floor, the busbar, the manifold, and the cabling at once.
Defining the rack as the IT integration unit
The integration unit is the level at which a thing is designed, procured, tested, and warranted as one. For decades that level was the server: you bought a 2U box, the vendor warranted it, and the rack was just where it lived. The scale-up domain pushed the unit up a level. When 72 GPUs must behave as one coherent accelerator over a copper NVLink spine that tolerates only sub-metre reach, the 72 GPUs, the nine NVSwitch trays, the spine, the busbar, and the manifold cannot be assembled by independent vendors and hoped to fit — they must be co-designed against a single mechanical, electrical, and thermal contract. That contract is the rack. So the rack is now the smallest independently meaningful object: you cannot buy a third of an NVL72, and you cannot service one GPU without respecting the loop and the spine that bind the whole.
That is why the rack is the natural home for an integration chapter that points outward. The rack is where the four facility spines — power, cooling, network, and data — converge, terminate, and must coexist in the rear plane of a frame roughly 600 mm wide externally and more than 1,200 mm deep. Each spine has a canonical chapter of its own, and none can be designed without knowing the geometry of the frame it lands in. The civil scope — how a 1.5–3 t object distributes load into the slab and resists seismic events — is its own discipline and lives in Chapter 6.7; this chapter owns everything from the frame inward to the point where each spine becomes canonical elsewhere.
Rack anatomy and form factors
Three frame lineages dominate in 2026, and the differences are not cosmetic — they propagate into every spine. The legacy line is 19" EIA-310: the universal enterprise rack, ~600 mm external width, nominal 19-inch equipment mounting width; clear aperture and external dimensions require the frame drawing, measured in 1.75" rack-units (U). It is the default for air-cooled and modest-density inference, and it is what most existing colocation halls are striped for. The OCP line is 21" Open Rack V3: a wider IT-usable aperture (~537 mm) measured in OpenU (OU, 48 mm — taller than a U), a 48 V DC blade busbar running the full height of the rear, and side-pocket volume for power shelves and BBUs. ORV3 was designed for hyperscale serviceability and is the baseline that the ±400 VDC Diablo 400 and 800 VDC sidecar designs extend. The newest line is Open Rack Wide (ORW) — Meta's 2025 contribution, a wider system standard with its own dimensions and assembly documents — built specifically to give rack-scale accelerators like AMD's 72-GPU Helios the lateral room for trays, manifolds, and a scale-up fabric that a single-wide frame cannot hold.
The second axis is depth, and it has grown faster than width. Enterprise racks lived comfortably at ~1,070–1,200 mm. AI racks have pushed past ~1,200 mm and are heading deeper because the rear of the rack is now contested territory: liquid manifolds, the named OEM’s quick-disconnect count, blind-mate power connectors, and dense cabling all compete for the same plane. NVIDIA's NVL72 MGX rack and its 44RU/48RU manifold variants exist precisely to package that rear congestion. Depth growth is not free — it consumes white-space floor area per rack, lengthens cable and hose runs, and complicates rear-aisle service clearance. The third axis is weight, which is a civil problem (Chapter 6.7) but originates in the frame: a populated MGX GB200 and its transport package have different mass and point-load records, and the selected OEM must supply both, which is why the frame, the slab, and the anchoring must be co-specified.
| Frame standard | Width (ext / IT-usable) | Vertical unit & busbar | Native IT / best-fit | Strands if mismatched |
|---|---|---|---|---|
| 19" EIA-310 | ~600 mm / 450 mm (17.72") | U (1.75"); AC whips or in-rack PDU; no native DC busbar | Air-cooled & modest inference; most existing colo halls; HGX OEM boxes | Caps density at the air/RDHx ceiling; no room for a deep DC busbar or wide manifold |
| 21" OCP ORV3 | ~600 mm / ~537 mm | OU (48 mm); 48 V rear blade busbar (~33 kW/shelf); side pockets for BBU/PSU | Hyperscale self-design; the baseline ±400/800 VDC sidecar designs extend | Different rail geometry and a 48 mm OU pitch against 19"'s 44.45 mm U — 19" rails, trays and shelves do not carry over; needs OCP-spec power & trays |
| MGX (19"-derived deep) | ~600 mm / 19"-class, deep frame >1,200 mm | 1RU (44.45 mm) liquid-cooled compute and switch trays; 1,400 A liquid-adjacent busbar; blind-mate manifold | NVIDIA GB200/GB300 NVL72 (the integrated, factory-built unit) | Vendor-coupled; rear plane fully consumed by manifold + QDs + NVLink spine |
| Open Rack Wide (ORW) | Double-wide (~2x ORV3) | OU; wide DC busbar; lateral room for trays + manifold + fabric | AMD Helios 72-GPU; Meta 2025 rack-scale; UALink scale-up | Doubles aisle footprint; a single-wide hall cannot absorb it without re-striping |
The rightmost column is the price of getting the frame wrong. A hall striped for 19-inch equipment needs an ORW swept-volume and route check before deciding whether containment or aisle changes are necessary; an ORV3 power shelf does not establish compatibility with an arbitrary 19-inch rack; an MGX rack's rear plane is so fully consumed by its manifold and NVLink spine that a cooling or cabling retrofit needs a measured service-envelope and routing check. The frame is the first irreversible interior decision, and the three lineages do not interoperate at the subsystem level even where their external footprint looks similar.
In-rack power delivery and the busbar
The power spine is where the frame standard bites first. In the legacy 19" world, power arrives as AC whips to in-rack PDUs — fine to ~20–40 kW, hopeless above it. The OCP world replaced whips with a vertical blade busbar running the full height of the rear: ORV3 operates at 47.5–50.5 V, delivers ~660 A and ~33 kW per power shelf, and a GB200-class rack stacks six to eight shelves. Blind-mate connectors let a tray drop into the busbar without a single bolted cable — a serviceability win that only exists because the busbar geometry is part of the frame standard. The NVL72 MGX busbar is rated ~1,400 A — NVIDIA's OCP contribution keeps the ORV3 width and deepens the profile to roughly double the ampacity. Do not derive rack power by adding supply and return: a conductor pair carries one circuit current, so 1,400 A at 48 V is ~67 kW delivered, and a ~132 kW continuous rack needs on the order of 2,750 A of aggregate circuit current — from parallel feeds or distributed injection whose topology and per-segment ratings you take from the OEM power-distribution drawing, not by back-solving the rack's nameplate wattage. This entire power chain — busbar ampacity, shelf ratings, PSU efficiency, the AC-to-DC topology — is engineered canonically in Chapter 4.6 (LV distribution, busway, PDUs) and Chapter 4.7 (the ±400/800 VDC and disaggregated-sidecar revolution). What lives here is the geometry: the busbar occupies a fixed plane of the rear, and it competes with the manifold and the cabling for that plane.
The density ramp is what forces the frame to anticipate a power architecture it does not yet run. At 54 V, a 1 MW rack implies ~18,500 A and ~200 kg of copper busbar per rack — a physical impossibility to scale, which is the entire reason the industry is moving to ±400/800 VDC, where NVIDIA's three-wire-DC versus four-conductor-AC comparison puts 800 V at ~157% more power in the same copper than 415 V. A frame cut today for a 48 V blade busbar may have to host an 800 VDC busbar fed from a sidecar power rack in two generations — and the transition has a purchasable on-ramp: NVIDIA's 800 V power rack ships as an optional Vera Rubin configuration from 3Q26 (TrendForce, Jun 2026), one generation before Kyber makes it mandatory. The BBU shelf — rack-level battery ride-through that triggers in milliseconds to absorb the synchronized GPU load step — also lives in the frame's side pockets and is made canonical in Chapter 4.5. The integration question the rack owns: does the frame reserve the volume and the busbar plane for the next voltage class, or does it lock you to 48 V?
Liquid distribution at the rack
The cooling spine is the most congested object in the rear of the rack and the one most coupled to the frame. A direct-to-chip liquid-cooled rack carries a vertical manifold pair (supply and return) feeding cold plates on every GPU and switch, connected through ~150–200 quick-disconnects (QDs) per rack — each a potential leak point, each a place where a tray must blind-mate into the loop without an operator coupling a hose by hand. The HPE GB200 NVL72 worked example derives ~165–236 L/min from its 115 kW liquid load at a 7–10 °C design rise; Clariant PG25 gives ~172–246 L/min at the same rise, roughly 4% more. The separate QCT GB200 NVL72 reference permits up to 45 °C liquid inlet and 65 °C liquid return; these are independent maxima, not the flow-sizing delta-T. ASHRAE W45 describes FWS supply capability, not a TCS operating pair. Colder supported supply buys thermal headroom at the cost of chiller capex. The dripless, blind-mate UQD/UQDB couplings that make this serviceable are an OCP-standardized interface, which is precisely why the manifold and QD geometry must conform to the frame standard — a manifold designed for an MGX rear plane will not fit an ORW one. The full thermal contract (CDU sizing, flow, delta-T, coolant chemistry, the technology-cooling vs facility-water loop split) is canonical in Chapter 5.4; the density wall that makes liquid mandatory is Chapter 5.1.
The integration decisions the rack owns are mechanical and operational. Leak detection — rope sensors at the manifold, drip trays, and per-tray flow monitoring — is a rack-level instrumentation layer, and its telemetry feeds the DCIM model. The emerging negative-pressure (sub-atmospheric) design is a rack-architecture choice with a real consequence: running the loop below ambient pressure means a breach pulls air in rather than pushing coolant out, so a leak becomes a detectable, contained event instead of coolant escaping under pressure across the loop. That safety property is bought with more complex pumping and tighter CDU control, and it is a decision made at the rack/CDU boundary before the loop is plumbed. With ~150–200 coupling points per rack and thousands of racks, expected leak events scale with the joint count — negative pressure is the lever that converts a catastrophic failure mode into a maintenance event.
In-rack network cabling: copper spine vs structured fiber
The network spine resolves a decision made canonical elsewhere but physically settled here, at the rack. Inside a scale-up domain, the choice is copper vs optics. NVIDIA's NVL72 keeps the entire 72-GPU NVLink scale-up fabric on copper — more than 5,000 copper cables in NVIDIA’s October 2024 OCP description. Short passive links avoid optical conversion at their endpoints; the engineering benefit is fewer powered optical components, with their heat, cost and failure exposure. The counterfactual rack saving requires a matched optical BOM and cannot be inferred from cable count. The catch is reach: the channel’s lane rate, insertion loss and host equalization set passive reach; a rack-cluster crossing needs its own qualified copper, active-copper or optical design. The copper-vs-optics economics and the NVLink fabric itself are canonical in Chapter 8.2; the physical-layer interconnect taxonomy in Chapter 8.9.
The scale-out (back-end) and front-end fabrics are a different physical animal: structured fiber, pre-terminated MPO trunks running from the rack to a row or end-of-row manifold, polarity-disciplined and loss-budgeted. The integration decision the rack owns is the boundary: where does the copper spine end and the fiber plant begin, and how does the rear plane accommodate both without the dense copper backplane fouling the fiber breakout. Pre-terminated trunk cabling is the primary install-velocity lever — it is what lets a rack land and link in hours instead of days — and that velocity story is owned in Chapter 7.15. The rack-level consequence: a frame that does not reserve cable-management volume for both a copper spine and a fiber breakout forces a field re-work that no amount of factory integration can recover.
Scope & caveats
Exact HPE product profile. Keep this 132 kW / 115 kW liquid / 17 kW air record separate from the OCP MGX Rev. 1.1 reference profile of 120 kW / approximately 102 kW liquid / 18 kW air.
Scope & caveats
NVIDIA's OCP contribution states a 1,400 A bus bar and 1RU liquid-cooled compute and switch trays. It does not establish a two-feed circuit: a supply and its return carry the same current, so 1,400 A at 48 V is ~67 kW delivered. A ~132 kW continuous rack implies ~2,750 A of aggregate circuit current from parallel feeds or distributed injection; take that topology from the OEM power-distribution drawing.
Scope & caveats
Guide water-property heat balance for the HPE 115 kW liquid load at a 10–7 K operating rise: density 1.00 kg/L and heat capacity 4.18 kJ/(kg·K). A derived design illustration, not an OEM flow requirement. The separate QCT 45 °C inlet and 65 °C return maxima do not prescribe this rise.
The operating flow must also satisfy the selected rack and CDU pressure/flow envelope.
Scope & caveats
NVIDIA's published figure (GTC 2025) is 600 kW per Rubin Ultra Kyber rack and GTC 2026 did not revise it. SemiAnalysis (2026-05-26) reports Kyber Ultra 'approaching 660 kW' — a single-source analyst estimate for a 2027 part, recorded here rather than adopted, since the vendor primary figure still stands.
The physical DCIM layer: the rack's digital twin
The fourth spine is data, and the rack is where it becomes physical. A rack record can use drawings, asset tables and port maps or a 3D digital twin of the rack and the hall, with every device mapped to a U/OU position, every port mapped to its peer, and a physical-layer telemetry feed (power per phase, coolant flow and temperature, leak-sensor state, busbar current, per-tray draw) streaming from the rack into the model. The integration discipline the rack owns is making the physical object legible: asset and port mapping accurate enough that a remote-hands technician can be directed to the exact tray and the exact QD, and a twin faithful enough that capacity, thermal, and cabling decisions can be simulated before steel moves. For a hall with thousands of liquid-cooled racks and hundreds of thousands of QDs, accurate geometry, connectivity and telemetry mappings are required to reason about where the next leak, the next thermal margin breach, or the next stranded U of capacity will appear.
The line between this chapter and operations is clean: the physical-layer twin — geometry, asset/port mapping, the sensor wiring — is a rack-integration deliverable and belongs here. The operational DCIM — alerting, predictive maintenance, agentic ops, the run-time observability that consumes this telemetry — is canonical in Chapter 14.2, and the failure-rate data that the telemetry exposes is Chapter 14.3. The rack's job is to be instrumented and mapped at integration time; operations' job is to act on what that instrumentation reports. A rack that ships without an accurate physical record and a wired telemetry plane is a rack the fleet cannot see, and at its contracted rack power and coupling count that is where the next undiagnosed leak or thermal breach lands.
Deep dive: factory L11 integration vs field integration — the wet-vs-dry shipping fork
Because the rack is the integration unit, the question of where it gets integrated is a first-class decision with consequences that ripple through install velocity, serviceability, and logistics. The manufacturing-level model (L1 component → L6 board → L10 server → L11 integrated rack → L12 multi-rack/pod, made canonical in Chapter 7.14) puts the fork at L11: do you ship trays and a bare frame and integrate on the floor, or do you ship a fully-built, cabled, and plumbed rack from the factory?
Factory L11 integration is the modern default for rack-scale systems because it compresses floor time from days of skilled cabling-and-plumbing labor to a crane lift and a few interface connections — the velocity lever behind Chapter 7.15. But it forces a sub-decision: ship wet or dry? A wet rack arrives with coolant already in the loop — fastest to commission, but it ships the populated rack plus its retained liquid inventory, with shock-and-tilt damage risk and freight-handling constraints that a dry rack avoids. A dry rack ships lighter and safer but defers the fill-and-leak-test to the floor, re-introducing some of the labor and risk that factory integration was meant to remove. The serviceability consequence is the mirror image: a deeply factory-integrated rack with a fully-consumed rear plane is fast to install but harder to service in the field, because the blind-mate density that makes assembly clean makes a single-component swap intricate.
The decision is therefore a function of distance-to-site, the maturity of the field integration crew, and how much the workload values time-to-goodput over field serviceability. Hyperscalers with disciplined logistics and short hauls lean wet-and-factory; operators with long international freight legs or thin field crews often ship dry and accept the floor labor. The right choice depends on your logistics envelope — model the freight and the field-crew maturity before committing, because picking blind risks either shock damage in transit or floor labor you did not budget.
Deep dive: why the frame outlives the silicon — designing the reserve you cannot retrofit
The rack standard endures because of a timescale mismatch. GPU roadmaps advance roughly annually; fleet replacement follows useful-service economics; a frame, a busbar plane, an aisle pitch, and a hall's containment scheme can outlast several silicon purchases. Possible future profiles include a density ramp from ~132 kW (GB200 NVL72, 2024) through 330 kW (Vera Rubin NVL72 facility design basis, 2026) to ~600 kW (Rubin Ultra Kyber, 2027) and a power architecture transition from 48 V blade busbar to 800 VDC sidecar. A frame modification needs structural, electrical and cooling requalification, so the integration discipline is to identify which interfaces are reservable and reserve them, and which are locked and must be chosen for the endpoint.
Reservable: rear-plane volume for a deeper manifold and more QDs; busbar-plane depth for a higher-voltage bus; cable-management volume for both a copper spine and a fiber breakout; side-pocket space for a larger BBU or a sidecar power feed; floor loading and anchoring for 2–3 t (the civil reserve of Chapter 6.7). Locked: the external width and the aisle pitch — an ORW conversion depends on measured external width, route, aisle pitch and service sweep. The doctrine that follows is the same one the economics chapters call build upgrade-ready: populate today's silicon, but price a reserved interface against its enabling works, expiry and demand trigger in Chapter 1.8 — because the alternative is a 'concrete husk,' a hall whose frame standard stranded it one generation early. The reserve costs an option premium now; skipping it can strand the building one generation early.
The integration discipline, summarized
The rack is integrated, not assembled: the four spines are co-designed against one frame rather than independently specified and hoped to fit. The frame is the first decision and the most durable, chosen against the workload and the density ramp, not against the existing hall. The rear plane is the contest a single integrator must arbitrate; the wet-vs-dry shipping fork trades commissioning speed against freight risk and field serviceability; and the digital twin must be wired and mapped at integration time so the fleet can see a 132 kW object with 150–200 coupling points. A frame chosen well lets every downstream subsystem land cleanly; a frame chosen wrong strands the floor, the power, the cooling, and the cabling together.
Scope & caveats
Hypothetical rack and hall, separate air/liquid duties. Unsupported limits and dates are explained in the opening callout; Chapters 5.1, 5.7, 6.7 and 2.1 own the methods.
Check each heat-removal path independently. Coolant temperature rise is liquid duty divided by mass flow and specific heat; add inlet temperature for return temperature. Subtract required service clearance from available clearance. Release date is the latest required interface date plus acceptance duration; nominal schedule float cannot waive a failed thermal or access interface.
Scope & caveats
Electrical margin = 132−120 = 12 kW; liquid-duty margin = 120−100 = 20 kW; room-air deficit = 20−18 = 2.0 kW. Coolant rise = 100 kW/(2.0 kg/s × 4.18 kJ/(kg·K)) = about 12 K, so return is about 52°C, below 55°C at the assumed flow. Rear-service deficit = 0.90−0.80 = 0.10 m. The nominal release is max(9,11,10,12) + 2 = week 14, but air and access failures block it. Both must be corrected by week 12 for the two-week test to finish by week 14; correcting them in week 15 instead gives week 17.
Reject an unconditional rack release despite spare electrical and liquid capacity. Add qualified room-air removal and clear the real service sweep, or obtain a different complete OEM profile; do not invent a lower-power mode by scaling heat splits. Record voltage/return/ground/protection, fluid chemistry/pressure/QDs, point loads, route, fabric and management ports alongside these arithmetic gates, then carry the dated enabling works into the integrated schedule.
Method: Open Compute Project Open Rack specifications. Chapter 2.1 owns the next handoff.
Cite this chapter
Fehn, J. (2026). The Rack as Integration Unit (Chapter 7.13). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-13-the-rack-as-integration-unit (accessed 2026-09-29).
@misc{aidc-7-13,
author = {Fehn, Jacob},
title = {The Rack as Integration Unit (Chapter 7.13)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-13-the-rack-as-integration-unit},
note = {Accessed 2026-09-29}
}