Chapter 7.14
In this chapter · 7 sections
Server & System Integration
Where you enter the integration pipeline — DGX, HGX-OEM, ODM-direct, or OCP self-design — sets who owns burn-in, the acceptance gate, and the RMA, and the days between a powered shell and a producing cluster.
What you'll decide here
- The integration model — DGX/turnkey vs HGX-from-an-OEM vs ODM-direct vs OCP self-design — which sets your margin stack, your serviceability terms, and how much systems-integration risk you are insourcing.
- Factory vs field integration, with deliverables and acceptance defined contractually because L-level numbering varies by vendor — where the rack actually gets built and tested, and therefore whether you ship wet or dry, how you move a 1.5–3 t rack, and how much install-day risk you carry.
- Which workload, quality, latency, power and recovery gates the contracted burn-in must prove, for the specified duration and signatory; a power-on smoke test alone cannot release the system.
- Which qualified CoWoS, HBM, board, integration or site milestone controls delivery, because an unfinished upstream allocation can leave the rack-assembly line idle and slip the build plan.
- Which spares, RMA ownership, firmware support and tray-repair commitments to buy: every hour to restore the tray costs useful output, so price repair time against this service’s failure and recovery evidence.
A modern AI rack is not a product you buy off a shelf and bolt to the floor. It is the output of a manufacturing pipeline that starts at silicon and ends at a benchmarked cluster, and what most shapes its cost and risk is where the operator enters it. Enter at the top, and a vendor hands you a turnkey, validated NVL72 with a warranty and a phone number. Enter at the bottom, and you are the systems integrator: you own the bill of materials, the firmware matrix, the burn-in scripts, the acceptance gate, and every tray that fails at 3 a.m. The two ends differ in quoted price, time-to-goodput and contractual liability when the cluster misses its acceptance target; a vendor logo does not assign that liability. That entry decision sets the margin stack, the time-to-goodput, and who owns the useful-output acceptance gate.
It starts with the guide's L1–L12 integration shorthand; vendor and integrator numbering varies, so contracts must define deliverables and acceptance rather than rely on a level number alone to say who builds what — and the ODM / OEM / systems-integrator roles mapped onto it. From there: the build-vs-buy fork (DGX vs HGX-OEM vs ODM-direct vs OCP self-design), the OCP Open Rack standards and the 2026 reference systems (HGX, MGX, GB200/GB300 NVL72, AMD Helios), the factory-vs-field integration question and the logistics of shipping a wet 1.5–3 t rack, the goodput-oriented acceptance gate, the separate CoWoS/HBM delivery milestones that can gate the system, and the commissioning handoff to operations. The rack as a physical integration unit is treated in Chapter 7.13; this chapter is about integrating it.
Guide shorthand for integration stages (vendor numbering varies)
The industry talks about hardware integration in numbered "levels," and getting fluent in them is the precondition for every contract you sign. The scale runs from raw components to a benchmarked cluster, and the level at which a vendor delivers is exactly the line that separates "you bought a server" from "you bought a working AI factory." The numbering varies slightly by vendor, but the structure is stable: L1–L5 are component and PCB assembly (bare board, SMT placement, the GPU baseboard or UBB); L6 is the populated motherboard/baseboard; L10 is a fully assembled server that boots an OS; L11 is a fully cabled rack — compute trays, NVSwitch trays, busbar, manifolds, top-of-rack switches, in-rack network and power cabling, tested as a unit; and L12 is a multi-rack cluster with cross-rack cabling, the customer's software loaded, and rack-scale benchmarks run to prove the thing actually performs (DCD; AMAX; Hyperscalers, 2025).
The classic division of labor: ODMs (Foxconn/Hon Hai, Quanta/QCT, Wistron/Wiwynn, Inventec, Supermicro on its ODM side) own roughly L1–L6 and increasingly push up into L10–L11; OEMs (Dell, HPE, Lenovo, Supermicro on its brand side) take L6–L10 and add brand, warranty, supply assurance, and a global service organization; the systems integrator — which can be the OEM, a specialist, or the operator itself — owns L11–L12, the part where validated servers become a producing cluster. The L11/L12 boundary is where time-to-goodput is won or lost: whoever owns it owns the burn-in, the acceptance gate, and the install-day risk.
| Level | What it produces | Typical owner | What you are buying | Where the risk sits |
|---|---|---|---|---|
| L1–L5 | Bare PCB → SMT-populated board → GPU baseboard (UBB/SXM) | ODM / contract manufacturer | Components and sub-assemblies | Yield, HBM/CoWoS supply |
| L6 | Populated motherboard / GPU baseboard | ODM (handed to OEM) | A tested board | Firmware, board-level defects |
| L10 | Fully assembled server that boots an OS | OEM / ODM | A working node | Node burn-in, thermal validation |
| L11 | Fully cabled, tested rack (trays, busbar, manifolds, ToR, cabling) | OEM / systems integrator | A deployable rack | Mis-cabling, leak test, rack-level burn-in |
| L12 | Multi-rack cluster, cross-rack cabling, software loaded, benchmarked | Systems integrator / operator | A producing cluster | Goodput/MFU acceptance, fabric validation |
Build vs buy: the four entry points
Map the entry decision onto four archetypes, ordered from most-bought to most-built. Each trades margin paid against integration risk insourced and control gained. There is no universally right answer — the right cell is a function of your scale, your engineering depth, and how much of the systems-integration burden you can actually carry.
DGX / turnkey (NVIDIA DGX, the GB-series "NVL72" sold as a system). You buy a fully-integrated, factory-validated, single-throat-to-choke supercomputer with NVIDIA's software stack, reference fabric, and warranty. A single accountable supplier can reduce integration work; price, acceptance scope and delivery still depend on the contract; the software and fabric still carry switching costs (Chapter 7.9). This is the right call for an enterprise standing up its first cluster or anyone who values a single accountable vendor over unit economics.
HGX-from-an-OEM (Dell, HPE, Lenovo, Supermicro building on the NVIDIA HGX baseboard or MGX modular server reference architecture). The middle path and the volume of the market. NVIDIA sells the HGX 8-GPU baseboard (or an MGX modular server reference architecture); the OEM does L6–L11 integration, adds its own chassis, thermals, BMC, service, and supply assurance. You get brand-name support and a global RMA org while escaping the full DGX premium. The cost: you inherit the OEM's firmware/validation cadence and pay an integration margin the ODM-direct buyer skips.
ODM-direct (buying L10/L11 straight from Quanta, Wiwynn, Foxconn, Supermicro's ODM arm). You strip out the OEM brand margin and contract the integrator directly, often to your own spec. Potentially lower purchase price and more control over BOM and firmware — but you are now closer to owning the acceptance gate and the RMA logistics yourself. This is the hyperscaler and large-neocloud default once volume justifies an in-house hardware team.
OCP self-design (you specify the rack against Open Compute standards and have ODMs build to it). Maximum design control, with design and sustaining costs inside your organization — and you are the systems integrator. You own the design, the BOM, the firmware matrix, the burn-in scripts, the acceptance criteria, and every serviceability decision. Justified when contracted volume and service benefits cover the design and sustaining organization, where a point of efficiency across hundreds of thousands of GPUs dwarfs the cost of an in-house infrastructure org. Meta, Microsoft, Google, and Amazon live here.
| Entry point | Who integrates L11/L12 | Relative unit cost | Integration risk you own | Best-fit buyer |
|---|---|---|---|---|
| DGX / turnkey NVL72 | Vendor (factory-validated) | Highest (full premium) | Minimal — vendor owns the gate | First cluster; single-vendor accountability |
| HGX-from-an-OEM | OEM (Dell/HPE/Lenovo/SMCI) | High (brand + integration margin) | Low — OEM warranty & RMA | Enterprise/mid-scale wanting brand support |
| ODM-direct | ODM, to your spec | Low (no brand margin) | Moderate — you co-own acceptance | Large neoclouds; in-house HW team |
| OCP self-design | You (the operator) | Lowest at scale | Full — you are the integrator | Hyperscalers; fleets >100k GPUs |
OCP, Open Rack & the 2026 reference systems
The Open Compute Project is the standards substrate that makes ODM-direct and self-design viable: it turns proprietary rack designs into shared, multi-vendor specifications so an operator can qualify a second source for the same rack from Quanta, Wiwynn, or Foxconn instead of being captive to one builder. The relevant standards for AI in 2026 are the Open Rack family. ORV3 (Open Rack v3, Meta-led, published 2022) moved the industry to a 21-inch rack with a vertical DC busbar, blind-mate power, native 48 V distribution, and provisions for direct liquid cooling — the form factor most current high-density AI racks descend from (OCP, 2025). At OCP 2025, Meta introduced Open Rack Wide (ORW), a double-wide standard explicitly designed for the power, cooling, and serviceability demands of next-generation rack-scale AI — the spec AMD's Helios is built on (OCP / Meta; AMD, 2025).
The reference systems are where these standards meet silicon. HGX is NVIDIA's 8-GPU baseboard reference — the building block OEMs integrate into air- or liquid-cooled servers; it is the enterprise and on-prem workhorse for inference and small training, while frontier serving deploys on NVL72-class racks alongside training. MGX is NVIDIA's modular server reference architecture that lets partners mix CPUs, GPUs, and DPUs into standardized rack designs. GB200 and GB300 NVL72 each join 72 GPUs and 36 Grace CPUs in a liquid-cooled NVLink domain. Their facility records remain separate: HPE GB200 specifies 132 kW nominal TDP split into 115 kW liquid and 17 kW air; Lenovo GB300 specifies 135 kW nominal and 155 kW peak. NVIDIA’s GB300 facility allowance is a separate 142 kW basis, not Lenovo’s operating mode. Mass, coolant inventory and the shipping/operating envelope must come from the chosen OEM configuration in Chapter 7.13. AMD Helios is the open challenger: an ORW double-wide rack carrying up to 72 MI455X (MI450-series) GPUs, ~1.4 EF FP8 / 2.9 EF FP4 and 31 TB HBM4 at rack scale, described against OCP, UALoE scale-up and Ultra Ethernet scale-out interfaces; implementation and interoperability require separate evidence — the open-standards answer to a single-vendor NVL72 (AMD; NextPlatform; DCD, 2025–2026).
The OCP catalogue lists Open Rack V3 Base 1.1 (December 2023), Meta Frame V3 1.3 (June 2024), ORW Base 1.0 and ORW Meta Design 1.0 (April 2026) as separate documents. The ORW Canister preliminary guide 0.1 remains under review. Preserve the named manifold, power-shelf, BBU, frame and canister revisions in the assembly schedule; a base-specification badge cannot replace them.
Use the ASHRAE–PNNL–NEMA AI Data Center Energy Performance Framework, released June 10, 2026, to coordinate IT and facility performance boundaries. It is guidance for integrated design and operation, not a rack certification or a substitute for applicable codes. Record which workload, energy, water and recovery evidence the project’s design basis requires. Appendix A owns the standards index, with Chapter 0.4 as intake; Chapter 13.1 opens the commissioning program. This chapter carries their requirements into the system integrator’s acceptance and handoff contract.
Scope & caveats
ORV3 base document only; Meta Frame V3 is separate. Revision as listed on the publisher page, which states no publication date.
Scope & caveats
Frame document, separate from ORV3 Base 1.1 and ORW documents.
Scope & caveats
ORW base document only; Meta Design and preliminary canister guide are separate. Revision as listed on the publisher page, which states no publication date.
Scope & caveats
Meta Design document; distinct from the ORW base specification and preliminary canister guide. Revision as listed on the publisher page, which states no publication date.
Scope & caveats
Preliminary canister document under review; distinct from the issued base and Meta Design documents. Revision as listed on the publisher page, which states no publication date.
Scope & caveats
Lifecycle design, commissioning, operation and retrofit guidance; does not supersede codes or standards.
Scope & caveats
Official specification history lists 1.0.3 as current; no public 1.1 established. Initial 1.0 release retained separately.
| System | Unit of integration | Accelerators | Power / weight | Fabric & standards posture |
|---|---|---|---|---|
| NVIDIA HGX (B200/B300) | 8-GPU server baseboard | 8 Blackwell/Ultra | ~30–60 kW/rack (air or liquid) | NVLink in-board; vendor-proprietary |
| NVIDIA MGX | Modular rack reference | Mix-and-match GPU/CPU/DPU | Density by configuration | NVLink/NVSwitch; NVIDIA reference |
| GB200 NVL72 | The rack (72-GPU NVLink domain) | 72 Blackwell + 36 Grace | HPE GB200: 132 kW = 115 liquid + 17 air; separately, 3,245 lb fully loaded with PGW | NVLink5 (130 TB/s rack); proprietary |
| GB300 NVL72 | The rack (Blackwell Ultra) | 72 Blackwell Ultra + 36 Grace | Lenovo 135 kW TDP / up to 155 kW peak; NVIDIA facility design basis up to 142 kW | NVLink5; ~90% liquid / ~10% air |
| Vera Rubin NVL72 | The rack (72-GPU NVLink domain) | 72 Rubin + 36 Vera | 188 kW Max Q / 228 kW Max P; 330 kW facility design basis (cabinet TDP); ~3.6 EF FP4 | NVLink6 (~260 TB/s rack); proprietary — 1st rack validated Jun 2026; partner availability planned H2 2026 |
| AMD Helios (ORW) | Double-wide rack | Up to 72 MI455X (MI450-series) | ORW double-wide; weight spread across 2 bays | UALink + Ultra Ethernet; OCP-open |
Read the last column as the real strategic axis. NVL72 is a vertically-integrated, single-vendor unit: you get a validated NVLink domain and a proprietary scale-up fabric, and you accept the lock-in. Helios is the open bet: a double-wide ORW rack on UALoE scale-up and Ultra Ethernet scale-out, designed to spread weight and improve serviceability by going wide rather than tall, with each OCP implementation separately qualified. Going wide changes load distribution and service volume; the actual point loads and swept volumes must resolve the floor-loading and field-serviceability problems that the selected single-bay NVL72 profile creates, which is the next section's subject. The scale-up fabric choices behind NVLink vs UALink are treated in Chapter 8.2; the merchant-vs-captive silicon framing in Chapter 7.1.
Factory vs field integration: where the rack gets built
Once you know who integrates, the next fork is where: is the rack built and tested at the factory (rack integration and factory acceptance completed before shipment; site/cluster acceptance follows installation) or assembled in the field at your site? This is the central velocity decision of the deployment, and it pivots on a hard physical fact: a populated NVL72 is a 1.5–3 t object that holds an OEM-specific internal coolant inventory and carries thousands of in-rack cables. Take its operating mass and point loads from the selected OEM configuration and its actual feet/wheel and support geometry — a footprint average is not a point-load rating — and do not infer the coolant inventory from a flow-rate figure.
Factory integration (ship the rack whole) is the 2026 default for dense liquid-cooled systems precisely because mis-cabling and leak risk are too high to absorb on the install floor. The integrator assembles trays, busbar, manifolds, ToR switches, and in-rack cabling in a controlled environment, runs rack-level burn-in, and ships a tested unit. NVIDIA's rack-scale partners explicitly factory-integrate the liquid loop and re-test at the rack level so the rack can be deployed directly at the customer site. Dell says its factory-validated PowerRack with PowerEdge XE9812 servers can move customers from delivery to production in under 6.5 hours; CoreWeave's June 2026 release separately confirms bring-up and system-level validation of its Vera Rubin NVL72 rack. The cost is logistics: you are now moving the named rack’s populated shipping mass, and the question becomes whether it ships wet (coolant already in the loop, factory-tested as-shipped) or dry (drained for transit, then filled and leak-tested in the field). Shipping wet preserves the factory test state and shaves field commissioning time but adds weight, freeze/spill risk, and stricter handling; shipping dry is lighter and safer in transit but reintroduces a fill-and-leak-test step on the critical path. Most high-density racks ship dry-of-coolant for transit and are filled on site (though hyperscalers with short, disciplined logistics hauls lean the other way and ship wet to preserve the factory test state — see Chapter 7.13), with the factory loop integrity certified separately — but the choice is contractual and worth pinning down explicitly.
Field integration (build the rack on site) — populating an empty rack with trays and cabling it in the data hall — survives only for lower-density, air-cooled, or 19-inch-EIA configurations where the weight and cabling risk are manageable. For NVL72-class systems it is an anti-pattern: you are doing precision liquid plumbing and thousands of cable terminations in an uncontrolled environment, against an install clock, with mis-cabling as the dominant acceptance failure (the velocity and cabling discipline this demands is the whole subject of Chapter 7.15).
Burn-in, validation & the goodput acceptance gate
The most expensive mistake in system integration is accepting a cluster on a power-on smoke test — "it boots, it pings, sign here." AI clusters fail in ways a smoke test never sees: a GPU that trains fine for an hour and then throttles on a thermal excursion, an HBM stack with marginal bit-error rates, an optic that flaps under load, a single mis-cabled link that quietly halves bisection bandwidth. The acceptance gate that catches these is goodput-oriented: a multi-day burn-in that drives the cluster at full power and measures whether it sustains its target goodput / MFU, not merely whether it powers on.
The empirical case is for a measured, evidence-bearing gate—not one portable duration. Meta's Llama 3 405B run logged 419 unplanned interruptions over 54 days on 16,384 H100s, about one every three hours, with 78% hardware-caused (Meta, 2024). SemiAnalysis's October 2024 operator playbook separately recommends at least 3–4 weeks of factory high-temperature burn-in before deployment and reports about 7 days MTBF for one 512-H100 cluster at a top-tier operator. Those are source-specific recommendation and observation, not a universal settling period or per-GPU rate. Contract the test intensity, fleet exposure, measured discovery curve, and statistical stopping rule so marginal components surface before production.
A defensible acceptance program therefore layers tests at each level: L10 node burn-in (thermal soak, memory test, per-GPU stress); L11 rack-level validation (leak test on the liquid loop, power-sequencing, in-rack link integrity, mis-cabling verification); and L12 cluster-level acceptance (collective-communication benchmarks like all-reduce bandwidth, a representative training run held to a target MFU, and a sustained multi-day goodput soak). The gate is contractual: it defines the number the integrator must hit, the duration the cluster must hold it, and the remedy if it does not. This connects directly to the formal commissioning levels in Part 13 — the integrated-systems and rack-scale acceptance machinery is built out in Chapter 13.1 and the cooling/electrical acceptance specifics in their respective chapters there.
Deep dive: what a goodput acceptance gate actually measures (and the failures it catches)
A goodput gate is not one test — it is a sequence designed so each layer catches a class of defect the layer below misses. Run them in order, because a fabric benchmark on a cluster with a thermally-marginal GPU just gives you a confusing number.
1. Component & node (L10). Per-GPU stress (compute + memory bandwidth at full TDP), HBM bit-error screening, and a thermal soak that holds the node at its power limit long enough to surface throttling. This is where the bulk of infant mortality — the faulty GPUs and marginal HBM stacks that dominated Meta's failure breakdown — is supposed to die before the rack is sealed.
2. Rack (L11). Liquid-loop leak test and pressure-hold; power-sequencing and busbar integrity; and the one that catches the most acceptance failures — cabling verification. A single transposed or under-seated link can pass a ping and still cripple collective bandwidth; automated link-map verification against the intended topology is the only reliable catch. The NVL72 packs thousands of in-rack copper NVLink cables, so the failure surface is large.
3. Cluster (L12). Collective benchmarks (all-reduce / all-gather bandwidth at scale, the operations a real training step is dominated by—see Chapter 8.2) prove the fabric against its signed traffic and performance requirements; then a sustained representative workload measures goodput for the named fleet, job, observation window, event taxonomy, checkpoint and recovery policy, and committed-progress rule. A 90%-versus-96% pair is an illustrative sensitivity, not an industry baseline or portable acceptance target. The output is a benchmarked, signed-off cluster—the L12 deliverable—with measured acceptance criteria, not merely a rack of servers that boots.
The real lead-time gate is upstream: CoWoS & HBM
It is tempting to plan the build around rack assembly — the visible, schedulable step. That is a mistake when the binding constraint sits two tiers upstream in advanced packaging: in 2026 TSMC’s CoWoS capacity is ramping toward ~120–140k wafers/month, against roughly a million wafers of reported allocation for the year, ~60% of it NVIDIA’s (Chapter 7.7 owns the series), and HBM is the companion gate (Chapter 7.6). The binding dependency can also be the logic wafer, board build or system acceptance, and TSMC, SK hynix, Samsung and Micron capacity disclosures are not a dated allocation to this rack. Record each predecessor’s committed lot, quantity, release evidence and substitute-qualification owner; Chapters 7.6 and 7.7 keep memory and packaging gates distinct.
The consequence for system integration is a sequencing rule: the latest unfinished dependency sets your delivery date. A flawless L11/L12 integration line standing idle waiting for accelerators is the default failure mode of a build plan that scheduled against the wrong constraint. The operators who deploy fastest secure dated CoWoS/HBM allocations before dependent commitments, treat the accelerator delivery curve as the master schedule, and stage the powered shell, cooling plant, and integration capacity to be waiting on silicon rather than the reverse. Procurement and allocation strategy belongs in Chapter 2.3; Chapter 2.1 owns the integrated schedule.
Scope & caveats
One HPE GB200 configuration, not GB300 or a generic NVL72 shipping mass. Guide unit conversion uses 0.45359237 kg/lb; about 1.47 t. Packaging, shipping state and foot/wheel/rigging reactions require the selected OEM inventory.
Scope & caveats
Size irreversible infrastructure to the facility design basis; run energy and TCO models on the operating profile.
Scope & caveats
Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.
Scope & caveats
Helios OEM/ODM reference design; the appendix uses the advertised configuration and separately records the delivery forecast.
Scope & caveats
Single validated rack; fleet general availability not yet. Dell factory-integration ('under 6.5 hours' delivery-to-production) per Dell; ~3.6 EFLOPS FP4 rack figure per DataCentre Magazine.
Deployment, commissioning & the operations handoff
The last act of integration is the handoff to operations — the moment the cluster stops being the integrator's project and becomes the operator's producing asset. A clean handoff is itself an acceptance gate: it transfers not just hardware but the documentation that makes the hardware operable — the as-built rack and link maps, the firmware/driver baseline, the burn-in and acceptance results, the asset-and-port inventory that feeds DCIM (Chapter 14.2), and the spares and RMA terms. Skip the documentation transfer and you have a cluster nobody can service without reverse-engineering it.
Spares, RMA, and serviceability are where the goodput thread closes the loop. In a cluster where one tray failure can interrupt synchronous work, mean-time-to-repair on that tray is not an operational footnote — it is a direct multiplier on effective availability and therefore on goodput. The serviceability decisions made at integration time govern it: front-serviceable trays vs racks you must pull from the aisle; blind-mate power and liquid quick-disconnects that let you swap a tray without draining the loop; an on-site spares depot sized to the fleet's failure rate rather than a vendor's standard SLA; and an RMA path whose turnaround you actually measured. The build-vs-buy fork resurfaces here: a turnkey buyer inherits the vendor's RMA org and SLA, while the OCP self-designer owns the spares pool and the repair logistics outright — another reason the entry decision is a multi-year operational commitment, not a one-time purchase. The reliability math behind why repair time dominates goodput is developed in Chapter 12.2; checkpointing, the software complement that bounds the cost of each failure, in Chapter 9.4.
Scope & caveats
Fixed manifest; unsupported synthetic outcomes and owner gates. Rationale is in the opening callout. Chapter 13.1 owns acceptance.
Useful throughput = 3,780,000 tokens/3,600 s = 1,050 tokens/s. The normal run clears 1,000 tokens/s, 89.0% quality and 50 ms p99. Recovery excess = 150−120 = 30 s, so the release is HOLD even though the normal run passes. The flip requires recovery ≤120 s, the same performance/quality gates, zero unresolved critical defects and successful inventory, authorized update, failed-update recovery and rollback checks on the restored manifest. A faster normal run alone cannot reverse the decision.
Scope & caveats
Same fixed manifest and unsupported owner gates as the input ledger. The worked trace states the arithmetic and every condition required to reverse HOLD.
Contract diagnosis, repair and full retest before transferring the accepted manifest, logs, spares and support contacts to operations.
Method: DMTF management and device-trust specifications; handoff: Chapter 13.1.
Operations handoff: firmware, telemetry and trust are tested interfaces
Ship a signed inventory and supported firmware bundle with recovery images, component identities, software bill of materials, vulnerability/support contact, update/rollback procedure and end-of-support date. Demonstrate an authorized update, a failed-update recovery and a component replacement without losing the configuration record. Chapter 14.8 owns the continuing firmware lifecycle; Chapter 14.6 owns the spare/RMA policy.
The DMTF publication catalogue lists Redfish DSP0266 v1.24.0, Redfish Data Model DSP0268 release 2026.1, SPDM DSP0274 v1.4.1 and PLDM Firmware Update DSP0267 v1.3.0. Record implemented versions/features and sensor-unit, timestamp, fault-state and identity mappings; exercise Redfish, SPDM and PLDM behavior against this manifest, not a latest-version badge.
The integrator owns fault diagnosis, corrected firmware evidence, the manifest and event logs; the operator imports them, reads power/thermal/error telemetry, exercises one supported firmware rollback and restores a replaced node to the approved image. The gate fails if the inventory no longer identifies the replacement or if the supported recovery path cannot restore it. Device authentication/attestation is evidence to verify against the platform’s implemented trust chain, not a property conferred by a rack logo. Use Appendix A for the applicable security and facility standards index, with Chapter 0.4 as intake.
Cite this chapter
Fehn, J. (2026). Server & System Integration (Chapter 7.14). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-14-server-and-system-integration (accessed 2026-09-29).
@misc{aidc-7-14,
author = {Fehn, Jacob},
title = {Server & System Integration (Chapter 7.14)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-14-server-and-system-integration},
note = {Accessed 2026-09-29}
}