Appendix F
In this chapter · 6 sections
Failure-Mode / FMEA Catalog
Random failures and adversarial scenarios need distinct initiating evidence and controls; link them only where a validated topology shows that they converge on a shared top event or protective response.
What you'll decide here
- Use this appendix as a lookup, not a narrative: find your failure mode in the master table, then jump to the owning chapter (last column) for the engineering derivation — this catalog summarizes, it does not replace, the canonical treatment.
- Keep random and adversarial initiators, likelihoods, authorities, observability, persistence, controls, and recovery evidence distinct; cross-link only the validated shared top event or protection function.
- Treat the blast-radius column as a candidate input to your fault-domain and RBD work (Chapter 12.1 / 12.5): a mode that strands one rack and a mode that trips the campus POI impose different service consequences and deserve different redundancy spend.
- Walk the propagation column for cascade interlocks. In a cascade, the first fault can be a single fault that defeated a shared mitigation (one CDU, one fuel header, one protection setting) and took an entire fault domain with it.
- Pair detection latency against propagation speed for each mode. Where the fault propagates faster than your detection-plus-actuation loop (load-step→grid trip, HBM overtemperature), prevent the trigger, add inertia, or qualify a faster independent reactive path — then demonstrate its margin.
This appendix consolidates the failure modes scattered across the engineering chapters into one uniform FMEA register. It is referenced from the resilience-standards chapter (Chapter 12.1), the reliability rethink (Chapter 12.2), the component-failure-rate chapter (Chapter 14.3), and the integrated-systems-test chapter (Chapter 13.6); and it supplies candidate top events and common-cause couplings to the quantitative availability model in Chapter 12.5, after the project validates equipment, operating state, fault domain and evidence; these rows are not executable model inputs. The canonical engineering of each mode lives in the chapter named in the right-most column — this catalog is the index and the cross-walk, not the derivation.
Some random and adversarial scenarios can converge on a similar top event or protective response, but they do not automatically share one record. Keep initiating evidence, likelihood, authority, observability, persistence, propagation, security controls, and recovery distinct. Where the validated topology joins the paths, link the records to the shared top event and reuse the proven protective response; do not claim every random failure has an attacker twin. The cyber-physical analysis is in Chapter 11.10.
How to read a row
Each failure mode is recorded against six fields, applied identically across the three master tables so the catalog is sortable and comparable:
- Trigger — the initiating event (random cause first; adversarial initiators stay in separate Part 11 records linked to a proven shared top event).
- Propagation path — how the fault spreads, and critically which shared mitigation it defeats to escalate from a local fault to a cascade.
- Detection — the sensing modality and its characteristic latency relative to propagation speed (the decisive ratio).
- Blast radius — the fault domain affected at full propagation: node, rack, row, hall, or campus/POI.
- Mitigation — the preventive or containing control, classified as preventive (stops the trigger), inertial (buys ride-through time), or reactive (acts after detection).
- Recovery — the path back to service, kept as two separate numbers: service-restoration time (detect, isolate, transfer, repair or replace, restart, validate) and, where the mode costs work, the replayed work plus any integrity look-back to a known-clean state. Sum them only where they occupy distinct, non-overlapping portions of the job timeline; they are not interchangeable.
Where propagation outruns detection-plus-actuation, that reactive loop is too slow: prevent the trigger, add inertia, or qualify a faster independent protection path. The ratio matters because neither an operator nor an untested standby earns timing credit.
Master FMEA catalog — thermal & mechanical (cooling) modes
| Failure mode | Trigger | Propagation path | Detection | Blast radius | Mitigation | Recovery | Owner |
|---|---|---|---|---|---|---|---|
| F-R01 · Coolant-leak cascade | QD/manifold/cold-plate breach, hose chafe, gasket creep | Local drip → conductive coolant on busbar/PDU → arc/short → de-rate or trip of the powered branch; if it reaches a shared CDU controller, the whole CDU loop and its row drop together | Floor/leak-rope sensors + CDU flow/pressure-decay; latency seconds-to-minutes; propagation can outrun it on a high-flow breach | Rack → row (if the CDU is shared); hall if isolation valves are absent | Negative-pressure loops (leak draws air in, not coolant out); dripless UQDs; zoned isolation valves; per-rack leak detection; N+1 CDU with independent controllers (preventive + reactive) Test: Safe leak-sensor/valve simulation and spray-path inspection; verify the branch contains the event without wetting shared controls. | Isolate the branch, drain/flush the loop, replace the failed coupling, re-pressure-test, re-fill, re-commission worst-case branch; TTR hours per rack | 5.11 |
| F-R02 · CDU / pump failure | Pump bearing/VFD failure, seal loss, filter blockage, control-board fault | Flow drops or stops on the affected zone → local heat accumulates and GPUs reach throttle/trip at the selected rack’s qualified limit; a shared CDU controller takes every rack on the loop together | Pump state, differential pressure, flow, coolant temperatures, valve position and GPU thermal telemetry; automatic protection must beat the demonstrated throttle/trip limit | The hydraulic zone the CDU serves — one rack for an in-rack unit, up to eight NVL72 racks for an in-row unit (HPE: one 1.3 MW CDU per eight) | Duty/standby or N+1 pumps with automatic switchover on UPS-backed power; N+1 CDUs with independent controllers; a pump-drop test at full load; GPU power capping as the last inertial layer (preventive + inertial) Test: Approved pump-drop test at qualified heat load; record worst-branch flow and detection-to-restored-flow time, including standby power/control independence. | Standby takes over inside the throttle window if the switchover was validated; otherwise isolate and repair the failed train and re-qualify flow — TTR minutes with a hot standby, hours for a train repair | 5.11 |
| F-R03 · Cooling-controls transient excursion | Synchronized GPU load drop (job ends / checkpoint pause) → loop heat input collapses faster than valves/VFDs can slew; setpoint hunt / control-loop oscillation | On a rapid load drop, supply-coolant temp overshoots downward → transient dew-point excursion → condensation risk on cold surfaces; or anti-hunting failure drives sustained oscillation that fatigues actuators and destabilizes neighboring loops | Coolant supply-temp rate-of-change, dew-point margin sensor, valve-position hunting; detectable but the excursion window is brief | Row → hall (controls coupling); condensation risk is local to cold surfaces | Slew-rate limits on control valves and pump VFDs; anti-hunting tuning; dew-point margin floor; thermal buffering/loop inertia and feed-forward cooling controls coordinated with workload telemetry (preventive + inertial) Test: Commission the specified load step; log supply temperature, dew point and valve motion. Require the OEM envelope without sustained hunting. | Re-tune control loops, restore dew-point margin, dry/inspect any condensation; TTR minutes, no hardware loss if caught | 5.12 |
| F-R04 · HBM overtemperature / thermal damage | Cold-plate contact loss, TIM pump-out, local flow starvation, or sustained over-temp on a stacked-DRAM site | HBM junction temp climbs → ECC error rate rises → uncorrectable error / package damage; on a tightly-coupled training step the failed device stalls the synchronous collective and the whole job stalls behind the straggler | Per-die thermal telemetry, ECC/CE rate trend, GPU throttle flags; trend-detectable early, but temperature rise can be fast once contact is lost | Node (the GPU/HBM package) → job (synchronous training stalls on the straggler) | Thermal screening/burn-in pre-deployment; ECC-rate alarming with proactive drain; flow-failure throttle floor; hot-spare nodes so the scheduler evicts and replaces the straggler (preventive + reactive) Test: Approved thermal/fault emulation and scheduler drain; show protection and checkpoint restart. Never damage HBM to demonstrate coverage. | Evict the node, fail the job over to a hot spare, RMA the package; training resumes from the last checkpoint; restoration is detect + evict + reschedule + reload, and the replayed work averages about half the checkpoint interval only for uniformly distributed failure timing and a valid checkpoint | 14.3 |
A direct-to-chip loop has two clocks. With flow maintained, the coolant and metal can buffer a heat step for tens of seconds where the measured mass and load support it. With flow lost, local mass remains but circulation no longer transports heat away; the rack's throttle and trip states can arrive in seconds, on a window the OEM's transient data and a commissioning pump-drop test set for the specific product. The leak and CDU/pump rows are engineered against the second clock: detection and switchover must complete inside it, which is why standby pumps switch over automatically — on UPS-backed feeds where the ride-through study shows the source transfer outlasts the window — and why GPU power capping, not operator action, is the last layer. The cause-and-effect is engineered and tested in Chapter 5.11; the measured response feeds the reliability model in Chapter 12.2.
Master FMEA catalog — electrical & power modes
| Failure mode | Trigger | Propagation path | Detection | Blast radius | Mitigation | Recovery | Owner |
|---|---|---|---|---|---|---|---|
| F-R05 · Simultaneous-GPU-load-step grid trip | Thousands of GPUs ramp in lockstep at job start/stop/checkpoint; di/dt event on every step | GW-scale aggregate ramp >1,000 MW/s presented to the POI → voltage/frequency disturbance → if the load-smoothing spine is absent, upstream protection or generators see a step they cannot follow → trip | Power-quality metering at the POI, PMU/PQM; fast, but the di/dt event is potentially faster than the installed reactive loop — milliseconds | Campus (POI) → contributes to wide-area grid disturbance | The chip→BBU→BESS smoothing spine (on-package capacitance → rack BBU → facility BESS); software ramp-rate limits and regulated wind-downs; grid-forming inverters (preventive + inertial) Test: Approved staged load step; record POI ramp, storage state and protection response against the utility-agreed envelope. | Re-energize per utility ride-through procedure; restore smoothing controls; no hardware loss if the spine held; TTR minutes if ride-through succeeded | 4.5 |
| F-R06 · Utility ride-through / voltage-disturbance event | Grid fault (e.g. 230 kV line fault) causes a voltage sag at the POI; sensitive customer-side protection drops the load | Undervoltage response across facility loads → cumulative ~1,500 MW customer-side load reduction during the July 2024 six-fault NoVA sequence → the reduction itself destabilizes the grid, a self-reinforcing reliability problem NERC flagged at Level 3 | POI relays, undervoltage/under-frequency elements, PMU; the disturbance is sub-cycle to cycles | Campus (full load drop) → wide-area grid | Fault-ride-through settings tuned to stay online through the sag (SEL/relay, UPS, undervoltage-load-retention); reactive/voltage support toward the POI; ride-through posture aligned to the NERC 2026 Level-3 alert's recommended large-load envelope — guidance, not penalty-backed; TPL-001/PRC standards bind transmission planners and generators, not data-center loads (preventive) Test: Relay secondary injection and approved ride-through simulation; verify load retention, source-transfer behavior and sequenced recovery. | Auto-recover as the sag clears if ride-through held; if tripped, sequenced re-energization and load ramp; TTR minutes | 4.10 |
| F-R07 · BESS thermal runaway | Cell defect, overcharge, internal short, or cooling loss in an LFP facility battery | Single cell vents → exothermic chain to adjacent cells → module-level runaway → fire/off-gas if pack-level isolation and venting fail; loss of the BESS also removes the ride-through and load-smoothing it was providing | Cell voltage/temp telemetry, off-gas (H2/CO) detection, BMS fault flags; warning lead time depends on chemistry, placement and qualified test evidence | BESS enclosure → adjacent enclosures/room if propagation isolation fails | LFP chemistry (higher thermal-runaway threshold than NMC); module-level thermal isolation and dedicated venting/deflagration paths; off-gas detection with pre-emptive isolation; physical separation of BESS from IT (preventive + reactive) Test: Review installed UL 9540A evidence and fire strategy; exercise alarm/isolation logic safely. No live runaway initiation in an operating hall. | Execute the pre-incident response plan: isolate electrically, keep ventilation/deflagration paths clear, monitor off-gas, coordinate fire-service response on the UL 9540A / NFPA 855 basis of design — never open or manually suppress the enclosure ad hoc; replace module/pack; re-commission; TTR hours-to-days; ride-through reverts to UPS/BBU meanwhile | 4.5 |
| F-R08 · Fuel-supply interruption (on-site generation) | Firm pipeline curtailment (correlated cold-snap), valve/compressor failure, or fuel-quality (Wobbe/dew-point) excursion; LNG/CNG storage depletion | Loss of fuel → on-site turbines/engines de-rate or trip → if the site is islanded or grid-import is constrained, generation cannot meet IT load → controlled load-shed or outage | Fuel header pressure, Wobbe-index/dew-point analyzers, tank level, generator load; minutes of warning on slow depletion, immediate on a hard cut | Campus (islanded sites) → partial if grid-import backstops | 'Synthetic-firm' fuel structure (multiple pipelines + interruptible + on-site LNG/CNG storage); dual-fuel switching; fuel conditioning to spec; sized on-site storage for correlated-curtailment duration (preventive) Test: Approved alternate-fuel/curtailment exercise; verify fuel quality, usable inventory and load shed before the modeled energy reserve is exhausted. | Switch fuel source / draw down on-site storage, restore generation; coordinate curtailment with curtailable-load agreement; TTR depends on storage sizing vs outage duration | 4.9 |
Scope & caveats
Load loss as seen by the grid. NERC's incident review ('Load Details') found the affected data centers transferred their loads to backup power — static UPS, decentralized rack UPS, or DRUPS — in response to the disturbance. The figure is a loss of demand at the interconnection, not evidence that IT power was interrupted or that training jobs restarted.
The approximately 1,500 MW is the total customer-side load reduction coincident with the six-fault sequence; NERC reports approximately 1,260 MW as the sustained drop at the third voltage depression. The NERC-investigated canonical case. A second, larger occurrence followed on 2026-07-22: ~3.8 GW dropped on a single normally-cleared Ashburn 230 kV fault (see companion key number). Two vintages of the same failure mode, not a replacement figure.
Scope & caveats
One 54-day pre-training window on a 16,384-GPU H100 cluster — a single fleet, generation, and workload, not an industry rate.
Guide arithmetic on the paper's own counts: 148/419 = 35.3% and 72/419 = 17.2%. Table 5 prints 30.1% for the faulty-GPU row, which no denominator in the paper reproduces; its 18 rows do sum to 419, but the printed percentages sum to 94.9%. Counts are auditable; treat the printed percentages as paper-printed.
Scope & caveats
Exact QCT reference. The 45 °C liquid-inlet maximum and 65 °C liquid-return maximum are separate limits, not a prescribed 20 K operating rise. Select a supported operating point, approved fluid, liquid heat load, and design ΔT; ASHRAE W45 describes FWS supply capability, not this product's setpoint.
Separate acceptance maxima, not a prescribed 20 K operating rise; do not attribute these limits to HPE without an HPE document that states them.
Scope & caveats
The near-zero endpoint requires dry, non-evaporative heat rejection; loop closure alone does not determine WUE-site.
Master FMEA catalog — connectivity, compute & data-integrity modes
| Failure mode | Trigger | Propagation path | Detection | Blast radius | Mitigation | Recovery | Owner |
|---|---|---|---|---|---|---|---|
| F-R09 · Fiber cut (inter-/intra-DC) | Backhoe/construction strike, conduit failure, or DCI route cut | Loss of a fiber path → if scale-across/DCI is single-routed, the affected campus or training partition is severed → distributed training stalls or a metro site loses connectivity | Optical LOS/LOF alarms, OTDR, BER collapse; immediate at the physical layer | Link → partition/campus (distributed training) or metro site reach (inference) | Physically diverse, geographically separated fiber routes; protected DCI (ZR/ZR+ with restoration); for training, topology that degrades gracefully on a partition loss (preventive + reactive) Test: Controlled protect-path transfer and end-to-end route survey; verify the declared cut leaves service intact with no common conduit. | Restore over the diverse path automatically; physical splice repair on the cut route (TTR hours-to-days for the splice, seconds for protected failover) | 3.6 |
| F-R10 · Optics flap storm | Marginal transceiver, dirty/over-bent connector, thermal cycling, or firmware bug causing repeated link up/down | One flapping link → routing reconvergence churn / ECMP rehashing → packet loss and tail-latency spikes propagate across the fabric → on a tightly-coupled collective, the flapping link gates the whole all-reduce and tanks MFU | Per-port link-flap counters, BER/FEC-error trend, CRC errors; trend-detectable but a storm builds in seconds-to-minutes | Link → fabric pod (reconvergence churn) → job (collective stalls) | Pre-install optics burn-in / BER screening off the critical path; FEC margin headroom; auto-quarantine of flapping ports; CPO removes some pluggable interfaces but changes the optical-engine replacement fault domain (preventive + reactive) Test: Approved link-flap injection; verify quarantine and collective recovery, measuring loss, tail latency and the surviving cut. | Quarantine and replace the offending transceiver; let routing reconverge; TTR minutes to swap, plus reconvergence | 8.9 |
| F-R11 · SDC corruption event (silent data corruption) | Marginal/defective silicon (a 'mercurial core'), aging, voltage/thermal margin loss producing a wrong result with no error flag | A miscomputed value flows silently into gradients/activations → corrupts the model state or an inference result → undetected for hours-to-days, potentially poisoning a checkpoint and forcing a roll-back of all work since the last clean checkpoint | No native hardware flag — requires a dedicated detection program (Fleetscanner periodic, Ripple in-fleet, Hardware Sentinel runtime); latency hours unless instrumented | Node (the mercurial core) → job/model (silent corruption of state) → potentially every downstream consumer of a poisoned checkpoint | Proactive SDC-hunting at scale; redundant/checksummed compute on critical paths; PVF-aware placement; roll-back to a verified-clean checkpoint; quarantine the device (preventive + reactive) Test: Known-answer fault injection in an isolated workload; verify detection, quarantine and verified-clean checkpoint recovery, including look-back. | Identify and quarantine the mercurial core, roll back to last verified-clean checkpoint, re-run; restoration is the quarantine and reschedule; the loss is the replay back to the last verified-clean checkpoint, and how far back that sits is set by the integrity look-back, not by checkpoint frequency — cadence alone cannot make a checkpoint clean | 14.3 |
| F-R12 · Water-curtailment event (cross-domain — canonical home Ch 3.7) | Drought-stage ordinance, basin/withdrawal-permit curtailment, or reclaimed-water supply interruption on an evaporative-cooled site (grid-curtailment orders like ERCOT's SB6 regime are an electrical event, not a water one) | Loss/limit of make-up water → evaporative/cooling-tower capacity drops → heat-rejection ceiling falls below IT load → forced de-rate of compute or a curtailment-driven load-shed | Make-up water flow/level, WUE telemetry, basin level, curtailment-order receipt; advance notice on scheduled curtailment, immediate on a hard order | Hall → campus (heat-rejection limited) | A dry, non-evaporative rejection path qualified for the design conditions (loop closure alone does not remove make-up-water dependence); reclaimed/non-potable sourcing; on-site water storage; curtailment-tolerant workload scheduling (batch defers); thermal-storage buffer (preventive) Test: Approved make-up-water-loss exercise or validated simulation; show heat rejection and curtailed IT stay inside the agreed envelope. | Shift heat rejection to the closed-loop/dry path or draw down water storage; defer curtailable batch load; restoration follows the curtailment duration; the structural fix is a qualified dry rejection path, not loop closure | 3.7 |
| Mode / operating state | Initiator → service effect | Test / evidence | Restoration criterion | Owner |
|---|---|---|---|---|
| F-R21 · shared-room operation | Fire/smoke defeats nominally separate paths. | Compartment/dependency survey; approved fire-strategy review and drill. | Authority release; contamination and both paths assessed. | 6.5 · fire/operations |
| F-R22 · severe weather | Flood/water ingress defeats common elevation, routes or access. | Elevation/drainage/access model; safe inspection and exercise. | Safe access, dry equipment and utilities independently released. | 3.8 · civil/operations |
Adversarial initiators: prove the shared physical event. Keep each security scenario’s access, command authority and persistence separate from random-fault rates. F-R01: spoofing a dew-point setpoint can force condensation where the command reaches exposed surfaces below dew point; the coolant-leak trace below retains that mechanism. F-R02: malicious VFD firmware or controller denial of service can stop circulation only where the compromised path can defeat the required flow and its independent protection. F-R04: CDU disablement that holds flow at zero links to HBM overtemperature only through the selected package’s demonstrated thermal response. F-R05: malicious power-cap firmware can force a synchronized load step only over the endpoints it controls; use the measured MW step and timing. F-R07: BMS spoofing that disables cell balancing or thermal protection reaches the battery hazard only where independent cell/pack protection does not veto it. Chapter 11.10’s OT destructive-primitive and control-dependency assessment owns these five initiators.
F-R09: a deliberate cut of an undiverse fiber route belongs to Chapter 11.2’s fiber-vault and sole-path sabotage assessment; the shared physical event is the lost route, not a new random cut rate. F-R10: a thermal attack on the optics environment is a candidate 11.10 OT scenario only if cooling authority or physical access can drive the named transceiver outside its qualified environment and link/error tests reproduce the flap mechanism. F-R11: fault injection on a known-marginal device is a candidate scenario only if the attacker controls a relevant voltage, clock or thermal path and isolated known-answer tests establish undetected wrong results; a crash or corrected ECC event is not SDC. 11.5 bounds the confidential-computing threat model; 11.10 assesses reachable controls and 14.3 owns SDC evidence. No attack frequency is inferred from component failure data. Credit protection only after it survives the assessed attacker authority.
Detection differs by mode: an active fiber-path cut alarms immediately at the physical layer, while an untested route-diversity defect and SDC remain latent; optics degradation is trend-detectable, and water curtailment is detected through make-up-water telemetry or receipt of the curtailment order. An SDC event can poison a checkpoint and sit undetected for days; an optics flap can erode MFU before tail latency is traced to a marginal transceiver; an untested protection path is exposed when the active route is cut. Match the control to the mode: a dedicated SDC-hunting program, per-port flap/FEC telemetry, verified-clean checkpointing, and periodic protection-path and physical-diversity verification bound the respective blast radii. The empirical fault taxonomy and the AFRs that feed these rows into the availability model are the canonical content of Chapter 14.3; the SDC detection program is treated there and demonstrated at commissioning in Chapter 13.6.
Common-cause couplings & cascade interlocks
For the availability model, the catalog's most useful output is the couplings between modes — the shared resources whose failure makes two independent-looking modes fail together. These are the common-cause terms an RBD or fault tree must capture, or it will badly over-state availability. The table below names the cross-mode interlocks worth modeling explicitly.
| Candidate shared condition | Modes it may couple | Condition to verify | Design / model response |
|---|---|---|---|
| CDU, controller, loop, or power source shared by a hydraulic zone | Leak escalation, loss of flow, thermal protection | Which racks actually share flow, control, isolation and power; whether one fault defeats standby | Model one shared event only for the confirmed fault domain; separate or isolate trains where required |
| One storage system used for smoothing and ride-through | Load-step response, ride-through, storage isolation | Whether the same cells, inverter, controls, protection or bus provide both functions | Separate functions or model the confirmed common dependency and its protection states |
| Fuel source or delivery path exposed to one regional event | Fuel interruption and generation availability | Contract, pipeline, storage and transport independence during the named event duration | Represent the named supply event explicitly; do not infer independence from two contracts |
| Fiber routes sharing conduit, crossing, right-of-way, building entry or carrier equipment | Route cut and loss of protect path | Surveyed physical diversity end to end, including meet-me and power dependencies | Use an explicit route-loss event where sharing is confirmed; correct false diversity |
| Checkpoint policy and last-known-good state | Restart and rollback recovery times | Integrity validation, cadence, write path and corruption horizon | Treat as a recovery-time dependency, not a hardware common-cause beta |
| Shared BMS/DCIM/SCADA authority or communications | Cooling setpoint, flow and load-control excursions | Which commands, sensors, interlocks and overrides share one control plane | Model only confirmed shared authority; provide independent protection where the hazard basis requires |
Using this catalog in the availability model
The rows above are inputs to the reliability work, and each field maps to a model term. Top events come from the project's loss criteria and confirmed fault domains — a row becomes a top event when its blast radius reaches a domain the contract prices. Basic-event rates may take the fleet evidence in Chapter 14.3 where component, population, maturity and duty match; otherwise use project evidence and carry the uncertainty. Shared events come from the coupling table, confirmed against the as-built topology — a confirmed coupling is an explicit shared event, never an assumed beta. Recovery distributions use measured detection, protection, repair, restart and validation times; checkpoint cadence couples the training rows' recovery. IST cases are the rows the design basis names for demonstration, by the evidence method Chapter 13.6 sets; the catalog does not itself authorize a hazardous live event. The model mechanics are in Chapter 12.5.
F-R02.1, normal duty: a pump bearing fails while standby control and power survive independently. Flow loss exposes four racks to throttle/trip and job interruption. The budget is 0.50 + 0.50 + 3.0 + 1.0 = 5.0 s, leaving 1.0 s before the assumed 6.0 s limit. HOLD hardware acceptance: obtain the selected rack’s allowable interruption and qualify the whole response; the assumed six-second screen authorizes no live pump-drop test. The commissioning lead records approved test/equipment IDs, heat load, inlet state, minimum worst-branch flow, timestamps, uncertainty and restart evidence. Accept only with restored qualified flow before the limit and temperatures inside the OEM envelope; report lost work separately from service restoration.
At an exact nominal 4.0 s standby recovery the budget becomes 6.0 s: equality fails. Shorten recovery or change the protection/thermal design and retest. F-R02.2, shared controller lost: the same pump count earns no standby credit when one controller disables both trains; the four-rack consequence remains. The design lead owns independence correction; the commissioning lead keeps acceptance open until signed evidence demonstrates the protective function. Chapter 5.11 owns hydraulic protection, Chapter 13.6 test authorization, and Chapter 12.5 the evidence-fed model.
Trace a coolant leak through the as-built pressure and isolation state
Start with a quick-disconnect leak and trace it through the as-built topology. A dripless coupling that opens inside a qualified negative-pressure loop can remain a rack-level ticket if the tested operating state holds: the loop draws air in rather than pushing coolant out, the CDU must demonstrate pressure/level detection within the qualified seconds-scale window, the branch isolates, and the rack is back after a drain, a coupling swap and a re-pressure test — hours, one rack. The same leak on a positive-pressure loop without zoned isolation is a row event: coolant reaches the busbar, the branch trips on a short, and if the shared CDU controller is in the spray path the whole loop drops. The difference between those two outcomes depends on pressure state, spray paths, fluid behavior and the demonstrated mitigation, which is why the model must take its fault domain from the built loop, not from a generic leak rate. A spoofed dew-point command that forces condensation is a separate adversarial scenario with the same physical consequence and different evidence, detection and persistence — modeled as its own initiating event on the shared top event. OT controls in Chapter 11.10; leak detection and isolation in Chapter 5.11.
Cite this chapter
Fehn, J. (2026). Failure-Mode / FMEA Catalog (Chapter F). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/appendix-appendices-and-reference-data/f-failure-mode-fmea-catalog (accessed 2026-09-29).
@misc{aidc-F,
author = {Fehn, Jacob},
title = {Failure-Mode / FMEA Catalog (Chapter F)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/appendix-appendices-and-reference-data/f-failure-mode-fmea-catalog},
note = {Accessed 2026-09-29}
}