Chapter 11.10
In this chapter · 7 sections
Cyber-Physical & Destructive Attacks on OT/Facility Systems
A compromised control plane can create destructive states, so allocate each protection function from the project hazard and risk analysis, then engineer the required independence, integrity, bypass control, diagnostics, proof testing, and defense in depth.
What you'll decide here
- Whether your OT estate (BMS, EPMS, SCADA, CDU controllers, BESS management, the power-cap/firmware planes) is genuinely segmented from IT per IEC 62443 zones-and-conduits, or merely firewalled on paper while exposed to the internet through a vendor jump-host nobody owns.
- Which destructive actions are physically possible from the control plane — a forced synchronized load step, CDU disablement, BESS abuse, or malicious firmware — and which independent or diverse protective layers limit each consequence.
- How the project hazard and risk analysis allocates protection functions among process controls, OEM protections, electrical/fire/life-safety systems, mechanical relief, dedicated relays, and any SIS/SIF — including required independence, integrity, bypass governance, diagnostics, proof testing, and secure maintenance.
- Where random failures and adversarial abuse share a physical consequence, and where the adversarial case needs additional authorization, integrity, detection, containment, persistence, and recovery evidence rather than an assumed identical twin.
- Who owns OT detection and response — a converged cyber-physical SOC with the escalation trigger wired in, or two teams (facilities and security) who each assume the other is watching the controllers.
Most of Part 11 defends the information in the building — model weights, tenant data, credentials. This chapter defends the building itself. An AI factory is a cyber-physical system: tens of megawatts of synchronized silicon sitting on a control plane of programmable logic controllers, building-management and electrical-power-management systems, coolant-distribution-unit firmware, battery-management systems, and a power-capping layer that the GPUs themselves obey. Every one of those is a computer, and reachability must be established from actual wiring, networks and maintenance paths. Alongside information loss, an attack on the OT plane can physically destroy the asset — cause equipment damage, trip the grid interconnection, contribute to battery hazards, or disable accelerators through an update path if credited protections fail. The reachable control scope and remaining protection layers determine blast radius; equipment damage can put long-lead replacement on the recovery path.
The sections that follow enumerate the OT assets and their security levels, the destructive primitives the control plane exposes and what each one costs, the segmentation boundary (the Purdue model and IEC 62443 zones-and-conduits), and how the project hazard and risk analysis allocates each protection function — process control, OEM protection, dedicated relays, mechanical relief, fire and life safety, and where the safety lifecycle requires it a safety instrumented function — so that no single compromised control domain can defeat every credited layer. Assess each Appendix-F failure mode for adversarial reachability: the transient you engineer against as a random fault is also the attacker's payload, and the attacker picks the timing.
The OT/ICS threat model: what is actually reachable
Start by enumerating the control planes, because each is a distinct attack surface with a distinct destructive payload. The BMS (building-management system) can command cooling setpoints, valve and pump commands, air handlers, leak detection, and interfaces to fire/EPO systems, whose authority and independent protection must be verified. The EPMS (electrical-power-management system) monitors electrical state and may command breakers, transfer switches, generator start/stop and UPS modes where the design grants it that authority. SCADA/PLC layers sit under both, executing the actual I/O. CDU controllers regulate the technology-cooling loop — flow, inlet temperature, secondary-loop pressure — that keeps a GB300 NVL72-class rack at its up-to-142 kW facility design basis from cooking itself. BESS management (battery-management systems and their EMS) governs charge/discharge, cell balancing, and the thermal-runaway protections on grid-scale storage. And the power-cap and firmware planes — the GPU/BMC layer that throttles and updates accelerators at fleet scale — need the same command-authorization and failure-containment analysis as the older controllers.
Published exposure studies justify testing these paths. Claroty's Team82 analyzed roughly 467,000 building-management devices across 529 organizations in 2025 and found 75% of organizations ran BMS affected by known-exploited vulnerabilities (KEVs), with 51% exposed to KEVs that are both ransomware-linked and insecurely connected to the internet (Claroty Team82, State of CPS Security 2025: BMS Exposures). Those are cross-industry organization-level results; the data-center-specific cut followed on 2026-07-29, when Team82 analyzed more than 750,000 CPS assets across the world's largest data-center facilities and found nearly 1 in 5 infrastructure assets (18%, 32,000+ of 174,577) one hop from a system making risky outbound connections to the public internet, 88% of BMS communicating over insecure protocols and 40% running outdated firmware, with PDUs (41%) and HVAC/cooling (32%) the most-exposed classes and more than 80% of OT control, power-monitoring and IoT systems speaking BACnet/Modbus-class legacy protocols (Claroty Team82, State of CPS Security: Data Center Exposures). The same team disclosed, in June 2026, two CVSS 9.8 flaws in Vertiv Liebert IS-UNITY-DP and RDU101 UPS network cards (CVE-2025-46412, CVE-2025-41426) and a five-CVE chain in the Trane Tracer SC+ HVAC controller up to v5.20.1362, fixed in v6.3 — named data-center equipment, not generic BMS. The population is the studied facilities, not this campus's measured exposure.
The power-cap and firmware as a weapon
This is the part of the threat model unique to AI factories, and it is where the destructive primitives concentrate. Four are worth naming explicitly, each with its downstream cost.
The forced synchronized load step → grid trip. A gigawatt-class training cluster already swings load by hundreds of megawatts when a job starts, stops, or stalls — that is the random-fault problem Chapter 4.5 engineers ride-through against. Now make it the payload. An attacker who controls the power-cap plane (or simply the scheduler) can command thousands of racks to drop or pick up load simultaneously, deliberately, at the moment of peak grid stress. NERC has already flagged the non-malicious version as a reliability crisis: in the July 2024 Northern Virginia incident, a permanent 230-kV transmission-line fault and staggered auto-reclosing produced six successive system faults over 82 seconds. Roughly 1,500 MW of data-center load reduced coincident with the disturbance; about 1,260 MW dropped simultaneously with the third voltage depression and remained off for hours (NERC). This was not an isolated event: NERC has documented repeated 1,000+ MW computational-load losses since 2022, and on May 4 2026 issued its rare Level 3 "Essential Actions" alert over the pattern — an implementation directive, not a penalty-backed standard, and issued ~22 months after this Virginia event rather than in immediate response to it. Weaponized, the same primitive is a tool to trip the interconnection or destabilize the local grid, and the facility's own protection relays may disconnect it — taking down the campus to save the grid. The downstream cost is not a reboot; it is a black-start and a damaged relationship with the utility that holds your interconnection.
CDU disablement → thermal runaway. An NVL72-class liquid-cooled rack's response after CDU flow loss (throttle, controlled shutdown, emergency shutdown or trip) must be established from OEM transient data, controls tests or a professional-engineering calculation; do not infer a portable seconds-to-tens-of-seconds value from chilled-water inertia. An attacker who stops the CDU pumps, closes a facility-water valve, or falsifies the inlet-temperature reading so the controller never opens the valve, drives the rack straight into thermal runaway. Done across a hall, that takes cold plates and silicon past their thermal limits, with a recovery bounded by component lead times.
BESS-induced runaway. Grid-scale and UPS-class battery storage protects against thermal runaway through the BMS: cell balancing, over-temperature trips, and isolation. Compromise that management layer and you can disable the very protection that prevents a runaway, force an over-charge or over-discharge, or suppress the thermal alarm — turning the energy-storage asset into an ignition source inside the building.
GPU-bricking via malicious firmware. The fleet firmware-update path (BMC, GPU VBIOS, NVLink-switch firmware) is a fleet-scale destructive primitive: a single malicious image, pushed through the legitimate update channel, can render thousands of accelerators non-functional or subtly mis-behaving. This is why firmware integrity (signed images, a hardware root of trust, measured boot) is the gate on the most expensive single payload in the building. The canonical treatment of that integrity chain is in Chapter 11.4; here it is one of four ways the control plane bricks the asset.
| Destructive primitive | Control plane abused | Physical consequence | Recovery scale | Independent protection layer (allocated from the hazard analysis) |
|---|---|---|---|---|
| Forced synchronized load step | Power-cap plane / scheduler / EPMS | Grid trip; interconnection disconnect; possible black-start | Hours to days (grid + black-start) | Utility protection relays; on-site ride-through (BBU/supercap) sized for the swing |
| CDU disablement / falsified inlet temp | CDU controller / BMS valve & pump I/O | Thermal runaway of liquid-cooled racks; silicon damage | Weeks to months (cold plates, GPUs) | Hardwired high-temp trip and flow-loss interlock independent of the BMS |
| BESS over-charge / alarm suppression | BESS BMS / energy-management system | Battery thermal runaway; fire; structural loss | Months (rack + structure + remediation) | Independent cell-level protection and gas/thermal detection on a separate logic solver |
| Malicious firmware push | BMC / GPU VBIOS / switch firmware plane | Fleet-scale bricking or covert mis-operation | Weeks to months (re-image or RMA fleet) | Hardware root of trust; signed measured boot; firmware governance (11.4) |
| EPO / breaker mis-operation | EPMS / SCADA | Unplanned full or partial campus shutdown | Hours (restart) to days (equipment stress) | Hardwired EPO logic; mechanical breaker interlocks; key-locked overrides |
OT/IT segmentation: the Purdue model and IEC 62443
The single highest-leverage control against all of the above is the oldest one: keep the control plane off the corporate network and off the internet. The Purdue model stratifies the estate into levels — physical process (0), basic control / PLCs (1), supervisory / SCADA-HMI (2), operations / MES (3), then the DMZ (3.5) and enterprise IT (4–5) — with the principle that traffic crosses level boundaries only through controlled, inspected conduits, never directly. Use IEC 62443-2-1:2024 for the asset-owner program and IEC 62443-3-2:2020 for system-design security risk assessment. The latter scopes zones and conduits and their target security levels; a network layout alone does not establish conformance. The standard's SL scale is defined by attacker capability — SL1 casual/coincidental, SL2 intentional with simple means, SL3 sophisticated with moderate resources, SL4 sophisticated with extended resources (i.e. a nation-state) (ISA/IEC 62443). The destructive primitives above sit at the IEC 62443 SL3–SL4 end of that scale — 62443's levels, not the RAND Weights Security Levels of Chapter 11.1 — and an internet-reachable BMS is evidence that a zone's achieved protection has drifted from its assessed target, not an automatic level score.
Two architectures pull against each other here. Convergence — running OT on the same IP fabric as IT, with shared identity and shared monitoring — is operationally attractive: one network team, one observability stack, remote vendor access for the cooling and electrical contractors who actually maintain the gear. Isolation — an air-gapped or diode-separated OT network with its own identity, jump-hosts, and one-way telemetry export — is far harder to operate and far harder to compromise. For an AI campus the practical middle is a hardened DMZ with brokered, recorded, time-boxed vendor access and unidirectional telemetry out, sized so that an IT compromise cannot reach a Level-1 controller and a controller cannot reach the internet. The microsegmentation and zero-trust mechanics that enforce this are the subject of Chapter 11.7; this chapter's contribution is the consequence of getting it wrong, which is the destructive-primitive table above.
Assign the IEC 62443 evidence to the party and assertion it covers. 62443-3-3:2013 supplies system security requirements; 62443-2-4:2023 addresses integration and maintenance service-provider processes. 62443-4-1:2018 addresses the product development lifecycle; 62443-4-2:2019, including its 2022 corrigendum, addresses component technical capabilities. Keep the target security requirements (SL-T), demonstrated system result and component capability (SL-C) separate. A component certificate does not establish the target or achieved protection of the assembled cooling or power system. The owner retains the system risk and integration evidence, even when suppliers provide certified components.
| Dimension | Full convergence (OT on IT fabric) | Brokered DMZ (recommended) | Hard isolation (air-gap / diode) |
|---|---|---|---|
| Attack surface | Largest — IT compromise reaches PLCs | Bounded — conduit + broker only | Smallest — no inbound path |
| Vendor/remote maintenance | Trivial but dangerous | Recorded, time-boxed, brokered | On-site or sneakernet only |
| Telemetry / observability | Native, unified SOC view | Unidirectional export to SOC | Manual or one-way diode export |
| Operational cost | Lowest | Moderate | Highest |
| Residual access paths into Levels 0–1 | IT compromise reaches PLCs directly | Broker, conduit, and any standing bypass account | Physical access, maintenance laptops, removable media |
Scope & caveats
Prevalence across the 2025 commercial-building study population (roughly 467,000 BMS devices, 529 organizations); Team82's July 2026 data-center report is a separate, data-center-specific population.
Scope & caveats
Population: more than 750,000 cyber-physical-system assets and core infrastructure across the world's largest data-center facilities in Team82's telemetry; the one-hop denominator is 174,577 data-center infrastructure assets. Studied facilities, not this campus's measured exposure.
Scope & caveats
Population: more than 750,000 cyber-physical-system assets and core infrastructure across the world's largest data-center facilities in Team82's telemetry; the one-hop denominator is 174,577 data-center infrastructure assets. Studied facilities, not this campus's measured exposure.
Scope & caveats
Population: more than 750,000 cyber-physical-system assets and core infrastructure across the world's largest data-center facilities in Team82's telemetry; the one-hop denominator is 174,577 data-center infrastructure assets. Studied facilities, not this campus's measured exposure.
Scope & caveats
Population: more than 750,000 cyber-physical-system assets and core infrastructure across the world's largest data-center facilities in Team82's telemetry; the one-hop denominator is 174,577 data-center infrastructure assets. Studied facilities, not this campus's measured exposure.
Scope & caveats
Population: more than 750,000 cyber-physical-system assets and core infrastructure across the world's largest data-center facilities in Team82's telemetry; the one-hop denominator is 174,577 data-center infrastructure assets. Studied facilities, not this campus's measured exposure.
Scope & caveats
Load loss as seen by the grid. NERC's incident review ('Load Details') found the affected data centers transferred their loads to backup power — static UPS, decentralized rack UPS, or DRUPS — in response to the disturbance. The figure is a loss of demand at the interconnection, not evidence that IT power was interrupted or that training jobs restarted.
The approximately 1,500 MW is the total customer-side load reduction coincident with the six-fault sequence; NERC reports approximately 1,260 MW as the sustained drop at the third voltage depression. The NERC-investigated canonical case. A second, larger occurrence followed on 2026-07-22: ~3.8 GW dropped on a single normally-cleared Ashburn 230 kV fault (see companion key number). Two vintages of the same failure mode, not a replacement figure.
Safety-instrumented systems: the control software cannot be allowed to win
Start with the project hazard and risk analysis and applicable process, electrical, fire/life-safety, equipment, and cybersecurity requirements. Allocate each protective function to the architecture that meets its required integrity and consequence: basic/process control, an OEM protective function, a dedicated relay, mechanical relief, fire/life-safety control, or — where the safety lifecycle requires it — a safety instrumented function (SIF) implemented by an SIS. An SIS is not the default home for every equipment trip. For an allocated SIF, IEC 61511 / ISA-84 lifecycle principles require the necessary independence, integrity, diagnostics, proof testing, management of change, and controlled bypasses; 'independent' is a designed and verified claim, not a promise that software can never defeat the function.
The 2017 Triton/Trisis attack showed that a safety layer can itself be targeted and reprogrammed. Use the project cyber/process hazard analysis to identify common dependencies, communications and maintenance paths, privileged access, sensor/final-element diversity, power, bypasses, and proof-test coverage. Separate or diversify protection where the required risk reduction demands it, secure and monitor authorized maintenance paths, and retain other independent or passive layers where one digital protection system cannot carry the consequence alone. The correct boundary is project-specific; neither 'hardwire every trip' nor 'the SIS is impossible to defeat' is a defensible universal rule.
Every Appendix-F failure mode is dual-use
The reframing that should run through your entire FMEA: each failure mode in the Appendix-F catalog has two possible causes, and the second one needs a cyber-reachability assessment. One is randomness — a pump fails, a valve sticks, a relay chatters, a load step happens because a job crashed. The other is an attacker who may induce a similar physical consequence through a different command or persistence path, with three advantages the random fault does not have: timing (at peak grid stress, during a maintenance window, when redundancy is degraded), correlation (across many units at once, defeating the N+1 that assumed independent failures), and persistence (re-triggering as fast as you recover).
This breaks a core assumption of classical reliability engineering. Simple independence calculations in Chapter 0.5 and Part 12 must be supplemented by common-cause models. A cyber-physical attacker deliberately violates independence: a single compromised controller can trip every CDU in a hall at once, and the 'redundant' second pump on the same compromised PLC is no redundancy at all. The design response is that fault domains must also be security domains: identify whether the controller that can fail unit A also reaches unit B or their shared protection; allocate the required independence and diversity from the project hazard and cyber-risk analysis.
Deep dive: worked example — the worst-timed attack on a 100 MW liquid-cooled hall
Walk the chain an adversary at the IEC 62443 SL3–SL4 capability end would actually run, to see where each control either holds or fails. Step 1 — access. The attacker reaches the OT DMZ through a contractor's jump-host left standing after a cooling-vendor maintenance window (the FrostyGoop pattern: a forgotten edge device, not a zero-day). If the DMZ is brokered and time-boxed, the jump-host is gone and this step fails; if OT is converged with IT, an earlier IT phishing foothold already put them here. Step 2 — reconnaissance. Modbus and BACnet are unauthenticated, so they map the CDU controllers, chiller PLCs, and EPMS without exploiting anything — they simply read.
Step 3 — the payload, timed. They wait for a maintenance window when the hall is on N (no spare cooling unit), then simultaneously falsify inlet-temperature readings on every CDU so the controllers hold their valves closed, and stop the pumps. With no chilled-water inertia, the rack response must be established by the selected OEM's transient data, controls tests or engineering calculation; do not infer a universal trip time. This is where the validated protection architecture changes the outcome. If all credited sensing, logic, actuation, power, communications, and bypass paths share the compromised control domain, the protection claim can fail. Where the project analysis requires independent or diverse layers, their tested response should drive the affected scope toward the defined safe state. Actual propagation, equipment damage, shutdown scope, and recovery time remain topology-, control-, inertia-, and attack-path-dependent; they must come from the project model and evidence, not a universal months-versus-hours promise. Step 4 — the multiplier. Simultaneously they command a synchronized load drop to stress the interconnection, so the facility is fighting a grid event and a thermal event at once, degrading the human response. The defenses that turned a catastrophe into a bad day were all decided at design time: segmentation that bounded step 1, an independent and verified protection layer that bounded step 3, and ride-through sized for step 4 (Chapter 4.5). None could be added in the moment.
Detection, response, and the converged escalation trigger
Prevention is the priority, but the OT plane also needs detection that a converged SOC can act on. The hard problem is that OT telemetry and IT telemetry live in different worlds: a falsified Modbus write looks like a legitimate engineering command, and the signal that distinguishes attack from operation is often physical — a setpoint that contradicts the measured process state, a valve commanded closed while temperature climbs, a load step with no scheduling event behind it. So the detection that matters here is cyber-physical correlation: cross-checking commanded state against sensed state against expected workload, and treating divergence as an incident.
The organizational failure mode is facilities and security each assuming the other is watching the controllers — the gap through which most OT incidents walk. The answer is a converged cyber-physical escalation trigger: a single condition (anomalous OT command + adverse physical trend) that pages both the SOC and facilities engineering simultaneously and invokes a joint playbook, because neither team can diagnose a CDU-disablement-into-thermal-runaway alone. The unified incident-command model that this escalation feeds into is canonical in Chapter 14.11, and the OT/cyber-physical IR playbook itself is built out in Chapter 11.12; this chapter's job is to insist the trigger exists and is wired to the physics, not just to the logs.
What to decide, and when
The decisions in this chapter sort sharply into irreversible-at-design-time and operational. The irreversible ones — which protection functions belong to process/OEM, electrical/fire, mechanical or SIS layers and what independence each requires, how the OT network is segmented from IT, whether fault domains are also security domains — must be made before you commission, because retrofitting an independent safety layer or re-segmenting a live OT network mid-life is brutal and sometimes impossible. The operational ones — vendor access brokering, firmware governance cadence, OT detection content, the converged escalation drill — are ongoing and improvable. The recurring mistake is treating an irreversible decision as if it were operational: assuming you can 'add a safety system later' to a hall whose trips already live inside the hackable BMS. Retrofitting may be possible, but its outage, redesign and safety-validation cost can dominate the original design choice. Assign the protection functions and verify common dependencies before commissioning; maintain their bypass and recovery evidence in service.
Cite this chapter
Fehn, J. (2026). Cyber-Physical & Destructive Attacks on OT/Facility Systems (Chapter 11.10). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-11-security/11-10-cyber-physical-and-destructive-attacks-on-ot-facility-systems (accessed 2026-09-29).
@misc{aidc-11-10,
author = {Fehn, Jacob},
title = {Cyber-Physical & Destructive Attacks on OT/Facility Systems (Chapter 11.10)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-11-security/11-10-cyber-physical-and-destructive-attacks-on-ot-facility-systems},
note = {Accessed 2026-09-29}
}