The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Security › 11.10

Chapter 11.10

In this chapter · 7 sections
Term help

Cyber-Physical & Destructive Attacks on OT/Facility Systems

A compromised control plane can create destructive states, so allocate each protection function from the project hazard and risk analysis, then engineer the required independence, integrity, bypass control, diagnostics, proof testing, and defense in depth.

POWER-BOUNDGOODPUT

What you'll decide here

  1. Whether your OT estate (BMS, EPMS, SCADA, CDU controllers, BESS management, the power-cap/firmware planes) is genuinely segmented from IT per IEC 62443 zones-and-conduits, or merely firewalled on paper while exposed to the internet through a vendor jump-host nobody owns.
  2. Which destructive actions are physically possible from the control plane — a forced synchronized load step, CDU disablement, BESS abuse, or malicious firmware — and which independent or diverse protective layers limit each consequence.
  3. How the project hazard and risk analysis allocates protection functions among process controls, OEM protections, electrical/fire/life-safety systems, mechanical relief, dedicated relays, and any SIS/SIF — including required independence, integrity, bypass governance, diagnostics, proof testing, and secure maintenance.
  4. Where random failures and adversarial abuse share a physical consequence, and where the adversarial case needs additional authorization, integrity, detection, containment, persistence, and recovery evidence rather than an assumed identical twin.
  5. Who owns OT detection and response — a converged cyber-physical SOC with the escalation trigger wired in, or two teams (facilities and security) who each assume the other is watching the controllers.

Most of Part 11 defends the information in the building — model weights, tenant data, credentials. This chapter defends the building itself. An AI factory is a cyber-physical system: tens of megawatts of synchronized silicon sitting on a control plane of programmable logic controllers, building-management and electrical-power-management systems, coolant-distribution-unit firmware, battery-management systems, and a power-capping layer that the GPUs themselves obey. Every one of those is a computer, and reachability must be established from actual wiring, networks and maintenance paths. Alongside information loss, an attack on the OT plane can physically destroy the asset — cause equipment damage, trip the grid interconnection, contribute to battery hazards, or disable accelerators through an update path if credited protections fail. The reachable control scope and remaining protection layers determine blast radius; equipment damage can put long-lead replacement on the recovery path.

The sections that follow enumerate the OT assets and their security levels, the destructive primitives the control plane exposes and what each one costs, the segmentation boundary (the Purdue model and IEC 62443 zones-and-conduits), and how the project hazard and risk analysis allocates each protection function — process control, OEM protection, dedicated relays, mechanical relief, fire and life safety, and where the safety lifecycle requires it a safety instrumented function — so that no single compromised control domain can defeat every credited layer. Assess each Appendix-F failure mode for adversarial reachability: the transient you engineer against as a random fault is also the attacker's payload, and the attacker picks the timing.

The OT/ICS threat model: what is actually reachable

Start by enumerating the control planes, because each is a distinct attack surface with a distinct destructive payload. The BMS (building-management system) can command cooling setpoints, valve and pump commands, air handlers, leak detection, and interfaces to fire/EPO systems, whose authority and independent protection must be verified. The EPMS (electrical-power-management system) monitors electrical state and may command breakers, transfer switches, generator start/stop and UPS modes where the design grants it that authority. SCADA/PLC layers sit under both, executing the actual I/O. CDU controllers regulate the technology-cooling loop — flow, inlet temperature, secondary-loop pressure — that keeps a GB300 NVL72-class rack at its up-to-142 kW facility design basis from cooking itself. BESS management (battery-management systems and their EMS) governs charge/discharge, cell balancing, and the thermal-runaway protections on grid-scale storage. And the power-cap and firmware planes — the GPU/BMC layer that throttles and updates accelerators at fleet scale — need the same command-authorization and failure-containment analysis as the older controllers.

Published exposure studies justify testing these paths. Claroty's Team82 analyzed roughly 467,000 building-management devices across 529 organizations in 2025 and found 75% of organizations ran BMS affected by known-exploited vulnerabilities (KEVs), with 51% exposed to KEVs that are both ransomware-linked and insecurely connected to the internet (Claroty Team82, State of CPS Security 2025: BMS Exposures). Those are cross-industry organization-level results; the data-center-specific cut followed on 2026-07-29, when Team82 analyzed more than 750,000 CPS assets across the world's largest data-center facilities and found nearly 1 in 5 infrastructure assets (18%, 32,000+ of 174,577) one hop from a system making risky outbound connections to the public internet, 88% of BMS communicating over insecure protocols and 40% running outdated firmware, with PDUs (41%) and HVAC/cooling (32%) the most-exposed classes and more than 80% of OT control, power-monitoring and IoT systems speaking BACnet/Modbus-class legacy protocols (Claroty Team82, State of CPS Security: Data Center Exposures). The same team disclosed, in June 2026, two CVSS 9.8 flaws in Vertiv Liebert IS-UNITY-DP and RDU101 UPS network cards (CVE-2025-46412, CVE-2025-41426) and a five-CVE chain in the Trane Tracer SC+ HVAC controller up to v5.20.1362, fixed in v6.3 — named data-center equipment, not generic BMS. The population is the studied facilities, not this campus's measured exposure.

The power-cap and firmware as a weapon

This is the part of the threat model unique to AI factories, and it is where the destructive primitives concentrate. Four are worth naming explicitly, each with its downstream cost.

The forced synchronized load step → grid trip. A gigawatt-class training cluster already swings load by hundreds of megawatts when a job starts, stops, or stalls — that is the random-fault problem Chapter 4.5 engineers ride-through against. Now make it the payload. An attacker who controls the power-cap plane (or simply the scheduler) can command thousands of racks to drop or pick up load simultaneously, deliberately, at the moment of peak grid stress. NERC has already flagged the non-malicious version as a reliability crisis: in the July 2024 Northern Virginia incident, a permanent 230-kV transmission-line fault and staggered auto-reclosing produced six successive system faults over 82 seconds. Roughly 1,500 MW of data-center load reduced coincident with the disturbance; about 1,260 MW dropped simultaneously with the third voltage depression and remained off for hours (NERC). This was not an isolated event: NERC has documented repeated 1,000+ MW computational-load losses since 2022, and on May 4 2026 issued its rare Level 3 "Essential Actions" alert over the pattern — an implementation directive, not a penalty-backed standard, and issued ~22 months after this Virginia event rather than in immediate response to it. Weaponized, the same primitive is a tool to trip the interconnection or destabilize the local grid, and the facility's own protection relays may disconnect it — taking down the campus to save the grid. The downstream cost is not a reboot; it is a black-start and a damaged relationship with the utility that holds your interconnection.

CDU disablement → thermal runaway. An NVL72-class liquid-cooled rack's response after CDU flow loss (throttle, controlled shutdown, emergency shutdown or trip) must be established from OEM transient data, controls tests or a professional-engineering calculation; do not infer a portable seconds-to-tens-of-seconds value from chilled-water inertia. An attacker who stops the CDU pumps, closes a facility-water valve, or falsifies the inlet-temperature reading so the controller never opens the valve, drives the rack straight into thermal runaway. Done across a hall, that takes cold plates and silicon past their thermal limits, with a recovery bounded by component lead times.

BESS-induced runaway. Grid-scale and UPS-class battery storage protects against thermal runaway through the BMS: cell balancing, over-temperature trips, and isolation. Compromise that management layer and you can disable the very protection that prevents a runaway, force an over-charge or over-discharge, or suppress the thermal alarm — turning the energy-storage asset into an ignition source inside the building.

GPU-bricking via malicious firmware. The fleet firmware-update path (BMC, GPU VBIOS, NVLink-switch firmware) is a fleet-scale destructive primitive: a single malicious image, pushed through the legitimate update channel, can render thousands of accelerators non-functional or subtly mis-behaving. This is why firmware integrity (signed images, a hardware root of trust, measured boot) is the gate on the most expensive single payload in the building. The canonical treatment of that integrity chain is in Chapter 11.4; here it is one of four ways the control plane bricks the asset.

Destructive primitive → control plane → physical consequence → the interlock that should veto it
Destructive primitiveControl plane abusedPhysical consequenceRecovery scaleIndependent protection layer (allocated from the hazard analysis)
Forced synchronized load stepPower-cap plane / scheduler / EPMSGrid trip; interconnection disconnect; possible black-startHours to days (grid + black-start)Utility protection relays; on-site ride-through (BBU/supercap) sized for the swing
CDU disablement / falsified inlet tempCDU controller / BMS valve & pump I/OThermal runaway of liquid-cooled racks; silicon damageWeeks to months (cold plates, GPUs)Hardwired high-temp trip and flow-loss interlock independent of the BMS
BESS over-charge / alarm suppressionBESS BMS / energy-management systemBattery thermal runaway; fire; structural lossMonths (rack + structure + remediation)Independent cell-level protection and gas/thermal detection on a separate logic solver
Malicious firmware pushBMC / GPU VBIOS / switch firmware planeFleet-scale bricking or covert mis-operationWeeks to months (re-image or RMA fleet)Hardware root of trust; signed measured boot; firmware governance (11.4)
EPO / breaker mis-operationEPMS / SCADAUnplanned full or partial campus shutdownHours (restart) to days (equipment stress)Hardwired EPO logic; mechanical breaker interlocks; key-locked overrides
Each row is a dual-use failure mode: the same event appears in Appendix F as a random fault. The right column is the hardwired or independent control that must hold even when the control plane is fully compromised.

OT/IT segmentation: the Purdue model and IEC 62443

The single highest-leverage control against all of the above is the oldest one: keep the control plane off the corporate network and off the internet. The Purdue model stratifies the estate into levels — physical process (0), basic control / PLCs (1), supervisory / SCADA-HMI (2), operations / MES (3), then the DMZ (3.5) and enterprise IT (4–5) — with the principle that traffic crosses level boundaries only through controlled, inspected conduits, never directly. Use IEC 62443-2-1:2024 for the asset-owner program and IEC 62443-3-2:2020 for system-design security risk assessment. The latter scopes zones and conduits and their target security levels; a network layout alone does not establish conformance. The standard's SL scale is defined by attacker capability — SL1 casual/coincidental, SL2 intentional with simple means, SL3 sophisticated with moderate resources, SL4 sophisticated with extended resources (i.e. a nation-state) (ISA/IEC 62443). The destructive primitives above sit at the IEC 62443 SL3–SL4 end of that scale — 62443's levels, not the RAND Weights Security Levels of Chapter 11.1 — and an internet-reachable BMS is evidence that a zone's achieved protection has drifted from its assessed target, not an automatic level score.

Two architectures pull against each other here. Convergence — running OT on the same IP fabric as IT, with shared identity and shared monitoring — is operationally attractive: one network team, one observability stack, remote vendor access for the cooling and electrical contractors who actually maintain the gear. Isolation — an air-gapped or diode-separated OT network with its own identity, jump-hosts, and one-way telemetry export — is far harder to operate and far harder to compromise. For an AI campus the practical middle is a hardened DMZ with brokered, recorded, time-boxed vendor access and unidirectional telemetry out, sized so that an IT compromise cannot reach a Level-1 controller and a controller cannot reach the internet. The microsegmentation and zero-trust mechanics that enforce this are the subject of Chapter 11.7; this chapter's contribution is the consequence of getting it wrong, which is the destructive-primitive table above.

Assign the IEC 62443 evidence to the party and assertion it covers. 62443-3-3:2013 supplies system security requirements; 62443-2-4:2023 addresses integration and maintenance service-provider processes. 62443-4-1:2018 addresses the product development lifecycle; 62443-4-2:2019, including its 2022 corrigendum, addresses component technical capabilities. Keep the target security requirements (SL-T), demonstrated system result and component capability (SL-C) separate. A component certificate does not establish the target or achieved protection of the assembled cooling or power system. The owner retains the system risk and integration evidence, even when suppliers provide certified components.

Convergence vs isolation for the OT plane
DimensionFull convergence (OT on IT fabric)Brokered DMZ (recommended)Hard isolation (air-gap / diode)
Attack surfaceLargest — IT compromise reaches PLCsBounded — conduit + broker onlySmallest — no inbound path
Vendor/remote maintenanceTrivial but dangerousRecorded, time-boxed, brokeredOn-site or sneakernet only
Telemetry / observabilityNative, unified SOC viewUnidirectional export to SOCManual or one-way diode export
Operational costLowestModerateHighest
Residual access paths into Levels 0–1IT compromise reaches PLCs directlyBroker, conduit, and any standing bypass accountPhysical access, maintenance laptops, removable media
The middle column — brokered DMZ — is the practical answer for most AI campuses; the question is which way you lean from there. Architecture does not award a security level: 62443 target SLs come out of the zone-and-conduit risk assessment, and target, capability, and achieved levels are three separate claims needing three separate bodies of evidence.
75%
of organizations with BMS affected by known-exploited vulnerabilities; 2025 commercial-building population (529 organizations)
Scope & caveats

Prevalence across the 2025 commercial-building study population (roughly 467,000 BMS devices, 529 organizations); Team82's July 2026 data-center report is a separate, data-center-specific population.

51%
of organizations exposed to KEVs that are ransomware-linked AND insecurely internet-connected
18% (nearly 1 in 5)
of data-center infrastructure assets one hop from a risky outbound internet connection (Team82 data-center report, 174,577 assets)
Scope & caveats

Population: more than 750,000 cyber-physical-system assets and core infrastructure across the world's largest data-center facilities in Team82's telemetry; the one-hop denominator is 174,577 data-center infrastructure assets. Studied facilities, not this campus's measured exposure.

41%
of PDUs one hop from a risky internet connection — the most-exposed data-center asset class
Scope & caveats

Population: more than 750,000 cyber-physical-system assets and core infrastructure across the world's largest data-center facilities in Team82's telemetry; the one-hop denominator is 174,577 data-center infrastructure assets. Studied facilities, not this campus's measured exposure.

32%
of HVAC/cooling systems one hop from a risky internet connection — second most-exposed class
Scope & caveats

Population: more than 750,000 cyber-physical-system assets and core infrastructure across the world's largest data-center facilities in Team82's telemetry; the one-hop denominator is 174,577 data-center infrastructure assets. Studied facilities, not this campus's measured exposure.

88%
of data-center BMS communicating over insecure protocols (66,395 devices)
Scope & caveats

Population: more than 750,000 cyber-physical-system assets and core infrastructure across the world's largest data-center facilities in Team82's telemetry; the one-hop denominator is 174,577 data-center infrastructure assets. Studied facilities, not this campus's measured exposure.

40%
of data-center BMS running outdated firmware
Scope & caveats

Population: more than 750,000 cyber-physical-system assets and core infrastructure across the world's largest data-center facilities in Team82's telemetry; the one-hop denominator is 174,577 data-center infrastructure assets. Studied facilities, not this campus's measured exposure.

600 buildings / ~2 days
FrostyGoop district-heating incident: malicious Modbus commands; distinguish the associated firmware action
~1,500 MW
total data-center load reduction coincident with six system faults over 82 s
Scope & caveats

Load loss as seen by the grid. NERC's incident review ('Load Details') found the affected data centers transferred their loads to backup power — static UPS, decentralized rack UPS, or DRUPS — in response to the disturbance. The figure is a loss of demand at the interconnection, not evidence that IT power was interrupted or that training jobs restarted.

The approximately 1,500 MW is the total customer-side load reduction coincident with the six-fault sequence; NERC reports approximately 1,260 MW as the sustained drop at the third voltage depression. The NERC-investigated canonical case. A second, larger occurrence followed on 2026-07-22: ~3.8 GW dropped on a single normally-cleared Ashburn 230 kV fault (see companion key number). Two vintages of the same failure mode, not a replacement figure.

SL1-SL4
IEC 62443 security requirements derive from risk and zone/conduit assessment; topology or destructive authority does not award a level
Port 502
Modbus TCP — unauthenticated by design; the protocol FrostyGoop and most BMS/CDU controllers speak

Safety-instrumented systems: the control software cannot be allowed to win

Start with the project hazard and risk analysis and applicable process, electrical, fire/life-safety, equipment, and cybersecurity requirements. Allocate each protective function to the architecture that meets its required integrity and consequence: basic/process control, an OEM protective function, a dedicated relay, mechanical relief, fire/life-safety control, or — where the safety lifecycle requires it — a safety instrumented function (SIF) implemented by an SIS. An SIS is not the default home for every equipment trip. For an allocated SIF, IEC 61511 / ISA-84 lifecycle principles require the necessary independence, integrity, diagnostics, proof testing, management of change, and controlled bypasses; 'independent' is a designed and verified claim, not a promise that software can never defeat the function.

The 2017 Triton/Trisis attack showed that a safety layer can itself be targeted and reprogrammed. Use the project cyber/process hazard analysis to identify common dependencies, communications and maintenance paths, privileged access, sensor/final-element diversity, power, bypasses, and proof-test coverage. Separate or diversify protection where the required risk reduction demands it, secure and monitor authorized maintenance paths, and retain other independent or passive layers where one digital protection system cannot carry the consequence alone. The correct boundary is project-specific; neither 'hardwire every trip' nor 'the SIS is impossible to defeat' is a defensible universal rule.

Every Appendix-F failure mode is dual-use

The reframing that should run through your entire FMEA: each failure mode in the Appendix-F catalog has two possible causes, and the second one needs a cyber-reachability assessment. One is randomness — a pump fails, a valve sticks, a relay chatters, a load step happens because a job crashed. The other is an attacker who may induce a similar physical consequence through a different command or persistence path, with three advantages the random fault does not have: timing (at peak grid stress, during a maintenance window, when redundancy is degraded), correlation (across many units at once, defeating the N+1 that assumed independent failures), and persistence (re-triggering as fast as you recover).

This breaks a core assumption of classical reliability engineering. Simple independence calculations in Chapter 0.5 and Part 12 must be supplemented by common-cause models. A cyber-physical attacker deliberately violates independence: a single compromised controller can trip every CDU in a hall at once, and the 'redundant' second pump on the same compromised PLC is no redundancy at all. The design response is that fault domains must also be security domains: identify whether the controller that can fail unit A also reaches unit B or their shared protection; allocate the required independence and diversity from the project hazard and cyber-risk analysis.

Deep dive: worked example — the worst-timed attack on a 100 MW liquid-cooled hall

Walk the chain an adversary at the IEC 62443 SL3–SL4 capability end would actually run, to see where each control either holds or fails. Step 1 — access. The attacker reaches the OT DMZ through a contractor's jump-host left standing after a cooling-vendor maintenance window (the FrostyGoop pattern: a forgotten edge device, not a zero-day). If the DMZ is brokered and time-boxed, the jump-host is gone and this step fails; if OT is converged with IT, an earlier IT phishing foothold already put them here. Step 2 — reconnaissance. Modbus and BACnet are unauthenticated, so they map the CDU controllers, chiller PLCs, and EPMS without exploiting anything — they simply read.

Step 3 — the payload, timed. They wait for a maintenance window when the hall is on N (no spare cooling unit), then simultaneously falsify inlet-temperature readings on every CDU so the controllers hold their valves closed, and stop the pumps. With no chilled-water inertia, the rack response must be established by the selected OEM's transient data, controls tests or engineering calculation; do not infer a universal trip time. This is where the validated protection architecture changes the outcome. If all credited sensing, logic, actuation, power, communications, and bypass paths share the compromised control domain, the protection claim can fail. Where the project analysis requires independent or diverse layers, their tested response should drive the affected scope toward the defined safe state. Actual propagation, equipment damage, shutdown scope, and recovery time remain topology-, control-, inertia-, and attack-path-dependent; they must come from the project model and evidence, not a universal months-versus-hours promise. Step 4 — the multiplier. Simultaneously they command a synchronized load drop to stress the interconnection, so the facility is fighting a grid event and a thermal event at once, degrading the human response. The defenses that turned a catastrophe into a bad day were all decided at design time: segmentation that bounded step 1, an independent and verified protection layer that bounded step 3, and ride-through sized for step 4 (Chapter 4.5). None could be added in the moment.

Detection, response, and the converged escalation trigger

Prevention is the priority, but the OT plane also needs detection that a converged SOC can act on. The hard problem is that OT telemetry and IT telemetry live in different worlds: a falsified Modbus write looks like a legitimate engineering command, and the signal that distinguishes attack from operation is often physical — a setpoint that contradicts the measured process state, a valve commanded closed while temperature climbs, a load step with no scheduling event behind it. So the detection that matters here is cyber-physical correlation: cross-checking commanded state against sensed state against expected workload, and treating divergence as an incident.

The organizational failure mode is facilities and security each assuming the other is watching the controllers — the gap through which most OT incidents walk. The answer is a converged cyber-physical escalation trigger: a single condition (anomalous OT command + adverse physical trend) that pages both the SOC and facilities engineering simultaneously and invokes a joint playbook, because neither team can diagnose a CDU-disablement-into-thermal-runaway alone. The unified incident-command model that this escalation feeds into is canonical in Chapter 14.11, and the OT/cyber-physical IR playbook itself is built out in Chapter 11.12; this chapter's job is to insist the trigger exists and is wired to the physics, not just to the logs.

What to decide, and when

The decisions in this chapter sort sharply into irreversible-at-design-time and operational. The irreversible ones — which protection functions belong to process/OEM, electrical/fire, mechanical or SIS layers and what independence each requires, how the OT network is segmented from IT, whether fault domains are also security domains — must be made before you commission, because retrofitting an independent safety layer or re-segmenting a live OT network mid-life is brutal and sometimes impossible. The operational ones — vendor access brokering, firmware governance cadence, OT detection content, the converged escalation drill — are ongoing and improvable. The recurring mistake is treating an irreversible decision as if it were operational: assuming you can 'add a safety system later' to a hall whose trips already live inside the hackable BMS. Retrofitting may be possible, but its outage, redesign and safety-validation cost can dominate the original design choice. Assign the protection functions and verify common dependencies before commissioning; maintain their bypass and recovery evidence in service.

The transient physics that make CDU disablement and load steps destructive on product- and implementation-specific timescales are engineered in Chapter 4.5 (ride-through and transient absorption) and Chapter 5.12 (cooling-controls stability). The grid-coupling side of the synchronized-load-step primitive is in Chapter 4.10, with NERC compliance for transmission-connected loads in Chapter 4.3. Firmware integrity — the gate on the GPU-bricking primitive — is canonical in Chapter 11.4; the segmentation and zero-trust mechanics that enforce the Purdue/62443 boundary are in Chapter 11.7. This chapter sits inside the broader threat model of Chapter 11.1; the insider who already has OT access is in Chapter 11.9; and the OT incident-response playbook plus converged SOC are in Chapter 11.12, escalating into the incident-command model of Chapter 14.11.
Cite this chapter
Fehn, J. (2026). Cyber-Physical & Destructive Attacks on OT/Facility Systems (Chapter 11.10). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-11-security/11-10-cyber-physical-and-destructive-attacks-on-ot-facility-systems (accessed 2026-09-29).
@misc{aidc-11-10,
  author       = {Fehn, Jacob},
  title        = {Cyber-Physical & Destructive Attacks on OT/Facility Systems (Chapter 11.10)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-11-security/11-10-cyber-physical-and-destructive-attacks-on-ot-facility-systems},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit