The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
GuidePart 14

Part 14

Day-2 Operations, Upgrades & Lifecycle

14 chapters

14.1
Operational KPIs, Goodput & the Reliability Economics of AI Factories
Goodput — accelerator-hours doing useful work on the critical path — is the number that governs an AI factory's day-2 economics; every reliability dollar must be justified by the goodput it buys.
14.2
DCIM, Telemetry & Observability for GPU-Dense, Liquid-Cooled Facilities
Sub-second IT health and minute-scale OT plant telemetry only explain a goodput drop if they share a clock and an asset namespace — a decision made before the DCIM purchase.
14.3
Component Failure Modes, Failure Rates & Fleet Reliability Data
A GPU fleet fails constantly — hard, transient, and silent — so measure per-component rates, expect cluster interruptions every few hours, and engineer detection and recovery around the failures you cannot prevent.
14.4
Reliability Engineering for Training (Operational)
At frontier scale a training job is interrupted every few hours, so operations turns on detecting, isolating, and recovering faster than failures accumulate — every recovery minute costs goodput.
14.5
Predictive & Preventive Maintenance of Power and Cooling Plant
Whether you can service the plant without dropping a synchronous job is set years earlier, by buying concurrent maintainability and building the condition-based program that intervenes on the equipment's schedule.
14.6
Spares Strategy, RMA Logistics & Repair Operations
On-site spares depth and hot-swap speed, more than the OEM warranty, set how much goodput a fleet of hundreds of thousands of accelerators earns back from every failure.
14.7
Capacity, Power & Thermal Management in Operation
The megawatts you energized are a fixed, capital-intensive ceiling; steady-state operations trades how full you run that budget against how violently the workload can swing it.
14.8
Firmware & Software Lifecycle Management at Fleet Scale
Firmware and software change thousands of times a year on a synchronized estate; the craft is rolling that change without sacrificing goodput or shipping a bad bit to 100,000 GPUs at once.
14.9
Hardware Refresh, Depreciation Strategy, Decommissioning & ITAD
Refresh turns the Chapter 1.8 depreciation assumption into a physical operation; pull timing, cascade routing, sanitization, and the ITAD channel decide whether the underwritten residual value is actually realized.
14.10
Facility Decommissioning, Repowering & Site Remediation
The facility outlives its silicon by decades, and end-of-life is decided years early among repower, demolish, or convert — in a power-bound market the energized shell is often the most valuable asset.
14.11
Operations Organization, Workforce, Talent & Incident Command
The org chart, talent bench, and incident-command model are first-order reliability infrastructure: human action drives most outages, so who is on shift and what procedure they hold decides goodput.
14.12
Operational Procedures, Change Management & Human-Error Control
Most preventable outages trace to procedure design rather than carelessness, so availability is moved by engineering the MOP/SOP/EOP regime and the change process that put people in front of live plant.
14.13
Agentic Ops, RL Control & the Autonomy Ladder
An agent may act unsupervised on a control loop only as far as you have bounded its action space, tested its override path, and assigned the liability for when it is wrong.
14.14
Continuous & Re-Commissioning on a Live Campus
On a live campus, commissioning evidence decays with every change ticket and density step; re-commissioning is a standing program, fired by drift and run against a revenue factory you cannot turn off.