The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Part 14

Part 14

Day-2 Operations, Upgrades & Lifecycle

14 chapters

14.1
Operational KPIs, Goodput & the Reliability Economics of AI Factories
Goodput — accelerator-hours doing useful work on the critical path — governs an AI factory's day-2 economics; justify each reliability dollar by the useful output it buys while meeting the contracted service-availability obligation.
14.2
DCIM, Telemetry & Observability for GPU-Dense, Liquid-Cooled Facilities
Sub-second IT health and minute-scale OT plant telemetry can explain a goodput drop when timestamps, data quality and asset identities reconcile — specify that correlation contract before the DCIM purchase.
14.3
Component Failure Modes, Failure Rates & Fleet Reliability Data
A large GPU fleet accumulates hard, transient and silent failures; Meta’s every-few-hours job interruptions illustrate one operating regime, so measure each component’s exposure and the whole-job trace before engineering detection, spares and recovery.
14.4
Reliability Engineering for Training (Operational)
Where a frontier training job is interrupted every few hours, as in Meta’s named run, operations turns on detecting, isolating and recovering faster than failures accumulate — every recovery minute costs goodput on the declared job ledger.
14.5
Predictive & Preventive Maintenance of Power and Cooling Plant
Whether you can service the plant without dropping a synchronous job is set years earlier, by buying concurrent maintainability and building the condition-based program that intervenes on the equipment's schedule.
14.6
Spares Strategy, RMA Logistics & Repair Operations
On-site spares depth and supported swap speed set how much goodput a fleet of hundreds of thousands of accelerators recovers when replacement delay binds; the OEM warranty and replenishment channel determine whether that local pool stays usable.
14.7
Capacity, Power & Thermal Management in Operation
The megawatts you energized are a fixed, capital-intensive ceiling; steady-state operations trades how full you run that budget against how violently the workload can swing it.
14.8
Firmware & Software Lifecycle Management at Fleet Scale
Firmware and software changes can reach 100,000 GPUs on a synchronized estate; the craft is rolling a validated tuple through bounded cohorts without sacrificing goodput or shipping a bad bit beyond the tested recovery boundary.
14.9
Hardware Refresh, Depreciation Strategy, Decommissioning & ITAD
Refresh turns the Chapter 1.8 depreciation assumption into a physical operation; pull timing, cascade routing, sanitization, and the ITAD channel decide whether the underwritten residual value is actually realized.
14.10
Facility Decommissioning, Repowering & Site Remediation
The facility outlives its silicon by decades, and end-of-life is decided years early among repower, demolish, or convert — in a power-bound market the energized shell is often the most valuable asset.
14.11
Operations Organization, Workforce, Talent & Incident Command
The org chart, talent bench and incident-command model are reliability infrastructure: procedures and human factors create controllable failure paths, so qualified shift coverage and executable procedures decide whether operators can preserve goodput under stress.
14.12
Operational Procedures, Change Management & Human-Error Control
Procedure design and qualification turn the MOP/SOP/EOP regime into a reliability control: prove the intended safe state before intervention on operating plant and the restored service before closing the change.
14.13
Agentic Ops, RL Control & the Autonomy Ladder
An agent may act unsupervised on a control loop only as far as you have bounded its action space, tested its override path, and assigned the liability for when it is wrong.
14.14
Continuous & Re-Commissioning on a Live Campus
On a live campus, each change ticket and density step can invalidate part of the commissioning evidence; a standing re-commissioning program retains unaffected proof and retests changed behavior while protecting the revenue-bearing load.