Guide › Part 14
Part 14
Day-2 Operations, Upgrades & Lifecycle
14 chapters
14.114.214.314.414.514.614.714.814.914.1014.1114.1214.1314.14
Operational KPIs, Goodput & the Reliability Economics of AI Factories
Goodput — accelerator-hours doing useful work on the critical path — governs an AI factory's day-2 economics; justify each reliability dollar by the useful output it buys while meeting the contracted service-availability obligation.
DCIM, Telemetry & Observability for GPU-Dense, Liquid-Cooled Facilities
Sub-second IT health and minute-scale OT plant telemetry can explain a goodput drop when timestamps, data quality and asset identities reconcile — specify that correlation contract before the DCIM purchase.
Component Failure Modes, Failure Rates & Fleet Reliability Data
A large GPU fleet accumulates hard, transient and silent failures; Meta’s every-few-hours job interruptions illustrate one operating regime, so measure each component’s exposure and the whole-job trace before engineering detection, spares and recovery.
Reliability Engineering for Training (Operational)
Where a frontier training job is interrupted every few hours, as in Meta’s named run, operations turns on detecting, isolating and recovering faster than failures accumulate — every recovery minute costs goodput on the declared job ledger.
Predictive & Preventive Maintenance of Power and Cooling Plant
Whether you can service the plant without dropping a synchronous job is set years earlier, by buying concurrent maintainability and building the condition-based program that intervenes on the equipment's schedule.
Spares Strategy, RMA Logistics & Repair Operations
On-site spares depth and supported swap speed set how much goodput a fleet of hundreds of thousands of accelerators recovers when replacement delay binds; the OEM warranty and replenishment channel determine whether that local pool stays usable.
Capacity, Power & Thermal Management in Operation
The megawatts you energized are a fixed, capital-intensive ceiling; steady-state operations trades how full you run that budget against how violently the workload can swing it.
Firmware & Software Lifecycle Management at Fleet Scale
Firmware and software changes can reach 100,000 GPUs on a synchronized estate; the craft is rolling a validated tuple through bounded cohorts without sacrificing goodput or shipping a bad bit beyond the tested recovery boundary.
Hardware Refresh, Depreciation Strategy, Decommissioning & ITAD
Refresh turns the Chapter 1.8 depreciation assumption into a physical operation; pull timing, cascade routing, sanitization, and the ITAD channel decide whether the underwritten residual value is actually realized.
Facility Decommissioning, Repowering & Site Remediation
The facility outlives its silicon by decades, and end-of-life is decided years early among repower, demolish, or convert — in a power-bound market the energized shell is often the most valuable asset.
Operations Organization, Workforce, Talent & Incident Command
The org chart, talent bench and incident-command model are reliability infrastructure: procedures and human factors create controllable failure paths, so qualified shift coverage and executable procedures decide whether operators can preserve goodput under stress.
Operational Procedures, Change Management & Human-Error Control
Procedure design and qualification turn the MOP/SOP/EOP regime into a reliability control: prove the intended safe state before intervention on operating plant and the restored service before closing the change.
Agentic Ops, RL Control & the Autonomy Ladder
An agent may act unsupervised on a control loop only as far as you have bounded its action space, tested its override path, and assigned the liability for when it is wrong.
Continuous & Re-Commissioning on a Live Campus
On a live campus, each change ticket and density step can invalidate part of the commissioning evidence; a standing re-commissioning program retains unaffected proof and retests changed behavior while protecting the revenue-bearing load.