The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Commissioning & Go-Live › 13.9

Chapter 13.9

In this chapter · 7 sections
Term help

Cluster-Scale Benchmarking, Reference Training & Storage/Scheduler Validation

A cluster's performance is accepted through the production scheduler and storage: training needs ML Productivity Goodput measured over a sustained reference training run, inference needs correct responses within its latency contract, and neither can average away a failed safety, correctness, isolation or recovery gate.

GOODPUTPOWER-BOUNDDENSITY-RAMP

What you'll decide here

  1. What busbw and completion-time bounds each collective, message size and rank layout must clear — and how you isolate nodes that pass burn-in but drag the collective, without assuming a fixed defective fraction.
  2. Whether acceptance stops at synthetic NCCL/RCCL/OSU/MLPerf results, or includes the contracted training, checkpoint/restore and inference service through production scheduler and storage for an evidence-selected window.
  3. The contractual goodput SLA — which goodput the operator commits to: ML Productivity Goodput (scheduling × runtime × program/MFU) or the runtime layer alone, since the two differ by the whole MFU factor — and the badput taxonomy that decomposes every percentage point you lose below it.
  4. Where storage acceptance sets its bar: sustained checkpoint write bandwidth (the failure-recovery path), data-loader read throughput (the steady-state path), and the metadata/small-file rate that LOSF workloads quietly destroy.
  5. Whether the scheduler is validated for gang/topology-aware placement and preemption-without-corruption before handover — because a scheduler that fragments NVLink domains strands bandwidth no benchmark will catch.

By the time a cluster reaches this chapter it has been heavily tested in pieces. The power chain demonstrated its failure modes under load (Chapter 13.6), the fabric was commissioned link-by-link with point-to-point bandwidth and PTP time-sync gates (Chapter 13.7), and every node completed the signed burn-in campaign, diagnostics, and silent-data-corruption screen selected from test intensity, fleet exposure, and the measured defect-discovery stopping rule (Chapter 13.8). All of it passed. And none of it tells you whether the machine you actually bought — a single, tightly-coupled supercomputer — works. This chapter closes that gap: cluster-scale validation, where the unit under test is the whole, and the acceptance criterion is no longer "does each part meet spec" but "does the assembled system deliver the goodput the contract assumes."

Validation proceeds in three widening rings, and a ring does not open until the prior one clears. Ring one is synthetic collective benchmarking — NCCL and OSU at scale — which isolates the fabric and the communication libraries from any application noise and gives you a hard, repeatable busbw number to gate on. Ring two is subsystem acceptance — storage (checkpoint write, data-loader read, metadata) and the scheduler/orchestrator (gang placement, topology-awareness, preemption) — each validated against its own bandwidth and correctness bar. Ring three is a reference (proxy) training run that fuses all of it: a real model, real optimizer state, real checkpoints to real storage, scheduled the way production will schedule, run long enough to surface stragglers, thermal drift, and the slow leaks that a five-minute benchmark never sees. For training, ring three supplies measured goodput for the declared SLA layer. For inference, it supplies correct request throughput and tail latency for the agreed traffic mix. Earlier correctness, safety, isolation and recovery gates remain independently binding.

Ring one: NCCL collective benchmarking and the busbw gate

The fabric was already proven point-to-point in Chapter 13.7 — each accepted path meets its named physical and payload gates, every cable reconciles to the map, and FEC/counter evidence is inside its profile. That proves the wires. It does not prove the collective, and the collective is what training actually runs: all-reduce, all-gather, reduce-scatter, and broadcast across the full job on every step. A fabric that is flawless link-by-link can still collapse on an all-reduce because of a single mis-seated transceiver three hops away, a routing imbalance that congests one rail, an adaptive-routing misconfiguration, or a SHARP/NVLS offload that failed to engage. The collective benchmark is the only test that exercises the fabric the way the workload will.

The canonical tool is the NCCL tests suite (all_reduce_perf, all_gather_perf, and relatives), run as a multi-node MPI/SLURM job across the full cluster and a sweep of message sizes from a few KB to multiple GB. The metric that matters is bus bandwidth (busbw) — not algorithm bandwidth — because busbw normalizes for the collective's data-movement factor; comparison to a hardware ceiling still requires the same operation, rank count and offload convention. You read it two ways: the plateau busbw at large message sizes (does the fabric reach its bandwidth ceiling) and the small-message latency floor (does the collective's fixed cost stay bounded, which governs strong-scaling efficiency). On a healthy GB200 NVL72, in-domain all-reduce busbw runs roughly 870-928 GB/s across large buffers, effectively saturating the 900 GB/s/GPU NVLink5 unidirectional rate — busbw follows nccl-tests' ring convention (algbw x 2(n-1)/n), so with NVLink SHARP doing in-switch reduction it can read slightly above the physical per-link rate (NCCL tests on GB200 NVL72, 2025); the scale-out, multi-rack number over InfiniBand or Spectrum-X is a separate gate derived from the named NICs, active rails, topology and oversubscription, protocol efficiency, message-size plateau, and collective data-movement model.

Two companion suites round out ring one. OSU Micro-Benchmarks (osu_bw, osu_latency, osu_allreduce) give a vendor-neutral, MPI-level cross-check on NCCL's numbers, useful precisely because they are a different code path. If the two disagree, compare transport, message size, rank placement and routing before assigning the fault: the suites take different paths through the same fabric, so a healthy NCCL sweep does not clear the fabric and a slow OSU run does not convict the collective library. MLPerf Training (and, for storage, MLPerf Storage, below) provides an industry-comparable, full-workload reference (the current round is MLPerf Training v6.0, June 2026 — the first with mixture-of-experts benchmarks including DeepSeek-V3 671B; CoreWeave trained it in ~2 minutes on 8,192 GB300 NVL72 GPUs) if the operator wants a number that benchmarks against published results rather than only against the cluster's own reference. MLPerf is heavyweight to run correctly, and whether it earns its engineering cost depends on the operator: for a neocloud selling capacity, a clean MLPerf submission is a marketing-grade external attestation that an internal NCCL number is not. A third, vendor-side comparable arrived in 2026: NVIDIA Exemplar Cloud, a production-recipe validation suite (Slurm-based, llmb-run exemplar) with a published pass floor of 95% of the NVIDIA reference architecture. NVIDIA's own data shows two clusters of identical iron routinely differing 8–12% on the same model from configuration alone — Grace-VM command-queue virtualization, C-states/NUMA, NCCL_IB_QPS_PER_CONNECTION, a topology file never mounted into the container — with worked cases running from a 12–14% GB200 MoE gap to a 31% iteration-time gap on a 512-GPU GB300/ConnectX-8 cluster (AllGather ~28 → ~61 GB/s once QPS went 1→4). It is a vendor RA gate, not a substitute for ring three's owner-workload soak or the goodput SLA; neoclouds now collect the Exemplar stamp as a go-live trophy (IREN's Horizon 1 was Microsoft-accepted and GB300 Exemplar-validated in the same week, Aug 2026). → external-rating context in Chapter 12.2. All three of those suites are NVIDIA-path tools; the table below carries the same device, workload and recovery evidence onto AMD Instinct, Google Cloud TPU and AWS Trainium, where the diagnostic suite differs and, on a managed service, the provider rather than the customer owns physical acceptance.

Platform acceptance paths
Platform / boundaryDevice and interconnect evidenceWorkload and recovery evidenceAccess / responsibility
NVIDIA / owner-operatedInstalled DCGM/OEM manifest; supported NCCL tests, topology and countersPinned framework, correctness, checkpoint and training/inference profileOwner accepts installed hardware; missing required tests remain HOLD
AMD Instinct / owner-operatedApplicable AGFHC/RVS/CVS procedures; AMD SMI, TransferBench and RCCL testsSupported ROCm/framework tuple; same outcome gates and restore recordAcquire gated OEM diagnostics; no assumed DCGM equivalence
Google Cloud TPU / managed serviceExposed slice/worker health and ICI/DCN topologyPinned provisioning route, distributed workload, checkpoint and slice-replacement behaviorProvider owns physical acceptance; customer proves service/recovery and escalation
AWS Trainium / managed serviceNeuron tooling and exposed device/error telemetryPinned Neuron/framework tuple; checkpoint and instance/job replacementProvider owns facility/device repair; customer proves exposed service and recovery

Methods: AMD Customer Acceptance Guide; NVIDIA DCGM; Google TPU multislice; AWS Neuron Monitor. Freeze the release and service route; tool names do not prove coverage.

Cluster-scale validation rings — what each tests, the tool, and the gate
Ring / targetPrimary toolMetric gated onAcceptance gateWhat it catches that lower rings miss
Ring 1 — fabric collectiveNCCL tests; OSU; (opt.) MLPerf TrainingBus bandwidth (busbw) + small-msg latencyIllustrative contract: ≥0.90 × matched reference busbw; no rail >5% below matched peer mean; separate completion-tail gateRouting imbalance, congestion, failed SHARP/NVLS, straggler nodes
Ring 2 — storage (write path)MLPerf Storage (checkpoint); fio; mdtestSustained checkpoint write GB/s; drain timeUnique bytes drain and commit before the protected-state deadline; pause/reload budgets passAggregate write collapse, metadata bottleneck, LOSF
Ring 2 — storage (read path)MLPerf Storage (training); fio; loader replaySustained data-loader read GB/s; per-GPU targetMeets per-GPU read target at full node countData-loader starvation, cache thrash, small-file IOPS wall
Ring 2 — scheduler / orchestrationSLURM block / topology.yaml; KAI/Run:ai; gang testsGang placement correctness; topology fidelityJobs land NVLink-domain-aligned; preempt w/o corruptionFragmented scale-up domains, broken gang semantics, quota leaks
Ring 3 — contracted training/inference runReal/proxy model (e.g. GPT/Llama-class) at scaleTraining: ML Productivity Goodput (scheduling × runtime × MFU); inference: correctness, throughput, latency tailsSustained training goodput or inference latency/throughput meets the contractual SLA over its declared windowThermal drift, slow leaks, real badput, end-to-end interactions
The discipline is to advance rings only when the prior gate clears. Busbw and bandwidth figures are GB200 NVL72-class reference points of 2025 vintage. GB300 NVL72’s nominal NVLink 5 rating is 1.8 TB/s per GPU bidirectionally, or 900 GB/s in one direction. That component rating does not validate a GB300 collective acceptance result; qualify its rank layout, software tuple, buffer sizes and NCCL normalization separately from the GB200 observation.

Ring two: storage acceptance — three different bars, not one

Storage is where cluster acceptance most often goes wrong, because operators test it as one thing when it is three, with three different failure modes and three different bandwidth bars. → architecture in Chapter 9.5 and Chapter 9.6; checkpoint math in Chapter 9.4.

The write path is checkpointing, and it gates failure recovery. When a synchronous training job checkpoints, every rank drains optimizer and model state to the parallel filesystem in a burst; the whole job stalls until the slowest writer finishes (unless async/multi-tier checkpointing overlaps it). The acceptance question is sustained aggregate write under a realistic concurrent burst, not peak write bandwidth, because that is what sets how long the job is frozen every checkpoint interval. The reference points from VAST's survey of 40 production runs are a sizing assumption of ~14 bytes per parameter of checkpoint state — an assumed format, not a universal one — while the survey's asynchronous overlap comparison is scoped to total training time, not a universal acceptance cap. Size the write path from the quantities kept separate: checkpoint interval, serialized bytes, the synchronous pause every rank takes, the asynchronous drain that follows, the point of durable completion, restore time after a failure, and the training overhead the contract permits. Keep pause, background drain and protected commit separate, as in Chapter 9.4: a write subsystem that benchmarks beautifully on a single stream but collapses when 128 nodes write simultaneously fails the test that matters. MLPerf Storage v2.0 (Aug 2025) added a dedicated checkpoint benchmark precisely because, in a 100,000-accelerator cluster at full utilization, failures can land roughly every 30 minutes (MLCommons), and checkpoint write/restore bandwidth is what converts those failures from minutes of lost work into hours.

The read path is the data loader, and it gates steady-state goodput. If the loader cannot feed the GPUs, they idle — invisible to a fabric benchmark, fatal to MFU. The bar is a sustained per-GPU read throughput target met at full node count (not a four-node extrapolation), validated by replaying the actual loader or by MLPerf Storage's training workloads. The acceptance run must hit the actual loader's per-GPU read target at full concurrency as well as the independent checkpoint write and restore budgets from Chapter 9.4; an aggregate read number or a write-to-read ratio hides hot-spotting and cannot select the recovery deadline.

The metadata path is the silent killer. Lots-of-small-files (LOSF) workloads — image datasets, sharded tokenized corpora — stress metadata operations and small-file IOPS, not bandwidth, and a parallel filesystem that delivers terabytes-per-second of sequential read can still fall over on a million-file-per-second metadata storm. mdtest and a representative LOSF replay belong in acceptance, because LOSF degrades scaling in a way no large-block fio run will reveal.

Deep dive: why you put the storage fabric where the rebuild traffic cannot collide with the collective

A subtle acceptance failure that only a full-cluster run exposes: storage traffic and collective traffic sharing a fabric. The reference architectures differ on where storage rides — the DGX SuperPOD H100 architecture defines a dedicated InfiniBand storage fabric, while the B300 architecture supports InfiniBand or Ethernet for it — so the acceptance question is not which protocol carries checkpoints but whether storage traffic and collective traffic share a contention and failure domain. The reason it matters is a correlated-failure interaction. When a node fails mid-run, two things happen at once: the parallel filesystem may begin rebuilding/rebalancing (a large, sustained read/write storm), and the surviving training job restarts from checkpoint (a large read storm to reload state). If both of those land on the same fabric the collective uses, the restart all-reduce contends with rebuild traffic exactly when the cluster is trying to recover, and goodput craters at the worst possible moment.

The acceptance implication is that you cannot validate storage and fabric in isolation and call it done. The ring-three reference run must include at least one injected node failure under checkpoint load, so you observe what the restart actually costs when storage and fabric interact under stress. A cluster that passes every static benchmark but has co-mingled storage and collective traffic will show a clean acceptance report and a punishing real-world recovery profile. This is the argument for isolating storage traffic from collective traffic — a separate fabric, or demonstrated isolation and service guarantees on a shared one — and the reason ring three exists. Identify the architecture's storage fabric, then prove its behaviour under concurrent checkpoint, restart and rebuild load; a dedicated InfiniBand storage fabric that follows its OEM reference is not a defect. → fabric topology in Chapter 8.5; checkpoint/restart economics in Chapter 9.4 and operational tuning in Chapter 14.4.

W-01 trace and policy choice. A snapshots t = 0, releases at 2 s, drains at 42 s and commits at 45 s. Drain throughput is 80/40 = 2.0 GB/s; multiplying by eight ranks would count state twice. Under the initial 120 s period, B snapshots at 120, releases at 122 and commits at 165 s. Loss at 130 s rejects uncommitted B and restores A: age 130 − 0 = 130 s passes. But loss at 164.9 s still restores A, now 164.9 s old: FAIL. A favorable event does not accept the policy.

Apply 9.4's bounds: period S + 2 + 40 + 3 ≤150 s requires S ≤105 s; 2/S ≤0.020 requires S ≥100 s. Select S =100 s: B snapshots at 100, releases at 102, drains at 142 and commits at 145 s. At loss at 130 s, B is still uncommitted and A remains recoverable outside the failed node; the worst pre-commit age approaches 145 s, within 150 s, while pause fraction is exactly 2.0%. The 45 s completion lag is below the period, so drains do not backlog. This spends all pause allowance for 5 s of age margin; faster capture is required if measurement or interference raises pause. A 105 s period uses all age margin; 106 s fails age, while 99 s fails pause.

Restore every shard from A, validate hashes, optimizer/RNG/loader state and the next step. Restart = 5 + 10 + 20 + 5 =40 s ≤45 s; reload =80/20 =4.0 GB/s. A 25 s reload reaches 45 s; 26 s makes 46 s and fails. Only the selected policy's simulated timing passes. Installed release remains HOLD for durable placement/manifests, kill/requeue timestamps, correctness and sustained interference tests; 13.7's hardware/path HOLD still blocks the cluster. Storage owns durability; scheduling owns placement; workload owns correctness. On corrupt state, quarantine output and restore the known-good generation. Replay with concurrent rebuild/collective traffic and restore baseline. PyTorch DCP distinguishes staging from asynchronous completion; acquire the actual store's atomic publication and failure-domain evidence.

Ring two: scheduler and orchestration acceptance

The scheduler is the subsystem most likely to be hand-waved at acceptance — "SLURM works, ship it" — and the most likely to silently strand the fabric you just spent a chapter validating. On rack-scale, NVLink-domain hardware, a scheduler that places a job's ranks without respecting the topology will scatter a tightly-coupled job across NVLink domains and force collectives that should have stayed on the 130 TB/s in-rack fabric out onto the comparatively narrow scale-out fabric. The benchmark busbw is fine; the delivered bandwidth to a real job is not, because the scheduler fragmented the domain. Topology-aware, gang-scheduled placement is therefore an acceptance criterion rather than a tuning knob deferred to operations.

The 2026 choice forks along a well-worn line. SLURM dominates dedicated training clusters (the 70% trade-press estimate has no defined population denominator) and brings parallel-job allocation, optional time-sliced gang scheduling (qualify it separately), fair-share/QoS priority, preemption, and — critically for Blackwell — block scheduling that allocates whole NVLink domains via a topology.yaml so coherent-memory jobs land aligned. Kubernetes (with KAI Scheduler, Run:ai, Volcano, or the SLURM-on-K8s bridge) brings multi-tenancy, fractional GPUs, and a cloud-native control plane, at the cost of needing an explicit gang-scheduler bolted on because vanilla K8s will happily start half a gang and deadlock. Which control plane to validate against is itself an acceptance decision — and for a multi-tenant neocloud the answer is increasingly both, converged. → the build-out treatment is in Chapter 10.1 and Chapter 10.2; here the bar is correctness, not architecture.

Scheduler acceptance — SLURM vs Kubernetes-native, what each must demonstrate
Scheduler modelGang / topologyMulti-tenancyAcceptance must proveFailure mode if skipped
SLURM (block scheduling)Native parallel-job allocation; topology.yaml domain blocks; time-sliced gang scheduling configured separatelyFair-share, QoS, accounts; coarser tenancyJobs land domain-aligned; Preempt/requeue from a complete committed generation within the agreed lost-work budgetFragmented NVLink domains; stranded scale-up bandwidth
Kubernetes + gang scheduler (KAI/Volcano/Run:ai)Gang via add-on; topology-aware via DRA/labelsStrong: namespaces, quotas, fractional GPUAll-or-nothing gang admission; no partial-gang deadlockHalf-started gangs deadlock; quota leaks across tenants
Converged (SLURM-on-K8s bridge)SLURM semantics over K8s podsK8s tenancy + SLURM job modelBoth paths schedule the same hardware without conflictTwo control planes fight over the same GPUs
Acceptance is about correctness under load, not feature lists. Share figures are 2026 practitioner estimates (HPCwire).

Ring three: the reference (proxy) training run

This is the acceptance test that the whole chapter builds toward, because it is the only one that runs the machine as a machine. A reference training run — a real model of representative scale (a GPT- or Llama-class transformer, often a deliberately-sized proxy rather than a frontier model) — advances through the canary and partial-fleet promotion gates of Chapter 13.10; only after full-fleet promotion does its sustained acceptance window begin across the full cluster, through the production scheduler, checkpointing real optimizer state to production storage. It is not run to convergence; it is run to measure goodput and surface the slow failures that short benchmarks structurally cannot see: thermal drift as the room heats, a single GPU that clocks down after six hours, a memory leak in the loader, a checkpoint that gets slower as the filesystem fills, a straggler that only manifests under sustained collective pressure.

The training output includes a defensible performance number: ML Productivity Goodput, using Chapter 14.1’s ledger over the declared wall-clock window and Chapter 12.2’s metric boundaries. That number — not the busbw, not the storage GB/s, not the green scheduler dashboard — is what the operator signs into the SLA, because it captures the training time/compute interactions across those layers, while correctness, isolation, recovery and safety retain separate pass/fail gates. A 90%-versus-96% pair may be shown as a declared sensitivity, but the cited ClusterMAX/CoreWeave material does not establish an industry average or portable acceptance boundary. The acceptance gate is whether the sustained goodput over the reference window clears the contractual floor, and whether the residual badput decomposes into causes you understand and can attribute.

Goodput / badput accounting and the contractual SLA

Assign acceptance shortfalls with 14.1's useful-output ledger and 12.2's metric boundaries: scheduling owns admission; reliability/storage owns interruption and restore; platform owns model/kernel/fabric performance. Attach workload, placement, precision, window and event attribution to the contractual floor so the remedy follows the failed boundary.

SemiAnalysis's roughly seven-day MTBF describes one 512-H100 cluster, not a per-GPU rate. Meta's July 2024 Llama 3 Table 4 reports 41% BF16 MFU for 405B pre-training on 16,384 H100s at 8K sequence length, not a Hopper population band or GB300 gate. Use the accepted fleet's own interruption distribution and model configuration. Without attribution, a missed SLA becomes a dispute among scheduler, storage and platform teams; operations needs both the result and its cause.

870-928 GB/s
in-domain all-reduce busbw on GB200 NVL72 (~saturates the 900 GB/s/GPU NVLink5 unidirectional rate; NVLS lets ring-convention busbw read slightly above it); scale-out gate derived independently from NICs, rails, topology, protocol efficiency, message size, and collective model
~14 bytes/paramderived
14 bytes/parameter is the declared mixed-precision Adam recipe; count serialized unique state and accept pause, drain, commit and restore separately
Scope & caveats

An arithmetic recipe for the standard BF16-compute / FP32-master / Adam mixed-precision configuration, not a measured population value. VAST Data assumed 14 bytes/param to infer model sizes from checkpoint sizes (its own footnote puts the resulting uncertainty at about ±15%); the survey did not establish the recipe. 8-bit Adam moments give 8 bytes/param; optimizers that drop the second moment change it again. Verify the tensors the chosen framework actually serializes.

~every 30 minderived
MLPerf Storage model for 100k accelerators at full utilization: ~one failure per 30 min; illustrative checkpoint workload, not an observed fleet rate
Scope & caveats

MLCommons MLPerf Storage v2.0 illustrative checkpoint benchmark model for 100,000 accelerators at full utilization, not an observed fleet failure rate. Size a real cluster's checkpoint interval from its own measured interruption distribution.

Illustrative benchmark workload, not a fleet forecast. Common-mode events, software, repair, job membership, censoring, detection policy, and non-constant hazard are outside this calculation.

90% vs 96% scenariomodeled
training-goodput sensitivity scenario: 90% vs 96% (illustrative — state which goodput layer it prices, and replace it with the named fleet's measured goodput)
Sep 2026Guide analysis — stipulated sensitivity scenario; no claim of an industry measurement.register ↗
Scope & caveats

Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.

~7 days / one 512-H100 cluster
reported MTBF for one 512-H100 cluster at a top-tier operator
Scope & caveats

SemiAnalysis reported this for one 512-H100 cluster at a top-tier operator in October 2024. It is not a per-GPU rate, a scaling law, or a portable fleet baseline; do not extrapolate it by accelerator count.

~70% / ~20% / ~10%estimate
AI/HPC scheduler share: Slurm / Kubernetes / in-house — a trade-press rule of thumb with no stated sampled population, not a fleet census
Scope & caveats

Trade-press rule of thumb with no stated sampled population or method; not a fleet census or a scheduler selection criterion.

~41%
Meta Llama 3 405B: BF16 MFU on 16,384 H100s at 8K sequence; Table 4, July 2024
Scope & caveats

Named 405B pre-training configuration; other Table 4 layouts report 38% or 43%. No Hopper fleet-wide MFU band or universal acceptance threshold follows.

The acceptance package: what ring three produces

Cluster-scale validation delivers a signed acceptance package — the contractual baseline and the seed of the day-2 reliability program. It contains: the full-scale NCCL/OSU busbw sweep with the straggler-hunt log and the final node-exclusion list; the storage acceptance results across write, read, and metadata paths with the per-GPU targets met; the scheduler validation including the preempt-checkpoint-resume trace; and the reference training run's sustained goodput curve with its badput decomposition, or the inference run's offered/accepted traffic, errors and TTFT/inter-token/completion latency tails. It also captures the fabric and node baseline — the busbw and per-node performance fingerprints — so that day-2 operations can detect drift against a known-good reference rather than guessing. → the baseline handoff into operations and the operational-readiness gate live in Chapter 13.10; the failure-rate data this seeds feeds Chapter 14.3.

W-02 trace, action and flip. Reconcile 9,950 + 50 = 10,000 offered requests; completion fraction = 9,950 / 10,000 = 99.5%, above 99.0%. Uncertainty-bounded tails are 191 ms, 48 ms and 7.801 s (report 7.80 s), within their separate gates. Normal-service simulation passes. Release stays HOLD for client/server traces and replica-loss/restore tests at the same traffic, with bounded retries, correct output, restored queue/cache state and tenant isolation. Service owns request accounting; platform owns replica/device events; operations witnesses restoration. Incorrect output or failed isolation aborts the canary; restore the qualified serving build. A 199 ms indicated TTFT reaches the 200 ms bound; at 200 ms indicated, the 201 ms upper bound fails even with 99.5% completion. Reduce admitted concurrency or repair the measured prefill bottleneck and repeat the same offered trace; do not hide rejected requests by changing the denominator. Hand the service profile to Chapter 13.10; the method follows MLCommons' separation of workload scenarios, correctness and latency constraints, with these owner-assumed gates kept distinct.

Anti-patterns

The recurring cluster-acceptance failures all share a root cause: validating components instead of the system.

  • Summing green checkmarks. Every node passed burn-in, every link passed point-to-point, the storage passed single-stream fio, the scheduler started a hello-world job — and the cluster is declared accepted without a single full-scale collective or a reference training run. The interactions (straggler-on-collective, storage-rebuild-vs-restart, scheduler-fragmenting-domains) are exactly the failures that only appear at the system level, and they are the ones that bite in production.
  • Peak bandwidth instead of sustained-under-burst. Accepting storage on a hero single-stream number, then watching the checkpoint freeze the whole job for minutes when 128 nodes write concurrently. The write path must be validated under realistic concurrent burst, not best-case streaming.
  • A goodput SLA with no badput decomposition. Committing to a goodput floor without the three-layer attribution model, so when the number misses there is no way to assign the shortfall to scheduling, runtime, or program causes — and the SLA becomes unenforceable.
  • Throwing away the baseline. Treating acceptance as a one-time gate and capturing no instrumented full-scale reference, leaving day-2 operations with no known-good to detect drift against. → Chapter 14.3.

Choose the workload profile whose correctness, durable recovery, isolation and sustained performance all pass through the production stack. A training goodput average can hide a torn checkpoint, and a fast inference tail can hide dropped requests. Holding those failures costs qualification time; releasing them transfers that cost to customer progress and trust.

This chapter is the convergence point of Part 13's earlier validation: power-chain failure demonstration in Chapter 13.6, link-level fabric commissioning and PTP time-sync in Chapter 13.7, and node burn-in / SDC hunting in Chapter 13.8 all feed the cluster-scale rings here; the accepted baseline then hands off through the staged ramp and operational-readiness gate in Chapter 13.10. The fabric topology and oversubscription that ring-one busbw validates is engineered in Chapter 8.5; the storage architecture behind ring two lives in Chapter 9.5 and Chapter 9.6, with checkpoint math in Chapter 9.4; the goodput-vs-availability reliability rethink that justifies a goodput SLA is in Chapter 12.2; quantitative availability modeling in Chapter 12.5; and the operational goodput stack, failure-rate data, and training-reliability tuning that this acceptance baseline seeds are in Chapter 14.1, Chapter 14.3, and Chapter 14.4.
Cite this chapter
Fehn, J. (2026). Cluster-Scale Benchmarking, Reference Training & Storage/Scheduler Validation (Chapter 13.9). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-13-commissioning-and-go-live/13-9-cluster-scale-benchmarking-reference-training-and-storage-scheduler-validat (accessed 2026-09-29).
@misc{aidc-13-9,
  author       = {Fehn, Jacob},
  title        = {Cluster-Scale Benchmarking, Reference Training & Storage/Scheduler Validation (Chapter 13.9)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-13-commissioning-and-go-live/13-9-cluster-scale-benchmarking-reference-training-and-storage-scheduler-validat},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit