Chapter 12.4
In this chapter · 6 sections
SLAs, Goodput Contracts & Availability Commitments
Promise facility availability to a tenant who is paying for goodput and the SLA tracks neither the customer's pain nor the provider's control; match the contracted metric to the workload.
What you'll decide here
- Whether you are committing to availability (the facility is energized and reachable) or to goodput (the customer's job makes effective forward progress) — the two diverge sharply for AI workloads and govern entirely different penalty mechanics.
- The measurement basis and attribution rules for any goodput or job-success commitment: what counts as badput, who owns each badput class (provider vs tenant vs force majeure), and from what baseline the shortfall is computed.
- The shape of the service-credit ladder — linear vs stepped, capped vs uncapped, credits-only vs termination rights — and the maximum monthly exposure you are willing to underwrite against your own failure environment.
- Which acceptance/commissioning gate establishes the contractual goodput baseline, so the SLA is measured against a number both parties signed at go-live rather than against marketing.
- How the customer-facing commitment is reconciled against the redundancy design-basis and the modeled availability — never promise a tier of continuity the physical plant and the failure environment cannot deliver at a profit.
An SLA is a reliability model with money attached. Everything in Part 12 up to this point — the goodput-vs-availability rethink (Chapter 12.2), the redundancy primer (Chapter 0.5), the facility tier standards (Chapter 12.1) — exists in the engineering domain, where being wrong costs an outage. The SLA drags all of it into the commercial domain, where being wrong costs a service credit, a churned tenant, or a take-or-pay dispute. What follows is the translation: how a failure environment becomes a promise, how that promise is measured and attributed, and how the penalty structure is shaped so that it disciplines the provider without bankrupting them on a bad month.
The recurring fork is the one Part 12 has been building toward. Availability asks: was the facility energized, cooled, and reachable? Goodput asks: did the customer's job make effective forward progress? For a web service those two questions have nearly the same answer. For an AI workload they diverge violently — a cluster can be 100% available at the facility meter and delivering 70% goodput because a single GPU's silent data corruption is poisoning every synchronous step, or keep a clean checkpoint through a ninety-second cooling interruption yet still lose those service seconds plus detection, restore and replay time. The SLA that measures the wrong one of these is worse than no SLA: it gives the customer a number that does not track their pain and gives the provider an exposure that does not track their control.
What is actually being promised: availability vs goodput
The legacy data-center SLA promises availability — a Monthly Uptime Percentage measured at the facility or the instance, with a service-credit ladder if it falls short. This is the world of the Uptime tiers — Tier III concurrently maintainable, Tier IV fault-tolerant — where each "nine" of the negotiated SLA is a downtime budget the operator commits not to exceed (Tier itself certifies topology, not an availability percentage; see Chapter 12.1). It is a clean, well-understood, auditable promise, and for a facility or instance service it is a useful boundary metric. An unreachable managed inference endpoint loses requests and revenue, but surviving replicas can mask facility downtime; contract correct, on-time logical completion under the agreed load envelope as well. → serving-engineering SLOs in Chapter 10.11.
The training SLA is a different animal, because the customer is not buying uptime — they are buying effective compute. For this contract, Provider Goodput = Scheduling Goodput × Runtime Goodput: retained forward-progress time divided by contracted wall-clock accelerator time. Provider-attributable badput includes unavailable resources, failed steps that get rolled back, checkpoint/restore, stragglers, rescheduling, and work discarded because it ran on a silently corrupting device; Program Goodput/MFU remains tenant-owned unless the provider accepts a reference-program target. Model 90% versus 96% only as a declared sensitivity; the cited provider-tier and marketed examples do not establish an industry average or portable best-in-class boundary. For an illustrative contracted service, value measured badput reductions against the actual facility-interruption and recovery consequences rather than assumed availability nines.
| Dimension | Availability commitment | Goodput contract |
|---|---|---|
| What is promised | Facility/node/instance energized & reachable | Effective forward progress of the customer's job |
| Right for | Facility/instance hosting and reachable infrastructure | Managed training; separately defined managed-inference request success |
| Unit of measure | Monthly Uptime % (downtime minutes) | Provider Goodput % = retained forward-progress time / contracted wall-clock accelerator time |
| Example figures (not market norms) | 99% node / 95% rack (ClusterMAX 2.0 example, 2025); 99.99% region-level Monthly Uptime (Amazon EC2 SLA) | Illustrative sensitivity: 90% vs 96%; no market baseline established |
| Measurement point | Meter, hypervisor, or endpoint probe — auditable | Inside the job: telemetry, checkpoint logs, step counters |
| Attribution difficulty | Probes still need clock, scope and exclusion rules | High — badput must be classed by owner (provider/tenant/force majeure) |
| Failure that breaches it | Power/cooling loss, network partition, node-down | SDC, straggler throttling, slow recovery, fabric BER, checkpoint stalls |
| Provider's lever to defend it | Redundancy (2N power, N+1 cooling) | Health-checking, hot spares, fast checkpoint/restore, node drain |
The two columns reward different capital. An availability commitment is defended with redundancy — 2N power, N+1 cooling, dual fabric — and its cost lands in the MEP-scope construction budget (2N duplicates the full power path; quantify the topology-specific BOM/TCO premium and operating posture). A goodput contract is defended with operational reliability — passive health-checks every few seconds, automatic node drain on a degradation-versus-golden-reference trigger, a hot-spare pool, and multi-tier checkpointing. None of those settings port between fleets, and the numbers are where the money is: the drain threshold follows from the reference workload, the cost of a false positive and the measured degradation distribution; the spare fraction from failure arrivals, repair and logistics time, topology-compatible placement and the depletion risk you accept; the swap and restore times from the named SKU, fabric and checkpoint tier. Derive each from your own acceptance data, and keep them apart — a performance-degradation trigger bounds throughput loss; it does not establish correctness or catch every SDC. The mistake operators make is buying the first kind of insurance for a workload that needs the second. A training tenant values required facility continuity and also the 40 minutes of work a slow restart can force it to replay after a checkpoint interval of that length; compare the contracted interruption budget with complete restart and actual lost-progress age. → the goodput-vs-availability tradeoff curve in Chapter 12.2, quantified by the model in Chapter 12.5.
Penalty and credit structures: the service-credit ladder
A commitment without a penalty is marketing. The penalty in a standard SLA is the service credit: a percentage of the monthly bill refunded (as credit, almost never cash) when the measured metric falls below a threshold. The ladder is a step function — miss the target band and you owe a credit; miss it badly and you owe more. The reference shape in Amazon EC2’s Region-Level SLA (May 25, 2022) is a three-rung ladder: 10% credit below 99.99% but at least 99.0% monthly uptime, 30% below 99.0% but at least 95.0%, and 100% below 95.0%, on the policy’s affected-fee basis. Those are EC2 uptime terms; a GPU agreement needs its own node, rack or goodput boundary before those rungs price a shortfall. Exhibit T below supplies an assumed training agreement with exact inequalities and an event ledger.
The design decisions inside the ladder are where the money and the disputes live. Stepped vs linear: a stepped ladder is simple to administer but creates cliff incentives — a provider one minute the wrong side of a threshold owes the same as one an hour past it, which can perversely make them stop fighting an outage once a rung is lost. A linear (or finely-stepped) ladder tracks pain more honestly at the cost of administrative complexity. Capped vs uncapped: a cap limits how much of the defined monthly fee is exposed; without it, a bad month can create a liability larger than that fee, so price collateral, insurance availability and the provider’s ability to pay before agreeing to an uncapped or multiplied remedy. Credits vs termination: the customer's real remedy for chronic underperformance is not the credit — which rarely exceeds a month's fee — but a termination-for-repeated-breach right (e.g., three breaches in a rolling quarter), which is the clause that actually disciplines a provider, and the one to negotiate hardest on either side.
| Ladder design | How a shortfall is paid | Who it favors | Failure mode to watch |
|---|---|---|---|
| Stepped (3-rung: ~10% / ~30% / ~100%) | Fixed credit % per uptime band missed | Provider — simple, predictable exposure | Credit cliff: recovery still limits deeper bands, user harm and repeat-breach exposure |
| Finely-stepped / linear | Credit scales with downtime minutes or goodput gap | Customer — tracks actual harm | Administratively heavy; needs trusted measurement |
| Capped at 100% of monthly fee | Total credits never exceed the period's bill | Provider — bounds catastrophic months | Under-compensates a customer whose loss dwarfs the fee |
| Uncapped / multiplied credits | Penalty can exceed the fee | Customer — real teeth | Model collateral, liability, insurance availability and price for the actual agreement |
| Credits-only | Future service credit; cash remedy and exit rights only as agreed | Provider — retains revenue & tenant | Toothless against chronic underperformance |
| Credits + termination-for-repeated-breach | Credit ladder plus exit right after N breaches/quarter | Customer — escape from a bad provider | Provider churn risk; the clause both sides fight over |
Choose the fee basis and overlap rule alongside the ladder. Two obligations measured on the same reservation can breach during the same event; the agreement must say whether credits accumulate or whether the greater applicable remedy settles that overlap. Keep the raw event evidence even when the settlement pays only once.
Scope & caveats
Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.
Scope & caveats
Illustrative provider contract structure reported by ClusterMAX 2.0; not a market baseline or recommendation.
Uptime definition, exclusions, measurement window, spare entitlement, service credits, and penalties remain contract-specific.
Scope & caveats
AWS’s affected-region eligible EC2 fee basis, exclusions, claim procedure and nonstacking rules apply; these are not generic GPU-cloud goodput bands.
Scope & caveats
SemiAnalysis reported this for one 512-H100 cluster at a top-tier operator in October 2024. It is not a per-GPU rate, a scaling law, or a portable fleet baseline; do not extrapolate it by accelerator count.
Scope & caveats
Scenario outputs for representative provider profiles, not sampled provider loss rates. Goodput expense prices extra rental time or reduced output; it includes fault-tolerance performance overhead and spare capacity as modeled. The rounded range and exact triplet describe this same scenario.
Measuring and attributing the shortfall: goodput accounting and badput
An availability metric is observable at a boundary — the meter trips, the endpoint stops answering, and the parties can reconcile probe results, clocks, measurement intervals and exclusions. A goodput metric is the opposite: it is measured inside the customer's job, where the provider and the tenant share the failure surface, and the entire commercial value of the contract turns on attribution — deciding, for every minute of badput, whose fault it was. This is the hardest part of an AI SLA and the part most contracts get dangerously vague on.
The contractual measurement basis is goodput accounting: instrument the job to produce a defensible ledger of total GPU-time, productive GPU-time, and each class of badput. The badput taxonomy is the heart of it, and each class has a natural owner. Hardware badput — a failed GPU, an HBM error, an SDC event auto-drained by the health-checker — is the provider's: it is their silicon and their fleet management. Recovery badput — checkpoint/restore latency, node-swap time, re-scheduling delay — is shared, and the split depends on whether the provider supplied the checkpointing stack or the tenant did. Workload badput — a tenant's inefficient parallelism, a bad hyperparameter that diverges, a job that simply ran slow — is the tenant's, and the provider must be able to fence it out or they are underwriting the customer's ML engineering. The contract must name these classes, name the attribution method (whose telemetry, what arbitration if the logs disagree), and name the baseline.
Deep dive: badput attribution and the silent-data-corruption problem
The clean cases are easy. A node hard-fails, the health-checker drains it, the job restarts from the last checkpoint — that is provider hardware badput, the minutes are logged, and the credit is owed. The pathological case, and the one that makes goodput contracts genuinely hard to write, is silent data corruption: a GPU that produces wrong results without erroring, poisoning gradients across a synchronous run until someone notices the loss curve has gone strange. Two attribution problems collide here. First, detection lag: the corruption may have been retained into checkpoints for hours before discovery, so the badput is not the few minutes to swap the bad device but the entire window of poisoned work that must be rolled back. Second, blame ambiguity: a diverging loss curve looks identical whether the cause is the provider's faulty silicon or the tenant's unstable training recipe, so attribution needs the agreed combination of device diagnostics, reproducible reference checks, workload evidence and a dispute procedure; a golden-reference comparison alone does not settle every cause.
This is why the operational practices and the contract are inseparable. The SLA's goodput floor is only defensible if the provider runs the machinery that produces clean attribution: passive health-checks every few seconds, periodic deep node diagnostics (DCGM-class), SDC-detection via golden-reference comparison, and an auto-drain trigger calibrated on the fleet's own measured degradation distribution and false-positive cost. Without that instrumentation, every badput dispute degenerates into a finger-pointing exercise the provider loses (because they cannot prove it was the tenant) or the tenant loses (because they cannot prove it was the provider). The acceptance-gate fingerprint — the all-reduce busbw the fabric actually achieved for the named SKU, topology, collective, message-size range and software tuple, the per-node nvbandwidth numbers, the fabric BER floor — is what both sides point back to when they disagree, and it is only usable in a dispute if every one of those conditions is recorded alongside the number. → commissioning fingerprint capture in Chapter 13.2; the failure-rate inputs that set the badput baseline in Chapter 14.3.
Exhibit T turns attribution into a payable decision: determine which reserved GPU-hours remain eligible, count overlapping failures once, then select the credit band using the exact fraction. The assumptions and ledger supply every operand.
| Event / half-open hours from window start | Ownership / precedence | Hours | GPU-hours |
|---|---|---|---|
| H: [0,40) | Hardware interruption; provider | 40 | 40 × 1,024 = 40,960 |
| R: [40,46) | Restore; provider-owned stack | 6 | 6,144 |
| U: [46,50) | Missing telemetry; provisional provider loss | 4 | 4,096 |
| C: [50,60) | Checkpoint overhead; provider-owned stack | 10 | 10,240 |
| T: [60,80) | Verified tenant-only delay; excluded | 20 | 20,480 |
| P: [80,720) | Retained reference-job progress | 640 | 655,360 |
| Raw tenant event T-overlap: [10,18) | Already inside H; retain raw evidence, add no second interval | 0 additional | 0 additional |
| Disjoint totals | H + R + U + C + T + P | 720 | 737,280 |
Contracted time is 1,024×720 = 737,280 GPU-hours. The raw tenant event overlaps provider hardware loss on [10,18), so only [60,80) is tenant-only: exclude 20×1,024 = 20,480 GPU-hours. Eligible time is 716,800 GPU-hours; retained progress is 640×1,024 = 655,360 GPU-hours. Eligible G = 655,360/716,800 = 32/35, displayed as 91.4%. Unadjusted resource productivity is 655,360/737,280, or 88.9%; it uses a different denominator.
The floor is max(92%, 0.95×96%) = 92%. Use the unrounded fraction to select the band: the month fails, and the 10% remedy is 0.10×$100,000 = $10,000. Pay max($10,000, $8,000) = $10,000, not their sum; the fee cap is not reached. Issue the future service credit and open the agreed attribution process with event IDs, timestamps, resource IDs, progress counters and device evidence. If logs conflict, use the agreed independent expert; unresolved provider evidence keeps the provisional treatment.
Flip: if evidence proves all four missing-telemetry hours were retained reference work, the numerator becomes 644×1,024 = 659,456 GPU-hours, and 644/700 = 92% exactly. The goodput credit becomes zero; the separate $8,000 availability credit remains payable. If only x of those hours are proven productive, the threshold is (640+x)/700 ≥ 0.92, hence x ≥ 4 hours. An exclusion needs its own evidence and a different denominator calculation. AWS’s published policy illustrates why fee scope and nonstacking must be explicit; this exhibit’s bespoke arithmetic belongs here. Operational accounting remains in Chapter 14.1.
Scope & caveats
1,024 GPUs,720h;20h tenant-only excluded;640h retained;655360/716800=32/35. Floor max(.92,.95*.96)=.92;10% credit on assumed USD100000 affected fee. No provider contract uses these illustrative terms by implication.
Tying the SLA to the commissioning baseline
A goodput contract measured against "the industry says ~90%" is a contract measured against nothing — it invites a dispute the day the first shortfall is claimed. What makes it enforceable is anchoring the SLA to a baseline captured at go-live: the commissioning process produces a quantitative fingerprint of the as-built cluster, and that fingerprint becomes the contractual reference the SLA is measured against. The acceptance gate thereby does double duty — it is also the moment the SLA's denominator is fixed.
Concretely, the go-live fingerprint that the SLA should cite includes the NCCL all-reduce busbw the fabric actually achieved (measured for the named SKU, topology, collective, message-size range, node count, software tuple, and direction against OEM qualification and the pre-production baseline), the per-node intra-node bandwidth, the fabric bit-error-rate floor, the burn-in evidence package (declared stress intensity and exposure, failures by mode and node-hours, pass/re-soak dispositions, and the contracted statistical stopping rule), and the measured goodput on a representative reference job. Writing the SLA against this number — "goodput shall not fall below 95% of the commissioned baseline reference" — converts an unfalsifiable marketing claim into an auditable commitment with an agreed starting point — provided you also state what it resolves to absolutely and set a floor beside it: Exhibit T's relative test resolves to 91.2% (95% of its 96% baseline), which its 92% absolute floor then overrides, and without that floor a weak acceptance baseline legitimizes weak service for the life of the contract. It also protects the provider: a tenant who later runs a pathological workload cannot claim the cluster regressed, because the baseline was established on a known-good reference job both parties signed. → acceptance scripts and pass/fail gates in Chapter 13.2.
For Exhibit T, attach the hardware SKU/count, topology, storage tier, software versions, checkpoint policy, program and dataset hashes to the assumed 96% baseline. Requalify a material change using the same retained-progress definition and acceptance records. Freeze the existing floor during a regression investigation; a slower reference cannot unilaterally lower the promise. Chapter 13.7 supplies fabric evidence and Chapter 13.8 supplies workload acceptance.
Mapping commitments to productization and serving SLOs
The SLA is the contractual face of two things engineered elsewhere. Upstream of the customer, it is the productization of capacity: the service tiers, the pricing, the reserved-vs-on-demand structure, and the onboarding commitments that turn a cluster into a sellable product. A reserved or take-or-pay commitment justifies a stronger SLA (the customer is locked in, so the provider can underwrite more); an on-demand spot tier carries little or no availability promise by design. The SLA tier and the commercial tier must be co-designed or they contradict each other. → customer delivery and productization in Chapter 10.9.
Downstream, for inference, the availability SLA is only the outer envelope; the metric the customer actually experiences is the serving SLO — time-to-first-token, time-per-output-token, p99 latency under load. A 99.99% availability commitment is worthless to an inference tenant if the endpoint is technically "up" but blowing its latency budget during every traffic peak. The contract must therefore reconcile the facility-availability layer with the serving-engineering layer: availability is necessary but not sufficient, and the agreement can contract deadline-qualified request success directly, or separate uptime and latency obligations with an explicit credit-overlap rule. → serving-engineering SLOs and latency budgets in Chapter 10.11.
Negotiating realistic commitments against the failure environment
One rule protects the provider from their own sales team: never promise a tier of reliability the physical plant and the failure environment cannot deliver at a profit. The SLA is the output of the reliability model. You start from the design-basis — the redundancy topology (Chapter 0.5), the facility tier (Chapter 12.1), the measured component AFRs (Chapter 14.3) — run the availability-and-goodput model (Chapter 12.5), and only then write a commitment with margin between the modeled number and the promised number. Promise the modeled number with no margin and a single bad-luck month — well within the variance the Monte-Carlo predicts — turns into a service-credit hit you did not price.
The failure environment is harsher than the marketing instinct assumes. At cluster scale, interruptions are routine, but populations and event definitions matter: Meta's 16,384-H100 Llama 3 run logged 419 unexpected interruptions over 54 days (roughly one every three hours, 78% hardware-attributed), while SemiAnalysis separately reported about seven days of MTBF for one 512-H100 cluster at a top-tier operator in October 2024. Neither observation is a per-GPU scaling law; the contract model must use the named service's measured interruption distribution. Set the goodput floor at the number the named service's interruption model supports with margin and its operational reliability machinery can defend. Offer a premium tier only when extra hot spares and dedicated checkpointing measurably support it, and price the difference. The move for the customer is to demand the measurement and attribution regime — because a high number with weak attribution is worth less than a modest number with airtight badput accounting.
Choose a service metric whose inputs both parties can reconstruct, an absolute floor supported by the accepted workload, and a remedy with a named fee basis and overlap rule. A higher headline target with weak attribution shifts the dispute without recovering work; a lower supported target exposes the service actually being sold. Price monthly breach exposure with Chapter 12.5 before accepting the obligation.
Cite this chapter
Fehn, J. (2026). SLAs, Goodput Contracts & Availability Commitments (Chapter 12.4). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-12-reliability-resilience-and-standards/12-4-slas-goodput-contracts-and-availability-commitments (accessed 2026-09-29).
@misc{aidc-12-4,
author = {Fehn, Jacob},
title = {SLAs, Goodput Contracts & Availability Commitments (Chapter 12.4)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-12-reliability-resilience-and-standards/12-4-slas-goodput-contracts-and-availability-commitments},
note = {Accessed 2026-09-29}
}