The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 10.9

In this chapter · 7 sections
Term help

Customer Onboarding, Delivery & Productization

A cluster becomes a product when customers can buy it, run a job, be metered, and leave; the rung you sell on the bare-metal-to-serverless ladder sets isolation, SLA exposure, billing, and margin.

GOODPUTPOWER-BOUND

What you'll decide here

  1. Choose bare-metal capacity, GPU-as-a-service, managed training or serverless inference; each moves metering, failure credits and margin risk to a different service boundary.
  2. Choose physical-node, VM, MIG or time-shared tenancy against the required isolation and noisy-neighbor controls, then bill the capacity actually reserved.
  3. What you actually commit to in the SLA — facility availability (the easy promise) versus goodput / effective-training-time (the promise the customer actually values) — and the credit schedule that prices a miss.
  4. Your capacity-and-billing model across the on-demand / spot / reserved / take-or-pay spectrum — the split that sets your revenue predictability, your debt capacity, and how much utilization risk you have transferred to the customer.
  5. Your time-to-first-job target and the onboarding automation that hits it — because TTFJ is the conversion metric of the whole business, and the offboarding/portability story is what a sophisticated buyer checks before they sign.

Every chapter in Part 10 builds the machine; this is the chapter that turns the machine into a product someone can buy. The orchestration plane, the multi-tenancy isolation, the provisioning automation, the health telemetry, the fault-tolerance loop, the training frameworks — all of it is internal plumbing until it is wrapped in a consumption model, an API, an SLA, a billing meter, and an onboarding path. Productization is the act of drawing a clean line between what the operator owns and what the customer owns, and then pricing the responsibility on each side of that line. Draw the line in the wrong place and you either give away margin you should have kept or carry risk you should have transferred.

We lay out the value-stack ladder from bare-metal to serverless and show how each rung moves the boundary; we work the multi-tenancy and isolation fork as a joint security/noisy-neighbor/billing decision; we build the control plane, API, and provisioning surface that makes the product self-serve; we separate facility availability, usable capacity and managed useful output, then identify which SLA the operator controls and the customer buys; we map the billing and capacity spectrum from spot to take-or-pay; and we close on onboarding, time-to-first-job, and offboarding/portability — the metrics that decide conversion and the exit story that decides trust. The rung changes the work you operate; the actual contract, isolation boundary and measured cost determine margin and liability.

Assumed: $2/active GPU-h; $1/standby GPU-h; $12 full cost; $2 creditmodeled
C7 exact assumed billing rates, hourly full cost and failure credit
Teaching assumptions, not a provider quote; retain the complete provisioned-capacity boundary in the calculation.
Scope & caveats

All case inputs and traces are assumed, with their rationale/bounds stated in the case. No provider performance, actual incident, procurement approval or achieved security level is asserted.

The value-stack ladder: four altitudes for the same GPUs

The same physical GPUs can be sold at four very different altitudes, and picking the altitude is the commercial decision everything else in this chapter follows from. At the bottom is bare-metal: you hand the customer a dedicated, single-tenant node (or a whole cluster) with raw access to the hardware, an OS image, and a network fabric — and almost nothing else. You keep a thin margin on raw capacity, you carry no software risk, and the customer owns goodput, drivers, schedulers, and uptime above the metal. One rung up is GPUaaS (the virtualized or containerized IaaS tier): VMs or Kubernetes with the GPU passed through, multi-tenant fabric, managed networking and storage, a self-serve API. Higher still is managed: the operator runs the orchestration plane for the customer — managed Slurm, managed Kubernetes, validated training stacks, active health-checks, automatic node draining — and prices the operational labor. At the top is serverless: the customer never sees a node at all; they send a request or a function, the platform scales GPUs up from zero and back down, and the meter runs in milliseconds.

Each step up the ladder is a systematic trade of margin for responsibility. Each rung adds operator-owned software and operator-carried risk, and each rung adds price you can charge for it. A bare-metal hour and a serverless second of the same H100 can differ by an order of magnitude in effective $/GPU-hr, and the difference is not arbitrage — it is the operator absorbing scheduling, idle-time, cold-start, and goodput risk that the bare-metal customer would otherwise carry themselves. The question is which risks you want to own and charge for, and which you want to push to the customer at a lower price. → multi-tenancy mechanics in Chapter 10.3; provisioning automation in Chapter 10.5.

The value-stack ladder — what each rung owns, sells, and risks
RungOperator ownsCustomer ownsBilling unitGoodput ownerMargin posture
Bare-metalPower, cooling, fabric, OS imageDrivers, scheduler, uptime, jobsGPU-hour (reserved/dedicated)CustomerThinnest; pure capacity
GPUaaS (IaaS)Hypervisor/K8s, network, storage, APIWorkloads, orchestration logicGPU-hour / GPU-secondSharedThin-to-moderate; volume tier
ManagedOrchestration, health-checks, validated stacksModel, data, hyperparametersGPU-hour + managed-service feeOperatorModerate; priced operational labor
ServerlessWhole stack, autoscale, cold-startThe request / function onlyPer-second / per-token / per-requestOperatorHighest per-unit; absorbs idle+cold-start
Synthesis of SemiAnalysis ClusterMAX 2.0, Crusoe managed-Slurm, Rafay GPU-PaaS, 2025-2026. 'Goodput owner' = who is contractually on the hook for effective useful compute, not just powered hardware.

Multi-tenancy and isolation: one decision, three boundaries

The moment you sell anything above dedicated bare-metal, you have to decide how tenants share hardware — and the isolation model you pick is simultaneously a security boundary, a noisy-neighbor boundary, and a billing-granularity boundary. These three pull in different directions, which is why the decision is harder than it looks. The taxonomy runs from hard to soft. Physical / bare-metal isolation gives each tenant a whole node or cluster — the strongest boundary, the coarsest billing unit (you cannot sell a fraction of a node), and no on-node noisy-neighbor risk — shared storage, the fabric, and the control plane still contend, so a dedicated node is a compute-exclusivity claim, not a contention-free one. VM-level isolation with GPU passthrough multiplexes tenants per host behind a hypervisor; the boundary is strong but the east-west fabric and shared storage become the contended surface. Fractional isolation — NVIDIA MIG (hardware-partitioned) or time-slicing/MPS (software-shared) — lets you sell a slice of a single GPU and expose fine billing granularity when workloads cannot saturate a device. Continuous request batching and multi-adapter serving on full GPUs provide another utilization path, while frontier models can consume whole multi-GPU domains; choose fractional isolation when the finer billing unit outweighs its weaker security and noisy-neighbor boundary.

Fractional sharing is a billing enabler and a security liability at the same time. MIG partitions are hardware-enforced and the strongest of the sub-GPU options, but time-slicing and MPS are not robust confidentiality boundaries — documented covert and side channels bypass MPS/MIG isolation, and real vGPU CVEs (e.g. the 2025 NVIDIA vGPU advisories) have shown that partitioning is not, by itself, a security boundary you can sell to an adversarial multi-tenant workload. So if your tenants are mutually untrusting (a public neocloud), you owe them either bare-metal/VM isolation or hardware-attested confidential computing, and you give up the finest billing granularity. If your tenants are one organization (an internal platform), fractional sharing is the right efficiency lever and the security objection mostly evaporates. → the full isolation engineering and its failure modes are Chapter 10.3; model/weight-in-use protection is Chapter 11.8.

Isolation models scored on the three boundaries they straddle
Isolation modelSecurity boundaryNoisy-neighborBilling granularityAdversarial-safe?
Bare-metal / dedicated nodeStrongest (physical)NoneWhole node/clusterYes
VM passthroughStrong (hypervisor)Fabric/storage contentionPer-GPU / per-VMYes, with fabric isolation
MIG (hardware partition)Moderate (HW-enforced)Low (partitioned)GPU fraction (fixed slices)Caution — side channels demonstrated
Time-slicing / MPSWeakest (software)High (shared SMs)Finest (arbitrary share)No — not a confidentiality boundary
Confidential computing (TEE)Strong (attested, in-use)Per underlying modelPer attested instanceYes — designed for it
Synthesis of NVIDIA MIG/Confidential Computing docs, Aarna multi-tenancy taxonomies, and NVIDIA vGPU security advisories, 2025-2026. 'Adversarial-safe' = defensible boundary for mutually-untrusting public tenants.

The control plane, API and provisioning surface

Productization is, concretely, the act of putting an API in front of the cluster so a customer can self-serve the lifecycle — request capacity, provision an environment, run a job, observe it, tear it down — without a human in the loop. The control plane that backs that API is the difference between a colo full of GPUs and a cloud. Below the API sits the bring-up automation from Chapter 10.5 (Redfish/IPMI, PXE, image pipelines, infrastructure-as-code), the scheduler from Chapter 10.1 (Slurm or Kubernetes, with fair-share, quota, and preemption), and the telemetry from Chapter 10.6 feeding the meter and the health dashboard. The product decision is how much of that surface you expose: a bare-metal vendor exposes node lifecycle; a managed vendor exposes a job-and-cluster abstraction and hides the nodes; a serverless vendor exposes only a function or an inference endpoint.

The consequence of getting the control plane wrong is measured in operational labor that should have been software. Every manual step between 'customer clicks provision' and 'job is running' is a margin leak and a TTFJ penalty; the maturity benchmark in the market (SemiAnalysis ClusterMAX) scores providers in part on exactly this — orchestration, lifecycle automation, and observability are explicit rated dimensions. A neocloud that bring-ups nodes by hand cannot hit a serverless cold-start budget or a managed-Slurm SLA, and it cannot scale its sales without scaling its NOC headcount linearly. The control plane is where the productization either compounds (software that serves the next 1,000 customers at near-zero marginal cost) or fails to (a services business whose headcount grows with revenue).

Deep dive: the API surface a credible GPU product must expose

A productized control plane is not one API but a layered set, and a buyer evaluates each layer. Identity and tenancy: per-tenant projects/namespaces, RBAC, quotas, and budget caps — the substrate of multi-tenancy and the thing that stops one tenant from spending another's capacity. Capacity: request, reserve, and release GPUs against on-demand, spot, and reserved pools, with the reservation model exposed as a first-class object (Google Cloud, instructively, makes capacity reservations distinct API objects from committed-use discounts — separating 'is the hardware held for me' from 'have I committed to pay'). Lifecycle: provision an environment (image, drivers, NCCL, scheduler), run/checkpoint/resume a job, drain/replace a faulty node — ideally declaratively via Terraform-style IaC so the customer can version their cluster. Observability and metering: per-job and per-tenant utilization, goodput/health signals, and a usage feed that the billing engine consumes. Egress and data: object and high-throughput storage, dataset staging, and — critically for the offboarding story — bulk export.

The design tension is abstraction versus control. Training customers running 3D-parallel jobs want low-level control: topology-aware placement, specific NCCL/fabric tuning, bare-metal performance with no hypervisor tax. Inference and app-layer customers want the opposite: a high abstraction that hides nodes entirely. A single product cannot be maximally low-level and maximally abstracted at once, which is the deeper reason the value-stack ladder exists as separate rungs rather than a single dial. Most credible operators ship two or three rungs (e.g. bare-metal/Slurm for training tenants, serverless endpoints for inference tenants) on one underlying fleet, and let the customer pick the abstraction that matches the workload. → scheduling plane in Chapter 10.1; bring-up IaC in Chapter 10.5.

Service interfaces: name the contract and test its boundary
InterfaceNamed referenceFilled acceptance action for C7
Login / application identityOpenID Connect Core 1.0, including published errataC7 login maps to C7 only; wrong issuer/audience is rejected. Authorization policy remains separate.
Enterprise user lifecycleSCIM 2.0 protocol, RFC 7644, when this integration is offeredDisable the test user; verify new login and resource actions are denied under the declared session/token policy.
Provision / reserve / cancelVersioned provider API with idempotency and allocation lifecycleRepeat C7’s create key; return the original allocation. Cancel after a partial failure and reconcile the same allocation ID.
Usage / invoiceCloudEvents 1.0.2 identity plus the C7 billing-event schemaReplay (active,a1) and late credit (credit,f1); the final balance stays $12.
Artifact exportDeclared checkpoint/model format, manifest and content digestsC7 verifies exported digests and an accepted restore before exit closes.
Primary references: OIDC Core, SCIM protocol, and CloudEvents. Security protocol mechanics remain in Part 11; implement only interfaces the product promises.

SLAs and reliability commitments: facility uptime vs goodput

The easy SLA to write is facility availability — a monthly powered-and-reachable commitment — because the retired Tier availability table Uptime disavows remains familiar in the market (Chapter 12.1). The SLA the customer actually values is goodput — the fraction of contracted GPU-time that produced useful work. For a training tenant, a node that is 'available' but flapping, throttling, or failing NCCL all-reduce is worthless; what they bought was effective training time, measured against the named job account in Chapter 14.1. The guide’s stipulated 90%-versus-96% pair tests sensitivity; site measurements set the contract baseline. The gap between those two SLA shapes is the gap between selling capacity and selling outcomes.

The trade is direct: the SLA shape you choose dictates which failures you pay for and how much you must invest to avoid them. Commit only to facility availability and your obligation is met by redundant power and cooling; the customer eats the goodput loss from a bad GPU or a fabric brownout, and you compete on price. Commit to goodput and you must build the whole fault-tolerance loop — active health-checks, fast node draining, automatic restart-from-checkpoint, burn-in and acceptance testing — because every percentage point of badput is now a credit you owe. The April 2026 ClusterMAX model’s ~6–21% goodput expense includes assumed recovery and fault-tolerance overhead; the engineering that shrinks it is optional spend for a capacity vendor and mandatory for a goodput vendor. The credit schedule is where this gets priced: tiered service credits (e.g. 10% of the affected charges below 99.9%, escalating with the miss) must be calibrated against your actual goodput distribution, or a single bad month of badput erases a quarter of margin. → goodput-vs-availability engineering in Chapter 12.2; the contract mechanics in Chapter 12.4; the fault-tolerance loop in Chapter 10.7.

Billing, metering and capacity reservation

The billing model is where the consumption decision becomes cash, and it spans a spectrum from pure-spot to take-or-pay that mirrors the value-stack ladder in a second dimension. On-demand bills per second or per hour at the highest unit rate, with no commitment — the customer pays for optionality and you carry the utilization risk of an unfilled fleet. Spot / preemptible sells interruptible capacity at a steep discount (60–91% off on-demand at Google's documented band) to drain idle inventory; the customer carries the interruption risk, which checkpoint-tolerant training and batch absorb directly and online inference can absorb only as a supplemental tier backed by standard or capacity-reserved headroom, rapid ejection, and admission/fallback policy. A term/price commitment commits the customer to months or years of spend for a discount; it does not by itself guarantee capacity. A capacity reservation makes the separate capacity promise: holding hardware for a tenant ('is it mine?') is a different promise from committing to pay for it ('have I agreed the spend?'), and mature platforms expose them as distinct objects. At the far end, take-or-pay obligates the customer to pay for a contracted block whether or not they use it — the structure that converts merchant utilization risk into contracted revenue, and the one that underwrites the debt capacity behind the build.

The on-demand/spot/committed-price/take-or-pay split is a risk-transfer dial, and where you set it determines your debt capacity and your survival of a downturn. A fleet sold entirely on-demand is maximally flexible for customers and maximally exposed for you: fewer sold GPU-hours or a lower realized rental price reduce contribution while fixed operating cash and scheduled debt service remain due. Compare the capacity commitment using Chapter 1.8’s matched cash-flow case, then stress merchant receipts against the debt schedule in Chapter 2.5’s DSCR case. A fleet anchored by take-or-pay has transferred that utilization risk to tenants, which is exactly why lenders price GPU-backed debt against contracted backlog rather than merchant hope. Two 2026 developments extend the dial past the tenant: CME and Silicon Data list cash-settled H100 and B200 rental-index futures (target 2026-10-05, pending regulatory review) — the first exchange-traded hedge on the $/GPU-hr this chapter prices in — and SemiAnalysis describes an NVIDIA neocloud rental backstop (a multi-year minimum-revenue guarantee on GPU capacity, upside shared above the floor), which puts the chip vendor itself into the offtake column that lenders underwrite. A backlog that includes the vendor of the collateral is leverage and concentration at once (→ the circular-financing analysis in Chapter 2.5). Take-or-pay carries its own exposure, concentration: a backlog that is take-or-pay but lives in two or three anchor tenants converts utilization risk into counterparty risk, and one non-renewal can strand a campus. The metering substrate underneath all of this — per-second granularity, per-token billing for inference, transparent egress, no surprise fees — is itself a competitive axis; ClusterMAX scores pricing transparency as a rated dimension precisely because hidden egress and rounding games erode trust. → the firm-level unit economics of this revenue (metered $/token and $/GPU-hr → margin) are worked in Chapter 1.8.

The capacity-and-billing spectrum — who carries which risk
ModelCommitmentIndicative discount vs on-demandUtilization riskBest fit
On-demandNone0% (baseline)OperatorBursty, unpredictable demand
Spot / preemptibleNone (interruptible)~60–91% offCustomer (interruption)Fault-tolerant training or batch; supplemental fault-tolerant inference
Term/price commitmentMonths-years~30-60% offSharedSteady, forecastable workloads
Take-or-payPay-regardless blockLargest, plus capacity guaranteeCustomer (transferred)Anchor tenants; debt-underwriting
Synthesis of Google Cloud GPU pricing/reservations docs, Compute Exchange reserved-vs-on-demand, and SemiAnalysis market-structure analysis, 2025-2026. Discounts are indicative ranges, highly volatile.

Onboarding, time-to-first-job and the productization lifecycle

Time-to-first-job (TTFJ) is the conversion metric of the entire business: the wall-clock from 'customer signs up' to 'customer's first useful GPU job is running'. On a bare-metal rung, TTFJ is dominated by node provisioning, image deployment, driver/NCCL setup, and acceptance burn-in; a manual operator measures this in days, an automated one in hours, and the gap is pure product quality. On a serverless rung, keep the nested endpoints separate. End-to-end TTFJ still starts at signup/entitlement/configuration and ends when the first useful job is running. Image or model cold start starts when an eligible request needs an unready worker and ends when that worker is ready. Request TTFT starts at the inference request and ends at the first token, with warm/cold state declared. Published seconds-to-token or cold-start figures belong to those narrower populations; they do not replace onboarding TTFJ.

Onboarding automation is the marginal-cost structure of the product. Each manual onboarding step caps how many customers one engineer can land, which caps growth at the speed of hiring. The fix is the same automation that backs the control plane — IaC-driven provisioning, self-serve quota/RBAC, pre-validated stacks, and acceptance tests that run without a human — so that the 1,000th customer onboards as cheaply as the first. The productization lifecycle that follows TTFJ is the long tail: usage growth, expansion to reserved/committed capacity, support tiers, and the metering that bills it all. But the lifecycle never starts if TTFJ is bad, which is why it is the first thing a sophisticated buyer tests with a trial workload before committing budget.

And then there is the decision most operators underweight and the best buyers check first: offboarding and portability. A customer evaluating a multi-year, take-or-pay GPU commitment is implicitly asking 'how hard is it to leave?' — because the answer prices their lock-in risk into the deal. The portability story has three parts: data egress (can I bulk-export my datasets and checkpoints without a punitive egress bill or a multi-week throughput bottleneck?), stack portability (is my orchestration standard — vanilla Slurm/Kubernetes, OCI containers — or a proprietary control plane that traps my workflows?), and commitment exit (does my reserved/take-or-pay block have real termination economics or a secondary market, or am I locked to a depreciating asset for the full term?). The conclusion is counterintuitive: a credible, low-friction exit story is a sales asset, not a risk. The operator who builds on open standards and clean egress wins the cautious enterprise buyer precisely because the lack of lock-in lowers the buyer's perceived risk of signing — and a secondary market for reserved blocks (an emerging structure for reserved-capacity trading) can offer a transfer route only where the contract permits assignment and an eligible buyer exists. Verify the exported checkpoint and manifest by restoring them under the agreed format, then close access, retention and deletion under Chapters 11.9 and 10.10. A GPU resale market does not establish customer portability.

90% vs 96% scenariomodeled
training-goodput sensitivity scenario: 90% vs 96% (illustrative — replace with the named fleet's measured goodput)
Sep 2026Guide analysis — stipulated sensitivity scenario; no claim of an industry measurement.register ↗
Scope & caveats

Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.

10 dimensions / 5 tiers
ClusterMAX 2.0 GPU-cloud rating: Security, Lifecycle, Orchestration, Storage, Networking, Reliability, Monitoring, Pricing, Partnerships, Availability — Platinum to UnderPerform
Scope & caveats

ClusterMAX 2.1 (2026-04-20) is the current published rating (adds providers + a priced 'Grand Unifying Theory of Goodput'); it names TorchFT (Meta, >10% per-iteration overhead routing collectives through CPU), AWS SageMaker HyperPod Checkpointless (Dec 2025, ~1m45s recovery vs ~15 min restart), and Clockwork TorchPass (licensed live migration) as the 2026 fault-tolerance fork. ClusterMAX 3.0 is not yet published.

~60–91% off
Google Cloud stated Spot discount band for most machine types and GPUs
Scope & caveats

Google Cloud stated Spot discount band for most machine types and GPUs; not a cross-cloud guarantee or capacity reservation.

The unit-economics tie-back

Everything in this chapter resolves into a single revenue line that Chapter 1.8 turns into a return. The value-stack rung helps define your unit; the isolation model constrains its sharing boundary; the SLA shape sets your reliability spend and your credit exposure; the capacity model sets how much of your revenue is contracted versus merchant; and TTFJ affects conversion. Together they determine metered revenue per token and per GPU-hour, which — net of the cost stack from Part 5 through Part 9 and the depreciation debate from Chapter 1.8 — contributes to the margin the service earns; keep paid standby, support, retries and credits in the same period as revenue. The clean way to read this chapter is as the revenue-side complement to 1.8's cost-side analysis: 1.8 asks whether the asset earns its cost of capital; this chapter decides how the revenue that feeds that question actually gets priced, packaged, metered, and collected. The recurring warning from 1.8 applies in full here: underwrite the metered revenue with explicit price and demand sensitivities over the contract term; a historical inference-price series does not prescribe this service’s annual decline. → inference serving efficiency, the governor of $/token, is engineered in Chapter 10.11; data-handling obligations that ride alongside the commercial contract are Chapter 10.10.

The workload archetypes that determine which rung of the ladder a customer needs are framed in Chapter 1.1, and the firm-level economics this chapter feeds are Chapter 1.8. The internal machinery being productized lives across Part 10: the scheduling plane in Chapter 10.1, multi-tenancy and isolation in Chapter 10.3, provisioning and bring-up in Chapter 10.5, observability and health in Chapter 10.6, fault tolerance in Chapter 10.7, and inference serving in Chapter 10.11. The SLA shape introduced here is engineered as goodput-vs-availability in Chapter 12.2 and contracted in Chapter 12.4. Weight/model-in-use protection behind the isolation boundary is Chapter 11.8; the customer-data governance riding alongside the commercial contract is Chapter 10.10; data residency and sovereignty constraints on where the product can be sold are Chapter 3.12.
Cite this chapter
Fehn, J. (2026). Customer Onboarding, Delivery & Productization (Chapter 10.9). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-10-software-orchestration-and-service-delivery/10-9-customer-onboarding-delivery-and-productization (accessed 2026-09-29).
@misc{aidc-10-9,
  author       = {Fehn, Jacob},
  title        = {Customer Onboarding, Delivery & Productization (Chapter 10.9)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-10-software-orchestration-and-service-delivery/10-9-customer-onboarding-delivery-and-productization},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit