The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 10.3

In this chapter · 5 sections
Term help

Multi-Tenancy, Isolation & Resource Sharing

Sharing a GPU means choosing performance isolation (whole-GPU, MIG, MPS, time-slicing, fractional) and security isolation (process, container, VM, confidential VM) separately; conflating them sells a partition as a boundary.

GOODPUTDENSITY-RAMP

What you'll decide here

  1. Where on the sharing spectrum each tenant class sits — whole-GPU for tightly-coupled training, hardware-partitioned MIG for SLO-bound serving, MPS or time-slicing for dev/notebook fleets — and therefore how much of your fleet's idle silicon you can actually reclaim.
  2. Choose dedicated GPUs, MIG or MPS/time-slicing against the tenant threat model; compare their distinct memory, fault and shared-host boundaries.
  3. The quota and fairness model — hard partitions vs hierarchical fair-share with preemption vs deadline/priority QoS — and which one your scheduler (Slurm Fair Tree, Kubernetes + KAI/Run:ai) can actually enforce.
  4. Choose a shared kernel, isolated VM or supported confidential VM/GPU path against the adversary, then test management-plane and device-assignment exposure.
  5. How you bound the noisy neighbor — at the hardware (MIG), the scheduler (limits, priorities, gang admission), or the fabric (per-tenant QoS) — before a single greedy job silently taxes everyone else's goodput.

Multi-tenancy is the lever that turns a depreciating pile of accelerators into a utility. A frontier pre-training run wants the whole machine to itself; smaller fine-tuning, batch, interactive-serving, notebook and CI workloads can under-fill a modern GPU. B200 SXM carries 180 GB of HBM and a GB200 Blackwell GPU 186 GB usable. Blackwell Ultra (B300 / GB300 NVL72) is a newer platform profile and raises HBM per GPU again; Hopper profiles instead span 80–141 GB. Recompute every slice size in this chapter against the part you actually bought. A 7B-parameter serving replica or a data-science notebook can fit within part of those resources; concurrency, context length and its latency target decide whether there is useful room left. Left whole, an under-filled GPU still consumes capital while the guide's contested 2–3-year bear-case economic clock runs (Chapter 1.8). This chapter is about reclaiming that stranded silicon without letting one tenant's work corrupt, starve, or spy on another's.

A specific error recurs: operators pick a sharing mechanism for a utilization reason, then quietly assume it also delivers isolation, and the two are different axes that happen to share the word "partition." MIG partitions documented GPU-local compute, memory and memory paths; shared host privilege, management and residual side channels still require a tenant threat model in Chapter 11.6. A container gives you a namespace but shares a kernel and a GPU driver with everyone else on the node. This chapter separates the two axes cleanly: the sharing spectrum (how the silicon is divided), quota and fairness (who gets how much, and when), and isolation (what one tenant can do to another). Get the mapping wrong and you either leave money on the floor or you put two adversarial tenants behind a wall that was only ever a speed bump.

The sharing spectrum: five ways to divide a GPU

There is no single "GPU sharing" feature. There are five mechanisms that sit at different points on a curve trading isolation strength against packing flexibility, and they compose — you can run MIG instances, each with MPS inside, scheduled by a fractional-GPU plugin. Read them as a spectrum from hardest partition to softest.

Whole-GPU is the degenerate case and the correct default for tightly-coupled training: one job owns the device, the NVLink domain, and the memory. No sharing overhead, no noisy neighbor, full bandwidth for collectives. The cost is utilization — anything that does not saturate the GPU wastes it.

MIG (Multi-Instance GPU) is the only mechanism that partitions the silicon in hardware: it carves the GPU into up to seven instances, each with a dedicated slice of SM compute, L2 cache, memory controllers, and a fenced region of HBM. GB200 supplies 186 GB usable HBM per GPU; that capacity alone does not specify a supported MIG layout. NVIDIA’s B200 profile table lists 180 GB total, two 3g.90gb instances, four 1g.45gb instances or seven 1g.23gb instances, with different compute fractions. Use the delivered SKU/driver’s profile table rather than dividing GB200 memory arithmetically. MIG dedicates documented GPU-local compute and memory resources, but the instances still share host, PCIe, fabric, storage and control-plane resources. It can reduce on-GPU interference; an end-to-end SLO still requires measured admission and QoS across every shared boundary. The cost is rigidity: you must drain and reconfigure the GPU to change the geometry, slices come in fixed sizes, and a B200 job that needs 30 GB cannot use a 1g.23gb slice and must take a supported 45 GB or larger profile and waste the rest.

MPS (Multi-Process Service) lets each client keep its own CUDA context and GPU address space while submitting work directly to shared GPU scheduling resources concurrently — true spatial sharing of the SMs, not just interleaving. It is the right tool when several small, cooperative, same-trust-domain workloads (e.g. a fleet of tiny inference replicas) can fill a GPU together. The cost is the absence of a hard wall: MPS provides limited resource controls and error containment, but a fatal fault can disrupt every co-client sharing the affected GPU set; clients confined to other GPUs remain unaffected, and with R610, Ampere-and-newer static SM partitions partially contain SM-triggered faults but not every GPU or system-level failure. It buys throughput by assuming the co-tenants trust each other.

Time-slicing is the GPU equivalent of a context switch: the scheduler round-robins the whole device between processes, each getting the full GPU for a slice of time. It needs no special hardware and works on any GPU, which is why it is the lowest-common-denominator sharing mode for dev, notebooks, and bursty low-priority work. The cost is latency jitter and zero memory isolation — every tenant sees the full device, can over-allocate HBM, and pays context-switch overhead; it is unusable for anything with a tight tail-latency SLO.

Fractional GPU is the scheduler-level abstraction (Run:ai, KAI, and similar) that lets you request "0.5 of a GPU" or "10 GB of a GPU" and have the platform place the workload via MPS or memory limits under the hood. It is the most flexible for bin-packing a mixed fleet and the friendliest developer experience, but its isolation is only as strong as the primitive it lands on — usually a soft MPS-class limit, occasionally MIG. Treat the fraction as a billing and packing construct, not a security boundary.

The GPU sharing spectrum — mechanism vs guarantee
MechanismLayerMemory isolationCompute isolationReconfig costBest fit
Whole-GPUDeviceTotal (sole owner)TotalNone — it is the whole deviceTightly-coupled training; max-throughput serving
MIGHardware partitionHard — fenced HBM per instanceHard — dedicated SMs + L2High — drain + reconfigure geometrySLO-bound multi-tenant inference; predictable QoS
MPSProcess / CUDA contextIsolated per-client address spaces; capacity soft unless the device memory limit is setConcurrent SMs; active-thread caps, or static SM partitioning on Ampere and newerLow — per-processSmall same-trust co-tenants filling one GPU
Time-slicingScheduler (temporal)None — full device per sliceNone — round-robin whole GPUNoneDev, notebooks, bursty low-priority batch
Fractional GPUCluster schedulerDepends on backing primitiveDepends on backing primitiveLow — logical requestBin-packing a mixed dev/serving fleet
Hardware partition = MIG only. MPS/time-slicing/fractional give utilization, not a trust boundary. Up to 7 MIG instances per GPU. The quoted MIG profiles are B200 profiles, separate from GB200’s 186 GB usable capacity. Read the supported profile table for the delivered B300/GB300 or GB200 SKU and driver. See keynumbers for sources.

Quota and fairness: who gets how much, and when

Dividing a single GPU is the easy half. The hard half is governing a 10,000-GPU fleet so that twenty teams, each convinced their work is most urgent, share it without a tragedy of the commons. Three models dominate, and the right one depends on whether your tenants distrust each other and whether your workloads are elastic.

Hard partitions (static quota). Each tenant gets a fixed slice of the cluster — N GPUs — and the guarantee has to be backed by dedicated nodes or a standing reservation. A Kubernetes ResourceQuota caps what a namespace may consume and Slurm partition membership routes jobs at a node set; neither, on its own, reserves matching GPUs, the topology they need, or free running capacity. Sell an admission cap as a reservation and you have contracted capacity nothing enforces. Dead simple, perfectly predictable, and the right model when tenants are external customers paying for guaranteed capacity (the neocloud and colo posture; Chapter 10.9). The cost is stranded capacity. When Tenant A is idle and Tenant B is queued, the idle GPUs sit dark because the partition forbids lending. Fleet utilization caps out well below what the silicon could deliver.

Hierarchical fair-share with preemption. The HPC-and-internal-platform model: tenants get a target share, not a hard cap, and the scheduler lets anyone borrow idle capacity, then preempts the borrower when the rightful owner returns. Slurm implements the share accounting as multifactor priority and Fair Tree — a tenant that has under-consumed its share recently floats to the top of the queue, an over-consumer sinks — but that only orders the queue. Evicting a borrower is a separate, explicitly configured mechanism (PreemptType/PreemptMode plus the QoS preemption rules), and it does nothing until you configure it. Kubernetes-native schedulers (KAI, Run:ai, Volcano) implement the same idea as hierarchical queues with reclaim. This is the model that actually drives high utilization, because idle silicon is always lendable. The cost is complexity and the preemption tax: a borrowed, preempted job must checkpoint and resume, so it is only sane for interruption-tolerant work (Chapter 9.4).

Priority / deadline QoS. Layered on top of either: jobs carry a priority or a deadline, and the scheduler admits, preempts, and orders accordingly — production-serving jobs outrank experiments, a deadline'd batch sweep gets escalated as its window closes. This is where you encode that an SLO-bound inference tenant must never be preempted by a speculative training run. The cost is that priority is only meaningful if it is scarce and policed; the moment every team sets priority=high, you are back to FIFO.

One non-negotiable couples all three to the prior chapter: GPU jobs are usually gang-scheduled — a 64-GPU job needs all 64 at once or none, or it deadlocks holding resources it cannot use. Quota and fairness must therefore admit and preempt at the granularity of the whole gang, topology-aware, or the fairness model fights the placement model (Chapter 10.2).

Quota / fairness models — utilization vs predictability
ModelUtilizationPredictabilityPreemptionEnforced byBest fit
Hard partition (static quota)Low — idle capacity strandedHighest — if dedicated or reserved capacity backs itNoneDedicated nodes or a Slurm reservation; namespace quota caps admissionExternal paying tenants; guaranteed capacity
Hierarchical fair-shareHigh — idle GPUs lendableSoft — target share, not a capRequired — configured preemption, not share priority aloneSlurm Fair Tree; KAI / Run:ai queuesInternal multi-team platforms; elastic work
Priority / deadline QoSTunableConditional on priority disciplinePriority-drivenJob priority + scheduler policyMixed prod-serving + research on one fleet
Choose by tenant trust and workload elasticity. Fair-share drives utilization; hard partitions strand it. Enforcement examples are 2026-current.
B200: up to 7 MIG instances
NVIDIA B200 MIG profiles; GB200 geometry requires its own SKU/driver table
Scope & caveats

B200 profiles include two 3g.90gb, four 1g.45gb or seven 1g.23gb instances; memory and compute fractions differ. This is not a GB200 profile specification.

B200: 180 GB; GB200: 186 GB usable
Usable HBM per GPU: B200 and GB200 are separate product records
Scope & caveats

GB200: 372 GB usable memory per two-GPU superchip / 2 = 186 GB per GPU. NVIDIA’s B200 MIG table identifies its full profile as 7g.180gb. These totals do not establish identical partition geometries.

CVSS 9.0
NVIDIAScape (CVE-2025-23266) Container Toolkit escape — container-to-host on shared GPU nodes
CVSS 2.5
CVE-2025-23290: vGPU Manager cross-VM GPU-metric information exposure
~2/3forecast
Deloitte November 2025 forecast: inference share of 2026 AI compute (½ in 2025, ⅓ in 2023)
Scope & caveats

Deloitte’s November 2025 prediction for inference as a share of AI compute in 2026. Not an observed fleet share, installed capacity, energy or instantaneous electrical draw. The forecast does not allocate an individual fleet.

A forecast for calendar 2026, not an observed 2026 outcome.

2–3 yr (bear case) vs 4–6 yr (GS)forecast
accelerated GPU economic life — contested: bear case 2–3 yr on obsolescence; Goldman Sachs estimates 4–6 yr useful life vs 5–6 yr book
Scope & caveats

Goldman Sachs characterizes four to six years as the estimated useful life of AI accelerators; it does not substantiate a general two-to-three-year economic life or a universal 20–40% three-year residual value.

Isolation models: the axis everyone conflates with sharing

The costliest error, and the one that shows up in real CVEs, is treating a sharing mechanism as an isolation boundary. Sharing answers "how is the silicon divided?" Isolation answers "what can one tenant do to another?" — corrupt their data, starve their compute, read their memory, or escape onto the host. The two axes are orthogonal, and the strength you need on the isolation axis is set by tenant trust, not by how you happen to be packing the GPU.

Stack the isolation tiers from weakest to strongest. Process isolation (bare processes, or MPS clients) shares a kernel, a driver, and frequently a CUDA context — fine for one team's own jobs, useless across trust domains. Container isolation adds namespaces and cgroups but still shares the host kernel and the GPU driver/runtime; the attack surface is the container runtime and the NVIDIA stack itself — exactly the surface the 2025 NVIDIAScape vulnerability (CVE-2025-23266, CVSS 9.0) punctured, letting a crafted container escape to the host on shared GPU nodes. VM isolation gives each tenant its own kernel behind a hypervisor; the GPU is passed through (full device) or virtualized (vGPU/MIG-backed). Stronger, though the vGPU manager is itself a shared component, and CVE-2025-23290 showed a guest reading global GPU metrics influenced by co-tenants, the first acknowledged cross-VM leakage through that layer. Confidential VM + GPU TEE is the top tier: the workload runs in a CPU trusted execution environment, the GPU runs in an attested, platform-specific confidential-compute mode: Hopper protects on-package HBM inside the package trust boundary but does not encrypt it, encrypts CPU↔GPU transfers, and leaves GPU↔GPU NVLink unencrypted in protected-PCIe mode; Blackwell multi-GPU pass-through also encrypts the NVLink path, so even the operator — hypervisor, host OS, cloud admin — cannot read tenant data or weights. This is the canonical home of Chapter 11.5 (GPU confidential computing & attestation); the multi-tenant security architecture that wraps it is Chapter 11.6, and the network half — per-tenant microsegmentation, east-west zero-trust on the storage and management planes — is Chapter 11.7.

The decision rule is blunt: match the isolation tier to the trust boundary, then choose any sharing mechanism that fits inside it. Two jobs from the same team can share an MPS context behind a single container — the trust domain is one, so weak isolation is fine and you maximize packing. Two mutually-distrusting external customers must never share a kernel: that is a VM boundary at minimum, and a confidential VM when your own platform must be untrusted (the regulated, sovereign-AI, and model-weight-protection cases). MIG and vGPU sit awkwardly in between — strong performance isolation, but documented side channels (uncore counters, shared metrics, microarchitectural leakage) mean you do not sell them as a confidentiality boundary between adversaries. They are a QoS boundary that happens to live in hardware.

Deep dive: three 2025–26 isolation failures and the specific boundary each crossed.

The "is partitioning a security boundary?" debate stopped being theoretical in 2025. Three classes of failure, each pointing at a different layer of the stack:

1. Container-escape (NVIDIAScape, CVE-2025-23266, CVSS 9.0). A flaw in the NVIDIA Container Toolkit let a malicious container break out to the host on shared GPU nodes — the worst outcome in multi-tenancy, because host compromise means access to every other tenant on the box. The lesson: the GPU runtime and toolkit are part of your trust boundary, not just the kernel. Container isolation across trust domains is only as strong as the GPU plumbing underneath it, and that plumbing gets CVEs like any other privileged software. Patch cadence on the GPU Operator and Container Toolkit is a multi-tenancy security control, not a hygiene nicety (Chapter 10.4) — and 2026 supplied the follow-on proof: CVE-2026-24260 (NVIDIA bulletin, Jun 2026; CVSS 8.5) is a TOCTOU race in the same Container Toolkit, a different bug class from NVIDIAScape's LD_PRELOAD hook on the same shared-kernel surface; fixed in Toolkit 1.19.1 / GPU Operator 26.3.2, so pin at or above those.

2. Cross-VM metric leakage (CVE-2025-23290). A guest VM could read global GPU metrics influenced by other VMs — low severity (CVSS 2.5) but conceptually important: it was the first acknowledged leak of co-tenant activity through the vGPU manager. The takeaway is that even a VM boundary leaks signal through shared observability surfaces. A side channel does not need to read your data to hurt you; reading your load can be enough in some threat models.

3. Microarchitectural side channels across MPS/MIG. Academic work ('Spy in the GPU-box' and uncore side-channel studies) demonstrated covert and timing channels that survive MPS and MIG partitioning. Performance partitioning is not confidentiality. If your tenants are adversarial, you climb to VM or confidential-VM isolation; if they are one trust domain, all of these are acceptable risks and you optimize for packing. Know which world you are in, and don't let a sales sheet blur the two.

Noisy neighbors and QoS guarantees

Even inside a single trust domain, where security is not the concern, multi-tenancy has a performance pathology: the noisy neighbor. One tenant's job saturates a shared resource and silently taxes everyone else's goodput — and on a shared GPU node the contended resource is rarely the one people watch. It is memory bandwidth (HBM is the bottleneck for most inference and the noisiest shared resource under MPS/time-slicing), PCIe/NVLink (a greedy host-to-device copy starves a neighbor's transfers), the shared fabric (a tenant's all-to-all floods the back-end and inflates everyone's collective time — congestion engineering in Chapter 8.6), and shared storage and the loader (one tenant's checkpoint write or dataset scan blows the cache for the rest).

You bound the noisy neighbor at one of three layers, hardest to softest. At the hardware: MIG hardware-partitions SM and L2 slices, bounding core and cache interference for tenants with a hard SLO; shared uncore paths remain observable and still need per-tenant telemetry. At the scheduler: resource limits and requests, priority and preemption so a low-priority neighbor yields, and gang-aware admission so a job either gets its full topology-clean allocation or waits rather than half-landing and contending (Chapter 10.2). At the fabric and storage: per-tenant QoS classes, rate limits, and bandwidth reservations so no one tenant monopolizes the network or the loader. The consequence of skipping all three is the failure mode that is hardest to debug because nothing errors: every tenant's p99 latency quietly degrades, the SLO dashboard goes amber, and the cause is invisible unless your telemetry attributes contention per tenant — which is why per-tenant goodput and contention metrics belong in the observability plane from day one (Chapter 10.6).

The QoS guarantee you can offer follows from the mechanism. No GPU partition alone guarantees an end-to-end tenant SLO regardless of neighbors. MIG or whole-GPU can bound GPU-local contention; the contract also needs host, PCIe, fabric, storage, control-plane and measurement boundaries. A soft guarantee ("this tenant gets a target share, best-effort, degrading gracefully under contention") is what MPS, time-slicing, and fractional GPUs can deliver. Selling a hard SLO on a soft mechanism works in the demo, when the GPU is half-empty, and breaches the first time the cluster fills.

Putting it together: a per-tenant-class decision

The mistake is to pick one sharing mechanism, one quota model, and one isolation tier for the whole fleet. The defensible design classifies tenants and maps each class down all three axes at once. A frontier training tenant: whole-GPU, hard partition or a top-priority gang reservation, container isolation (one trust domain), noisy neighbor bounded by sole ownership. An external SLO-bound serving tenant: MIG where measured GPU-local partitioning helps, reserved ready capacity behind the quota, a qualified VM or CVM boundary across trust domains, and end-to-end QoS plus telemetry for every shared path. An internal research/dev fleet: fractional GPU and time-slicing for packing, hierarchical fair-share with preemption for utilization, container isolation, noisy neighbor bounded by scheduler limits. A preemptible batch tier: whatever packs tightest, lowest priority with a measured reclaim deadline, container isolation, explicitly soft QoS. A qualified physical fleet can run all four — the art is keeping each class's mechanism, quota, and isolation tier internally consistent so the guarantees you sell have capacity and a tested recovery path behind them. Fair-share without preemption changes future queue order; it does not revoke a running allocation.

This chapter is the sharing-and-isolation framework; the pieces are engineered elsewhere. The scheduling plane that enforces quota and gang admission is Chapter 10.1; the topology-aware placement that fairness must respect is Chapter 10.2; the node software stack (driver, GPU Operator, Container Toolkit) whose CVEs define the container trust boundary is Chapter 10.4; per-tenant contention and goodput telemetry live in Chapter 10.6. The security axis is canonical downstream: GPU confidential computing, platform-specific memory protection, and attestation in Chapter 11.5; the multi-tenant workload-isolation security architecture in Chapter 11.6; network microsegmentation and zero-trust in Chapter 11.7. Fabric congestion that creates cross-tenant noise is Chapter 8.6; the checkpoint math that makes preemption affordable is Chapter 9.4; and the commercial terms that turn these guarantees into a product are Chapter 10.9.
Cite this chapter
Fehn, J. (2026). Multi-Tenancy, Isolation & Resource Sharing (Chapter 10.3). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-10-software-orchestration-and-service-delivery/10-3-multi-tenancy-isolation-and-resource-sharing (accessed 2026-09-29).
@misc{aidc-10-3,
  author       = {Fehn, Jacob},
  title        = {Multi-Tenancy, Isolation & Resource Sharing (Chapter 10.3)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-10-software-orchestration-and-service-delivery/10-3-multi-tenancy-isolation-and-resource-sharing},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit