The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Security › 11.6

Chapter 11.6

In this chapter · 5 sections
Term help

Multi-Tenant & Workload Isolation Security

Choose the sharing mode by the boundary it actually enforces, then qualify tenant handoff, privileged control paths and failure propagation before selling hostile tenants the same resource.

GOODPUTDENSITY-RAMP

What you'll decide here

  1. Where on the multi-tenancy spectrum each pod sits — bare-metal-per-tenant, VM-per-tenant, MIG-partitioned, vGPU-virtualized, or kernel-level time-sliced — and therefore which isolation failures you have signed up to own.
  2. Which memory, execution, fault, device-management and operator boundaries each proposed MIG, vGPU, MPS or time-slicing configuration enforces, using the exact supported mode rather than its product label.
  3. Your memory-hygiene policy on every reassignment: who scrubs GPU framebuffer, HBM, and local/shared memory between tenants, and whether that scrub is verified or assumed.
  4. Whether the tenant excludes the ordinary platform administrator from plaintext access and therefore needs a supported confidential execution path, or accepts that administrator within a qualified VM or dedicated-host trust boundary.
  5. Which container, GPU-operator and orchestration privileges reach tenant workloads, and which hardware confidentiality boundary still holds when one of those privileges is compromised.

Multi-tenancy is the business model of the entire neocloud and most enterprise platform teams: buy one expensive accelerator, rent fractions of it to many workloads, and amortize its paid capacity across tenants whose measured workloads leave shareable headroom. The economic gain depends on useful output after contention, reserve and recovery, and the security consequences of that sharing are routinely underpriced. The instant two tenants share a physical GPU, a memory controller, an L2 cache, an NVLink domain, or merely a host kernel and a container runtime, you have created a shared dependency to qualify — and the only question left is whether it is a strong boundary, a soft convenience dressed up as a boundary, or an outright leak. The work here is telling those apart before a tenant's KV-cache shows up in a neighbor's framebuffer.

The route: the multi-tenancy spectrum from bare-metal to time-slicing, with the documented isolation failures at each rung; the question every platform team mishandles — is MIG / vGPU / time-slicing a security boundary? — answered from the hardware reality rather than the marketing; memory hygiene on reassignment as an explicit, owned policy instead of an assumption; and the isolation-versus-utilization economics that govern the whole thing, with confidential multi-tenancy addressing a selected host threat and retaining its own sharing and side-channel limits. The canonical TEE and attestation machinery lives in Chapter 11.5; here we decide when you need it.

The multi-tenancy spectrum

Tenant isolation is a spectrum, and each step down it trades a stronger boundary for higher utilization. Read it top to bottom as decreasing isolation strength and increasing density: the moment you stop giving each tenant a whole physical GPU, you start sharing silicon that was never designed as a confidentiality boundary.

Bare-metal-per-tenant is the strongest posture: a tenant gets whole nodes, whole GPUs, and ideally a whole rack or NVLink domain, with no hypervisor and no co-tenant on the same silicon. The blast radius is the tenant's own. This is what serious training customers and sovereign deployments demand, and it is what the strongest neocloud SLAs implicitly promise. The cost is utilization: a tenant who under-uses the hardware strands it, and you cannot backfill without reintroducing a co-tenant.

VM-per-tenant (whole-GPU passthrough) keeps the GPU undivided but interposes a hypervisor, giving each tenant a strong CPU-side boundary (SR-IOV / PCIe passthrough) while the GPU itself is still owned end-to-end by one VM at a time. The boundary you now trust is the hypervisor and the IOMMU, not the GPU. This is the workhorse of enterprise multi-tenant clouds and a genuinely defensible boundary — the residual risk is hypervisor escapes and the host-side GPU management stack.

MIG (Multi-Instance GPU) partitions a single physical GPU into up to seven hardware-isolated instances with dedicated SM and memory-system slices. Copy and media-engine allocation is product- and profile-specific, and compute instances inside one GPU instance share their parent's memory and engines; select a +me profile where dedicated media engines are required. This is the only fractional-sharing mode with a hardware-enforced partition, which is why it is the one rung of fractional sharing that can credibly be called a security boundary. The cost is rigidity (fixed geometries, no oversubscription) and the residual uncore side-channels discussed below.

vGPU (mediated virtualization) time-multiplexes a GPU across guest VMs via a host-resident vGPU Manager that schedules and mediates access. The VM boundary is real, but the GPU is now shared through a software mediator with privileged host code — and that mediator has been the source of the most consequential 2025 multi-tenant CVEs. vGPU can be layered on MIG (MIG-backed vGPU) to combine hardware partitioning with the VM boundary, which is the strongest practical fractional mode.

Time-slicing is the bottom rung: processes share one GPU with neither memory nor fault isolation. It is a utilization tool for trusted workloads, not a tenant boundary.

MPS is a separate cooperative-workload service. Volta-and-later clients have isolated GPU address spaces, but scheduling hardware, memory bandwidth, caches, and capacity remain shared. Standard MPS provides no error isolation; R610 static SM partitioning contains some SM-triggered faults but is not a general fault-isolation guarantee. Neither mode is suitable as a hostile-tenant boundary.

The multi-tenancy spectrum: isolation strength vs utilization
Isolation modeBoundary trustedMemory isolationFault isolationSecurity boundary?UtilizationDocumented failure class
Bare-metal-per-tenantPhysical separationTotal (no co-tenant)TotalYes — strongestLowest (strands idle)Decommission / reassignment remanence
VM-per-tenant (passthrough)Hypervisor + IOMMUTotal (whole GPU per VM)TotalYesLow-moderateHypervisor / SR-IOV escapes
MIG (hardware partition)GPU hardware partitionHardware-enforced slicePer-instanceYes — strongest fractionalHigh (≤7 instances)Uncore side-channels; fixed geometry
vGPU (mediated)VM + host vGPU ManagerPer-VM via mediatorPer-VMConditional — see CVEsHigh (oversubscribable)vGPU Manager CVEs; cross-VM leakage
Time-slicingNone (shared context)NoneNoneNoHighest (oversubscribed)Cross-process memory exposure; noisy-neighbor DoS
Boundary classification reflects 2026 hardware reality, not vendor positioning. "Boundary trusted" names the component whose compromise breaks isolation. Utilization is qualitative — see economics section.

Documented isolation failures: the empirical record

The 2024–2025 disclosure record contains real, assigned CVEs at every layer of the stack, and reading them is the fastest way to calibrate which boundaries hold. The pattern is instructive: the catastrophic breaks are almost never in the GPU partitioning itself, but in the control plane around it.

The control-plane escape (the real catastrophe). NVIDIAScape — CVE-2025-23266, CVSS 9.0, disclosed by Wiz in July 2025 — is a three-line container escape in the NVIDIA Container Toolkit: a malicious container sets LD_PRELOAD against an OCI createContainer hook and gains root on the host, from which it can read, steal, or tamper with the models and data of every other tenant on the shared machine. It affected Container Toolkit up to 1.17.7 and GPU Operator up to 25.3.1. Multi-tenant security is a full-stack property, and the runtime is the weakest, most-exposed link.

The mediator leak (vGPU). NVIDIA's July 2025 vGPU bulletin disclosed cross-VM issues in the vGPU Manager: CVE-2025-23290 (cross-VM information disclosure — a guest reading global GPU metrics influenced by neighboring VMs, the first publicly acknowledged leakage of co-tenant activity through the mediator) and CVE-2025-23285 (a guest consuming global resources to deny service to neighbors). The same bulletin patched stack buffer overflows in the vGPU Manager (CVE-2025-23283/23284, CVSS 7.8) enabling guest-to-host code execution. The mediated-virtualization boundary is real but it is privileged software, and privileged software has bugs.

The memory-remanence leak (affected local-memory stacks). LeftoverLocals — CVE-2023-4969, Trail of Bits, 2024 — recovered another process's data from uncleared GPU local memory across process and container boundaries on affected Apple, Qualcomm, AMD, and some Imagination devices, enough to reconstruct an LLM's responses (≈181 MB recoverable per query against a 7B model on llama.cpp). Tested NVIDIA devices returned zeros, and the disclosure did not evaluate global HBM, MIG teardown, or vGPU release. The control belongs at the local/shared-memory layer: use patched drivers and kernels that clear scratchpad on affected stacks.

The uncore side-channels (even MIG). Spy in the GPU-box (arXiv 2203.15981) demonstrated cross-GPU L2/NVLink channels on an eight-P100 Pascal DGX-1; it did not test MIG. Veiled Pathways demonstrated uncore channels based on DRAM-frequency contention, media engines under MPS, and shared PCIe I/O between MIG instances. MIG's partition is genuinely hardware-enforced for memory and compute, but it is not a proof of side-channel resistance. For data where timing leakage is in-scope, MIG is necessary and not sufficient — and neither is a GPU TEE, whose deployment threat model excludes side channels. Pick the mitigation from the specific shared resource and leakage mechanism: stop sharing the contended engine, dedicate the allocation, or partition the path — and test the chosen control against the actual attack rather than assuming a TEE established timing noninterference.

9.0
CVSS of NVIDIAScape (CVE-2025-23266) — three-line container escape to host root in NVIDIA Container Toolkit
≤1.17.7
NVIDIA Container Toolkit versions vulnerable to CVE-2025-23266/NVIDIAScape — through 1.17.7; NVIDIA GPU Operator affected through 25.3.1
Scope & caveats

For Container Toolkit versions before 1.17.5 and GPU Operator versions before 25.3.1, CVE-2025-23266 applies only when CDI mode is used.

CVE-2025-23290
first publicly acknowledged cross-VM co-tenant information disclosure via the vGPU Manager
7
Supported NVIDIA GPUs offer up to seven MIG instances in specified profiles; qualify the actual model, mode and shared failure domain
≈181 MB
LLM-response data recoverable per query via LeftoverLocals (CVE-2023-4969) from un-scrubbed GPU local memory
10-dimension
ClusterMAX 2.0 operator-maturity rubric grades tenant/fabric isolation, health-checks, and goodput as first-class

Memory hygiene on reassignment: the owned policy

The most routine multi-tenant control is also the most overlooked: verifying the teardown path when a resource changes hands. Every time a whole GPU, MIG instance, vGPU slice, or container allocation moves between tenants, require the named mode's vendor-defined clearing or cryptographic teardown and test it on the deployed stack before reassignment. The decision you own is who produces that evidence and whether reassignment fails closed when verification does not pass.

Two postures compete: rely on the documented driver/firmware teardown or gate reassignment on an operator-controlled verification test. The former minimizes handoff latency but inherits the vendor implementation; the latter adds measured scheduler delay and produces evidence for the tenant. For confidential workloads, verify that the platform's attested teardown destroys the protecting keys before treating residual ciphertext as inaccessible (Chapter 11.5).

Deep dive: the three layers of GPU memory remanence and who scrubs each

"Scrub the GPU memory" is too coarse to be actionable, because GPU memory is at least three distinct regions with three distinct owners, and a policy that covers one and not the others leaks through the gap.

Global / HBM (framebuffer and device allocations). Establish teardown semantics separately for whole-GPU, MIG, and vGPU reassignment from the vendor's mode-specific documentation, then verify the deployed driver/firmware path with an acceptance test before enabling cross-tenant reuse. Measure that exact clear-or-erase and verification path: raw capacity divided by peak HBM bandwidth is not a teardown-latency result.

Local / shared memory (per-SM scratchpad). The small, fast, software-managed region that LeftoverLocals exploited. It is not automatically cleared between kernel launches on affected stacks, so a reader kernel could dump whatever a prior victim kernel left behind. Mitigation is a vendor driver fix plus, defensively, kernels that clear their own local memory on exit — an overhead to measure on the affected implementation, with the vendor fix and cleanup semantics checked.

Caches and registers (L2, register file). The hardest to reason about and the realm of side-channels rather than direct remanence: MIG dedicates L2 slices per instance, but contention on shared uncore paths is observable, which is the uncore side-channel result. Specify cleanup and residual observation channels by mode. Confidential computing does not erase every side channel; accept the remaining channels explicitly or choose a deployment without the implicated sharing. Record which threat and platform each conclusion covers.

Confidential multi-tenancy: the strongest, most expensive rung

When the data class genuinely cannot tolerate exposure to a co-tenant or to the operator — regulated health and financial data, sovereign workloads, frontier weights rented on infrastructure you do not own — the only sufficient answer is confidential computing: a per-tenant Trusted Execution Environment spanning the CPU (SEV-SNP / TDX) and the GPU (NVIDIA Hopper/Blackwell CC with an access-fenced HBM region (encrypted on Blackwell), the BAR0 decoupler, and encrypted transfers), gated by attestation so a tenant releases keys only to a measured, verified configuration. The full machinery — Compute Protected Region state, attestation via NRAS/RIM, TEE-I/O across NVLink, the residual attack surface — is the canonical subject of Chapter 11.5. The decision here is whether you need it, because it is not free.

Strong VM isolation + disciplined hygiene defends against an honest-but-curious neighbor and an accidental leak; it does not defend against a malicious co-tenant exploiting a side-channel, nor against a compromised or coerced operator, nor against an insider with host access. Confidential multi-tenancy moves the operator out of the trust boundary and makes the attestation the gate — at the cost of a measurable performance tax (encrypted PCIe/NVLink transfers and TEE overhead, larger for small-transfer chatty workloads than for large compute-bound ones), reduced GPU-sharing flexibility (CC modes constrain partitioning), and real operational complexity in key brokering and attestation-policy management. You pay goodput and ops to remove the operator from the trust model. If your tenants do not require that removal, you are buying nines of confidentiality the data class does not value — the security analog of over-provisioned redundancy.

Choosing a multi-tenant posture by adversary and data class
Adversary you must defend againstMinimum sufficient postureBoundary that must holdGoodput / cost penaltyResidual risk you accept
Accidental cross-tenant leakVM-per-tenant + verified scrubHypervisor + scrub policyLowHypervisor escape; side-channels
Noisy neighbor / DoSMIG (hardware partition)GPU hardware partitionLow (fixed geometry overhead)Uncore side-channels
Malicious co-tenant (data theft)MIG-backed vGPU + hygieneHW partition + VM mediatorLow-moderatevGPU Manager CVEs; timing leakage
Malicious co-tenant (timing/side-channel)Dedicated allocation of the leaking resource; a GPU TEE does not cover this threatNo sharing of the contended uncore pathModerate-high (stranded capacity)Architectural side channels the TEE threat model excludes; TEE implementation flaws
Compromised / coerced operatorConfidential computing + tenant-held keysAttestation-gated key releaseModerate-high (ops complexity)Supply-chain / firmware trust root
Read top to bottom as strengthening adversary model. Each row is the minimum sufficient posture for that threat — over-buying wastes goodput, under-buying is a breach waiting to happen.

The isolation-versus-utilization economics

Every rung you climb for stronger isolation costs utilization, and utilization is the entire economic premise of multi-tenancy. This is the central tension and it is quantifiable. A bare-metal tenant who drives a $30k+ accelerator at 40% leaves ~60% of it idle and unbackfillable. MIG recovers much of that by packing up to seven isolated tenants onto one die. Time-slicing recovers the most by oversubscribing — at the price of being no boundary at all. Confidential computing climbs back down the utilization ladder: it constrains partitioning, taxes transfers, and complicates scheduling, so the same hardware serves fewer effective tenant-hours.

The decision is therefore not "how secure can we be" but "what is the cheapest posture that is sufficient for this data class's adversary model" — which is why the table above is organized by adversary, not by feature. Over-isolating is a real cost: a commercial inference fleet serving public, non-sensitive content that runs everything under per-tenant confidential computing is burning goodput and capex to defend against a threat its data class does not face. Under-isolating is a breach: a regulated-data tenant time-sliced onto shared silicon with a stranger is one un-scrubbed allocation away from a disclosure. The ClusterMAX 2.0 rubric (SemiAnalysis) bakes this judgment into how the market grades neoclouds — tenant and fabric isolation are scored as first-class alongside goodput and health-checks, because buyers have learned to price the boundary, not just the FLOPs.

The pragmatic 2026 default for a serious multi-tenant operator is a tiered offering: bare-metal or VM-per-tenant for customers who pay for it and demand it; MIG-backed vGPU with verified hygiene as the standard fractional product; time-slicing reserved strictly for a single tenant's own co-operative workloads (never across trust boundaries); and confidential computing as a premium tier for regulated and sovereign demand. Selling time-slicing as a cross-tenant boundary, or selling "isolation" without naming which rung, is the misrepresentation that turns a CVE into a liability.

Deep dive: the control plane is the real multi-tenant attack surface

The in-GPU isolation mechanisms — MIG, vGPU, CC — get the attention, but the empirical breach record points relentlessly at the orchestration layer that sits above them. In a Kubernetes-based GPU platform the multi-tenant boundary is enforced by a stack of software components, each of which is a tenant-reachable attack surface: the container runtime and NVIDIA Container Toolkit (NVIDIAScape lived here), the GPU Operator and device plugin that advertise and allocate GPUs, the scheduler / fractional-GPU layer (e.g. Run:ai-style quota and policy) that decides which tenant lands on which slice, and the namespace / RBAC / network-policy fabric that is supposed to keep tenant A from reaching tenant B's pods and services.

The hard isolation design (the hard / soft / hybrid taxonomy) treats these as the primary controls: per-tenant namespaces with enforced RBAC and resource quotas, network policies that default-deny east-west traffic between tenants, admission controllers that block privileged containers and dangerous host mounts, and a patched, current Container Toolkit as table stakes. A neocloud that nails MIG partitioning but runs an unpatched Container Toolkit, permits privileged pods, or shares a flat tenant network has a strong GPU boundary wrapped in a soft control plane — and attackers, like water, find the soft part. In 2026 the CVE flood moved one layer up, to the inference-serving control plane: NVIDIA's August bulletin for Dynamo carried fifteen CVEs at once (CVE-2026-24254, an unauthenticated out-of-bounds write in the multimodal serving topology, CVSS 9.8 — all fifteen fixed only at v1.3.0), and Triton took CVE-2026-47627 (CVSS 9.8 path traversal, versions through 26.05). The disaggregated router and the inference server are tenant-reachable software with their own patch cadence — pin a patched-serving-version acceptance gate exactly as you pin the Container Toolkit. The corollary for buyers: when you diligence a multi-tenant provider, audit the control plane and the patch cadence before you ask about MIG. The fabric-side enforcement (DPU-VPC, per-tenant PKeys vs shared VLANs) is detailed in Chapter 11.7.

Confidential computing — the GPU TEE, the protected HBM region, attestation via NRAS/RIM, TEE-I/O over NVLink, and the residual attack surface — is the canonical subject of Chapter 11.5; this chapter decides when its cost is warranted. The container-runtime and orchestration hardening that NVIDIAScape exposes connects to firmware and BMC integrity in Chapter 11.4 and to the network/microsegmentation boundary that carries east-west tenant traffic in Chapter 11.7. Memory remanence on reassignment is the in-GPU analog of media sanitization and data remanence at decommission in Chapter 11.3. The model and weight protection that a tenant boundary ultimately exists to guarantee is treated in Chapter 11.8. The multi-tenant scheduling, quota, and fair-share mechanics that decide who shares what are engineered in Chapter 10.3; the isolation-vs-utilization economics tie back to the procurement and goodput framing of Chapter 1.6.
Cite this chapter
Fehn, J. (2026). Multi-Tenant & Workload Isolation Security (Chapter 11.6). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-11-security/11-6-multi-tenant-and-workload-isolation-security (accessed 2026-09-29).
@misc{aidc-11-6,
  author       = {Fehn, Jacob},
  title        = {Multi-Tenant & Workload Isolation Security (Chapter 11.6)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-11-security/11-6-multi-tenant-and-workload-isolation-security},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit