MIG · Multi-Instance GPU
An NVIDIA feature partitioning one GPU into isolated instances so multiple workloads share it securely.
Current numbers
~2/3inference share of AI compute in 2026 — the workload class that most rewards fractional/MIG sharing
up to 7MIG instances per GPU (B200/GB200): 2×~93GB, 4×~46GB, or 7×~23GB profiles
7max MIG instances per GPU — the only hardware-enforced fractional partition (dedicated SMs, L2 slice, memory controllers, HBM slice)