Expert parallelism · EP
Distributing a Mixture-of-Experts model's experts across GPUs, routing tokens to whichever GPU holds the chosen expert.
Current numbers
~7x / ~10xDynamo + wide-EP MoE throughput on GB200 NVL72 vs B200 (Dynamo 1.0 GA at GTC 2026); NIXL+GPUDirect Storage prefill speedup for long context
~18×NVLink5 scale-up (1.8 TB/s/GPU bidirectional, ~900 GB/s/dir; 130 TB/s NVL72 rack) over ~400G scale-out NIC (~50 GB/s/dir) — keep TP/EP inside scale-up