The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
GuideGlossaryExpert parallelism

Expert parallelism · EP

Distributing a Mixture-of-Experts model's experts across GPUs, routing tokens to whichever GPU holds the chosen expert.

Current numbers

~7x / ~10xDynamo + wide-EP MoE throughput on GB200 NVL72 vs B200 (Dynamo 1.0 GA at GTC 2026); NIXL+GPUDirect Storage prefill speedup for long contextas of 2026 · register ↗
~18×NVLink5 scale-up (1.8 TB/s/GPU bidirectional, ~900 GB/s/dir; 130 TB/s NVL72 rack) over ~400G scale-out NIC (~50 GB/s/dir) — keep TP/EP inside scale-upas of 2025 · register ↗

← All terms