The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Glossary › Expert parallelism

Expert parallelism · EP

Distributing a Mixture-of-Experts model's experts across GPUs, routing tokens to whichever GPU holds the chosen expert.

Current numbers

~18×NVLink5 scale-up (1.8 TB/s/GPU bidirectional, ~900 GB/s/dir; 130 TB/s NVL72 rack) over ~400G scale-out NIC (~50 GB/s/dir) — keep TP/EP inside scale-upas of 2025 · register ↗

← All terms