The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
GuideGlossaryMixture of Experts

Mixture of Experts · MoE

An architecture routing each token to a few specialized sub-networks, widening parallelism and reshaping fabric needs.

Current numbers

~1/4 the number (~75% fewer, ~4×)Vera Rubin GPUs to train an MoE model vs Blackwell (training metric)as of 2026-07 · register ↗
~10xtokens/watt advantage of Blackwell-class over Hopper on MoE inference — the power-limited leveras of 2026 · register ↗
~7x / ~10xDynamo + wide-EP MoE throughput on GB200 NVL72 vs B200 (Dynamo 1.0 GA at GTC 2026); NIXL+GPUDirect Storage prefill speedup for long contextas of 2026 · register ↗

← All terms