MoE · Mixture of Experts
A sparse model that activates only a subset of expert sub-networks per token, cutting compute per token at scale.
An architecture routing each token to a few specialized sub-networks, widening parallelism and reshaping fabric needs.
Also written as: Mixture of Experts
Current numbers
25T tokenslongest documented 4-bit pre-training run (Nemotron 3 Super, 120B-total / ~12B-active hybrid Mamba-Transformer MoE, NVFP4, 25T total seen tokens)