Meta Llama 3 405B pre-training: 41% BF16 MFU on 16,384 H100 GPUs at 8,192-token sequence length
~41%observed
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | Named 405B pre-training configuration; other Table 4 layouts report 38% or 43%. No Hopper fleet-wide MFU band or universal acceptance threshold follows. |
| As of | 2024-07 |
| Source | Meta, The Llama 3 Herd of Models, Table 4, July 2024 report. · Table 4: 16,384 GPUs; TP8, CP1, PP16, DP128; sequence 8,192; batch 16/DP; BF16 MFU 41%. |
| Review | checking…review by 2026-07-27 · standard cadence |
| Recorded changes | last 2026-09-16 · 2 revisions tracked |
| Claim id | bf16-mfu-achieved-pre-training-llama-3-on-16k |
Where the guide uses it
- 10.8 MLOps & Training Frameworks
- 13.9 Cluster-Scale Benchmarking, Reference Training & Storage/Scheduler Validation
← Full numbers register — every date-stamped figure in the guide, with revision history.