KV-cache per token of context, Llama 3.1 405B (327 KB for Qwen-2.5 72B)
504 KiB/token (516.096 kB/token)derived
| Value kind | derived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method. |
|---|---|
| Scope | Aggregate logical KV at stated precision; excludes weights, workspaces, allocator overhead and replication. Derive actual per-rank placement before testing fit. |
| As of | 2024-07 |
| Source | Meta model dimensions published 2024; primary paper. BF16 storage and ideal KV-head sharding are guide assumptions, not a serving measurement. · Table 3; 70B or 405B layers, model width, attention-head and KV-head counts |
| Derivation | 2 × 126 layers × 8 KV heads × 128 head dimension × 2 bytes = 516096 bytes/token = 504 KiB/token. |
| Review | checking…review by 2027-09-05 · standard cadence |
| Recorded changes | last 2026-09-16 · 3 revisions tracked |
| Claim id | kv-cache-per-token-of-context-llama-3-1-405b |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.