The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Numbers register › Claim

KV-cache per token of context, Llama 3.1 405B (327 KB for Qwen-2.5 72B)

504 KiB/token (516.096 kB/token)derived

Value kindderived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method.
ScopeAggregate logical KV at stated precision; excludes weights, workspaces, allocator overhead and replication. Derive actual per-rank placement before testing fit.
As of2024-07
SourceMeta model dimensions published 2024; primary paper. BF16 storage and ideal KV-head sharding are guide assumptions, not a serving measurement. · Table 3; 70B or 405B layers, model width, attention-head and KV-head counts
Derivation2 × 126 layers × 8 KV heads × 128 head dimension × 2 bytes = 516096 bytes/token = 504 KiB/token.
Reviewchecking…review by 2027-09-05 · standard cadence
Recorded changeslast 2026-09-16 · 3 revisions tracked
Claim idkv-cache-per-token-of-context-llama-3-1-405b

Where the guide uses it

← Full numbers register — every date-stamped figure in the guide, with revision history.