The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Numbers register › Claim

aggregate KV-cache state per sequence across participating ranks for Llama 3 70B at 128K context: 40 GiB aggregate and 5 GiB/rank in a named TP=8 ideal head-sharded illustration before overhead

40 GiB aggregate; 5 GiB/rank at ideal TP=8derived

Value kindderived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method.
ScopeAggregate logical KV at stated precision; excludes weights, workspaces, allocator overhead and replication. Derive actual per-rank placement before testing fit.
As of2024-07
SourceMeta model dimensions published 2024; primary paper. BF16 storage and ideal KV-head sharding are guide assumptions, not a serving measurement. · Table 3; 70B or 405B layers, model width, attention-head and KV-head counts
Derivation2 × 80 × 8 × 128 × 2 × 131072 = 42949672960 B = 40 GiB = 42.94967296 GB. Ideal eight-way KV-head sharding gives 5 GiB/rank, before overhead.
Reviewchecking…review by 2027-09-05 · standard cadence
Recorded changeslast 2026-09-16 · 3 revisions tracked
Claim idkv-cache-for-one-llama-3-70b-request-at-128k

Where the guide uses it

← Full numbers register — every date-stamped figure in the guide, with revision history.