aggregate KV-cache state per sequence across participating ranks for Llama 3 70B at 128K context: 40 GiB aggregate and 5 GiB/rank in a named TP=8 ideal head-sharded illustration before overhead
40 GiB aggregate; 5 GiB/rank at ideal TP=8derived
| Value kind | derived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method. |
|---|---|
| Scope | Aggregate logical KV at stated precision; excludes weights, workspaces, allocator overhead and replication. Derive actual per-rank placement before testing fit. |
| As of | 2024-07 |
| Source | Meta model dimensions published 2024; primary paper. BF16 storage and ideal KV-head sharding are guide assumptions, not a serving measurement. · Table 3; 70B or 405B layers, model width, attention-head and KV-head counts |
| Derivation | 2 × 80 × 8 × 128 × 2 × 131072 = 42949672960 B = 40 GiB = 42.94967296 GB. Ideal eight-way KV-head sharding gives 5 GiB/rank, before overhead. |
| Review | checking…review by 2027-09-05 · standard cadence |
| Recorded changes | last 2026-09-16 · 3 revisions tracked |
| Claim id | kv-cache-for-one-llama-3-70b-request-at-128k |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.