LLMflation: inference cost decline at fixed quality — Epoch AI puts it at ~9–900x/yr depending on the capability threshold, ~40x/yr for GPT-4-level GPQA Diamond
~9–900x/yr by thresholdobserved
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | The headline rate depends entirely on the basis. a16z's original 'LLMflation' series tracks GPT-3-level quality and gives ~10x/yr (~1,000x over three years, ~$60 to ~$0.06 per million tokens); Epoch AI's cross-benchmark median is nearer ~50x/yr, ~40x/yr at the GPT-4-level GPQA Diamond threshold, and ~9x to ~900x/yr across thresholds. Always state which basis a quoted rate uses — these are not restatements of one measurement. |
| As of | 2024-2025 |
| Source | Epoch AI, LLM inference price trends |
| Review | checking…review by 2026-09-22 · fast cadence |
| Recorded changes | last 2026-08-31 · 2 revisions tracked |
| Claim id | llmflation-inference-cost-decline-at-fixed |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.