LLMflation
The rapid collapse in the cost to serve a given level of AI capability as models and hardware improve.
Current numbers
~$1.90 → ~$2.50/M tokself-hosted vs market-avg inference cost per million tokens; ~10x/yr token-price deflation (LLMflation)
~10x/yrLLMflation: inference cost decline at fixed quality (Epoch Mar-2025: ~50x/yr median; ~200x/yr post-2024 models)
~10x/yrLLMflation: drop in cost to serve a fixed-quality token; ~1,000x over 3 yr (GPT-3 quality ~$60 to ~$0.06/M)