Test-time compute
Spending extra inference compute (e.g. chain-of-thought reasoning) to improve answers, shifting cost from training to serving.
Current numbers
5-100xmore tokens/test-time compute per query for reasoning models vs one-shot inference — workload-dependent (NVIDIA cites up to ~100x compute)