guide-derived self-hosted 70B inference-cost scenario using declared GPU price, measured throughput, utilization and batching assumptions
$1.90-$2.50/M tokens (guide scenario)derived
| Value kind | derived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method. |
|---|---|
| Scope | Scenario output, not a market benchmark; recompute with the selected instance price and workload-specific throughput. |
| As of | 2026-08-24 |
| Source | Guide derivation using public cloud pricing and NVIDIA NIM benchmark fixtures · Scenario inputs and formula are declared in Chapter 0.3 |
| Derivation | Guide scenario: instance or owned-GPU hourly cost divided by measured workload throughput and effective utilization, scaled to one million output tokens. Recompute rather than treating the range as observed. |
| Review | checking…review by 2026-09-23 · fast cadence |
| Recorded changes | last 2026-08-24 · 4 revisions tracked |
| Claim id | inference-cost-per-million-tokens-self-hosted |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.