The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Numbers register › Claim

guide-derived self-hosted 70B inference-cost scenario using declared GPU price, measured throughput, utilization and batching assumptions

$1.90-$2.50/M tokens (guide scenario)derived

Value kindderived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method.
ScopeScenario output, not a market benchmark; recompute with the selected instance price and workload-specific throughput.
As of2026-08-24
SourceGuide derivation using public cloud pricing and NVIDIA NIM benchmark fixtures · Scenario inputs and formula are declared in Chapter 0.3
DerivationGuide scenario: instance or owned-GPU hourly cost divided by measured workload throughput and effective utilization, scaled to one million output tokens. Recompute rather than treating the range as observed.
Reviewchecking…review by 2026-09-23 · fast cadence
Recorded changeslast 2026-08-24 · 4 revisions tracked
Claim idinference-cost-per-million-tokens-self-hosted

Where the guide uses it

← Full numbers register — every date-stamped figure in the guide, with revision history.