serverless GPU startup must separate infrastructure resume, model-ready time and end-to-end time to first token; each is model-, platform-, image- and cache-state specific
Measure resume, model-ready and TTFT separatelyrecommendation
| Value kind | recommendation — Recommendations are design guidance, not measured facts; validate them against the project boundary conditions. |
|---|---|
| Scope | Warm-state resume is not end-to-end first-token latency, and no portable H100 range is retained. |
| Caveat | Platform-published benchmark; validate on the selected model, image, region, cache state and serving stack. |
| As of | 2026-08-24 |
| Source | RunPod, controlled vLLM cold-start benchmark · Cold-start phases and controlled model-size benchmark |
| Review | checking…review by 2026-12-22 · standard cadence |
| Recorded changes | last 2026-08-24 · 3 revisions tracked |
| Claim id | serverless-gpu-time-to-first-token-h100-warm |
Where the guide uses it
Not quoted in a chapter yet; it is kept in the curated register.
← Full numbers register — every date-stamped figure in the guide, with revision history.