Blackwell (B200) confidential-computing overhead for a tuned LLM serving stack; stock stacks and eight-GPU training cost more
~1–3% tuned servingobserved
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | One physical host with Intel TDX and NVIDIA Blackwell CC; ~1–3% tuned serving, 30–40% stock serving and ~10–13% eight-GPU training in the evaluated configurations. Not a multi-node or all-workload bound. |
| As of | 2026-09 |
| Source | arXiv, independent paired-run B200 study arXiv:2608.26575v2; submitted 2026-08-27, revised 2026-09-01 · Version 2; abstract and workload evaluations |
| Review | checking…review by 2027-02-01 · standard cadence |
| Recorded changes | last 2026-09-16 · 3 revisions tracked |
| Claim id | blackwell-cc-overhead-on-large-matrix-ops |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.