Rust tokenizer speedup vs pure-Python; ~1 GB text encoded in tens of seconds per server, still CPU-bound at corpus scale
10–100x (Rust/HF tokenizers vs pure Python, 1 GB-text benchmark)observed
| Value kind | observed — Reported measurements, counts, and specifications keep the precision and scope stated by their source; an exact specification is not treated as a range. |
|---|---|
| Scope | tokenizer speedup |
| As of | 2025 |
| Source | Hugging Face |
| Review | checking…review by 2026-09-30 · standard cadence |
| Recorded changes | last 2026-09-01 · 3 revisions tracked |
| Claim id | rust-tokenizer-speedup-vs-pure-python-1-gb-text |
Where the guide uses it
Not quoted in a chapter yet; it is kept in the curated register.
← Full numbers register — every date-stamped figure in the guide, with revision history.