The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
Guide › Storage & Data › 9.9

Chapter 9.9

In this chapter · 8 sections
Term help

The Data-Prep Supercomputer: Offline Data Processing

Before any GPU sees a token, a second cluster must dedupe, filter, decontaminate, and tokenize trillions of tokens — undersize it and you starve the training fleet or waste GPU-hours on string processing.

GOODPUTPOWER-BOUND

What you'll decide here

  1. Whether data prep runs as a first-class, separately-sized cluster (its own CPU/storage/network profile) or is bolted onto idle GPU nodes — and what that does to your GPU goodput and your prep wall-clock.
  2. Where on the dedup spectrum you sit: exact-match only, fuzzy near-dedup (MinHash/LSH), or embedding-based semantic dedup — and how aggressive a similarity threshold you can afford before you delete the high-quality long tail.
  3. Whether quality and safety filtering is heuristic (cheap, CPU, interpretable) or classifier-based (expensive, often GPU, higher ceiling) — and who owns the false-positive rate that silently shapes your model.
  4. How rigorous your decontamination is against every eval and benchmark you will report — because the leakage you fail to remove is the leaderboard number you cannot trust.
  5. Whether tokenization and the rest of preprocessing are CPU-bound at your token volume — making cores and memory bandwidth the limiting purchase — or whether a measured GPU implementation moves the bottleneck to extraction, deduplication, classification, scratch or network before the preparation deadline.

There are two supercomputers in a frontier training program, and only one of them is GPU-dominant. The first — the one this part of the guide has been about — keeps thousands of accelerators saturated with a clean, shuffled, tokenized stream. The second runs before that, often weeks before, and its job is to manufacture the stream in the first place: ingest raw web crawls and licensed corpora, strip boilerplate, deduplicate at corpus scale, filter for quality and safety, scrub out anything that overlaps your evaluations, and tokenize the survivors into the exact binary format the loader will memory-map. This is the data-prep supercomputer, and it is the most consistently undersized, under-instrumented, and under-respected machine in the building.

The reason it gets disrespected is that it looks like ETL — and ETL is something every org thinks it already knows how to do. But web-scale data prep for a frontier run is not a nightly Spark job. It is a multi-petabyte, trillion-token batch pipeline whose stages have opposite hardware appetites — some embarrassingly parallel and CPU-bound, some metadata-storm small-file workloads, some GPU-accelerated classifier sweeps, some all-to-all shuffles that look like a network benchmark. Size it as one thing and you bottleneck on whichever stage you got wrong. This chapter walks the pipeline stage by stage and ends by sizing the cluster that runs it, which has a power, network, and storage profile deliberately unlike the GPU fleet next door.

Why prep is a goodput problem, not a side quest

The strategic case for taking prep seriously is a goodput case, and it cuts two ways. First, prep is on the critical path to the run. If your corpus takes six weeks to prepare and your GPU cluster is energized and idle for four of them, you have paid frontier-cluster lease or depreciation for a month of string processing. The prep cluster's wall-clock is a direct deduction from the training program's schedule, and the GPUs cannot start that training release until its authorized, versioned input is ready; incremental corpus releases need their own quality gates. Second, the quality of prep sets the ceiling on what the GPUs can achieve. The dominant lesson of the open-data era — FineWeb, RefinedWeb, DataComp-LM, Dolma, Nemotron-CC — is that careful filtering and deduplication beats raw scale: a smaller, cleaner corpus trains a better model than a larger, dirtier one at the same token budget. Every duplicate you fail to remove is a GPU-hour spent memorizing boilerplate; every leaked benchmark sample is a leaderboard number you have to caveat.

The decision is rarely whether to do prep at all; it is how much engineering to invest, and on what hardware. The expensive failure mode is the one nobody plans: discovering, two weeks before a launch, that the prep pipeline cannot tokenize the licensed corpus fast enough on the CPU pool you have, and either delaying the run or — far worse — borrowing GPU nodes to brute-force tokenization, which is an expensive way to run a Rust regex engine if the exact tokenizer and normalization pipeline still execute on the CPU while the accelerators sit idle.

The pipeline, stage by stage

A web-scale prep pipeline is a directed sequence of stages, each with a distinct hardware personality and a distinct fork. The canonical order — text extraction → language ID → quality/heuristic filtering → exact dedup → fuzzy dedup → classifier filtering → decontamination → tokenization → shuffle/shard — is not arbitrary: you filter cheaply before you filter expensively, and you dedup before you tokenize so you never tokenize a duplicate. The table below maps each stage to what binds it, so you can see where the prep cluster's bottleneck actually lives.

Prep pipeline stages → what binds each one
StageWhat it doesBinds onHardware fitConsequence of underspeccing
Text extractionWARC/HTML to clean text; strip boilerplateCPU + sequential read bandwidthCPU pool, streaming from object storeThroughput floor for the whole pipeline
Language ID + heuristic filtersfastText langID, C4-style rules, repetition/perplexity gatesCPU; cheap per-docCPU poolCheap to run; skipping it makes later stages pay
Exact dedupHash whole docs/lines; drop identicalMemory + hash-table I/OCPU, large RAM, fast scratchLeaves trivial dupes for fuzzy stage to find
Fuzzy near-dedupMinHash signatures + LSH banding; drop near-dupesAll-to-all shuffle + memoryCPU at scale, or GPU (RAPIDS/NeMo Curator)Signature generation, bucket skew or shuffle becomes the bottleneck
Classifier filteringModel-scored quality/safety/domain curationMeasured CPU/GPU inference throughputGPU or accelerated CPUQuality ceiling; false-positive rate set here
Decontaminationn-gram/substring overlap vs eval setsCPU + index lookupsCPU pool, in-memory eval indexBenchmark leakage; untrustworthy evals
TokenizationBPE/Unigram encode to token IDsMeasured end-to-end tokenizer throughputBenchmark CPU and GPU implementationsNormalization, regex, or encoding becomes the wall
Shuffle + shardGlobal shuffle, pack to loader formatNetwork + storage write bandwidthFast scratch + back-end networkLoader stalls or poor mixing at train time
Starting pipeline; measure the binding resource for the chosen implementation, corpus and release graph. CPU/GPU allocation follows those measurements.

Deduplication: the highest-leverage and most dangerous stage

The web mirrors, syndicates, and templates itself relentlessly. Removing those duplicates is the single biggest lever on token efficiency, because a model wastes capacity memorizing anything it sees too many times, and over-represented duplicates skew the distribution. But dedup is also a deletion from the derived training set; retained source IDs and an authorized immutable input make that exclusion reproducible and reversible, and the aggressiveness knob — how similar is "the same" — is the most consequential single number in the pipeline.

There are three tiers, and they compose. Exact dedup hashes whole documents or lines and drops identical copies; it is cheap, CPU-bound, embarrassingly parallel, and catches the trivial cases. Fuzzy / near-duplicate dedup is the real work: compute a MinHash signature per document (a compact sketch of its n-gram set), then use Locality-Sensitive Hashing (LSH) to bucket documents whose signatures collide in any band, so you only compare plausibly-similar pairs instead of the quadratic all-pairs explosion. The banding configuration sets a probability curve, not a hard Jaccard cutoff. FineWeb used word 5-grams and 14 bands of eight MinHash values; its method and per-snapshot comparison are documented by the authors in 2024. At similarity s, the candidate probability is 1 − (1 − s8)14; at s = 0.75 it is about 77%. Preserve normalization, candidate handling and survivor policy with the resulting corpus. Semantic dedup goes further: embed every document, cluster in embedding space, and drop near-neighbors that are paraphrases or translations a MinHash would miss — at the cost of an embedding-model inference pass over the entire corpus.

The deduplication fork: exact vs fuzzy vs semantic
TierMethodKillsCostWhen it is the right last stage
ExactSHA/xxHash on doc or line; drop identicalByte-identical copies and mirrored pagesLowest (CPU, near-linear)Tiny budgets; a pre-pass before fuzzy
FuzzyMinHash + LSH banding; probability curve plus explicit survivor policyNear-dupes: templated, lightly-edited, reformattedModerate; dominated by the all-to-all shuffleDefault for web-scale pretraining corpora
SemanticEmbed corpus, cluster, drop near-neighborsParaphrases, translations, semantic restatementsHighest (an embedding inference pass over everything)When a measured quality gain justifies the embedding pass and deletion risk
Compose them in order. The 'kills' column is what each tier removes that the prior tier cannot. Cost is per-token relative, at trillion-token scale.

The hardware fork inside dedup carries the biggest 2026 swing. Fuzzy dedup at trillion-token scale is an all-to-all shuffle of billions of signatures, and on a CPU cluster it is routinely the slowest stage in the entire pipeline — days of wall-clock on a large cluster. GPU-accelerated dedup (NVIDIA's NeMo Curator on RAPIDS, and the FED-class GPU dedup frameworks) collapses that. NVIDIA’s NeMo Curator deduplication stages put GPU implementations on the shortlist, but compare signature generation, bucketing, shuffle and survivor selection separately at fixed corpus, normalization and dedup policy. Keep end-to-end wall-clock, peak memory, spill and accepted output with the CPU/GPU comparison; a fast signature kernel does not time the complete corpus release. That is the rare case where the prep cluster legitimately wants a few GPUs — not to train, but to keep the dedup stage from being the long pole. You either pay for a large CPU pool to grind the shuffle, or borrow a small GPU partition to do dedup in hours and free the CPUs for tokenization.

Quality and safety filtering: heuristics vs classifiers

After dedup, you decide what is worth keeping. The fork is heuristic filtering vs classifier-based curation, and most serious pipelines use both — heuristics first because they are cheap and interpretable, classifiers where a measured quality gain justifies their cost.

Heuristic filters are the C4/Gopher/FineWeb lineage: drop documents by language-ID confidence, mean line length, fraction of alphabetic characters, repetition ratios, presence of boilerplate or blocklisted terms, perplexity under a reference model. They are CPU-cheap, fully auditable, and every threshold is a knob you can explain to a regulator or a model-behavior reviewer. Classifier filters score each document with a trained model — a fastText or transformer quality classifier (often trained to recognize "educational" or instruction-like text, as in FineWeb-Edu and the DataComp-LM curation), plus safety classifiers for toxicity, CSAM signals, and policy categories. These can capture learned quality signals beyond a hand-authored regex, but corpus-scale cost depends on the selected CPU/GPU implementation and their decisions can be harder to interpret than explicit rules; their false-positive rate silently shapes the model: a quality classifier that down-weights a dialect, a domain, or a non-English register is making a model-behavior decision, not a data-cleaning step.

For the technical stratum, A rejects 8/1,000 = 0.80% and passes this sample gate; B rejects 35/1,000 = 3.5% and fails. The crossover is 20 eligible rejections; 21 fails even if the pooled global score improves. Select A only for the illustrated sample comparison. Real corpus release remains HOLD until every required stratum, retained-corpus quality test and eligibility gate passes, with uncertainty and adjudicator disagreement reported.

Compare dedup scope and filter thresholds at fixed token budget, tokenizer and evaluation protocol; inspect source/language retention as well as model outcomes. A sample pass is not a population bound. Preserve excluded IDs and reason codes under authorized retention so the corpus can be regenerated after an approved policy change. FineWeb's controlled curation comparisons supply the experimental method; Chapter 10.10 owns policy and this chapter owns its execution.

Decontamination: the eval you can trust

Decontamination is the stage that protects the meaning of every benchmark you will ever report. If your pretraining corpus contains the test items from GSM8K, MMLU, HumanEval, or whatever you benchmark on, your model can memorize the answers and your reported scores are fiction. This is not hypothetical: MMLU questions, like other evaluation material, can enter a training corpus as exact copies or transformed questions and make a model's score reward leakage rather than generalization. Test both kinds of overlap against the versioned evaluation registry before interpreting the score. The fork is how aggressively you scrub, and against which eval sets.

The standard mechanism is n-gram / substring overlap: index every eval set you care about, then drop or flag any training document that shares a long enough contiguous span (commonly an 8- to 13-gram match) with a test item. It is CPU-bound, index-lookup-heavy, and conceptually simple. The catch is that it is necessary but not sufficient: paraphrase, translation, and reformatting slip straight past a literal n-gram match, so a corpus can pass decontamination and still be semantically contaminated. What actually works is procedural rather than algorithmic: maintain a frozen, versioned registry of every benchmark and decontaminate against the entire registry before training; when you add an eval after training, scan the frozen training corpus referenced by its manifest for overlap and, if overlap exists, mark the original score contaminated and use a clean mirror or held-out replacement; then add the eval to the registry and remove overlaps before the next training run. Get it wrong and the cost is a credibility event: an outside party reproduces your eval on a clean split and your numbers collapse.

Deep dive: why n-gram decontamination quietly fails, and what to add

The n-gram overlap method has two opposite failure modes, and a serious pipeline mitigates both. False negatives (contamination it misses): any transformation that breaks the literal token sequence — paraphrasing a question, translating it, swapping multiple-choice option order, reformatting a code prompt — defeats exact substring matching while preserving the information the model can memorize. Literal decontamination misses leakage that survives as paraphrased, translated or reformatted questions. Include those positives alongside exact copies in the overlap fixture, then report which literal, embedding-based or model-based mechanism detects each one and what it wrongly removes. False positives (clean data it wrongly drops): on multiple-choice benchmarks, unrelated questions share option boilerplate ("A) B) C) D)", common stems), so a naive n-gram match flags and deletes legitimate training documents — a quiet quality tax.

The defensible posture layers three things. (1) Long-n-gram exact as the cheap floor, tuned long enough (13-gram is a common choice) to avoid the false-positive flood from short shared spans. (2) Embedding / semantic overlap against the eval registry to catch paraphrase and translation — the same embedding infrastructure you already stood up for semantic dedup. (3) A held-out clean mirror for the benchmarks you most care about, so you can measure the contamination delta directly rather than trusting that the scrub worked. The standard n-gram script is the floor here, not the ceiling; running only it is not a defensible decontamination posture. → quality and provenance governance in Chapter 10.10.

Tokenization: benchmark the exact pipeline

Tokenization is where the prep cluster's hardware profile reveals itself, and where the most common sizing surprise lives. BPE applies ranked pair merges; Unigram selects the best path through a scored segmentation lattice. Widely deployed production tokenizers are often CPU-oriented and parallelize across documents, but GPU BPE implementations can also parallelize merge work within a document. Benchmark the exact tokenizer, normalization and regex path, document-size distribution, batch size, and available CPU/GPU implementation; those measurements — not an assumed CPU-only limit — size the tokenization fleet.

At frontier-corpus scale, tokenization is a large compute pass whose wall-clock depends on measured end-to-end throughput. The decision is whether to size a dedicated CPU pool, use a GPU implementation, or lean on a managed distributed framework (Spark, Ray, Dask, or NeMo Curator's streaming executor that overlaps CPU and GPU stages). Choose the hardware path by throughput, cost, and pipeline compatibility on the real corpus.

Sizing the prep cluster: a deliberately different machine

The prep supercomputer is sized against the opposite constraints from the GPU fleet, and the headline is that CPU-bound extraction and tokenization can make prep CPU-dense, RAM-heavy and storage-bandwidth-bound; GPU dedup or classification changes that mix, so size each stage’s compute, memory, storage and network before calling the hall power-light. A CPU-heavy prep rack can fit an air-cooled row where a dense training GPU rack needs liquid cooling; its populated CPU/GPU mix still sets the actual power and cooling envelope, and its scarce resources are CPU cores, DRAM capacity (dedup hash tables and shuffle buffers are memory-hungry), fast local scratch (NVMe for intermediate shuffle spill), and sustained sequential read bandwidth from the object/capacity tier where the raw corpus lives (→ Chapter 9.6). The network it stresses is the data-center fabric for all-to-all shuffles, not a non-blocking GPU back-end.

The two structural decisions are dedicated vs shared and co-located vs remote. A dedicated prep cluster (or a large pool of general CPU compute) decouples prep wall-clock from GPU availability — you prepare the next corpus while the current run trains — at the cost of standing up and powering a second cluster. Sharing idle GPU-node CPUs is capital-efficient but couples your prep schedule to GPU idleness and wastes accelerators. On location: prep wants to run where the data already is, because raw corpora are multi-petabyte and data gravity makes them economically immovable (cloud egress at ~$0.087–0.12/GB first-tier list price in 2026 makes moving a petabyte a ~$90–120k line item; negotiated rates lower) — so prep is the textbook case of move compute to data, the same gravity logic developed in Chapter 9.8.

Prep cluster vs GPU training fleet: the inverse profile
AxisGPU training fleetData-prep cluster
Binding resourceGPU FLOPs + scale-up bandwidthCPU cores + DRAM + read bandwidth
AcceleratorsThe entire pointMeasured benefit in dedup, classification or tokenization
Rack powerPopulated GPU rack power and cooling envelopePopulated CPU/GPU mix and qualified cooling envelope
NetworkNon-blocking back-end (IB/RoCE)Shuffle fabric; oversubscription must meet the stage deadline
Storage emphasisHot parallel FS, checkpoint write burstsCapacity/object read bandwidth + fast NVMe scratch
Job shapeSynchronous, restart-from-checkpointBatch, idempotent, restartable per-shard
Siting driverCheap firm power + cold climateWhere the corpus already lives (data gravity)
The point of the table is that almost every axis is opposite. Sizing the prep cluster with GPU-fleet instincts (liquid cooling, non-blocking fabric, accelerator-first) overspends on the wrong things and underspends on cores, RAM, and read bandwidth.
15T tokens
FineWeb corpus from 96 Common Crawl snapshots (~44 TB on disk as GPT-2 tokens)
112 hashes (14 × 8) over 5-grams
FineWeb 2024 MinHash configuration; probabilistic candidate selection and per-snapshot experiment, not a hard Jaccard cutoff
Scope & caveats

Word 5-grams and MinHash banding; 0.75 is a similarity target, not a hard cutoff. Collision probability 1−(1−s^8)^14 is about 77% at s=.75. Per-snapshot dedup is an experiment result for that corpus, not a universal policy.

~50%
of already-filtered RefinedWeb documents removed by fuzzy+exact dedup (~12% of original Common Crawl documents) — dedup runs after language/quality filtering
Scope & caveats

corpus dedup removal share

8–13-gram
typical contiguous-span threshold for n-gram benchmark decontamination overlap matching
~29%
estimated MMLU contamination across public web corpora; clean-mirror retests drop scores high-single to low-double digits
Deep dive: a worked prep budget, and where it goes wrong

Compute each stage as the maximum of its overlapped compute, read, shuffle and write times; then sum the sequential stages. Extraction offers 128 × 64 × 0.010 = 81.92 GB/s of assumed CPU capacity, so the 64 GB/s source is the first bottleneck. Tokenization is the longest stage at the assumed per-core rate. The table carries the trace instead of assuming every stage has the same hardware personality.

The six unrounded durations sum to about 26 h; adding 25% contingency yields about 33 h, inside the 36-hour release window. Reserve this resource mix for qualification. Flip: a 24-hour deadline fails even before contingency; the current configuration needs at least 33 h at the displayed precision. Re-profile and expand the limiting stages or change the release scope; assigning extra GPUs to an unaccelerated CPU stage does not meet the date.

Signatures use 109 × 112 × 8 B = 896 GB; ×2.5 for index allocation gives 2.24 TB. Per worker, 2,240/128 + 48 + 16 = about 82 GB, below 128 GB. Check the largest duplicate bucket: the uniform average is not a skew guarantee. Scratch needs 200 + 400 = 600 TB; after one worker loss, 127 × 8.0 × 0.75 = about 760 TB is available. About 160 TB of extra simultaneous spill would exhaust that surviving allocation; exceeding the memory or scratch bound requires repartitioning or a revised schedule.

Publish only an accepted output manifest: source snapshots and eligibility decisions, stage/container hashes, normalization and tokenizer versions, dedup/filter policy, eval-registry hash, object checksums, shard boundaries, token counts and shuffle seed. Rerun a failed partition and require the declared IDs and contents to reproduce, or publish an explicitly new version. The FineWeb experiments motivate measuring curation policy rather than assuming quality follows removal rate; they do not supply this cluster's rates. Hand the accepted manifest to Chapter 9.5 for real loader/restart qualification. Schedule and hardware remain HOLD until the assumed rates and worst buckets are tested.

The stage ledger below exposes which resource sets each duration. Its rates are qualification targets from the stated assumptions; a faster isolated kernel changes the release date only if it changes the limiting term.

Preparation case: stage-by-stage wall-clock trace
StageTrace and limiting resourceDuration
Extraction / early gates1,000,000 GB / 64 GB/s / 3,600 s/h; CPU ceiling 81.92 GB/s; output 200,000/32 sAbout 4.3 h; source binds
Dedupmax(6.0 h compute, 400,000 GB / 40 GB/s / 3,600, 200,000/32/3,600 read, 150,000/32/3,600 write)6.0 h; compute binds
Classification10^9 documents / (32 × 2,000 documents/s) / 3,600; 150 TB input and 100 TB output at 32 GB/s clear soonerAbout 4.3 h; inference binds
Decontaminationmax(2.0 h compute, 100,000/32/3,600 read, 99,000/32/3,600 write)2.0 h; compute binds
Tokenization10^13 tokens / (8,192 × 40,000 tokens/s) / 3,600; 99 TB read at 32 GB/s and 40 TB write at 20 GB/s clear soonerAbout 8.5 h; tokenizer binds
Shuffle / shardmax(1.0 h compute, 80,000/40/3,600 shuffle, 40,000/20/3,600 write, 40,000/32/3,600 read)1.0 h; compute binds
Use unrounded intermediate values when summing; add the single 25% schedule contingency afterward. Rates and yields are stated assumptions in this one case.
about 33 hderived
Release time in the stated sequential preparation case; illustrative assumptions
Scope & caveats

Illustrative inputs; qualify project rates and limits before selection.

Anti-patterns

The same prep failures recur, because each comes from treating prep as ETL rather than as a sized supercomputer:

  • Tokenizing without implementation benchmarks. Sizing a CPU pool or borrowing GPU nodes before benchmarking the exact normalization, regex, and tokenizer pipeline. The wrong hardware mix leaves either CPU cores or accelerators idle.
  • Treating a post-training scan as a repair. When a new eval overlaps the frozen training manifest, mark the original score contaminated and use a clean mirror or held-out replacement; add the eval to the registry and remove overlaps before the next training run.
  • Treating dedup threshold and classifier operating point as hygiene, not hyperparameters. Picking a similarity threshold and a quality-classifier cutoff because they sounded reasonable, never ablating them, and discovering a quality regression or a silent capability gap after a frontier run. Global dedup that deletes the good recurring web is the canonical case.
  • Sizing the prep cluster with GPU-fleet instincts. Liquid cooling, a non-blocking back-end, accelerator-first — buying that package before measuring the prep stages can overspend on idle GPUs while starving the cores, DRAM, scratch and read bandwidth the CPU-bound stages need. Derive each resource from the complete pipeline budget. → Chapter 9.8.

Choose the prep resource mix by the stage that controls delivery, and the curation policy by accepted output quality at fixed training budget. Keep corpus release behind both gates: finishing early with an untraceable or over-filtered dataset does not feed the intended run. Under-sizing memory or scratch delays the program; accepting an untested exclusion policy spends the GPU budget on a different corpus than the one the program meant to train.

Offline prep produces what the runtime loader streams — the online data path is engineered in Chapter 9.5, and the capacity/object tier the corpus lives on is Chapter 9.6. The data-gravity and move-compute-to-data logic that sites the prep cluster is developed in Chapter 9.8; the checkpoint write path that shares the prep cluster's restartable, idempotent batch discipline is Chapter 9.4. What you are legally permitted to put in the corpus — provenance, licensing, opt-outs, PII — is the regime enforced mechanically by these filters, treated in Chapter 10.10. The economics of running this second cluster as a distinct power and capital line item connect to the on-site generation strategy in Chapter 3.5.
Cite this chapter
Fehn, J. (2026). The Data-Prep Supercomputer: Offline Data Processing (Chapter 9.9). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-9-storage-and-data/9-9-the-data-prep-supercomputer-offline-data-processing (accessed 2026-09-29).
@misc{aidc-9-9,
  author       = {Fehn, Jacob},
  title        = {The Data-Prep Supercomputer: Offline Data Processing (Chapter 9.9)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-9-storage-and-data/9-9-the-data-prep-supercomputer-offline-data-processing},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit