Chapter 9.6
In this chapter · 7 sections
Object Storage, Data Lakes & the Capacity Tier
Object storage holds the whole corpus, every checkpoint lineage, and every shipped model; its cost turns less on vendor choice than on whether you build it as a flash-fronted serving layer the GPUs read from.
What you'll decide here
- Whether your capacity tier is on-prem object, cloud object, or a hybrid — and therefore who owns the egress bill, the designed durability and protection model, the separate availability SLA/service-credit terms, and the data-gravity consequences once the petabytes land.
- Whether the bucket is HDD-backed nearline (low media cost per TB, mechanical seek latency) or flash-fronted / all-QLC (a price premium for random reads and density) — the measured object throughput and tail latency decide whether it feeds training directly or stages through a cache, and reuse decides whether that cache earns its cost.
- Which lakehouse table format (Iceberg, Delta, Hudi) governs the data lake, given that the ecosystem has converged on Iceberg as the interop standard and the wrong bet strands your catalog and multi-engine access — pin the exact reader, writer and catalog releases, because every choice leaves someone owning compaction, snapshot expiry and catalog recovery.
- The durability and protection scheme — replication vs erasure coding, single-AZ vs multi-AZ, and the immutability/object-lock posture — because extra copies, cross-zone writes and independent recovery cost different amounts and protect against different losses; replica count alone does not stop corruption or deletion.
- Whether object is also your inference cold tier and model-distribution backbone — because if it is, its read latency and fan-out throughput stop being archive concerns and become cold-start and time-to-first-token concerns.
Every tier above this one — node-local NVMe scratch, the parallel file system, the checkpoint store — is sized to a working set. Object storage is sized to everything. It is the only layer in the building that holds the full training corpus, the complete checkpoint lineage across every run, the model registry, the eval artifacts, the inference logs, and the cold copies of weights waiting to be paged onto a GPU. Because it holds everything, it is the largest tier by capacity and usually the cheapest per terabyte — and that combination breeds a dangerous instinct: to treat it as a passive archive that sits behind the "real" storage and never touches the GPUs. In the 2026 AI data center that instinct is wrong, and acting on it is expensive.
This chapter is built around one shift: object storage stopped being the cold tier and became a serving tier. Modern data loaders stream Parquet and WebDataset shards directly from S3-compatible buckets into the training loop; model servers page multi-hundred-gigabyte weight files out of object storage on cold start; lakehouse engines query Iceberg tables in place without a copy. The moment a GPU is waiting on an object GET, the bucket's read latency and aggregate throughput are no longer archive metrics — they are goodput metrics, and an under-provisioned capacity tier idles the most expensive hardware in the building. The forks ahead are on-prem vs cloud, HDD vs flash, replication vs erasure, and archive vs serving, each with a downstream cost. The data-loader path that consumes this tier is Chapter 9.5; the file-system tier that sits above it is Chapter 9.2; the sizing math that ties them together is Chapter 9.8.
What object storage is for in an AI fleet
Object storage is defined by three properties that make it attractive as a capacity tier, while the implementation’s delivered reads and request tails decide whether it can also serve the performance tier: a flat key-value namespace (no directory tree, no POSIX semantics, no in-place updates — you replace an object whole, by key, though a reader can GET any byte range within it), massive horizontal scale (object count and namespace throughput scale within the chosen service’s documented limits), and an HTTP/REST access path (the S3 API is the de-facto wire protocol the entire ecosystem speaks). The cost of those properties is latency and small-object overhead: a GET carries request setup, and a workload of millions of tiny objects pays per-request cost that crushes throughput. Range GETs are what let a reader fetch only the bytes it needs — a Parquet reader pulls the footer, then just the column chunks in play — but each range is still a priced, latency-bearing request, which is why the data-loader layer packs many small samples into large sequential shards (WebDataset tar, Parquet, MDS) before they ever reach the bucket — the format choice in Chapter 9.5 exists largely to make object storage fast.
In the AI lifecycle, object storage carries five distinct payloads, and they have different access patterns even though they share a bucket. The training corpus is read-heavy, large-sequential, fan-out across thousands of GPUs. Checkpoints are write-burst (incast) then mostly cold, until a restart makes one of them suddenly latency-critical (Chapter 9.4). The model registry is small in object count but each object is huge and read on cold start. Eval and inference logs are append-heavy, queried analytically. And the data lake is the governed, queryable view over the raw corpus. Sizing the bucket to the average of these patterns guarantees you under-serve the peak of each one.
On-prem object vs cloud object
The first and least-reversible decision is where the petabytes physically live, because data has gravity: once a multi-petabyte corpus lands somewhere, moving it costs money, time, and egress fees, and every workload that touches it is pulled toward it. The choice is not really "AWS vs MinIO" — it is a choice about who owns the durability design and the separate legal availability/service commitment, who pays for reads, and how locked-in the data gravity makes you over the asset's life. The data-gravity economics are developed quantitatively in Chapter 9.8.
Cloud object (S3, GCS, Azure Blob, and the regional sovereign clouds) buys you a published durability design target distinct from the service-availability SLA, both specific to the chosen storage class, elastic capacity within service limits, and no operator-owned storage hardware to buy, at the price of a per-GB-month storage rate, per-request charges, and — the line that quietly dominates AI economics — egress fees when bytes leave the cloud on a billed route, plus the selected class’s request and retrieval charges. If your GPUs are also in that cloud, a same-region service path can avoid transfer charges; cross-region and other billed paths still incur transfer or retrieval costs, so price the exact service boundary; if your GPUs are in a colo or self-build and your data is in the cloud, you are paying to feed your own training loop, and the data gravity makes it progressively harder to leave. On-prem object (MinIO, Ceph RADOS Gateway, VAST, Pure, Cloudian, Dell ECS, and the parallel-FS vendors' native S3 endpoints) inverts the trade: you own the capex and the operational burden, you set your own durability scheme, and — critically — there is no egress fee to feed a co-located GPU fleet. For a durable, large, self-hosted training operation the on-prem bucket is usually the lower-TCO floor; for spiky, multi-region, or cloud-native inference the managed bucket wins on elasticity and reach.
| Dimension | On-prem object | Cloud object | Hybrid (cloud-DR / burst) |
|---|---|---|---|
| Capex / opex | High capex, low marginal opex | Zero capex, per-GB-month + per-request opex | Mixed; pay cloud only for the overflow/DR copy |
| Egress to GPUs | None if GPUs co-located | Internal-cheap in-cloud; expensive cross-cloud/colo | Egress only on the spillover path |
| Durability design / service commitment | You design it (EC + replication) | Published durability design target; availability SLA and credits are separate | Off-site copy’s replication lag, loss boundary and recovery service |
| Elasticity | Bounded by what you racked | Elastic within service quotas, request ramp and data-placement limits | Cloud absorbs the burst |
| Data-gravity lock-in | Locks workloads to your site | Locks workloads to that cloud + egress moat | Splits gravity; highest operational complexity |
| Best-fit | Durable, large, co-located self-build | Spiky, multi-region, cloud-native serving | On-prem primary + cloud archive/DR/overflow |
S3-over-flash: the capacity tier learns to serve
The second fork is the one that has changed most since 2024: what media sits behind the bucket. Classic object storage is HDD-backed — cheap per terabyte, dense, and slow, with ~4 ms seek latency and ~hundreds of MB/s per spindle. That is fine for a true cold archive and fatal for a serving tier whose loader deadline it misses; test aggregate sequential demand and request tails before choosing HDD, flash or a cache. The new pattern is S3-over-flash: an S3-compatible namespace whose hot data lives on NVMe SSD, delivering single-digit-millisecond GETs and an order-of-magnitude higher throughput, so the data loader can stream training shards directly from object storage at line rate instead of staging them through a separate file system first.
The cloud made this concrete with S3 Express One Zone: up to ~10x faster data access and up to ~80% lower request cost than S3 Standard, at single-digit-ms latency — but bought by collapsing to a single availability zone, which trades away the multi-AZ durability that defines S3 Standard. That is the flash fork in miniature: you pay for speed in dollars (storage roughly $0.11/GB-month, several times the Standard rate) and in durability blast-radius (one AZ, not three). On-prem, the same shift is driven by high-capacity QLC SSDs — the 245.76 TB class is shipping as of mid-2026 (Micron 6600 ION from May; Kioxia LC9 qualified in Dell ObjectScale at 9.83 PB raw per 2U; Solidigm's 245 TB still an end-2026 target; a 512 TB DapuStor prototype exists but is not a product), at ~13.7 GB/s sequential read on the newest Gen5 parts (the fleet-typical Gen4 figure in this chapter's table is ~7 GB/s) / ~3 GB/s write versus ~300 MB/s for an HDD and 20–100 µs latency versus ~4 ms — density and watts-per-terabyte that let an all-flash object tier out-pack an HDD rack several-fold. The economics, though, reversed during the NAND shortage: a 30 TB QLC SSD ran ~15× the price of a 30 TB nearline HDD in Aug 2026 (VDURA index), so the QLC-vs-HDD crossover story of 2025 is on hold — the fork is once again cold-on-HDD unless the latency or density case pays for the flash premium.
| Property | HDD nearline object | Flash-fronted / all-QLC object |
|---|---|---|
| Per-TB cost | Lowest ($/TB floor) | ~15x 30 TB-class nearline HDD (Aug 2026 VDURA index; supply-sensitive) |
| Read latency | ~4 ms (seek-bound) | 20–100 µs device; single-digit-ms at the bucket |
| Sequential throughput | ~300 MB/s per spindle | ~7 GB/s per drive (≈23x) |
| Density per rack | High (30+ TB HDDs) | Equal or higher (122–245 TB QLC) |
| Watts per TB | High — many spindles | Low — fewer, denser devices; less storage energy per usable TB (PUE is a facility ratio and need not improve with it) |
| Feeds GPUs directly? | Only if aggregate sequential throughput and request latency meet the loader's measured demand (1,024 GPUs × 20,000 tokens/s × ~4 B/token ≈ 82 MB/s); stage for random access, bursts, or higher rates | Yes — stream shards into the loop at line rate |
Lifecycle and tiering: paying archive prices for archive data
Not every object earns flash. The corpus shard read every epoch does; the checkpoint from a run that finished four months ago does not; the inference log from last quarter belongs in deep archive. Lifecycle and tiering policies are the mechanism that keeps the average cost low while the hot path stays fast — they automatically demote objects down a temperature ladder (hot flash → warm standard → cool infrequent-access → cold/archive) as access frequency drops, and the cloud's intelligent-tiering variants do this without an explicit policy by watching access patterns.
The subtlety is in retrieval. Deep-archive classes are cheap to store but slow and sometimes expensive to read back — fine for a compliance copy, disastrous for a checkpoint you might restart from. The anti-pattern is tiering by age alone: a six-month-old checkpoint looks archival until a model regression makes it the one you must restore now, and a multi-hour glacier retrieval becomes the long pole in your recovery. Tier by access likelihood and recovery-time requirement, not by calendar age — and keep anything on a recovery path in an instant-retrieval class even if it is cold.
Data lakes and the lakehouse: why this guide chooses Iceberg
A bucket full of Parquet files is a data lake; a bucket full of Parquet files with a transactional metadata layer that gives you schema evolution, time travel, ACID commits, and partition evolution is a lakehouse. The metadata layer is the open table format, and for AI fleets it matters because the same governed corpus must be queryable by the data-prep pipeline, the training loader, the eval harness, and the analytics engine — without copying it four times. The format is the contract that lets many engines read one copy of the data in place.
This guide chooses Apache Iceberg as the interoperability standard for a new AI fleet. AWS gives Iceberg first-class support, and S3 Tables use Iceberg; Athena's general CTAS default remains Hive. Snowflake added native Iceberg tables; Databricks acquired Tabular (Iceberg's creators) and uses Delta Lake UniForm to expose Iceberg-compatible metadata so a single Delta table reads as either format. Iceberg v3 deliberately aligns deletion semantics, file layout, and row tracking with Delta so one copy of data serves both. So the call is straightforward: pick Iceberg unless you are all-in on the Databricks/Delta stack, in which case Delta-with-UniForm gives you interop anyway; Hudi earns the pick where streaming upserts and CDC merges dominate the write path and you are prepared to own its compaction and table-services operations. The one condition on the choice: name the writer, reader and catalog releases before committing the corpus, because UniForm exposes read metadata without authorizing arbitrary Iceberg writers to mutate a Delta table, and schema evolution, deletes and partition changes must work through every required engine or the shared copy becomes another conversion pipeline. Betting on a proprietary or orphaned format is the expensive mistake — it strands your catalog and forecloses multi-engine access, which is the entire reason to run a lakehouse. The governance and lineage regime that sits on top of the lake is Chapter 10.10.
| Format | Select when | Qualify before adoption | Maintenance consequence |
|---|---|---|---|
| Iceberg | Default for a new AI fleet: read in place by AWS (S3 Tables, Athena, EMR), Snowflake native tables and Databricks via UniForm; v3 aligns deletes and row lineage with Delta | Concurrent commit, schema/partition evolution, delete handling and catalog restore on the pinned engine releases | Expire snapshots and remove orphan files only after retention and reader checks |
| Delta | You are all-in on the Databricks/Delta stack; UniForm publishes Iceberg-readable metadata, so the interop comes anyway (reads only, not arbitrary Iceberg writers) | Protocol features and every UniForm reader; keep writers on the authorized transaction path | Coordinate log retention, cleanup and generated read metadata |
| Hudi | Streaming upserts and CDC merges dominate the write path and you will own compaction and table services | Selected copy-on-write or merge-on-read query behavior, concurrency and delete propagation | Budget compaction/cleaning against foreground ingest and query latency |
Before enabling cleanup, restore a catalog snapshot together with the data and delete files it references into an isolated namespace. Verify the same record IDs through each authorized reader. Compaction needs space for old and replacement files to coexist; snapshot retention can keep both live. Expiring a snapshot and removing orphan files are different operations, and an age threshold shorter than an active writer or reader can destroy required data. For a required deletion, trace the record through table snapshots, source objects, training shards, retrieval indexes, replicas and caches; record the completion watermark required by the data owner. Hold release when any required consumer still returns it. Method sources: Iceberg maintenance, Delta UniForm and Hudi table types.
Durability and protection: the same eleven nines, bought different ways
"Eleven nines" is a marketing number until you ask how it is achieved, because the same durability target costs very differently depending on the protection scheme. The two primitives are replication (store N full copies; simple, fast to repair, but N-x the raw capacity) and erasure coding (split each object into k data + m parity shards spread across failure domains; survives m losses at a fraction of the overhead of replication). Erasure coding can buy substantial media savings at large corpus scale: triple replication of an illustrative 10 PB corpus needs 30 PB before reserve, with 20 PB buying protection rather than new logical capacity. EC reduces that overhead according to its code, but spends CPU on encoding and read reconstruction, and repairs touch multiple nodes. Compare the saving with foreground service: the extra replica capacity is not automatically wasted, and EC earns the choice only if foreground service still passes during the required failure. The failure domain the shards span is the real durability lever: count shards lost to each drive, node, rack or zone event. Reconstruction requires at least k surviving shards of the object; merely spanning zones does not establish zone survival. S3 Standard's eleven nines come from multi-AZ erasure coding plus continuous background integrity scanning (checksums, auditors, automated re-replication); S3 Express One Zone deliberately gives up the multi-AZ span for latency, which is why its durability blast-radius is a single AZ.
The two failure modes durability marketing hides are silent data corruption and operator error. Bit rot is caught by end-to-end checksums and background scrubbing — verify your object store does both, because a corrupted training shard poisons a run silently. And eleven nines of durability is no protection against an rm, a buggy lifecycle rule, or ransomware: that is what object lock / immutability (WORM) and versioning are for. For checkpoints and the model registry, an immutability window plus versioning is the difference between a recoverable incident and a destroyed lineage; the SDC-detection discipline connects to fleet health in Chapter 10.6.
Worked decision: test the object layout against the lost zone
An (8,3) code stores 11/8 = 1.375 times the logical bytes and reconstructs from any eight shards. The illustrative 1.0 PB corpus therefore needs about 1.4 PB encoded; allowing 25% free raw space requires 1.0 × 11/8 / 0.75 ≈ 1.8 PB raw. Select a supported hardware configuration above the unrounded requirement. That reserve is not extra promised data capacity. Triple replication would instead require 1.0 × 3 / 0.75 = 4.0 PB raw, with different placement and repair behavior.
Test the listed shards, not the zone count. Losing A removes s0–s3: seven survive, so reconstruction fails. Losing B also leaves seven; losing C leaves eight. The proposed 4/4/3 layout is rejected for any-one-zone recovery even though it spans three zones. Flip: after moving s3 and s7 to D, losing A, B or C leaves eight shards; losing D leaves nine. The crossover is a maximum of three shards in any lost zone. This revised layout passes the mathematical read-survival test, with no extra parity margin after a three-shard loss.
New writes still fail when only eight remain because the assumed write gate requires nine. If uninterrupted writes are required after every zone loss, the 3/3/3/2 proposal remains HOLD; redesign code and placement with the implementation owner instead of silently lowering the write threshold. Use the mapping tool on the deployed object set, confirm the declared host/rack/zone hierarchy, and replay reads with repair running and the chosen zone absent. A repair cannot restore the promised zone placement into a full or disallowed destination. Choose the least costly placement that passes the required read/write state, then qualify its repair traffic through Chapter 9.8; this chapter owns the placement test.
Object as the inference cold tier and model-distribution backbone
The last role retroactively justifies treating object storage as a serving tier rather than an archive: object storage is where models live between requests and how they reach every GPU in the fleet. A model server on cold start pages a multi-hundred-gigabyte weight file out of the bucket onto the GPU; an autoscaler spinning up a new replica reads the same weights; a fleet-wide model rollout fans the same objects out to thousands of nodes at once. When that happens, the bucket's read latency becomes cold-start latency and its aggregate fan-out throughput becomes the ceiling on how fast you can scale a model up to meet a traffic spike — both of which are user-facing inference SLO concerns, not archive concerns.
This is why the inference-era capacity tier is increasingly flash-fronted and why the new inference memory hierarchy (Chapter 9.7) has an object/Ethernet-flash layer explicitly at its base: the KV-cache and weight tiers above it page to and from object storage, and a slow bucket throttles cold-start and cache rehydration. If object storage is your model-distribution backbone, you size it for the fan-out burst of a fleet-wide rollout — thousands of simultaneous reads of the same large objects — not for steady-state archive throughput. Under-provision it and your cold-start time-to-first-token and your ability to absorb a traffic surge both degrade. The model-serving and cold-start mechanics live in Chapter 9.7; the read-amplification of a synchronized fan-out is a fabric co-design problem shared with Chapter 8.5.
Deep dive: when to stage through a file system vs read object directly
The recurring architecture question is whether the training loop reads directly from the object bucket or whether object storage feeds a faster file-system or NVMe cache tier that the GPUs actually read. The answer turns on three variables. First, object media and service: an HDD-nearline bucket that cannot feed the loader’s read rate and request tails must stage through a faster tier; a flash-fronted / S3-over-flash bucket can be read directly when it passes those same limits. An HDD pool can also feed a modest packed-token stream directly. Second, reuse: if the same shards are read every epoch for weeks, a one-time copy into a node-local NVMe or parallel-FS cache amortizes cheaply and removes the bucket from the steady-state hot path; if data is touched once (streaming, single-pass), staging is wasted copy and direct read wins. Third, working-set fit: if the corpus fits in the cache tier you stage once and forget the bucket; if it dwarfs the cache you stream, and the bucket's sustained throughput is the binding constraint.
The 2026 default for large pre-training is a caching architecture: object storage as the durable, capacity-priced source of truth, fronted by a flash cache (parallel FS or distributed cache like the Meta RSC's 46 PB flash cache) that absorbs the per-epoch reads, with the loader prefetching ahead of the GPU. For single-pass streaming and for inference cold-start, the flash-fronted bucket is read directly. The wrong call — staging a single-pass stream, or reading an HDD bucket directly into a training loop whose delivery budget it cannot meet — shows up immediately as data-loader stall and collapsed GPU utilization. The loader-and-cache mechanics are Chapter 9.5; the NVMe and GPUDirect path is Chapter 9.3.
Choose the bucket, media and table format that meet both the active read path and the retained-data recovery contract. Keep cold data on cheaper media when it meets the service window; buy flash where measured request tails or fan-out require it. An erasure-code ratio saves raw capacity only after the actual placement, quorum and repair test passes. A shared lakehouse saves copies only when every required engine sees the same committed records and deletions.
Scope & caveats
Illustrative inputs; qualify project rates and limits before selection.
Cite this chapter
Fehn, J. (2026). Object Storage, Data Lakes & the Capacity Tier (Chapter 9.6). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-9-storage-and-data/9-6-object-storage-data-lakes-and-the-capacity-tier (accessed 2026-09-29).
@misc{aidc-9-6,
author = {Fehn, Jacob},
title = {Object Storage, Data Lakes & the Capacity Tier (Chapter 9.6)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-9-storage-and-data/9-6-object-storage-data-lakes-and-the-capacity-tier},
note = {Accessed 2026-09-29}
}