Chapter 9.3
In this chapter · 6 sections
NVMe Tiers, GPUDirect Storage & the CPU-Bypass Data Path
Whether a GPU stays fed can turn on per-request latency before aggregate bandwidth; a supported direct path routes bytes from flash straight into HBM without a host-memory bounce buffer, but the measured workload and surviving path decide whether that bypass earns its cost.
What you'll decide here
- Where each tier of the media hierarchy lives — node-local NVMe scratch vs networked all-flash vs an Ethernet-attached context tier — and therefore which I/O personality (ingest, checkpoint, KV-cache) each path is sized for.
- How far you bypass the CPU: legacy buffered POSIX, GPUDirect Storage (data path off the CPU), or 2026-class GPU/DPU-initiated I/O (SCADA, BlueField-4) that takes the control path off the CPU too.
- TLC vs QLC per tier — endurance and write-bandwidth headroom for scratch and checkpoint vs $/TB density for the read-mostly capacity and context tiers.
- The PCIe generation and SSD form factor (E1.S vs E3.S) you commit the storage node to — which fixes per-drive bandwidth, drive count, and whether the storage shelf needs liquid cooling.
- Where storage rides — dedicated links, a shared compute fabric, or a provisioned front-end/storage network — and which supported NVMe/RDMA over InfiniBand or RoCE, or NVMe/TCP over Ethernet, meets the CPU, I/O and recovery budgets. Keep OOB management separate.
Chapter 9.2 chose the file system. This chapter chooses the data path underneath it — the physical and software route a byte travels from flash to a GPU's HBM — and that path, not the file system's headline aggregate, is what determines whether the accelerators stay fed. The reason is structural. A parallel file system can advertise multiple terabytes per second across a cluster and still starve an individual GPU, because the classic I/O path routes every read through the host CPU: the NIC DMAs data into a kernel bounce buffer in system DRAM, the application stages it in host memory, and the GPU DMA engine transfers it into device memory. Each hop adds latency, consumes CPU cycles and DRAM bandwidth, and — at the small-request, high-concurrency access patterns that dominate AI — leaves the PCIe bus underutilized while the GPU spins. When the trace shows host staging limiting delivery, the file system is not the bottleneck; the bounce buffer is.
The engineering that follows is a sequence of bypasses, each deleting a piece of the host from the path and each a genuine fork with a downstream cost. First, the media hierarchy: how many tiers of flash, where they sit, and which I/O personality each absorbs. Then GPUDirect Storage (GDS), which takes the data path off the CPU. Then the 2026 frontier — SCADA and BlueField-4 — which take the control path off the CPU too, letting the GPU or a DPU initiate I/O with the host processor entirely out of the loop. Underneath all of it, the silicon decisions — TLC vs QLC, PCIe 5.0 vs 6.0, E1.S vs E3.S — and the fabric question of where the storage traffic rides. Each carries its cost in goodput, watts, or stranded capacity.
The media hierarchy: four tiers, four personalities
AI storage is not one medium but a stack, and the stack exists because the four I/O personalities of Chapter 9.1 — sequential ingest, bursty checkpoint writes, metadata-heavy small-file access, and latency-critical KV-cache retrieval — pull in incompatible directions. No single tier serves all four economically. The hierarchy, from hottest to coldest:
Node-local NVMe scratch is the closest persistent medium to the GPU: one to a handful of NVMe drives inside the GPU server itself, on the host PCIe complex, contributing no network traffic at all. It absorbs the staleness-tolerant, node-private workloads — the local checkpoint tier (write to local flash in seconds, drain to the global store asynchronously), data-loader shuffle buffers, and spill space. Its usable bandwidth is bounded by the populated drives, PCIe tree, controller and application concurrency, and its great virtue is that it never touches the storage fabric. The cost of skipping it: every checkpoint and every shuffle becomes networked traffic that collides with training collectives. → Chapter 9.4 treats the local-NVMe checkpoint tier in depth.
Networked all-flash (the hot tier) is the parallel/distributed file system itself — WEKA, VAST, DDN/Lustre, IBM Storage Scale — built on TLC NVMe when its write endurance and sustained service fit the job, and reached over RDMA when the filesystem and client support that path. This is the tier sized to the per-GPU read target and the checkpoint write budget; it is where 9.2's file-system choice lives. QLC capacity flash (the warm tier) trades endurance and write speed for density: read-mostly datasets, model registries, and the all-flash data lake that replaces HDD for active data when its read rate and rack density justify the flash premium. The Ethernet-attached context tier is the 2026 newcomer — petabyte-scale flash addressed as inference KV-cache memory (NVIDIA CMX/ICMSP-class), a tier that did not exist in the training-only storage stack and that Chapter 9.7 treats as a memory-hierarchy problem rather than a file-system one.
| Tier | Primary personality | Medium | Reach / placement | Sized for |
|---|---|---|---|---|
| Node-local NVMe scratch | Local checkpoint drain, shuffle/spill | TLC NVMe (high-endurance) | Inside the GPU server (off-fabric) | Seconds-to-local write; async drain to global |
| Networked all-flash (hot) | Sequential ingest + checkpoint writes | Qualified NVMe media; supported RDMA or TCP path | Dedicated storage links or qualified shared fabric | Measured reads; unique state / drain window |
| QLC capacity (warm) | Read-mostly datasets, model registry | QLC NVMe + data reduction | Networked; same namespace or object | $/TB density; all-flash data lake |
| Ethernet-attached context | Inference KV-cache offload | QLC/TLC flash as 'memory' | Ethernet-attached, shared per GPU pod | KV reuse; complete transfer within TTFT budget |
| Object / archive (cold) | Durable corpus, cold checkpoints | QLC or HDD; erasure-coded | Object store; on-prem or cloud | Durability and $/TB; egress economics |
Map personality to tier, and resist the temptation to collapse them. Putting checkpoint writes on a QLC capacity tier whose rated host writes or full-drive drain rate are insufficient burns its endurance or saturates its write path; size and test the pool after a failed host; serving KV-cache from the cold object tier blows the latency budget. The hierarchy answers a sizing question, not a procurement convenience — and the placement column is where it intersects the fabric decision at the end of this chapter.
GPUDirect Storage: taking the data path off the CPU
GPUDirect Storage is the first and most established bypass. It establishes a direct DMA path between storage and GPU memory, eliminating the bounce buffer in host DRAM: the NIC (or local NVMe controller) DMAs data straight into HBM across the PCIe root complex, skipping the CPU-mediated copy entirely. The host CPU still initiates the transfer — it owns the control path, issuing the I/O and managing the file system — but it is no longer in the data path. The payoff is threefold: lower latency (one fewer copy), reclaimed CPU cycles and DRAM bandwidth, and far higher effective bandwidth on the small, concurrent requests that defeat the buffered path. Published GDS results are configuration aggregates, not a promise for one GPU behind one link. A 200 Gb/s link has a 25 GB/s decimal line-rate ceiling before protocol overhead; any aggregate result above one link's payload ceiling must name the drive, link, PCIe and concurrency boundary. Benchmark payload GB/s and GPU stall time end to end on the specified topology.
Adopting GDS is not free, and the failure mode is concrete. It requires an end-to-end supported stack — a GDS-aware file-system client, a supported NIC, a clean PCIe topology (the GPU and the NIC/NVMe should sit under the same PCIe switch or root complex so the DMA does not traverse the CPU's inter-socket link), and applications that issue I/O through the cuFile API rather than ordinary POSIX read(). Miss a required GDS condition and the stack can fall back to compatibility-mode host staging when configured to allow it — you pay for GDS-capable hardware and get a bounce buffer. Other unsupported paths fail instead; eligibility checks and per-process transfer counters must establish which path actually ran. The fork here is whether your file system and data loaders are GDS-native: a WEKA or VAST or GPFS deployment with a GDS client and a DALI/cuFile loader gets the bypass; a generic NFS mount with a PyTorch loader doing buffered reads does not, no matter what NICs you bought. → loader integration in Chapter 9.5.
Prove the bypass with the installed tools: retain gdscheck.py -p eligibility output, run a checksum-verified disposable read through GDSIO's direct and host-staged transfer modes, and collect gds_stats for that process. Confirm flags with the installed version's help. Inspect direct operations, compatibility/POSIX fallback, routing and errors, then repeat on the real loader; a GDSIO result does not prove a buffered application used cuFile. Compare delivered bytes and step wait at identical I/O sizes and concurrency. NVIDIA GDS diagnostics distinguish path support from path use.
The 2026 frontier: SCADA and BlueField-4 take the control path too
GDS removed the host from the data path but left it owning the control path — the CPU still decides what to read and when. The 2026 generation removes the control path as well, and it does so along two distinct routes that are easy to conflate but architecturally different.
SCADA (GPU-initiated I/O) pushes the storage control path onto the GPU. Where GDS kept the CPU as the initiator, SCADA lets the GPU itself issue and control storage I/O operations directly — the host processor is fully out of both the data and control loops. The motivating problem is inference: a GPU sustaining more than a thousand parallel threads cannot wait on a CPU to orchestrate millions of tiny reads against a multi-petabyte context dataset. Wiwynn's SCADA-class reference server pairs this with extreme media density — 96 liquid-cooled E3.S drives for roughly 2.9 PB in a single server on PCIe 6.0 — precisely because the GPU-initiated path can finally drive that many drives without a CPU bottleneck in the way.
BlueField-4 (DPU-offloaded storage) takes the opposite route: it moves the storage, networking, and security control plane onto a DPU sitting between the network and the GPU, so the host CPU is bypassed by offload rather than by GPU initiation. NVIDIA's BlueField-4 STX architecture, announced at GTC 2026, is built around a storage-optimized BlueField-4 DPU (announced at GTC Washington D.C., Oct 2025; 800 Gb/s and roughly 6x the compute of BlueField-3) plus a ConnectX-9 SuperNIC, routing data through a dedicated accelerated-storage layer via RDMA over Spectrum-X Ethernet. WEKA, VAST, and DDN are building platforms on it for H2 2026 availability. The same silicon underpins the CMX context-memory tier for inference KV-cache. → the inference memory hierarchy in Chapter 9.7; DPU security offload in Chapter 11.6.
| Path | Data path on CPU? | Control path on CPU? | Initiator | Availability |
|---|---|---|---|---|
| Buffered POSIX (legacy) | Yes (bounce buffer) | Yes | Host CPU | Universal |
| GPUDirect Storage (GDS) | No (direct DMA to HBM) | Yes | Host CPU | Mature (CUDA/cuFile stack) |
| SCADA (GPU-initiated) | No | No | GPU | Emerging 2026 |
| BlueField-4 / STX (DPU-offload) | No | No (on DPU) | DPU | H2 2026 (WEKA/VAST/DDN) |
The silicon: TLC vs QLC, PCIe generation, and form factor
Underneath the data-path software sit three hardware forks that fix the per-node economics, and each maps cleanly onto a tier.
TLC vs QLC is an endurance-and-write-bandwidth versus density tradeoff. TLC (3 bits/cell) has higher write endurance (DWPD) and better sustained write performance — the right medium for the scratch and hot/checkpoint tiers that absorb bursty, write-heavy traffic. QLC (4 bits/cell) packs ~33% more capacity per die at lower $/TB but with materially lower endurance and weaker write throughput, which is exactly tolerable for the read-mostly capacity, data-lake, and context tiers. The mistake is picking by NAND category instead of by measurement. Compare the drive's rated sustained write bandwidth and its endurance budget in absolute TB written per day — including write amplification and degraded-mode rebuild — against what the tier will actually do. Shuffle spill and small random writes are where QLC fails that test; large sequential checkpoint drains onto a wide enough enterprise-QLC pool can pass it. Pair the medium to the measured write pattern, and default to TLC wherever the proposed QLC design cannot clear both limits.
PCIe 5.0 vs 6.0 sets the per-drive ceiling. A PCIe 5.0 x4 enterprise NVMe drive tops out near 14 GB/s sequential read; PCIe 6.0 doubles the lane rate, and the first mass-production Gen6 enterprise drives (Micron's 9650, in volume in 2026) reach roughly 28 GB/s read / 14 GB/s write at around 5.5M IOPS. The consequence is drive count: a fixed per-node bandwidth target needs half as many Gen6 drives as Gen5 only if the qualified drives deliver twice the usable bandwidth and PCIe, NIC, endurance and capacity do not set a higher minimum count — and it is that drive count that reshapes the storage shelf. A proposed GPU-initiated server packing 96 drives needs qualification of the complete populated PCIe and cooling layout; drive count alone does not require Gen6. E1.S vs E3.S is the form-factor fork that rides along: E1.S (the ruler) is the dense, hot-swap, often liquid-cooled-friendly format favored inside compute and the densest flash servers; E3.S is the higher-power, higher-capacity format common in dedicated storage nodes. The 9650 ships in both, with E1.S the liquid-cooling target — a reminder that at Gen6 power and density, the storage shelf inherits a cooling decision of its own.
Scope & caveats
Advertised DPU interface capability, not an end-to-end application benchmark or a guarantee that all scale-out traffic traverses this DPU.
Scope & caveats
NVIDIA B300 SU: 32 nodes, 256 GPUs; distinct from the 576-GPU GB200 NVL72 rack-scale planning SU. The Enhanced tier is 250 / 124 GB/s per SU and 2,000 / 992 GB/s at eight SUs; the Standard tier is 80 / 40 and 640 / 320. NVIDIA states these assume a mix of workloads and that requirements must be characterized per workload — they are reference tiers, not a per-GPU floor.
Storage-fabric placement: where the I/O traffic rides
The last fork is topological: every byte from the networked tiers has to cross a fabric, and which fabric is a co-design decision with the back-end compute network of Part 8. Three placements, three sets of consequences.
A dedicated storage rail — separate switches and NICs for storage — gives the cleanest isolation: checkpoint-write incast and dataset-read bursts never collide with the all-reduce collectives that the training job lives or dies on. Select it when the coexistence test cannot protect both storage and collective deadlines on shared resources. The separate front-end/storage network in the cited NVIDIA reference illustrates this placement; it is not a universal goodput ranking. Extra NICs, switch ports, optics and their power buy network separation, while shared PCIe, host memory, storage controllers or power still require a failure check. Count those additions in Chapter 8.5 rather than charging isolation as a nominal allowance.
Converging storage onto the back-end compute fabric reuses the non-blocking InfiniBand/RoCE network you already paid for, saving the dedicated rail. A checkpoint write from 16k GPUs is a synchronized incast event; convergence saves ports only if the shared NIC, PCIe path, queues and target preserve both service contracts through that burst and the required failure. Separate classes and tuned congestion control protect allocation but create no physical capacity. Choose convergence when the concurrent test below passes; otherwise buy independent resources or reduce admitted traffic. → surviving-capacity ledger in Chapter 8.5 and congestion controls in Chapter 8.6. A high-performance front-end/storage fabric can carry the hot path when provisioned for it, as in NVIDIA's separate storage-network reference. An out-of-band management network carries BMC and recovery control; combining those names conceals different capacity and failure requirements. → NVIDIA network fabrics.
| Placement | Isolation from collectives | Capex | Best for | Failure mode if wrong |
|---|---|---|---|---|
| Dedicated storage links | Separate switches and NIC ports; host/controller dependencies still checked | Added ports, optics, power and support from 8.5 | Shared resources miss a required storage or collective deadline | Network isolation leaves PCIe, memory or controller contention unresolved |
| Converged compute/storage | Shared physical resources with a tested allocation | Saves only equipment actually removed from the 8.5 ledger | Concurrent foreground and degraded/rebuild contracts pass | Class separation passes while common capacity misses the workload deadline |
| Provisioned front-end/storage; independent OOB | Separated from back-end compute; check the front-end traffic mix | Count incremental storage capacity and separate OOB equipment | Hot storage fits beside ingress/client traffic under failure | An undersized front-end or shared management dependency blocks delivery/recovery |
On the rail, the transport itself is the final sub-fork: NVMe-oF over RDMA (RoCE or InfiniBand) and NVMe/TCP expose different latency, CPU, congestion-control and operational tradeoffs, but neither protocol supplies a portable tier result. Measure the named target and device at the workload's I/O size, queue depth, concurrency, cache state, CPU path, network and hop count, including tail latency under failure and checkpoint incast. Select RDMA or TCP from that end-to-end evidence and the governing throughput/latency SLO, not from the label ‘GPU hot tier.’ Between the two RDMA flavors, the InfiniBand-vs-RoCE decision is the same one Part 8 makes for the compute fabric, and the storage version has the same sting: untuned RoCE under checkpoint incast is where tail latency blows out. Converged Ethernet has caught up where NIC, target and congestion recipe are qualified together — Spectrum-X-class fabrics are sold on exactly that promise for GDS and KV-offload — so hold every binding to Chapter 8.4’s supported-operation and recovery contract: NVMe/RDMA over InfiniBand or RoCE, NVMe/TCP over the qualified IP path; an Ethernet NIC alone proves neither RDMA support nor NVMe target compatibility. Carry the eligible choice into Chapter 8.5’s ledger and Chapter 8.6’s congestion qualification. The connection trace below tests whether a surviving path restores the intended NVMe namespace and completes correct I/O before the application deadline.
A surviving Ethernet link is only the first step in storage recovery. Record the host identity, target subsystem NQN, transport addresses and namespace identifiers. Discovery identifies reachable subsystems; a fabric connection associates the host with a controller, then the host establishes the admin and I/O queues it needs. Identify the namespace and its access state before issuing work. An NVMe submission/completion queue pair is protocol state; an RDMA queue pair or TCP connection carries that state over its transport. Their identifiers and failure lifetimes are different. With TCP, a re-established byte stream alone does not restore an NVMe controller association; with RDMA, a live QP alone does not establish namespace access. Pin the supported base and transport revisions separately from the NIC's Ethernet speed or the SSD's PCIe generation. NVM Express specification boundaries.
Trace a read from submission to completion: select an accessible namespace path, submit the command, transfer the requested data through the transport binding, observe successful completion, then make that buffer available to the application. A successful write completion must also satisfy the configured persistence contract; ordinary completion and power-loss durability are not synonyms. A file checkpoint adds flush and publication requirements beyond one NVMe command, owned by Chapter 9.4. Record where buffers remain staged if the direct path fails so that a reconnect does not silently change CPU load, memory demand or the promised delivery time.
Multipath must preserve the namespace and its service deadline. Present paths to the same intended namespace and inspect controller and namespace identity, rather than treating any second target address as a replica. Linux native NVMe multipath uses Asymmetric Namespace Access (ANA) state when selecting paths; optimized, non-optimized and inaccessible paths are not equal candidates. Choose the supported policy for the installed kernel and target, verify it through nvme list-subsys and the namespace's ANA evidence, and verify independent NIC, switch, controller and power dependencies. Two addresses behind one controller do not cover controller loss. Linux NVMe multipath.
The client owns command timeout and reconnect behavior; the target owns controller state and recovery behavior; the application owns its request or checkpoint deadline. Record their ordered time budgets, including controller-loss timeout, reconnect delay, detection and any queue drain. The installed nvme connect options expose these controls, but a setting must match the target's supported recovery policy. Infinite retries can preserve a pending command while violating the training or serving SLO. Kill the selected path during reads and writes, then kill its controller. Require the surviving controller to complete correct I/O before the application deadline; when all valid paths disappear, require bounded failure reporting. Restore the path and check that failback neither reorders application publication nor overloads the survivor. Retain command status, timeouts, ANA transitions and data checksums. nvme-cli connection controls.
Coexistence is tested at the shared physical bottleneck. Run dataset reads, checkpoint drain, metadata operations and repair alongside the actual collective job. Identify the common NIC, PCIe uplink, switch queue or target service. A separate VLAN or priority cannot create capacity, and separate network switches still leave shared host memory or PCIe contention. Compare per-job delivery, commit delay, metadata tails and collective step time with matched solo runs; use the workload budgets from Chapter 9.2 and Chapter 9.8. Convergence earns its port saving only while every foreground contract survives the required failure and repair traffic.
MLPerf Storage exercises real storage with defined workload generators; its training harness emulates accelerator computation rather than running the proposed model on that many physical GPUs. Record benchmark version, workload, access API, client count, cache state and the emulated accelerator profile with every result. Neither simulated accelerator utilization nor an isolated storage score proves real collectives, decode, cuFile delivery or model-state restart. First use the published workload for comparison, then replay the real loader, model and fault sequence on the installed stack. The first test characterizes storage; the second qualifies the application. MLCommons benchmark implementation.
Deep dive: walking a single training read through the buffered path vs GDS vs SCADA
Trace one read of one data shard and the bypasses become concrete. Buffered POSIX: the data loader calls read(); the kernel issues an NVMe-oF request; the target's NIC DMAs the data into a kernel bounce buffer in the host's system DRAM; the CPU copies it from the bounce buffer across the PCIe root complex into a user buffer; the framework then copies it again into pinned memory and DMAs it to GPU HBM. Two or three copies, a context switch, and CPU + DRAM bandwidth consumed on every shard — fine for one stream, fatal for ten thousand concurrent small reads.
GPUDirect Storage: the loader calls cuFile instead; the GDS stack programs a DMA directly from the storage target (or local NVMe) into GPU HBM across the PCIe complex. The bounce buffer is gone, the CPU copy is gone, the DRAM round-trip is gone. The CPU still issued the request — it owns the control path — but it is no longer touched by the bytes. This removes a copy and may improve an I/O-bound configuration; measure the effect on the named workload. The direct DMA is also why PCIe topology matters: if the NIC and GPU sit on different root complexes, the DMA crosses the CPU's inter-socket link and you lose part of the win.
SCADA / DPU-offload: now even the request is gone from the CPU. The GPU itself (SCADA) or a BlueField-4 DPU (STX) initiates and manages the I/O; the host processor is out of both paths. For an inference engine running a thousand parallel threads against a 2.9 PB context store, this is the difference between feeding the GPU and starving it on CPU orchestration overhead — the host simply cannot dispatch millions of sub-4 KB reads fast enough, so you stop asking it to. Buffered staging retains the host-memory bounce; GDS can bypass that bounce while the CPU still submits I/O; SCADA/GPU initiation and DPU services can move submission or storage work off the CPU. Each rung adds software and hardware dependencies that cost portability, and buys goodput only if the removed work was on the critical path. Compare payload correctness, end-to-end step or request time and host resources on each supported path.
Deep dive: why QLC won the capacity tier and where it must not go
The all-flash-everywhere trend rests on QLC plus data reduction beating HDD on total cost for active data — not on QLC matching TLC — and as of Aug 2026 that economic premise is suspended: the NAND shortage pushed a 30 TB QLC SSD to roughly 15× the price of the equivalent nearline HDD (VDURA index; the 2–4× ratio that held through 2025 no longer does). The architecture below still holds where flash is already bought; the media split for new capacity tiers is a mixed-media decision again. QLC's economics come from packing 4 bits per cell (33% more capacity per die than TLC's 3), and layered with dedup, compression, and similarity reduction (VAST-style), it makes a petabyte-scale all-flash data lake cost-competitive while delivering flash read latency the corpus actually benefits from during training ingest. That is the right home for QLC: read-mostly datasets, model registries, the warm capacity tier, and the read-dominated context tier.
Writes are where QLC has to be proven rather than assumed. Its rated drive-writes-per-day is a fraction of TLC's, and on cache-fronted designs sustained write throughput drops once the SLC cache fills — but DWPD is a rate, not a budget: a large QLC pool can carry an adequate absolute TB-written-per-day despite the lower rating, and DPU-native designs that coalesce writes into large sequential blocks make a separate cache tier optional (Solidigm). Land a checkpoint stream — bursty, large, repeated every few minutes at scale — on a QLC pool too narrow for it and you both burn through its endurance budget and hit the write cliff exactly when the synchronous checkpoint needs full bandwidth, stalling the training job. So the split is a measurement, not a taxonomy: TLC wherever the writes are random, latency-sensitive, or exceed what the QLC pool can sustain (scratch, shuffle spill); QLC where reads dominate or the writes are large and sequential (capacity, lake, context, and checkpoint drains onto a pool sized for them). Hold the independently derived checkpoint drain, durable-commit and restore budgets from Chapter 9.4 against the pool's measured rates under concurrent load, whichever medium you chose; a write-to-read ratio does not select those budgets. → checkpoint bandwidth sizing in Chapter 9.4 and the capacity tier in Chapter 9.6.
Anti-patterns
The recurring mis-builds in this layer all come from optimizing the file system's headline number while ignoring the path:
- Buying GDS-capable hardware and running buffered I/O. A supported NIC and file system mean nothing if the data loader issues ordinary
read()and silently falls back to the bounce buffer. You pay for the bypass and get bounce-buffer latency. Validate the path, not the bill of materials. - QLC under the checkpoint stream. Putting write-heavy, bursty checkpoint traffic on a cheap capacity tier can burn through its endurance or hit a sustained-write cliff just as the checkpoint must drain. TLC is a candidate for write absorption and QLC for read-heavy capacity; a QLC pool earns checkpoint duty only when rated host writes, full-drive sustained service and the degraded pool pass the endurance and drain deadlines.
- Converging storage onto the collective fabric without traffic isolation. Saving a dedicated rail by dropping 16k-GPU checkpoint incast onto the same InfiniBand that carries all-reduce — converting a checkpoint into a training stall. The capex you saved is dwarfed by the goodput you lost.
- Treating the storage shelf as a thermal afterthought at Gen6. Speccing a 96-drive Gen6 flash server into an air-cooled storage row without checking the populated chassis’s power, inlet and airflow limits can strand the density behind a cooling wall. Qualify flow and service access as well if the selected chassis needs liquid cooling.
Choose the medium and transport that clear capacity, host-write endurance, payload delivery and recovery together. Preserve a qualified staged path when its operational simplicity outweighs a measured direct-DMA gain; select offload when it releases the resource that limits useful work. The purchase is conditional on the populated chassis and surviving connection passing those tests. A faster drive or wire that leaves the command, controller or application stalled adds cost without feeding the GPU.
Cite this chapter
Fehn, J. (2026). NVMe Tiers, GPUDirect Storage & the CPU-Bypass Data Path (Chapter 9.3). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-9-storage-and-data/9-3-nvme-tiers-gpudirect-storage-and-the-cpu-bypass-data-path (accessed 2026-09-29).
@misc{aidc-9-3,
author = {Fehn, Jacob},
title = {NVMe Tiers, GPUDirect Storage & the CPU-Bypass Data Path (Chapter 9.3)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-9-storage-and-data/9-3-nvme-tiers-gpudirect-storage-and-the-cpu-bypass-data-path},
note = {Accessed 2026-09-29}
}