The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 10.2

In this chapter · 7 sections
Term help

Topology-Aware & Rack-Scale Scheduling

For tightly coupled groups, placement is a performance contract: inside one qualified NVLink domain the job runs at scale-up bandwidth; split across the boundary it falls off a bandwidth cliff no tuning recovers, so price that loss against the idle capacity needed to keep the group local.

GOODPUTPOWER-BOUNDDENSITY-RAMP

What you'll decide here

  1. Choose node, subgroup or NVLink-domain allocation for the job's actual communication groups; whole-rack reservation is a policy choice whose unused GPUs still cost money.
  2. How you express topology to the scheduler: Slurm block/segment plus the IMEX switch plugin, or Kubernetes ComputeDomains via DRA with clique labels — and who owns the one-ComputeDomain-per-node constraint.
  3. Where TP, EP, PP, DP and context-parallel groups cross the fabric, using measured collective exposure and the layout’s locality constraints.
  4. Your fragmentation policy: whether you accept stranded GPUs to keep domains whole, defragment with preemption, or relax locality with segments — and what that costs in goodput versus utilization.
  5. Whether discovery is trusted (vendor labels, clique IDs) or verified (NCCL/nvbandwidth probes at admission), because a mislabeled or degraded NVLink link places jobs onto a cliff the scheduler thinks isn't there.
One job, two Blackwell placements: inside the NVLink domain at ~1.8 TB/s bidirectional per GPU versus full-duplex scale-out at ~0.1 TB/s on GB200/CX-7 (~18×) or ~0.2 TB/s on GB300/CX-8 (~9×). These are nominal endpoint ratios, not job slowdowns; locality need not mean whole-domain allocation.

Chapter 10.1 established the scheduling plane as the control loop that turns a fleet of accelerators into a service. This chapter is about the single hardest constraint that loop has to respect on 2026-era hardware: the network is not flat. Between any two GPUs in a modern AI cluster there are three radically different classes of link, and the bandwidth between them differs by roughly an order of magnitude at each step. Inside a node, GPUs talk over NVLink at terabytes per second. Inside a rack-scale NVLink domain — a GB300 NVL72, the 2026 volume rack and the same shape as the GB200 NVL72 it succeeds, is 72 GPUs across 18 compute trays sharing 130 TB/s of aggregate NVLink — they still talk at scale-up speeds as if they were on one giant board. The moment a collective has to leave that domain, it drops onto the scale-out fabric: a ~400G or 800G NIC, an order of magnitude less per-GPU bandwidth, and a different failure and congestion regime. A scheduler that ignores this hierarchy will cheerfully place a tightly-coupled job with eight GPUs in one rack and eight in the next, and the job will run — at a fraction of its rightful throughput, forever, with no error logged.

Topology-blind versus topology-aware placement is the choice, and choosing blind is paid for in goodput at every step of every job for the life of the run. We walk the bandwidth cliffs and how you discover them; the two production mechanisms for defending them, Slurm block scheduling with IMEX and Kubernetes ComputeDomains over Dynamic Resource Allocation; how to map a parallelism strategy onto the fabric so the right collectives stay local; and the fragmentation problem that makes the rack, not the node, the natural unit of allocation on rack-scale systems.

Why topology matters: the bandwidth cliffs

Hold three numbers in mind, because everything in this chapter is a consequence of the gaps between them. NVLink (scale-up, within the coherent domain): NVLink 5 on Blackwell delivers 1.8 TB/s bidirectional per GPU; an NVL72 rack aggregates to 130 TB/s. PCIe (host-local, GPU-to-CPU or GPU-to-NIC): Gen5 x16 is ~64 GB/s each way — the path KV-cache and host staging traverse, and a bottleneck the moment data must touch the CPU. Scale-out NIC (inter-node, over InfiniBand or RoCE Ethernet): GB200 assigns one 400 Gb/s ConnectX-7 path (50 GB/s per direction) per GPU; GB300 assigns one 800 Gb/s ConnectX-8 path (100 GB/s per direction) per GPU. The platform-scoped ratio is roughly 18× on GB200/CX-7 and 9× on GB300/CX-8 (900 GB/s per direction on NVLink 5 against 50 or 100 GB/s per-GPU scale-out), and the consequence is unambiguous: the most bandwidth-hungry collectives must be fit inside the scale-up domain, and the scheduler is the thing that decides whether they are.

Bandwidth does not degrade gracefully as a job spreads; it steps down discontinuously at each boundary. A tensor-parallel all-reduce that fits in one NVLink domain runs at NVLink speed; the same all-reduce split 9-and-9 across two domains runs at NIC speed for the cross-domain portion and stalls every other rank waiting on it. NVIDIA's own Slurm guidance states the case flatly: when a job crosses the NVLink-domain boundary, performance "drops sharply," which is why their block plugin treats the domain as a hard constraint rather than a best-effort preference. The failure mode is insidious because it is silent — the job completes, the loss curve descends, nothing alarms — but MFU sits 20–40% below where the hardware should deliver, and the bill for that gap compounds across the entire run.

Discovery: how the scheduler learns the fabric

A scheduler can only defend a topology it can see, and discovery is where most topology-aware deployments quietly fail. There are two postures, and they correspond to a trust decision. Trusted discovery reads the fabric from labels the platform asserts: NVIDIA exposes an NVLink clique identifier (the nvidia.com/gpu.clique node label in Kubernetes; equivalent block definitions in Slurm's topology.yaml) that tells the scheduler which nodes share a coherent NVLink partition. PCIe and NUMA affinity come from the driver and from the node's hardware topology. Scale-out topology — which leaf switch a node hangs off, which rail a NIC belongs to — comes from cabling databases, LLDP, or a topology file the operator maintains. Trusted discovery is fast and is how production schedulers run day-to-day.

Verified discovery distrusts the labels and measures. At node admission and after every repair you run point-to-point bandwidth probes (nvbandwidth, NCCL all-reduce/all-gather benchmarks, p2pBandwidthLatencyTest) and compare against the expected matrix for the asserted topology. The labels can be right while the physical link is degraded: a marginal NVLink lane that has fallen back to a lower width, a transceiver running hot, a cable seated badly. The label still says "one clique, full bandwidth"; the hardware delivers half. A scheduler trusting that label places a tightly-coupled job onto a cliff it believes isn't there. Verifying the topology the scheduler will trust — gating nodes into the schedulable pool only after they pass the bandwidth matrix, not merely after they boot — is what prevents topology-label drift from consuming the named fleet's measured goodput; use any 90%-versus-96% comparison only as an explicit sensitivity case. Verification belongs to the burn-in and health pipeline of Chapter 10.6; the scheduler consumes its verdict.

The interconnect hierarchy a scheduler must respect
TierReachNominal endpoint / aggregate bandwidth (basis stated)What rides itScheduling implication
NVLink (scale-up)Within the coherent NVLink domain (node → NVL72 rack)~1.8 TB/s bidirectional (~900 GB/s per direction; NVLink 5); 130 TB/s aggregate per NVL72TP and EP all-reduce / all-to-all; KV-cache over NVLinkKeep each required group within its qualified domain; smaller allocations remain valid.
PCIe (host-local)GPU ↔ CPU / GPU ↔ NIC inside a node~64 GB/s (Gen5 x16, each direction)Host staging, GPUDirect setup, KV spill to hostNUMA/PCIe affinity: pin GPU to its local NIC and CPU socket
Scale-out NIC (inter-node)Across the back-end fabric (IB / RoCE), domain to domain~50 GB/s per direction per GPU on GB200/CX-7/400G; ~100 GB/s per direction per GPU on GB300/CX-8/800GData-parallel all-reduce; pipeline-stage point-to-pointDP/PP or cross-node EP where measured traffic and overlap pass the workload contract. Honor rail alignment, leaf locality and each group’s declared boundary
Per-GPU bandwidths are platform-versioned Blackwell reference points; see keynumbers for sources and vintages. The scheduling implication column is the load-bearing one.

Slurm block scheduling: the rack as an atomic block

Slurm is an established scheduler for large training fleets, and its answer to rack-scale topology is block scheduling (the topology/block plugin, hardened for NVL72 in 2025). The model is simple and powerful: you declare, in topology.yaml, that a set of nodes forms a block corresponding to one NVLink domain — for the cited GB200 NVL72 profile that is 18 four-GPU compute nodes carrying 72 GPUs. The scheduler then uses the block as a locality unit; locality does not require reserving every GPU in it. A job requesting up to 18 nodes is consolidated into a single block, provided the required TP/EP groups are mapped within that block and the selected allocation policy preserves their locality. A job larger than one block spans whole blocks rather than being smeared arbitrarily, so the cross-domain traffic is confined to the parallelism dimensions that tolerate scale-out bandwidth.

Two refinements make this practical rather than rigid. The --segment argument lets a user declare the atomic group size their job needs — for example --segment=4 lets a 12-node job be split into three 4-node segments that the scheduler can pack flexibly across blocks while still keeping each segment's locality intact. This is the release valve for fragmentation: instead of demanding one contiguous 12-node hole, the scheduler can fill three smaller holes, lifting utilization without abandoning topology. The second refinement is the switch/nvidia_imex plugin (Slurm 24.05+), which manages IMEX channels — the driver-level memory-export/import access control that lets GPUs in a multi-node NVLink domain address each other's memory. Slurm provisions and tears down the IMEX channel per job, so authorized memory import/export follows the job; verify isolation with the exact driver and job teardown, rather than treating channel provisioning as whole-platform security. Block scheduling without IMEX management is a security and correctness gap on shared NVL72 hardware, not merely a tuning omission.

Deep dive: what topology.yaml actually encodes, and why segments beat raw node requests

The topology.yaml file is the scheduler's map of the bandwidth cliffs. For a two-rack NVL72 cluster it declares two blocks — say block01 and block02 — each binding a named range of 18 nodes to a block size of 18. That single declaration changes the scheduler's behavior from "find me 16 free nodes anywhere" to "find me 16 free nodes that share an NVLink domain, and if you can't, span whole domains." The plugin, introduced in Slurm 23.11 (the --segment allocation flag arriving in 24.05) and matured for rack-scale systems through 2025, enforces this as a hard constraint: jobs at or below the block size stay consolidated; larger jobs grow in block-sized increments.

The reason --segment matters is the difference between locality and contiguity. A naive topology scheduler will only place a job if it can find one contiguous hole large enough — which, on a busy fleet, strands capacity because the holes are the wrong shape. Segments decouple the two: the user tells the scheduler "my job needs groups of N that each stay NVLink-local, but the groups themselves can be anywhere." A 12-node job with --segment=4 becomes three independent 4-node locality requirements, far easier to satisfy from a fragmented fleet than one 12-node block. The cost is that cross-segment traffic now rides scale-out, so you only request segments smaller than your job when the parallelism crossing the segment boundary (typically DP or PP) can tolerate it. Choosing the segment size is choosing which parallelism dimension you are willing to push onto the scale-out fabric — the segment size is a mapping decision, not merely a scheduler flag. → parallelism mapping below; oversubscription budget in Chapter 8.5.

Kubernetes: ComputeDomains, DRA, and the IMEX-per-node constraint

Kubernetes arrived at rack-scale topology from the opposite direction. Its original device-plugin model exposed a GPU as a flat, countable, location-free resource — nvidia.com/gpu: 8 — which alone does not express cross-node NVLink domain membership. The fix is Dynamic Resource Allocation (DRA), the resource API that lets a workload request a structured, parameterized resource rather than an opaque count, and on top of it NVIDIA's ComputeDomains abstraction. A ComputeDomain represents reachability between the distributed workers of a multi-node job: when its pods are scheduled, the platform dynamically creates the IMEX domain that lets those pods' GPUs address one another's memory over NVLink, and tears it down when the job ends. The IMEX channel surfaces inside each container as a device file, so CUDA or NCCL can use the authorized memory-export path; device preparation and a real transfer still have to pass.

Two constraints define how you design around this. First, placement still needs an affinity rule: pods must be steered onto nodes that share an NVLink partition, expressed by matching the nvidia.com/gpu.clique node label, or the ComputeDomain spans a boundary it cannot bridge. The clique label is the Kubernetes analogue of the Slurm block. Second — and this is the sharp edge of the 2026 implementation — the selected ComputeDomain driver can impose one domain per node; that allocation constraint is not the number of channels IMEX can represent. If a ComputeDomain claims only part of a node's GPUs, the remaining GPUs on that node cannot join a different ComputeDomain and are unavailable to that second domain, though a compatible local workload may still use them. That constraint can favor whole-node allocation when different jobs need different domains; it does not require reserving the whole rack. Record domain ownership separately from the GPU count reserved and billed. The NVIDIA GPU Operator 26.7 managed workflow requires Kubernetes ≥1.34.2, NVIDIA driver R580+, CDI and a fresh GPUCluster installation without ClusterPolicy. It offers no in-place migration from standalone DRA; ComputeDomains additionally require supported Multi-Node NVLink hardware and correctly owned IMEX services. Missing preparation evidence blocks admission rather than authorizing a presumed fallback.

Decision fork: Slurm block scheduling vs Kubernetes ComputeDomains for rack-scale jobs
DimensionSlurm block schedulingKubernetes + ComputeDomains (DRA)
Topology unitBlock in topology.yaml = one NVLink domain (18 nodes / 72 GPUs)ComputeDomain over nodes sharing the gpu.clique label
Locality enforcementHard constraint; --segment relaxes contiguity, not localityPod affinity on gpu.clique; scheduler must keep pods in-partition
IMEX managementswitch/nvidia_imex plugin provisions channel per jobIMEX domain created/torn down dynamically per ComputeDomain
Partial-node sharingSupported within a node; gang semantics per jobBlocked: one ComputeDomain per node strands unused GPUs in a partial claim
Native gang schedulingYes — all-or-nothing allocation is intrinsicNeeds a gang-aware scheduler (KAI, Volcano, Kueue) atop DRA
Best fitLarge synchronous training; HPC-heritage fleetsMixed train+serve estates; cloud-native, multi-tenant platforms
Both are valid in 2026; the choice usually follows the existing control plane (HPC-heritage training fleet vs cloud-native platform). Convergence (Slurm-on-K8s) is blurring the line — see Chapter 10.1.

Mapping parallelism onto the fabric

Topology-aware scheduling is only half the bargain; the other half is the job declaring its parallelism so the scheduler can place the right dimension on the right tier. 3D (and 4D) parallelism uses four collective patterns with sharply different bandwidth appetites, and the art is matching appetite to tier. Tensor parallelism (TP) all-reduces activations on every layer — the most bandwidth-hungry pattern, and the one that must live inside the NVLink domain. Expert parallelism (EP) for MoE models does all-to-all token routing; wide EP (EP32 and beyond) is precisely what a 72-GPU NVLink domain unlocks, because the all-to-all stays on NVLink instead of hitting the scale-out fabric. Pipeline parallelism (PP) passes activations point-to-point between stages — low volume, latency-sensitive, tolerant of crossing the domain boundary. Data parallelism (DP) all-reduces gradients once per step — high volume but infrequent and overlappable, the canonical dimension to push onto scale-out.

The default placement rule that falls out: TP and EP inside the scale-up domain; PP and DP across the scale-out fabric. EP is the dimension where that default is a preference rather than a law — DeepSeek-V3 trained with 64-way expert parallelism spanning eight H800 nodes over InfiniBand, using node-limited routing and communication/compute overlap to keep the all-to-all off the critical path — so judge cross-node EP on measured routing volume, expert placement, overlap, and available network bandwidth rather than ruling it out. A correctly mapped GB300 NVL72 job sets TP+EP to fit within the 72-GPU domain, runs PP across racks, and lets DP all-reduce ride the rail-optimized fat-tree where SHARP in-network reduction can absorb it. VR200 NVL72, in production now, keeps that 72-GPU domain at higher scale-up bandwidth, and Rubin Ultra Kyber on the 2027 roadmap widens the domain itself, so the TP/EP fit is a mapping you recheck each generation rather than a constant. Get the mapping backwards — DP inside the domain, TP across racks — and you have inverted the bandwidth hierarchy: the cheap-to-cross dimension hogs NVLink while the expensive-to-cross dimension chokes on NIC bandwidth. This is why the scheduler needs more than a node count; it needs the parallelism shape, so block sizes and segment sizes line up with the TP/EP group rather than slicing through it. The framework that emits this shape is the subject of Chapter 10.8; the fabric that carries it is Chapter 8.5.

72 GPUs / 130 TB/s
GB200 NVL72 scale-up domain; topology membership does not require whole-domain allocation
1.8 TB/s bidirectional (~0.9 TB/s per direction)
NVLink 5 nominal per-GPU bandwidth, bidirectional; one direction is half this value
Scope & caveats

Named NVIDIA generation/platform, bidirectional aggregate convention; not delivered collective bandwidth, cache coherence or arbitrary mixed-platform performance.

~18× GB200 / ~9× GB300derived
scale-up (NVLink 5, 900 GB/s per direction) vs per-GPU scale-out — ~18× on GB200/CX-7/400G and ~9× on GB300/CX-8/800G
Scope & caveats

Nominal directional endpoint ratios: 900 GB/s divided by 50 GB/s for the specified GB200 profile; divided by 100 GB/s for the specified GB300 profile. Not collective or application speed ratios.

18 nodes
GB200 example: 18 four-GPU nodes per NVL72 block in topology.yaml
1 / node
one ComputeDomain per node in the Kubernetes DRA driver — a partial claim strands the node's other GPUs
Scope & caveats

A driver-imposed allocation constraint of one ComputeDomain per node, not a count of the channels the IMEX service can represent. Verify against the GPU Operator / DRA driver release actually deployed.

K8s 1.32+
legacy 2025 ComputeDomain path: Kubernetes 1.32 with DRA APIs enabled; GPU Operator 25.3+
8.4% (10.0% incl. NIC)
network share of Llama 3 interruptions — switch/cable 8.4%; 10.0% including NIC
Scope & caveats

Table 5 is internally inconsistent: its 17 listed counts sum to 417 although §3.3.4 states 419 unexpected interruptions; printed percentages sum to 94.4%; and 148/419 = 35.3%, not the printed 30.1%. Treat counts as auditable; label percentages as paper-printed. Network Switch/Cable is 35 (8.4% printed); NIC is classified Host and NCCL watchdog timeouts Unknown.

90% vs 96% scenariomodeled
training-goodput sensitivity scenario: 90% vs 96% (illustrative — replace with the named fleet's measured goodput)
Sep 2026Guide analysis — stipulated sensitivity scenario; no claim of an industry measurement.register ↗
Scope & caveats

Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.

Fragmentation and the rack as the scheduling unit

Topology-awareness creates a tension that does not exist in flat scheduling: the better you defend locality, the worse your bin-packing gets. If a job can only land inside one NVLink domain, then a domain with 6 of 72 GPUs busy has 66 GPUs that are useless to any job needing more than 66 NVLink-local — they are stranded by fragmentation, not by lack of capacity. Flat schedulers never see this problem because they will place a job anywhere; they pay for it instead in silent cliffs. Topology schedulers make the fragmentation visible and force a policy choice.

There are three levers, and a 2026 fleet uses all three. Accept the strand: keep domains whole, leave the 66 GPUs idle until a job that wants them arrives, and price the lost utilization as the cost of full-bandwidth placement — the right call when goodput per job dominates, i.e. large training. Defragment: use preemption and migration to compact small jobs out of partially-used domains, freeing whole domains for large ones — effective but expensive on synchronous jobs because preemption forces a checkpoint-and-restart. Relax locality with segments: let jobs declare smaller atomic groups (Slurm --segment, or sizing the ComputeDomain below a full rack) so they fit the holes you have, accepting scale-out traffic on the boundary the segment crosses. On rack-scale hardware the rack — the NVLink domain — becomes the natural unit of allocation, not the GPU and not the node. Placement is what genuinely wants the domain as its unit: any policy that subdivides one has to reckon with the IMEX-per-node constraint, the bandwidth cliff, and the fragmentation it creates. Billing, drain, and repair units are separate decisions — a job that fits inside a domain can take a subset of its nodes, and NVIDIA's own GB200 block-scheduling guidance places eight- and twelve-node jobs inside a single block — so reserve and price a whole rack where a workload or a capacity guarantee needs the whole rack, not by default.

Deep dive: why discovery errors are worse than placement errors

A placement error — putting a job on a cliff — is at least diagnosable: the bandwidth is wrong, a benchmark reveals it, and a better policy fixes the next run. A discovery error is worse, because it makes the scheduler confidently wrong. If a node's gpu.clique label or Slurm block assignment says it shares an NVLink domain with neighbors it does not actually reach at full bandwidth — a degraded NVLink lane that fell back to half-width, a partially-seated backplane connector, a transceiver throttling on heat — then the scheduler will place a tightly-coupled job into what it believes is one coherent domain, and the job will run on a hidden cliff that no policy can see. The scheduler did everything right against a map that was wrong.

This is why 2026 practice is to verify the topology before trusting it, and to make the schedulable pool a function of measured bandwidth, not asserted labels. The mechanics: at admission and after every repair, run an NCCL all-reduce or nvbandwidth probe across the asserted domain and compare against the expected matrix; if a link is below threshold, mark the node degraded and pull it from the topology the scheduler trusts, even though it boots and passes basic health. The Meta figure in the keynumbers — roughly a tenth of large-training job interruptions traced to network and configuration issues — is in large part this class of problem: the fabric was not what the control plane thought it was. The scheduler is only as good as its map, so the map has to be measured, not declared. The verification pipeline lives in Chapter 10.6; the fault-domain and recovery framing in Chapter 10.7.

Anti-patterns

The same mistakes recur because each comes from treating the GPU as fungible when the fabric says it is not:

  • Flat scheduling on rack-scale hardware. Exposing nvidia.com/gpu: 8 with no clique or block awareness, then wondering why a 16-GPU job that landed 8-and-8 across two racks runs at 60% MFU. The hardware is non-blocking inside the domain and an order of magnitude slower across it; a location-free resource model cannot see the difference and so cannot avoid the cliff.
  • Topology without gang scheduling. Placing the right pods on the right nodes but admitting them incrementally, so synchronous jobs deadlock holding partial allocations. Locality and all-or-nothing admission are a pair; shipping one without the other trades a silent cliff for an outright stall.
  • Inverted parallelism mapping. Putting data-parallel all-reduce inside the NVLink domain and tensor-parallel across racks — using the scarcest bandwidth on the dimension that least needs it. The mapping must follow the appetite: TP/EP local, PP/DP remote.
  • Trusting labels over measurements. Letting a node into the schedulable pool because it booted and its clique label is present, without verifying the NVLink bandwidth matrix — placing jobs onto a degraded domain the scheduler believes is healthy.
  • Subdividing domains under the one-ComputeDomain-per-node limit. Designing a Kubernetes multi-tenant scheme that hands partial nodes to distinct ComputeDomains, then discovering the one-ComputeDomain-per-node constraint strands the remaining GPUs. On rack-scale NVLink hardware, align allocations to domain boundaries and keep each ComputeDomain whole at node granularity; sell and bill the whole rack only where the workload or the capacity guarantee actually requires it.

Reserve the scale-up domain when the job’s measured TP/EP communication needs it; otherwise pack compatible jobs inside the isolation and SLO limits. Reserving every rack wastes paid capacity, while splitting the wrong collective buys nominal utilization with longer completion time. Keep the rank map and measured crossover beside the allocation policy.

This chapter sits inside the scheduling plane defined in Chapter 10.1 and feeds the multi-tenancy and isolation policy of Chapter 10.3, where the fragmentation-vs-utilization fork becomes a fairness and quota question. The bandwidth cliffs it defends are engineered in the fabric chapters: scale-out topology, sizing, and oversubscription in Chapter 8.5, transport and protocols in Chapter 8.4, and the congestion control that governs cross-domain traffic in Chapter 8.6. The parallelism strategy that the scheduler must be told about is the province of Chapter 10.8; the topology verification and goodput telemetry that decide which nodes the scheduler may trust live in Chapter 10.6; fault domains and recovery in Chapter 10.7; and the goodput-vs-availability reframe that underlies the whole fragmentation trade in Chapter 12.2. Why training is the archetype that most rewards full-bandwidth placement is established in Chapter 1.2.
Cite this chapter
Fehn, J. (2026). Topology-Aware & Rack-Scale Scheduling (Chapter 10.2). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-10-software-orchestration-and-service-delivery/10-2-topology-aware-and-rack-scale-scheduling (accessed 2026-09-29).
@misc{aidc-10-2,
  author       = {Fehn, Jacob},
  title        = {Topology-Aware & Rack-Scale Scheduling (Chapter 10.2)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-10-software-orchestration-and-service-delivery/10-2-topology-aware-and-rack-scale-scheduling},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit