The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 7.9

In this chapter · 5 sections
Term help

Software Ecosystems & Lock-In

Buying an accelerator commits you to CUDA, ROCm, XLA or Neuron; kernel ports, collective tuning and continuing releases are switching costs the datasheet never priced, so measure them with the useful output of the executable path.

POWER-BOUNDGOODPUTDENSITY-RAMP

What you'll decide here

  1. Which software regime you are committing to — CUDA, ROCm, XLA/JAX, or Neuron — and therefore which workloads you can run on day one versus which require a porting project before they earn revenue.
  2. Whether the quoted system’s measured useful output pays for acquisition, porting and continuing support, using MFU only with a named compute denominator.
  3. How much of your stack you write against a portability layer (Triton, MLIR, torch.compile, OpenAI-compatible serving APIs) versus a vendor-native one (CUDA C++, cuDNN, TensorRT-LLM) — the choice that sets your switching cost years before you exercise it.
  4. Whether the lock-in that binds you is the kernel language, the collective-communication library, the serving framework, or the scale-up fabric — each has a different exit cost and a different rate of erosion.
  5. Whether to run a single-vendor fleet for operational simplicity or a heterogeneous fleet for supply leverage and price discipline — and who on your team owns the second toolchain if you do.

Chapter 7.1 through Chapter 7.8 compared accelerators on hardware — HBM capacity, FP4 PFLOPS, NVLink bandwidth, dollars per GPU — and that comparison is incomplete in a way that costs real money. Buying an accelerator also commits you to a runtime, a kernel language, a collective library, a compiler, a serving stack, and a fabric, and that stack is what decides how much of the silicon's peak you ever see, how fast a new model lands on the box, and how much it costs to leave. Four stacks matter in 2026: CUDA (NVIDIA), ROCm (AMD), XLA (Google TPU, via JAX/TensorFlow), and Neuron (AWS Trainium/Inferentia). Each is a distinct lock-in regime.

Three questions follow — which regime, which portability posture, which fleet composition — and each carries a downstream cost: the porting project that delays revenue, the realized-MFU gap that erases a price advantage, the kernel you write yourself because no library shipped it, a model whose NVIDIA path ships before a required operation is supported on the alternative release. Ignore the software stack at purchase and you do not avoid the bill; you defer it to the day the cluster is live and the delivered goodput misses the acceptance target.

The four regimes, and what each one actually locks

These four ecosystems are not "CUDA, plus three challengers." Each locks a different layer of the stack, erodes at a different rate, and charges a different exit cost — so the useful way to sort them is by what each one actually holds captive.

CUDA is the deepest and broadest moat, and it is mostly a library and tooling moat, not merely a language one. The language (CUDA C++) is replaceable; what is hard to replace is two decades of accreted libraries — cuDNN, cuBLAS, CUTLASS, NCCL, TensorRT-LLM — plus the profilers, debuggers, and the simple fact that every new model, every new attention variant, every new quantization scheme ships and is tuned on NVIDIA first. The lock-in is subtler than 'you cannot leave': on NVIDIA the kernel you need already exists and is tuned, and everywhere else you write it or wait for it.

ROCm is the open challenger, and as of 2026 it has closed the gap from "unusable for production" to "feature-matched on the workloads AMD chooses to fund." ROCm mirrors CUDA layer-for-layer — HIP for the language, MIOpen/rocBLAS for the math, RCCL for collectives — and the porting tax is genuinely low for code written against PyTorch rather than raw CUDA. The residual lock-in is coverage: the long tail of custom kernels and bleeding-edge model architectures that are tuned on CUDA months before ROCm catches up. ROCm's exit cost is asymmetric — coming from CUDA-portable PyTorch is cheap; coming from hand-written CUDA C++ is not.

XLA (the compiler behind JAX and TensorFlow on TPU) inverts the model entirely. The lock is not the absence of a kernel language — JAX's Pallas does let you hand-write TPU kernels, as AWS's Neuron Kernel Interface does on Trainium. XLA is a graph compiler that takes a whole-program traced computation and emits TPU machine code, and the lock-in is that your code is written in JAX against XLA's compilation model, not in a portable imperative GPU style. The upside is that XLA is genuinely multi-target — it compiles to GPU too — so the regime locks the programming model more than the silicon. The downside: the ecosystem of pretrained checkpoints, third-party kernels, and community tooling around JAX/TPU is narrower than CUDA's, and TPUs are rentable only from Google.

Neuron (AWS's SDK for Trainium and Inferentia) is the most vertically integrated and the most single-tenant. The Neuron compiler ingests PyTorch/JAX and targets the Trainium ISA; the lock-in is the whole stack plus the fact that Trainium exists only inside AWS. The trade is explicit: external adopters cite roughly 50% cost savings versus comparable NVIDIA capacity, paid for with a Neuron porting project and the loss of any exit to another cloud. → Chapter 7.4.

The four software regimes — what each one locks, and the exit cost
RegimeVendor / siliconPrimary lock-in layerDay-0 model availabilityExit cost (dominant term)Erosion vector
CUDANVIDIA GPUsLibraries + tooling (cuDNN, NCCL, TensorRT-LLM, CUTLASS)Day 0 — models land and are tuned here firstRewriting hand-tuned CUDA C++ kernels with no portable equivalentTriton/MLIR/torch.compile abstract the kernel layer away
ROCmAMD Instinct GPUsKernel coverage (the custom-kernel long tail)Days-to-weeks on funded workloads; longer on the tailRe-tuning + closing the realized-MFU gap on your specific modelHIP source-compat + PyTorch upstreaming + InferenceMAX-style benchmarking
XLA / JAXGoogle TPU (also compiles to GPU)Programming model (whole-program graph compilation)Strong for Google/JAX-native models; thinner third-party tailRewriting imperative PyTorch into JAX/XLA tracing semanticsJAX-on-GPU + PyTorch/XLA bridge widen the target set
NeuronAWS Trainium / InferentiaFull vertical stack + single-cloud availabilityAnchor-tenant models day 0; general tail lagsNeuron port + total loss of cross-cloud portabilityPyTorch/JAX front-ends + managed model hubs reduce port friction
Vendor-neutral characterization current to mid-2026. "Day-0 model availability" means how quickly a newly-released open model runs in production on that stack without a porting project.

The realized-MFU gap: where paper FLOPS go to die

A costly misconception in accelerator procurement is that two chips with similar peak FLOPS deliver similar work. They do not, and the gap reflects software, memory, fabric and the model. Model-FLOPS-utilization is the ratio of the useful floating-point work your model actually performs to the chip's theoretical peak over the same wall-clock. For a fixed model, precision, FLOP-count convention and device count, compare useful model FLOPs with the matching dense peak over the same wall-clock. No universal 35–55% band classifies a stack as mature. Profile the difference in the kernels, the fused attention path, the collective-communication overlap, the compiler's ability to keep the tensor cores fed.

This is where ROCm's story turned in 2026, and the direction of travel matters more than any single snapshot. The historical knock on AMD was a severe realized gap: the MI300X carried roughly 1.3x the paper FLOPS of an H100 but, on real inference, showed workload-specific shortfalls on the tested kernels and serving stack at the reported model, precision and release (SemiAnalysis AMD-vs-NVIDIA benchmarking, 2025). By mid-2026 that gap had narrowed dramatically on the workloads AMD and the community chose to fund. On specific recent models, the MI355X running ROCm on SGLang reached feature parity with B200 running CUDA — same FP8 KV-cache path, same MLA kernels — and undercut the B200 on cost per million tokens across much of the single-node Pareto frontier (SemiAnalysis InferenceX, 2026). The lesson is not "AMD won" or "NVIDIA won." The lesson is that the realized-MFU gap is a moving target set by software investment, workload by workload, and a procurement decision made on last quarter's gap may be wrong this quarter.

The consequence for the buyer is structural. A hardware-price advantage requires matched quotations, then useful output, porting and recurring support on the same boundary. In the stated illustration, a chip priced 20% lower but delivering 40% less matched useful output costs 0.8/0.6 ≈ 1.3 times as much per unit of output before other costs. It breaks even only when its useful output reaches 80% of the reference; below that, the cheaper chip buys a more expensive cluster. Compare the full cash flows in Chapter 1.8 — cost per token or cost per useful FLOP. Flip it around and the same logic favors the challenger: on a workload where ROCm has reached parity and the complete service costs less, it wins on cost-per-token. The number you must produce before signing is not on any datasheet; it is the realized cost-per-token of your workload on the software you will run.

~40%
MI355X (ROCm/SGLang FP8) cost-per-Mtoken advantage vs B200 (SGLang) at points on the GLM-5 single-node frontier — gap reversed on funded workloads
Scope & caveats

MI355X cost/Mtoken gap

90% vs 96% scenariomodeled
training-goodput sensitivity scenario: 90% vs 96% (illustrative — replace with the named fleet's measured goodput)
Sep 2026Guide analysis — stipulated sensitivity scenario; no claim of an industry measurement.register ↗
Scope & caveats

Stipulated endpoints for sensitivity only. They are neither measured provider outcomes nor universal targets. Chapter 14.1 reconciles productive-time boundaries; a site measures its own baseline.

~50%
external-adopter cited cost savings on AWS Trainium vs comparable NVIDIA capacity — paid for with a Neuron porting project
Scope & caveats

Trainium adopter savings

27.8% forecast for 2026forecast
TrendForce’s January-2026 forecast of ASIC-based AI-server unit share for 2026
Scope & caveats

Forecast of server-unit shipments, not an observed accelerator, revenue, energy or workload share; retain the historical forecast date.

Default
Triton is the default GPU code generator inside torch.compile/TorchInductor (since PyTorch 2.0), now with NVIDIA, AMD and Intel backends

Portability layers: paying the option premium up front

If lock-in is a contract, portability layers are the clause that lets you exit — but you pay for the clause whether or not you exercise it. The strategic decision is how much of your stack to write against a portable abstraction versus a vendor-native one, and it is made years before you would ever switch. Three layers matter in 2026, at three different altitudes.

The framework layer (PyTorch / JAX). The single most important portability decision most teams make is simply to write models in PyTorch (or JAX) rather than in raw CUDA. PyTorch is the great equalizer: idiomatic PyTorch can target CUDA, ROCm and, through PyTorch/XLA, TPU, but operator coverage, numerical behavior, compilation, performance and distributed-runtime support must be validated per backend. This is why ROCm's porting tax is low for PyTorch code and high for hand-written CUDA: the abstraction already exists, and AMD invests heavily in keeping the AMD backend upstreamed. The lock-in you create by dropping into vendor-native kernels for the last 10% of performance is the lock-in you will pay to undo.

The kernel layer (Triton / MLIR). Below the framework sits the kernel, and this is where CUDA's deepest moat is being abstracted away. Triton — the kernel language built on MLIR — offers multiple backends, but a kernel is portable only across the specific devices, operators, layouts, numerical behavior and performance targets validated in CI, and since PyTorch 2.0 it has been the default code generator inside torch.compile's TorchInductor. Its governance moved from OpenAI to a community triton-lang project with contributions from NVIDIA, AMD, Intel, Meta, IBM and Red Hat — a deliberately multi-vendor structure. A custom Triton kernel can reduce porting work, though it is not automatically portable or performance-equivalent across vendors; the same kernel written in CUDA C++ with wgmma and tensor-memory-accelerator intrinsics is not portable at all. This is the layer where the CUDA moat is most actively eroding, and where your kernel-authoring policy directly sets your future switching cost.

The serving layer (OpenAI-compatible APIs). For inference, the cheapest portability you can buy is an OpenAI-compatible HTTP surface in front of whatever engine you run. vLLM, SGLang, TensorRT-LLM, and their ROCm equivalents all expose a near-identical request API, so the application above never learns which silicon it is talking to. This makes the inference fleet the easiest place to be heterogeneous: route traffic to whichever accelerator delivers the best cost-per-token this quarter, behind a stable API. The lock-in at this layer is shallow by design.

Portability posture — where to abstract, where to go native
LayerPortable choiceVendor-native choiceOption premium of portableWhen native is right
FrameworkPyTorch / JAX (multi-backend)CUDA C++ application codeModel- and backend-specific; test operator coverage, numerics, compilation and distributed behaviorAlmost never; reserve for a profiled hot path
KernelTriton / MLIR (PTX + AMDGCN + Intel)CUDA C++ + cuDNN / CUTLASSOften single-digit % on the hottest kernelsFrontier kernels where the last 10% pays for itself at scale
CollectivesNCCL-API-compatible (RCCL, oneCCL)NCCL tuned to one fabricLow at the API; tuning differs per fabricWhen fabric-specific tuning is the bottleneck
ServingOpenAI-compatible API over vLLM/SGLangVendor-locked inference microserviceAPI syntax may be small; semantics, tokenizer, batching, errors, observability and SLO behavior require validationWhen the native service's validated semantics or performance justify the switching cost
Scale-up fabricUALink / Ethernet-based open fabricNVLink (proprietary; NVIDIA-controlled, licensed via NVLink Fusion)Bandwidth/latency gap vs the leader, todayWhen you want the densest, fastest scale-up domain now
The trade is performance-now vs switching-cost-later. "Option premium" is the performance or effort you forfeit today to keep the exit cheap.

Quantifying the switching cost

Lock-in is only actionable once you put a number on it, and the number is the total cost of switching regimes — which is not one cost but four, and they are paid at different times.

The porting cost is engineering time to make your stack run at all on the new regime. For PyTorch-native training and inference it is genuinely modest — days to weeks, mostly dependency and container work. For a codebase laced with hand-written CUDA kernels, custom NCCL tuning, or TensorRT-LLM-specific serving paths, it is a quarters-long project staffed by scarce specialists. This cost scales with how much vendor-native code you wrote in the first place, which is why the kernel-authoring policy of two years ago is the switching cost of today.

The realized-MFU recovery cost is the larger and more often ignored term. Getting your model to run on the new regime is not getting it to run well. The gap between first-light and tuned-throughput is the period where you are paying for silicon at a fraction of its peak while engineers chase the missing 20–40% of MFU — re-tuning kernels, fixing collective overlap, matching the serving framework's batching to the new memory system. On an immature stack this recovery can take longer than the port itself, and during it your cost-per-token is underwater.

The day-0 coverage cost is the opportunity cost of not being on the regime where new models land first. New architectures, attention variants, and quantization schemes ship and are tuned on CUDA first; on other regimes they arrive days to weeks later, after community or vendor porting. For a frontier lab that must serve the newest model the hour it drops, that lag is lost revenue and a competitive gap — and it is a recurring tax, not a one-time switching cost.

The operational-complexity cost is the standing overhead of running a heterogeneous fleet: two toolchains, two driver/firmware matrices, two sets of failure modes, two on-call playbooks. It is paid every day, not once, and it is the reason many operators stay single-vendor despite a price advantage elsewhere — the second toolchain needs an owner, and that owner is a headcount.

12 custom operators × 8 h; collective work 40 h; precision qualification 40 h; CI/rollback work 24 h; loaded labor $150/h; sustaining 20 h/month; gross service saving $8,000/month for 12 months.modeled
Pay for a second software stack only after its release qualifies — input ledger
Scope & caveats

Matched qualified service on CUDA/ROCm. Unsupported labor, savings and engineering budgets are explained in the opening callout; ROCm supplies compatibility methods and Chapter 1.8 the full appraisal.

Add operator, collective, precision and release-engineering hours; multiply by the loaded hourly rate for the initial port. Annual sustaining is monthly hours × billing months × rate. First-year surplus is gross monthly saving × months minus port and sustaining. Dividing those two costs by billing months gives the gross-saving flip threshold; keep the untouched baseline and tuned manifest as separate runs.

$30,000 first-year surplus in the scenario; $5,500/month break-even gross saving; −$6,000 at $5,000/month.derived
Pay for a second software stack only after its release qualifies — result and flip threshold
Scope & caveats

Port hours = 12 × 8 + 40 + 40 + 24 = 200 h. Initial port = 200 h × $150/h = $30,000. Annual sustaining = 20 h/month × 12 months × $150/h = $36,000. First-year surplus = 12 × $8,000 − $30,000 − $36,000 = $30,000. The first-year gross-saving threshold is ($30,000 + $36,000)/12 = $5,500/month. At $5,000/month the result is −$6,000. Benchmark the untouched portable baseline and tuned manifest separately; count the hours between them rather than calling the gap a ROCm tax.

Fund the bounded port in this scenario, with acceptance at the complete tuple. Release requires the quality, p99 latency, collective scale, checkpoint/recovery and rollback gates to pass on that manifest; an API-compatible NCCL/RCCL call is not that evidence. Below the stated saving threshold, keep the incumbent unless a separately funded continuity requirement pays the difference.

Method: AMD ROCm documentation and compatibility guidance. Chapter 7.14 owns the next handoff.

Deep dive: why the CUDA moat is libraries, not the language — and what that means for erosion

The popular framing — "CUDA is a programming language and that is the moat" — gets the mechanism wrong and therefore mispredicts where the moat erodes. The CUDA language (a C++ dialect with a launch syntax) is the least defensible part of the stack; HIP source-translates it almost mechanically, and most production code never touches it directly because it lives one or two layers up in PyTorch. The actual moat is the library and tooling estate: cuDNN's hand-tuned convolution and attention kernels, cuBLAS and CUTLASS for the matrix math, NCCL for collectives that overlap with compute on NVLink, TensorRT-LLM for serving, plus Nsight profilers and a debugger ecosystem — and, above all, the network effect that every researcher prototypes on NVIDIA, so every new technique is born CUDA-tuned.

This diagnosis predicts where the moat erodes. It does not erode by someone cloning the CUDA language; it erodes by abstractions that make the library estate irrelevant. Two are doing exactly that. PyTorch hoists the application above the library layer, so a model is portable even when the kernels beneath it are not — and AMD/Intel invest to keep their backends upstreamed. Triton/MLIR attacks the kernel layer itself: a fused kernel written once compiles to PTX, AMDGCN, and Intel GPU code, and because Triton is the default inside torch.compile, ordinary PyTorch users generate portable kernels without knowing it. The moat does not fall; it gets routed around, one layer at a time, on the workloads the ecosystem chooses to fund. The practical reading for a buyer: the moat is strongest exactly where you are still writing vendor-native kernels, and weakest where you have already moved up to PyTorch + Triton. Your switching cost is, to first order, a measure of how much vendor-native code you let accumulate. → precision and quantization kernels in Chapter 7.10.

Single-vendor vs heterogeneous fleets

Fleet composition is a strategy decision that reaches you disguised as a procurement one. A single-vendor fleet buys operational simplicity: one toolchain, one driver matrix, one set of failure modes, one on-call playbook, and the deepest day-0 model coverage if that vendor is NVIDIA. It pays for that simplicity with price exposure — you are a captive buyer with no credible alternative to walk to — and with supply exposure, because a single allocation queue gates your entire build (→ Chapter 7.6 on the HBM/CoWoS upstream gate).

A heterogeneous fleet inverts both terms. Running NVIDIA for the bleeding edge and AMD, TPU, or Trainium for the steady-state workloads where they have reached parity buys price discipline — a real outside option in every negotiation — and supply diversity across multiple allocation queues. It pays for that leverage with the operational-complexity cost above: two toolchains that each need an owner, and a routing layer (best built at the serving API) that sends each workload to the silicon with the best current cost-per-token. The heterogeneous strategy is most defensible at the inference layer, where an OpenAI-compatible API keeps the calling application unchanged across the swap and the porting, throughput-recovery and operating costs separated above are what you actually pay. It is least defensible at the frontier-training layer, where day-0 coverage and the deepest scale-up domain still favor a single vendor.

What forces this question onto every roadmap is the forecast rise of custom ASICs toward roughly a quarter of 2026 AI-server shipments. Every Maia, MTIA, TPU, and Trainium is, by construction, a non-CUDA software regime — and the hyperscalers building them have already paid the heterogeneity tax internally because the cost-per-token and supply-control upside justified it at their scale. Whether everyone smaller can capture the same upside, or whether the operational-complexity cost eats it, comes down to their workload mix and the depth of their engineering bench. → the merchant-silicon disruption in Chapter 7.5; the qualified configuration from Chapter 7.11 enters the lifecycle cost model in Chapter 1.8.

Deep dive: the collective-communication library as a hidden lock-in (NCCL vs RCCL)

Buyers obsess over the kernel and the framework and forget the layer that most directly gates multi-node training MFU: the collective-communication library. NCCL (NVIDIA) implements all-reduce, all-gather, and reduce-scatter tuned to NVLink and the InfiniBand/Spectrum-X fabric, overlapping communication with compute to reduce exposed gradient-sync time; dependencies and network contention can still stall the tensor cores. On a large synchronous training run, a poorly-overlapped collective can leave substantial compute time exposed — and it is invisible on the datasheet.

The portability story here is real but incomplete. RCCL (AMD) is API-compatible with NCCL, so PyTorch's distributed layer calls it transparently — the supported call surface can remain unchanged; custom extensions and build dependencies still need verification. But tuning is fabric-specific: the topology-aware algorithms, the buffer sizes, the ring-vs-tree selection that NCCL has accreted over years on NVLink are not automatically optimal on AMD's Infinity Fabric or a UALink topology. API compatibility can reduce porting work; it does not establish realized collective bandwidth, whose tuning cost folds straight back into the software budget. This is why a multi-node training benchmark — not a single-GPU one — is the only honest way to compare regimes for training: the lock-in lives in the spaces between the GPUs as much as on them. → fabric topology and collectives in Chapter 8.2.

This chapter is the software contract that sits underneath the hardware choices in Chapter 7.1 (the accelerator taxonomy), Chapter 7.2 (NVIDIA), Chapter 7.3 (AMD/ROCm), and Chapter 7.4 (TPU/Trainium and their XLA/Neuron regimes). The merchant-silicon disruption that multiplies the number of regimes is in Chapter 7.5; the HBM/CoWoS supply gate that makes single-vendor fleets risky is in Chapter 7.6. The precision and quantization kernels whose portability this chapter assumes are engineered in Chapter 7.10. The realized-MFU benchmark is one acceptance input in Chapter 7.14; Chapter 14.1 defines useful output, Chapter 7.11 selects eligible equipment, and Chapter 1.8 prices that service and scores it economically. The scale-up fabric lock-in is engineered in Chapter 8.2 and taxonomized in Chapter 8.9; goodput as the metric that replaces paper FLOPS is reframed in Chapter 12.2.
Cite this chapter
Fehn, J. (2026). Software Ecosystems & Lock-In (Chapter 7.9). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-9-software-ecosystems-and-lock-in (accessed 2026-09-29).
@misc{aidc-7-9,
  author       = {Fehn, Jacob},
  title        = {Software Ecosystems & Lock-In (Chapter 7.9)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-9-software-ecosystems-and-lock-in},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit