The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Chapter 7.7

In this chapter · 8 sections
Term help

Advanced Packaging & the Integration Substrate

Packaging yield on the largest interposers still caps how many accelerators ship, so choose a package whose memory layout, die crossings, thermal path and yield can meet the workload and delivery plan; an interposer roadmap is not a qualified accelerator.

GOODPUTDENSITY-RAMP

What you'll decide here

  1. Which 2.5D packaging family (CoWoS-S silicon interposer, CoWoS-R RDL, or CoWoS-L stitched silicon bridge) your accelerator targets — because that single choice sets the reticle-multiple ceiling, the HBM-stack count per package, the yield curve, and which OSAT can build it.
  2. Whether to stay monolithic or disaggregate into chiplets across a UCIe die-to-die fabric — trading the yield and reuse upside of small dies against the area, power, and latency tax of every die crossing.
  3. How many HBM stacks the program needs (8, 12, or 16+) and therefore which reticle-multiple and interposer technology that demands — the package-area math, not the memory contract, is the real gate on capacity-per-package.
  4. Which silicon or RDL interconnect structure supplies the routing density and reach you need, separately from a glass-core substrate choice, and how routing, warpage and yield affect cost and the cooling envelope.
  5. Which logic, memory, packaging, assembly or test milestone controls delivery: a missing CoWoS allocation can leave an otherwise ready final-assembly line idle, so give that predecessor an owner and a recovery path.

"How fast is this chip?" used to be answered at the transistor: a smaller node, more transistors, a faster clock. That era is functionally over for AI silicon. A single reticle field — the largest area a lithography scanner can pattern in one shot — is about 858 mm² (roughly 26 × 33 mm), and the largest useful monolithic die is already pressed against that limit. An H100-class GPU is essentially a full-reticle die. You cannot make the logic meaningfully bigger by printing it; you have hit the reticle wall. Everything that has happened to AI accelerator performance since is, at the physical level, a packaging story: how to stitch multiple reticle-sized fields together, how to bolt a dozen HBM stacks beside the logic, and how to wire it all with enough bandwidth that the assembled system behaves like one chip. The package is the integration substrate, and the firm that controls the largest, highest-yielding package controls how much compute the world can ship.

What follows is the engineering of that substrate: the 2.5D/3D packaging taxonomy and why it is the most-cited binding constraint on AI compute through 2030; the CoWoS-S/R/L families and SoIC / hybrid bonding, and the reticle-stitching and per-package-area ceilings each imposes; the interposer fork (silicon vs RDL vs glass) and its reach/cost/yield triad; the way HBM-stack-count-per-package falls directly out of interposer area — the engineering driver behind the allocation view of Chapter 7.6; chiplet disaggregation and UCIe and the economics of every die-to-die crossing; and the thermal, warpage, and yield consequences of building packages the size of a coaster. Procurement and allocation of CoWoS capacity stay in Chapter 7.6 and Chapter 2.3; here the concern is the physics.

Why packaging became the binding constraint

The constraint moved upstream the same way power did in the facility. The industry assumed the gate on accelerator supply was wafer-out logic — how many GPUs a leading-edge fab could print. It is not. The gate is advanced packaging capacity, specifically TSMC’s CoWoS lines, which through 2026 are fully booked across both the CoWoS-S and CoWoS-L families. Reported allocation for the year runs to roughly a million wafers — NVIDIA alone ~595,000, about 60%, with Broadcom ~15% and AMD ~11% (Tom’s Hardware, May 2026) — while monthly TSMC capacity climbs from ~75–80k wafers toward ~120–140k by year-end, plus ~50–60k at OSAT partners (TrendForce, June 2026). TSMC discloses only relative growth, so every absolute figure is an analyst estimate, and none of them is this buyer’s monthly allocation. A logic wafer that cannot be packaged is not an accelerator: a TSMC/CoWoS design can hold N3/N5 logic allocation yet ship late if its package slot or qualified HBM lot is missing, and a total wafer count cannot be converted into delivered GPUs without package area, yield, test and product mix.

Choosing a more aggressive package — more reticles stitched together, more HBM stacks — buys performance, but it consumes more interposer area per accelerator (a 5.5×-reticle part eats far more interposer area, and yield risk, than a 2× part) and can put the program onto a newer process with less production-yield evidence. A conservative package preserves yield and supply but caps memory bandwidth and compute area. No axis is free: the package is where the GOODPUT ambition of the silicon team collides with the DENSITY-RAMP reality of the packaging line.

Compute and HBM memory dies, base die and TSVs sit above microbumps, interposer routes, C4 bumps, substrate and BGA. A conventional 26 × 33 mm reticle field is about 858 mm²; a larger patterned interposer needs stitching, adding alignment and yield constraints. With a top-mounted cold plate, buried-die heat crosses the overlying layers. Interposer area and substrate footprint are separate limits.

The 2.5D / 3D taxonomy

Advanced packaging splits cleanly into two regimes, and the fork between them is the first decision an architecture team makes. 2.5D places multiple dies side-by-side on a shared horizontal carrier — an interposer — that routes thousands of fine wires between them. This is how logic meets HBM today: the GPU die and its HBM stacks sit beside each other on a silicon or RDL interposer; a glass-core package substrate is a different layer, with HBM's wide thousand-plus-bit bus fanned out across that carrier. CoWoS is the canonical 2.5D flow. 3D stacks active dies vertically, face-to-face, bonded directly so that a die sits on top of another die rather than beside it. TSMC's SoIC and the hybrid-bonding processes underneath it are the 3D flow; AMD's MI300 stacks compute chiplets on base dies this way, and the HBM stack itself is a 3D structure internally.

The consequence of the fork is geometric. 2.5D buys you area — you spread components across a large flat carrier, and the limit is how large an interposer you can yield. 3D buys you proximity and density — wires between stacked dies are microns long instead of millimeters, slashing interconnect energy and latency, but you inherit a severe thermal problem (with a top-mounted cold plate, a buried lower die must send heat through overlying material toward the spreader and plate) and a yield-multiplication problem (a bad die anywhere in the stack can scrap the whole assembly). Some 2026 accelerators are hybrids: a 2.5D interposer carrying 3D-stacked logic and 3D-stacked HBM. The package is two integration technologies at once.

CoWoS family fork: which 2.5D flow, and what it ceilings
FamilyInterposer / carrierReticle-multiple ceiling (2026)HBM stacksBest fitCost & yield posture
CoWoS-SFull silicon interposer (TSV-bearing)~3.3× reticle (mature); pushing higher8–12Today's mainstream GPU + 8 HBM (H100/H200-class)Highest interposer cost; best signal integrity; mature yield
CoWoS-RRDL interposer (organic, no large Si)Lower than -S; smaller spansUp to ~8Cost-sensitive parts; fewer stacks; warpage-tolerantCheaper carrier; relaxed routing; warpage easier to manage
CoWoS-LStitched local-silicon-interconnect (LSI) bridges in RDL3.5× in production; 5.5× ramping 2026 (100×100 mm); ~9.5× in development (2027)Product-specific stack count and placementReticle-stitched multi-die GPUs + many HBM (B200/Rubin-class)Most area for the money; stitching = new yield-loss modes
TSMC CoWoS variants are the dominant 2.5D flows for AI accelerators in 2026. Reticle-multiple and HBM-stack ceilings are TSMC roadmap/symposium figures; OSAT alternatives (ASE, Amkor, Samsung) exist but trail on the largest reticle multiples. "Reach" is max routable die-to-die span.

CoWoS-S uses one big slab of TSV-bearing silicon as the interposer — offering dense silicon routing, but a large silicon interposer is expensive and its own reticle-stitching limits cap how big it can get. CoWoS-R drops the silicon for an organic RDL carrier — with different routing, warpage and assembly constraints; price, yield and enabled stack count require the selected flow. CoWoS-L is the 2026 frontier: it abandons the single monolithic silicon interposer for an RDL carrier with small local silicon interconnect bridges embedded only where dense die-to-die routing is needed — letting the package span far more than a single reticle (3.5× in production since 2024, 5.5× certified and ramping in 2026, and a ~9.5× node in development for ~2027 supporting 12+ HBM4-generation stacks) while spending silicon only where it earns its cost. The consequence of choosing -L is that you inherit a new family of yield-loss mechanisms — bridge-to-RDL alignment, stitching seams, larger-package warpage — in exchange for a larger qualified routing area; compare usable routes and package yield, not the family suffix.

SoIC, hybrid bonding, and the 3D ceiling

Where 2.5D spreads outward, SoIC (System-on-Integrated-Chips) stacks upward using hybrid bonding — a bumpless, copper-to-copper, dielectric-to-dielectric direct bond that eliminates the solder microbumps of older 3D stacking. The payoff is interconnect density: classic microbump pitch is ~40–50 µm; hybrid bonding takes the bond pitch into the single-digit-micron range and, on the roadmap, sub-micron. That is one-to-two orders of magnitude more vertical connections per unit area, which is what lets two dies behave electrically as if they were one — picojoule-per-bit interconnect, full-bandwidth die-to-die, latency measured in fractions of a clock. AMD's MI300-class parts stack compute chiplets on a base die this way; HBM's internal joining process is a separate question, and it has not settled in hybrid bonding's favour: SK hynix announced 16-layer HBM3E in 2024 and shipped 12-layer, 48 GB HBM4E samples in June 2026, both on Advanced MR-MUF. Treat stack height and joining process as independent attributes, per supplier and per product.

The 3D ceiling is thermal, and it is a hard one. When you put logic on top of logic, one die is always buried — and with the cold plate above the stack it is the lower die whose heat must traverse the upper die to escape, while power density at the bond interface can exceed what any external cooling can extract without throttling. Draw the cooler, the die order and the heat path for the actual assembly before naming the bottleneck: AMD puts second-generation 3D V-Cache below the CPU cores precisely to keep the hotter logic facing the cooler. This is why 3D-on-logic is selective — you stack the parts that benefit most from proximity (cache, certain compute tiles) and keep the hottest, highest-power logic where it can see the cold plate directly. The decision to go 3D is a decision to make the cooling problem of Part 5 harder in exchange for interconnect you cannot get any other way. The forward pointer is explicit: large-package hotspot and heat-flux behavior is engineered in Chapter 5.1 (the density wall) and removed by the cold plate in Chapter 5.4.

The interposer fork: silicon vs RDL vs glass

Underneath the CoWoS branding sits the real material decision: what is the carrier made of? Separate the interposer or redistribution layer from the package substrate core before comparing reach, cost and yield. Silicon interposers route the finest lines and carry TSVs for power and signal straight through — dense routing for HBM's wide bus — but a large silicon interposer is costly, is itself reticle-limited (you cannot pattern one bigger than a scanner field without stitching), and its CTE mismatch against the organic substrate drives warpage as it grows. RDL (redistribution-layer) carriers can be made large, but their routes are coarser than local silicon interconnect: CoWoS-L uses silicon bridges where the RDL carrier alone cannot supply the required density. Cost and warpage depend on that complete stack-up, rather than following automatically from the carrier material. Glass is the emerging third option: glass-core package substrates and glass interposers are separate structures; proposals for each target the dimensional stability and flatness of silicon with potential panel-scale manufacturing, with excellent high-frequency loss characteristics — but the ecosystem is immature, with handling, via-formation, and crack-propagation risks still being industrialized through 2026.

Interposer material triad: reach, cost, yield, and what it costs you downstream
CarrierRouting density / reachRelative costWarpage / yield risk2026 statusDownstream consequence
Silicon interposerHighest density; best SI; reticle-stitch limitedHighestCTE-mismatch warpage grows with areaMainstream (CoWoS-S)Largest CoWoS-capacity draw per part
RDL (organic)Coarser routing; large carriers feasibleLowestMost warpage-tolerant; relaxedMainstream for cost parts (CoWoS-R) & as CoWoS-L baseCaps stacks/bandwidth; cheaper supply
Glass coreSilicon-like density; panel-scale, low lossMid (promised); ecosystem premium todayFlat & stable, but crack/handling riskPre-volume / ramping 2026–2027Could relieve the area ceiling if it yields
Qualitative practitioner ranking, 2026. Glass is pre-volume for high-end AI accelerators; figures reflect the technology's promise and current maturity, not a shipped commodity.

HBM-stack-count is an interposer-area problem

This is the link between this chapter and the memory chapter, and it is close to an identity: HBM stacks per package is a function of interposer area, not of how much HBM you can buy. Each HBM stack occupies a fixed footprint beside the logic die and must be reached by a thousand-plus-wire bus fanned across the interposer. Read the platform sizes as envelopes, not as a law of stack count: TSMC's ~3.3–3.5×-reticle platforms accommodate eight HBM stacks, and the ~5.5× platform that CoWoS-L begins delivering in volume in 2026 accommodates up to twelve; 16+ stacks need the next reticle-multiple node still ramping. Interposer size follows the whole floorplan — logic-die area, I/O, and routing as well as stacks — not the HBM generation alone. So when Chapter 7.6 describes HBM as the top-3 BOM line and the binding allocation gate, the engineering reason a given accelerator carries the stack count it does lives here: the package area the program chose set the stack count, which set the memory capacity and bandwidth, which set where it lands on the memory roadmap.

Memory ambition and packaging ambition are therefore one decision. A team that commits to an 8-stack, ~288 GB HBM4 accelerator (Rubin-class, ~2.75 TB/s/stack — ~22 TB/s per package) has by that act committed to a large-reticle CoWoS-L package, its yield curve, and its slice of the scarcest packaging capacity on earth — though NVIDIA has not disclosed Rubin's exact reticle multiple, and eight stacks alone do not require the full 5.5×. You cannot buy your way to more bandwidth without buying your way to more interposer area, and interposer area is exactly the thing in shortest supply. The procurement reflex — "order more HBM" — is necessary but not sufficient; without the CoWoS slot to mount it on, the HBM is inventory, not bandwidth. That allocation logic lives in Chapter 2.3; the engineering driver lives here.

858 mm² (26 × 33 mm)derived
single reticle field (≈26×33 mm) — the area limit every advanced-packaging technique exists to defeat
Scope & caveats

A named lithography scanner exposure field, not a guarantee of usable die area, yield or a stitched interposer ceiling. Guide rectangle area: 26 × 33 = 858 mm².

3.5× → 5.5× (2026) → 9.5× (dev)forecast
TSMC package roadmap: distinguish interposer reticle multiple from 100 × 100 mm substrate footprint
<10 µm
hybrid-bonding bond pitch (vs ~40–50 µm microbump), heading sub-micron — the density behind SoIC 3D stacking
>20 Tbps/mm
UCIe 3.0 signaling reaches 64 GT/s; a third-party 64 Gbps PHY demonstration exceeds 20 Tbps/mm edge density, while UCIe-3D reaches up to ~300 TB/s/mm² at 1 µm pitch
~288 GB
HBM4 capacity per Rubin-class package at ~2.75 TB/s/stack (~22 TB/s per package) — the memory the package area enables
Scope & caveats

2.75 TB/s = Rubin config (22 TB/s ÷ 8× 36 GB stacks). JEDEC HBM4 base ~2.0 TB/s/stack; shipping 2026 HBM4 parts run ~2.56 (SK hynix) → ~2.8 (Micron) → up to ~3.3 (Samsung) TB/s/stack.

~120–140k wafers/monthestimate
TSMC CoWoS capacity in 2026, from ~75–80k wafers/month; +50–60k at OSAT partners (TrendForce, June 2026)
Scope & caveats

Analyst-reported capacity, not a buyer allocation; TSMC discloses only relative growth.

~1.0M wafersestimate
2026 CoWoS allocation implied by reported shares: NVIDIA ~60%, Broadcom ~15%, AMD ~11% (May 2026)
Scope & caveats

Reported allocation shares; allocation is contractual and the total is implied from the NVIDIA share.

~60%estimate
NVIDIA’s reported share of 2026 CoWoS allocation (~595k wafers)
Scope & caveats

Reported allocation share; allocation is contractual.

Chiplet disaggregation and UCIe

Once the reticle wall makes one big die impossible, the question becomes how to build a big system from small dies — and that is the chiplet decision. Disaggregation splits a would-be monolithic SoC into multiple dies ("chiplets") that are packaged together: compute tiles, I/O dies, cache dies, each potentially on a different process node tuned to its job. The upside is real and quantifiable. At a fixed defect density, smaller die area improves individual-die yield — defect density punishes large dies brutally, so four small dies can improve known-good-die supply, but assembly yield and testing decide the cost of a usable package. You reuse a chiplet across many products. You put I/O on a cheap mature node and spend leading-edge wafers only on the logic that needs them. AMD productized this years ago; it is now the default architecture for the largest accelerators.

The cost of disaggregation is paid at every die-to-die crossing, and this is where UCIe (Universal Chiplet Interconnect Express) enters. A monolithic SoC pays wire, buffering and switching energy on-die. A chiplet system must drive and receive across a physical die boundary, with serialization and encoding determined by its interface — spending area on PHYs, power on the link (picojoules per bit that add up across terabytes per second), and latency on the crossing. UCIe standardizes that interface to support interoperability within an implemented, qualified profile: UCIe-S targets standard 2D packages, UCIe-A targets advanced 2.5D packages, and UCIe-3D covers 3D integration; 64 GT/s arrived with UCIe 3.0, while the 3D interface was introduced in UCIe 2.0. UCIe's published catalogue lists 3.0 as the current revision; voltage, bump pitch, lanes, protocol, test access and thermal/mechanical compatibility still need agreement. The strategic opportunity is an open chiplet market: a buyer can, in principle, assemble an accelerator from best-of-breed dies rather than one vendor's monolith — the packaging-era analog of the open-system disaggregation Part 8 describes for the network.

Useful one-way payload 1.0 TB/s; transmitted-bit factor 1.10; combined PHY energy A 1.0 pJ/bit, B 3.0 pJ/bit; link allowance 20 W; buried-die power 600 W; path resistance 0.040 K/W; permitted rise 30 K.modeled
Choose a package only if its die crossing and heat path both close — input ledger
Scope & caveats

Hypothetical package, combined PHY endpoints and top cold plate. Unsupported energy/thermal inputs and rationale are in the opening callout; Chapters 5.1 and 5.7 supply methods, not numeric bounds.

Transmitted rate = 1.0×10¹² B/s × 8 bit/B × 1.10 = 8.8 Tb/s. Combined endpoint power A = rate × 1.0 pJ/bit = 8.8 W; B at 3.0 pJ/bit needs about 26 W and exceeds 20 W. Buried-die rise = 600 W × 0.040 K/W = 24 K, within 30 K. A’s payload threshold is 20 W/(8×1.10×1.0 pJ/bit), about 2.3 TB/s; 2.3 itself is above that exact bound. Independently, the thermal crossover is 30 K/600 W = 0.050 K/W; at 0.060 K/W the rise is 36 K and fails.

A: 8.8 W link and 24 K rise; B: about 26 W, rejected. A flips above about 2.3 TB/s or 0.050 K/W.derived
Choose a package only if its die crossing and heat path both close — result and flip threshold
Scope & caveats

Transmitted rate = 1.0×10¹² B/s × 8 bit/B × 1.10 = 8.8 Tb/s. Combined endpoint power A = rate × 1.0 pJ/bit = 8.8 W; B at 3.0 pJ/bit needs about 26 W and exceeds 20 W. Buried-die rise = 600 W × 0.040 K/W = 24 K, within 30 K. A’s payload threshold is 20 W/(8×1.10×1.0 pJ/bit), about 2.3 TB/s; 2.3 itself is above that exact bound. Independently, the thermal crossover is 30 K/600 W = 0.050 K/W; at 0.060 K/W the rise is 36 K and fails.

Qualify A’s actual bump pitch, voltage, lane width, protocol, test access and assembly process. A catalogue UCIe version alone cannot make chiplets interchangeable. Request the thermal and assembly-yield evidence before reserving its HBM placement and packaging capacity.

Method: UCIe Consortium specification catalogue. Chapter 7.14 owns the next handoff.

UCIe 3.0
UCIe catalogue; contract the implemented profile
Scope & caveats

Specify implemented package/PHY/protocol profile; not automatic chiplet interchangeability.

Deep dive: the die-to-die tax, quantified — when disaggregation stops paying

Disaggregation is not free; the break-even is an engineering calculation. Every signal that would have stayed on-die in a monolith now crosses a die boundary, and that crossing costs three things. Area: the D2D PHY (the SerDes or parallel interface) consumes silicon on both dies — beachfront along the die edge that could have been compute. Power: on-die wires cost a fraction of a picojoule per bit; even an excellent D2D link costs more, and at the multi-terabyte-per-second bandwidths a GPU needs internally, the aggregate link power is a real fraction of the package budget. Latency: serialize-cross-deserialize adds cycles that on-die routing never pays.

Hybrid bonding changes the arithmetic. Because it pushes bond pitch into the single-micron range, UCIe-3D gets the per-bit energy and latency close enough to on-die that the crossing nearly disappears — which is precisely why the most aggressive disaggregation (many small dies) pairs with hybrid bonding rather than microbump 2.5D. The decision rule that falls out: disaggregate when the yield-and-reuse savings on the dies exceed the area/power/latency tax of the links between them, and reach for hybrid bonding when the chiplet count is high enough that microbump D2D would eat the savings. A two-die split over a coarse interface can be a net loss; a many-chiplet design over hybrid bonding is the architecture of the frontier accelerator. The standards-war analog for the scale-up fabric (NVLink Fusion vs UALink) lives in Part 8; here the fork is purely about the package.

Thermal, warpage, and yield: the consequences of a coaster-sized package

Large packages punish you three ways, and each is a downstream cost of the area you bought to defeat the reticle wall. Warpage is the first: a large silicon interposer on an organic substrate has a coefficient-of-thermal-expansion mismatch, and as the assembly heats and cools it bows. Past a certain span the bow threatens the solder joints to the board and the bond integrity across the package — which is a major reason CoWoS-L's RDL-plus-bridge construction exists (the organic carrier is more compliant) and why glass cores are attractive (dimensionally stable and flat). Warpage is not a yield footnote; it is a first-order limit on how large a package you can build at all.

Yield is the second, and large packages multiply it sharply. A 12-HBM, multi-reticle, stitched package is the product of many independent yields (each die, each HBM stack, each bond, each stitch seam), and a defect anywhere late in the flow scraps an assembly carrying enormous accumulated value (a dozen HBM stacks and multiple known-good logic dies). This is why known-good-die testing before assembly is non-negotiable and why the largest packages carry the largest scrap-cost exposure: you are gambling the most expensive components on the last and least-reversible step. Thermal is the third: a large, dense package concentrates hundreds of watts into a small footprint with internal hotspots over the 3D-stacked regions, producing heat fluxes that air cannot remove and that even direct-to-chip liquid must be engineered for. The package's heat-flux map is the boundary condition the cold plate inherits — the explicit hand-off to Chapter 5.1 and Chapter 5.4.

How the package decision propagates

Walk the chain forward and the package sits at the center of it. The reticle-multiple you choose sets the interposer area; interposer area sets the HBM-stack count, which sets memory capacity and bandwidth (→ Chapter 7.6); the reticle-multiple also sets which CoWoS family and OSAT can build it, which sets yield and lead time (→ Chapter 2.3); the package's power and heat-flux map sets the cooling boundary condition (→ Chapter 5.1, Chapter 5.4); and the on-package power-delivery and di/dt behavior that a big, dense package demands is its own engineering problem (→ Chapter 7.12). The package is the substrate every other accelerator decision is mounted on. Get its reticle-multiple and family wrong and you have either an accelerator you cannot get capacity to build or one whose memory and compute ambition you under-shot — and unlike a board respin, a package family change is a multi-quarter, multi-million-dollar reset.

Deep dive: why a CoWoS-L respin is a worse mistake than a node respin

The package looks like a back-end detail you can adjust late. It is the opposite — in 2026 the package family is one of the least-reversible decisions in an accelerator program, for two reasons. First, allocation. CoWoS-S and CoWoS-L slots are booked quarters ahead and fully subscribed; changing reticle-multiple or family mid-program does not just mean re-engineering — it means re-queuing for capacity that is already someone else's, against a roadmap clock that does not stop. A node change at least keeps you in a wafer queue you may already hold; a package-family change can put you at the back of a line for the single scarcest manufacturing resource in AI hardware.

Second, co-design depth. The HBM stack count, the interposer routing, the chiplet floorplan, the power-delivery network, and the thermal solution are co-designed around the chosen package; pulling the reticle-multiple unwinds all of them. A team that scoped for 8 stacks on CoWoS-S and discovers it needs 12 is not editing a parameter — it is changing interposer family, re-floorplanning, re-doing power and thermal, and re-entering the allocation queue. The discipline is the same one Part 1 preaches for the facility: identify the irreversible decision (here, reticle-multiple and CoWoS family) and over-specify or hedge it at scoping time, because re-deciding it costs a generation. The procurement-side mitigation — deposits, allocation locks, design-for-substitution — is owned by Chapter 2.3; the engineering reason it is irreversible lives here.

The package is the substrate the rest of the silicon stack mounts on. HBM as a top-3 BOM line and the allocation gate it forms is owned by Chapter 7.6; this chapter is the engineering driver behind its stack-count-per-package view. CoWoS/HBM long-lead procurement, deposits, and design-for-substitution live in Chapter 2.3. The hyperscaler XPUs and custom ASICs that make these packaging choices are profiled in Chapter 7.4 and Chapter 7.5. The on-package power delivery and di/dt physics a large dense package demands are in Chapter 7.12; the host-attach and system composition around the package in Chapter 7.8. The large-package heat-flux problem this chapter hands off is engineered in Chapter 5.1 and removed by direct-to-chip liquid in Chapter 5.4. The package-vs-fabric boundary — when to build a bigger package vs wire more packages together — is set at the rack in Chapter 7.13. The consolidated packaging-and-memory roadmap through 2030 is in Chapter 16.2.
Cite this chapter
Fehn, J. (2026). Advanced Packaging & the Integration Substrate (Chapter 7.7). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-7-advanced-packaging-and-the-integration-substrate (accessed 2026-09-29).
@misc{aidc-7-7,
  author       = {Fehn, Jacob},
  title        = {Advanced Packaging & the Integration Substrate (Chapter 7.7)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/part-7-compute-silicon-and-system-integration/7-7-advanced-packaging-and-the-integration-substrate},
  note         = {Accessed 2026-09-29}
}
Spotted an error? Suggest an edit