Chapter 8.9
In this chapter · 7 sections
Physical-Layer & Interconnect Taxonomy
Every link bets on the cheapest, lowest-power medium that closes the required channel and error budget; faster lanes consume copper margin, so the selected host, module, temperature and repair path decide where that bet gives way to optics.
What you'll decide here
- Where the copper-to-optics boundary falls for the selected host and interface: passive DAC, active copper or optics, with verified channel loss, reach and error margin rather than a universal reach-halving rule.
- Which DSP placement you certify — fully retimed, LRO/RTLR (half-retimed), or LPO — trading module-retiming power against host-board co-design burden and multi-vendor interoperability risk. Compare total system power at equal delivered bandwidth on the named host/module/channel pair before selecting bids.
- Which pluggable form factor (QSFP-DD, OSFP, OSFP-XD) and thermal lid (flat-top vs finned) you standardize on — a cage decision that is mechanically locked into the switch and the rack airflow for the life of the hall.
- How much of the reliability and goodput budget you spend on optics — a single unstable link can stall a synchronous collective and waste the waiting GPUs’ work, while a surviving redundant path can absorb the same event. Price hard failures, flaps and repair disturbances through the actual routes and affected jobs, not component frequency alone.
- Which FEC, transmitter TDECQ and passive-channel discipline you commission to, because a link that becomes unstable at field temperature can turn purchased bandwidth into retries and stalled steps. Keep transmitter qualification, channel certification and loaded error tests on their own measurement boundaries and limits.
The fabric chapters before this one — scale-up (8.2), scale-out protocols and topology (8.4, 8.5), the switch and NIC silicon (8.3) — all assume bits arrive at the far end of every link intact. This chapter is about the wire and the glass that make that assumption true, and about the physical decision that propagates through all of them: for a given reach at a given lane rate, what is the cheapest, lowest-power medium that still closes the link budget? The answer is one of three — a passive copper cable, an electrically-boosted copper cable, or an optical link — and the boundary between them is set by signal-integrity physics, not taste. It moves toward the operator's wallet every time the SerDes ladder steps up.
The answer is a portfolio. An AI hall can use “copper inside, optics outside” as a starting layout — keeping qualified short in-rack links electrical and spending optics where reach, routing or density earns it. Where you place that inside/outside seam is the most consequential physical-layer fork in the building. Set it too far toward copper and you stretch DACs past their margin and eat link flaps; too far toward optics and you burn megawatts of transceiver DSP power and a fortune in pluggables on spans copper could have carried for free. What follows walks the link budget, the SerDes ladder, the FEC/TDECQ margin discipline, the interconnect medium taxonomy, the form factors, and the DSP-placement spectrum — then treats optics as what it is at scale: a recurring network failure and operating-cost exposure whose component events and repairs can interrupt useful work.
The link budget: the equation every interconnect decision answers
Strip away the marketing and a link is an accounting problem. The transmitter launches a signal with some power and some quality; the channel — copper trace, cable, connectors, or fiber and its splices — subtracts loss and adds distortion; the receiver needs the signal to arrive above a threshold with enough eye-opening that the forward-error-correction engine can clean up what remains and hit the target bit-error rate. The link budget is that ledger: launch power and equalization on one side, insertion loss and impairments in the middle, receiver sensitivity and FEC coding gain on the other. Every entry in the interconnect taxonomy is just a different way of balancing that ledger over a different distance.
Two physical facts dominate the ledger in 2026. First, modern datacom links are PAM4, not the older NRZ. PAM4 encodes two bits per symbol using four amplitude levels, doubling throughput at a given symbol rate — but it cuts the vertical eye opening to roughly a third of NRZ's, so the same channel loss costs far more margin and the link leans hard on equalization and FEC to survive. Second, copper loss rises steeply with frequency. For fixed modulation and coding, doubling the per-lane service rate (50G to 100G to 200G to 400G) doubles symbol rate and Nyquist frequency. Copper loss rises with frequency, while return loss, crosstalk and package/connector impairments consume receiver equalization budget. That pressure drives the copper-to-optics migration, but not a universal reach-halving law: trace geometry, cable gauge, host capability and interface limits decide which span closes.
The SerDes ladder and the FEC/TDECQ margin discipline
The SerDes (serializer/deserializer) is one interface in the complete link: switch or NIC package → host board → module electrical attachment → optical PMD → remote endpoint. An 800G service can use eight 100G or four 200G electrical lanes, while 1.6T can use eight 200G lanes; a gearbox can change that count before the optical PMD, so optical lanes and fibers need not match the host. Record the exact port mode, coded electrical rate, symbol rate and lane/FEC mapping; use Chapter 8.10 for the resulting fiber and polarity schedule. The P802.3dj status tile below distinguishes draft interface work from available products. A module expecting a 112G-class electrical attachment cannot simply accept a port forced into a 224G-class mode. Backward operation needs an explicitly supported mode: a newer ASIC does not make arbitrary modules interoperable.
What keeps a PAM4 link alive across the budget is forward error correction. Datacom uses named Reed-Solomon FEC profiles, including RS(544,514) KP4, but no one pre-FEC input maps universally to one post-FEC result. State the Ethernet generation, lane rate, code/FEC mode, measurement interface and counter semantics. The P802.3dj MAC/PLS service-interface objective is not a portable installed-link threshold, and older Ethernet or InfiniBand profiles must be evaluated against their own normative and OEM limits. FEC is not free: it adds latency and a small bandwidth tax, and — critically — it has a cliff. Below its correction threshold the link is pristine; cross the threshold and errors leak through catastrophically. That makes margin the real engineering quantity. For optical transmitters, TDECQ (Transmitter Dispersion Eye Closure Quaternary) quantifies a transmitter penalty under the specified reference receiver and dispersion test; it is not a field receiver margin reading. A transmitter near its permitted TDECQ penalty leaves less tolerance for connector contamination, temperature drift and aging elsewhere in the system; those impairments can push the assembled link into errors or flaps. Installed margin also depends on the passive channel and remote receiver, so keep transmitter qualification separate from the field error record. TDECQ discipline at acceptance test is therefore a goodput decision, not merely a metrology one.
Scope & caveats
Public task-force material; product 1.6T ports are not proof of IEEE final approval.
Deep dive: PAM4, TDECQ, and why temperature drift is a commissioning trap
PAM4's three stacked eyes are not equal, and they are not stable. Transmitter non-linearity squeezes the upper and lower eyes differently, and the laser/driver behavior shifts with junction temperature — so an optic that passes TDECQ on a cool bench at the factory can drift out of margin once it is jammed into a populated finned cage drawing hot inlet air at the switch host’s limiting airflow condition. This is why TDECQ is specified with temperature, and why field-relevant qualification matters more than a datasheet number. The failure mode is insidious: the link does not fail at install, it fails statistically. As the eye closes under thermal stress, pre-FEC errors creep up toward the RS-FEC threshold; for most of the day the code corrects them and the link looks healthy on the dashboard, but at the thermal peak it tips over the cliff and the link flaps. The operator sees an intermittent, load-correlated, time-of-day-correlated fault that is maddening to chase — and the fault could be a temperature-sensitive transmitter, contaminated connector or marginal receiver/host channel; isolate those candidates because a rising pre-FEC counter alone cannot identify the cause. The discipline that prevents it is boring and non-negotiable: clean every connector (a single fiber contamination event can cost more margin than the entire fiber span), qualify transmitter performance at the temperatures it will actually run and reject conformance failures at acceptance, then certify passive loss and run the assembled channel’s loaded error tests at field conditions rather than discovering it in production. The connector-level EMC and bonding that protect copper links from a different class of impairment are canonical in Chapter 4.11.
Qualify this channel by separating three records. The supplier supplies transmitter/receiver conformance, including the prescribed TDECQ test and operating envelope. The installer supplies every fiber’s continuity, insertion loss, reflectance and end-face record using the specified reference method and uncertainty. The system test supplies loaded per-lane pre-FEC statistics, corrected and uncorrectable codewords, frame loss, link resets and recovery at temperature. A clean post-FEC counter does not establish reserve margin, and optical receive-power telemetry is neither calibrated channel certification nor TDECQ.
Scope & caveats
Undated application overview checked on this date, describing IEEE 802.3df-2024 DR8: eight parallel optical lanes per direction, 16 SMF fibers, 1310 nm, OS2 reach to 500 m. The page’s apparent symbol-rate typo is not used. This is a passive channel limit, not a field pre-FEC BER threshold.
Worked decision A: 800G DR8 budget and loaded acceptance
Use the passive channel boundary from module interface to remote module interface. Fiber loss is 0.100 km × 0.40 dB/km = 0.040 dB. Connector loss is 4 × 0.50 dB = 2.0 dB. Total passive loss is 2.040 dB; with 0.40 dB reserve and 0.20 dB uncertainty, required budget is 2.640 dB, about 2.6 dB. Against the source-backed 3.0 dB application limit, unrounded remaining allowance is 0.360 dB, about 0.4 dB: the planned channel passes.
One extra mated pair adds 0.50 dB, making the reserved total 3.140 dB, about 3.1 dB: fails. Keep four pairs or procure/test a lower-loss arrangement; the fifth pair’s serviceability benefit otherwise spends the reserve. This calculation does not credit FEC gain a second time against a PMD channel limit.
Installed acceptance remains HOLD. Certify every strand in both directions at the specified wavelength, identify the worst path, and demonstrate continuity/polarity and reflectance before connecting equipment. Then run the stated loaded test at the declared populated-host thermal condition, recording the exact pre-FEC limit and lane counter semantics alongside corrected/uncorrectable errors and resets. Acquire missing conformance limits or test evidence rather than treating the arithmetic as a measurement. The DR8 application overview supplies the loss basis; Chapter 8.10 supplies the plant map and Chapter 13.7 executes the loaded test. This chapter is the optical-budget home reached from Chapter 4.11.
A next-generation proposal inherits the installed loss record, not the preceding application’s acceptance. Keep the current DR8 plant record immutable and attach the candidate’s new endpoint and PMD qualification alongside it.
Worked decision B: 1.6T candidate on HOLD
The installed passive arithmetic is unchanged: 0.100 × 0.40 + 4 × 0.50 = 2.040 dB; including reserve and uncertainty gives 2.640 dB, about 2.6 dB required. Call the candidate’s verified draft-version loss ceiling Bdj. Its passive gate is Bdj ≥ 2.640 dB before display rounding, together with the candidate’s reach, reflectance and conformance conditions. The accessible IEEE optics material establishes draft work, not a complete set of installed-link limits.
HOLD the migration. Acquire the exact draft clauses accepted in the purchase contract, both endpoint specifications and interoperation evidence. The decision flips to a passive-budget pass only when those limits admit this channel with reserve; a ceiling below the unrounded 2.640 dB requires fewer/lower-loss connections or a different route. Full release still requires candidate-specific transmitter, receiver, FEC and loaded-temperature acceptance. Retain the working 800G service until that evidence closes; a compatible fiber count saves a possible re-pull but cannot authorize a transceiver-only upgrade. The calculation stays here; lane/polarity and repair access belong to Chapter 8.10.
The interconnect medium taxonomy: where qualified copper gives way to optics
With the budget and the ladder established, the taxonomy is itself a reach ladder. As distance grows you climb from passive copper, to electrically-assisted copper, to optics — paying more power and more dollars at each rung for more reach. The useful principle behind copper inside, optics outside is to keep short, dense, latency-critical scale-up links electrical when channel and service access close, then spend optics where reach, routing or density makes the complete system better. The rack boundary helps lay out routes; it is not a physical cutoff.
Passive DAC (direct-attach copper) is a plain shielded twinax cable with no active electronics — zero added module power, lowest cost, lowest latency, and no active parts to fail, but the shortest reach: about 1–2 m at 800G-class rates, and shorter again on 224G/lane channels (SemiAnalysis, January 2026). ACC/AEC (active copper / active electrical cable) adds a small redriver or retimer IC in the connector to boost and re-clean the signal, buying a few more metres — Credo’s 800G ZeroFlap AECs run 1–7 m — at the cost of retimer power, firmware/interoperation constraints and another population of components that can fail; these are 2025–26 reference points, so verify the qualified host/cable channel at your lane rate. AOC (active optical cable) is a captive, factory-terminated optical assembly — optics on both ends, fiber in between, sold as a single fixed-length cable; it reaches tens to hundreds of metres with its end count and length fixed by the assembly; breakout requires a purpose-built, supported cable. Pluggable transceivers are the fully modular option: a transceiver in a cage on each end with field-installed structured fiber between them, the most flexible and the basis of every large scale-out fabric — and, at ~15–17 W per 800G DSP module, the most power-hungry and failure-prone rung on the ladder.
| Medium | Added electronics | Typical reach (2025–26 reference) | Added power / port | Failure profile | Where it lives in an AI cluster |
|---|---|---|---|---|---|
| Passive DAC | Twinax and connectors; no active conditioning | ~1–2 m at 800G-class rates; shorter at 224G/lane | ~0 W added; host power separate | Cable/contact/host faults still possible | Scale-up inside the rack; shortest in-row hops |
| ACC / AEC | Redriver or retimer in assembly | Up to ~7 m (Credo 800G ZeroFlap, 112G/lane) | Retimer power; Credo: up to 50% below an AOC | Electronics plus cable/contact faults | Rack-to-rack where copper still reaches; ToR to spine within a row |
| AOC | Captive optical ends and fiber | Tens to ~100+ m, fixed assembly length | Optics at both ends (assembly-specific) | Optics failure rate, but factory-terminated; replace the whole assembly | Fixed point-to-point runs where length is known |
| Pluggable optics + structured fiber | Module at each end and certified passive channel | 100 m to multi-km by PMD class | ~15–17 W (800G DSP module) | Dominant network failure/flap source; module, fiber and contamination | Scale-out leaf-spine and beyond; the modular default |
Pluggable form factors: the cage you are mechanically married to
Once you choose pluggable optics, the form factor is a long-lived commitment because the cage and connector belong to the host. QSFP-DD and OSFP provide eight electrical lanes; OSFP-XD expands to sixteen, as defined by the QSFP-DD hardware specification and OSFP/OSFP-XD MSA. Those lanes, cage dimensions and thermal surfaces constrain the switch you live with, but do not define the optical PMD. Check module width, electrical attachment, power class, flat or finned heat-removal arrangement and every supported backward mode before ordering. A 1.6T label alone cannot establish lane count or compatibility, and a module that fits mechanically can still exceed the host’s cooling budget. Package and FRU consequences belong to Chapter 8.10.
The thermal lid is a second fork. Finned-top OSFP modules carry their own heatsink and rely on the switch's airflow — standard in air-cooled halls. Flat-top OSFP modules omit the integral fin because the host supplies the heat-removal surface — a riding heat sink pressed onto the module, which can itself be air-cooled or liquid-cooled; NVIDIA's air-cooled DGX H100/H200 links run flat-top modules. Finned and flat-top are not freely interchangeable in a given cage, but what decides the fork is the host SKU's cage geometry and qualified thermal interface, not the hall's cooling architecture. Get the host-SKU/module-SKU compatibility matrix — cage, heat-sink interface, power class, and the airflow or coolant conditions each is qualified at — before you order optics; discovering the mismatch later is a faceplate redesign, not a swap.
A last mechanical nuance, an engineering implication of the packaging rather than a vendor spec: high-lane-count modules multiplex links, and links from one module share its fate as a unit. A switch-side twin-engine 800G module driving 2×400G serves two distinct endpoints from one field-replaceable unit — NVIDIA's twin-port OSFP modules explicitly feed up to two adapters, and SuperPOD compute fabrics reach all eight GPUs of a node through four such switch-side two-port transceivers (NVIDIA, 2025) — so removing, swapping, or wholly losing one module interrupts both endpoints' links simultaneously (a single-engine fault may take only one). Draw the breakout map so fate-shared pairs never back each other up (never both planes of one accelerator through one module), and size spares by module count, not link count.
The DSP-placement spectrum: retimed, TX-retimed/RX-linear, LPO
Inside a conventional pluggable transceiver sits a power-hungry DSP that re-times and re-equalizes the signal on both the host (electrical) and line (optical) sides. That DSP spends power and cost on retiming, which is why linear and partially retimed architectures remove some or all of the module’s retiming work. The saving moves channel responsibility toward the host; count its equalization power and qualification burden before claiming a lower total bill. It is a three-way fork, trading optics power against host-board co-design burden and interoperability risk.
Fully-retimed pluggable is the incumbent: a full DSP retimes both sides, isolating the optic from the host channel. It is the most robust, the most plug-and-play, the most multi-vendor-interoperable — and the most power-hungry, at roughly 15–17 W for an 800G module (SemiAnalysis, January 2026). LRO — linear receive optics, also sold as RTLR or half-retimed; the names describe the same architecture, so treat them as aliases rather than as separate purchases — retimes the transmit side and runs the receive path linear, recovering the receive-side retiming power while leaving its channel equalization and qualification burden with the host. LPO (linear pluggable optics) removes the DSP entirely, relying on the host ASIC's own SerDes equalization to drive the optic directly; it cuts 800G module power to about 8.5 W on a shipping DR8 part (FS, September 2025) — close to halving it — but demands host-board co-design and end-to-end qualification, because the optic no longer cleans up the host channel. Price and power belong to a named module at a named rate and reach, not to the category label. The further you move toward LPO, the more module retiming you remove and the more electrical margin the host must own. That can save power and cost, but compare equal delivered bandwidth, reach and temperature including host power; an alternate vendor’s module still has to pass the complete channel before it becomes a substitution option.
| Architecture | DSP placement | ~Power / 800G (2025–26 reference) | Host co-design burden | Interop / serviceability | Best fit |
|---|---|---|---|---|---|
| Fully-retimed pluggable | Retiming in both directions | ~15–17 W (800G DSP) | Low — optic isolates host | Highest; multi-vendor, field-swappable | Default scale-out; multi-vendor fabrics |
| LRO / RTLR / half-retimed | TX retimed; RX linear in this arrangement | Between LPO and fully retimed; module-specific | Moderate — linear receive channel qualified with the host | High; faceplate FRU | Pragmatic power recovery without full LPO risk |
| LPO | No module retiming DSP | ≤8.5 W (FS 800G DR8, 2025) | High — own the whole electrical path | Narrower vendor set; faceplate FRU | Hyperscale fabrics that control host design |
| Near-package / co-packaged engine | Placement, not a DSP-retiming definition | ~5.4 W engine + laser (Broadcom Bailly class) | Highest — switch/optics co-design | Lowest; laser, engine and board FRUs are product-specific | Highest-radix switches at the power wall (→ 8.10) |
Scope & caveats
Illustrative assumed rates, not measured fleet telemetry. Count modules, not duplex links; correlate common-cause faults separately. Event rate becomes job-interruption rate only through topology, recovery and affected-job exposure.
Scope & caveats
A configuration-specific modeled estimate (GB300 NVL72, 3-layer network, first-party LinkX optics), not a universal ratio; merchant-optics builds sit lower on cost share.
Scope & caveats
2025–26 reference points at 112G/lane-class electrical lanes; 224G/lane channels are shorter — verify the qualified host/cable channel at your lane rate.
Scope & caveats
Module-class reference points; price and power belong to the named module at its rate and reach.
Scope & caveats
One vendor’s DR8 module; host equalization power sits outside the module figure.
Scope & caveats
Analyst figures for a named switch; CPO serviceability is decided in Chapter 8.10.
Scope & caveats
Counterfactual vendor estimate for the named NVL72 design as reported; not measured per-module power and not a saving transferable to another fabric.
Optical flaps and failures: the operating cost of interrupted work
At cluster scale, optics are a population, and a component event rate is only the first line of the reliability model. Define hard failure as replacement-requiring loss, a flap as transient service loss, and a repair disturbance as an interruption caused while working on another component. Measure mean-time-to-flap separately from hard-failure MTBF: a link that drops and re-establishes can stall a collective without ever requiring replacement. Collect count, exposure in module-hours, temperature, firmware and lot identity separately. For an independent constant-rate model, expected events over time t are N × t / M for population N and per-component mean interval M. The illustrative cadence below uses declared assumptions; it is not observed fleet telemetry or a transferable job-failure rate.
One unstable link that forces an illustrative 32,000-GPU job to roll back a full hour wastes 32,000 GPU-hours of progress; the loss is conditional on that restart and checkpoint age. Trace each event into the installed topology before pricing it. A failed optical end can leave a job running over a surviving plane at reduced bandwidth; two nominally redundant links sharing a module, laser, tray or power source can disappear together. Join the affected port and route identities to the admitted jobs, then distinguish rerouted service, timeout/retry, communicator failure and restart. Correlated temperature/lot faults and maintenance disturbances need separate event groups, not independent-rate multiplication. The MRC production account illustrates why path and transport recovery matter; its rates do not describe an unmeasured fleet. Hand restart probability and affected jobs to Chapter 9.4 for checkpoint consequences, and useful-output loss to Chapter 14.1.
Deep dive: commissioning and operating the optical plant as a reliability program
Treating optics as a managed population changes what "done" means at every lifecycle stage. At factory and incoming inspection, obtain transmitter conformance at its specified hot and cold limits, not only on a cool bench; sample-test incoming lots with the prescribed reference method rather than treating the supplier’s coupon as the installed channel. At install, connector cleanliness is the single highest-leverage discipline — a contaminated ferrule can erase more link margin than the entire fiber run, and contamination is cumulative across every mate/de-mate, so inspect-and-clean before every connection is a hard rule, not a courtesy. At commissioning, run the fabric under synthetic all-reduce load at thermal soak and read each lane’s pre-FEC/error statistics against its named acceptance limit and record passive loss reserve, uncorrectable codewords and link resets. A link that barely passes one counter can still become unstable under field heat or contamination; accept the complete qualified record, not a generic BER cutoff. In operation, the win is telemetry-driven: trend per-link pre-FEC BER and flag the optic whose error rate is climbing toward the cliff before it stalls a job, so the swap can be scheduled into a maintenance window on a verified surviving route instead of being triggered by a flapping link that restarts the training job. Keep a hot-spare optic population sized to the statistical replacement rate, and design the cabling so a single module is swappable without disturbing its neighbors. These choices decide whether a bad optical end is a scheduled replacement or repeated lost training work. Identify the population and preserve failed-part evidence; counts without exposure cannot size spares. Structured cabling and polarity belong to Chapter 8.10; the optical-budget calculation stays here.
Putting the taxonomy to work: drawing your boundaries
The physical layer is three boundaries, each derived from the layer above it and each costly if drawn wrong. First, the copper-to-optics boundary, set by your SerDes generation's reach physics: keep scale-up and the shortest scale-out hops on passive copper (no module power, no active electronics to fail); cross to active copper only for the reach copper still closes; reserve optics for spans copper physically cannot carry. Draw it toward optics and you burn on the order of 20 kW/rack of needless DSP power on an NVL72-class design; draw it toward copper and you eat link flaps. Second, the DSP-placement boundary: how much optics power you cut versus how much host co-design and interop risk you absorb — fully-retimed is safe and power-hungry, LPO halves the module power but makes you owner of the whole electrical path. Third, the form-factor boundary: set by the host SKU's cage geometry and qualified thermal interface — flat-top where the host supplies the riding heat sink, finned where the module carries its own — and unwound only by a faceplate redesign.
Around all three sits the reliability discipline, because the physical layer is where goodput is silently won or lost. A fabric that looks identical on a topology diagram can lose very different amounts of training work when its optics have different field thermal margin, shared FRUs and repair access; commission and operate that physical population, not just its logical links. The taxonomy tells you which medium to use; commissioning discipline tells you whether it will carry your bits for the life of the cluster.
Cite this chapter
Fehn, J. (2026). Physical-Layer & Interconnect Taxonomy (Chapter 8.9). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-9-physical-layer-and-interconnect-taxonomy (accessed 2026-09-29).
@misc{aidc-8-9,
author = {Fehn, Jacob},
title = {Physical-Layer & Interconnect Taxonomy (Chapter 8.9)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/part-8-networking-fabrics-and-optics/8-9-physical-layer-and-interconnect-taxonomy},
note = {Accessed 2026-09-29}
}