Optics, cables and the power nobody budgeted
S2·E5Where are the optics? · A joint facilities and networking review, three weeks to the filing
Builds on: The other three networks: converged, storage, out-of-band
Before you read: what do you already know?
3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.
After this lesson you can
- Select a LinkX part class for a given link by reach plane count and fibre plant using five stated rules.
- Judge an optics line on a quote and name the specific line that is wrong and why.
- Compute the optics power of a fully populated switch and add it to the switch's own draw.
- Defend a transceiver choice against the switch's published thermal depth and airflow envelope.
Episode 5 — Where are the optics?
The BOM finally reaches facilities, and their power engineer comes back with one question. The network racks are drawn beside compute racks in a contained hot aisle, the switch lines list a chassis draw, and nothing on the sheet accounts for the transceivers. Where are the optics?
They are not on it, and they are not small. A fully populated SN5600 carries 64 cages, and the twin-port MMA4Z00-NS is listed at 15 Watts for all configurations, so the optics alone add about 960 W on top of the switch — published wattage, my multiplication.[2] On the InfiniBand side the cable class moves it by more: a QM9700 is listed at 747 W typical with passive cables and 1720 W maximum with active ones.[5]
Then there is where the box can sit: chassis depth, airflow direction and the temperature ceiling decide placement — SN5600 at 0–35 °C against SN5610 at 0–40 °C — not switching capacity,[4] and the compute racks these leaves live beside “exceed 50 kW” on the Ethernet DGX B300 RA.[11] The pluggable exists for a reason the GB300 RA states plainly: NVIDIA Ethernet interconnects “make extensive use of pluggable optical transceivers and detachable optical fibers for easier installation, inspection, and debugging”.[7] Field-serviceable on purpose — and every module clicked into a cage is a watt somebody has to find.
Optics are a power line item, not an accessory; a BOM that lists switches but no transceivers has never been lit.
Segment 1 starts with the part families, because reach eliminates most of them before power is a question.
1The catalogue is a decision tree, not a lookup table
LinkX spans “10, 25, 40, 50, 100, 200, and 400G” solutions in “SFP, QSFP, and OSFP form factors”, for reaches “from ½-meter-to 40-kilometers”, across DAC, AOC and optical-transceiver categories, for both Ethernet and InfiniBand.[3] That range is too large to memorise and too consequential to guess at, so learn the prefixes and the reach bands instead of the part numbers.
The prefixes carry the category: MCP is passive copper DAC, MCA is active copper, MFS is fibre AOC, MMA and MMS are transceivers, and MFP is optical fibre accessories.[1] The families sort by speed: 1600G OSFP twin-port (MMS4B10, MMS4C1X, MMS4A00); 800G twin-port OSFP 2x400 (MMS4X00, MMA4Z00); 400G OSFP and QSFP (MMS4X00, MMA4Z00, MMS1V00); 200G QSFP56 (MMS1W50, MFS1S00); 100G QSFP28 (MMS1V70, MMA1B00); 25G SFP28 (MMA2P00).[1]
Reach then does the elimination. DAC covers 1–3 m at the short end; long-reach transceivers extend to 10 km; many single-mode variants sit at “up to 500m” or 2 km.[1] Because reach eliminates whole families, it is the first question you ask, before price and before brand.
| Product | Speed | PCIe | Role | GPU generation |
|---|---|---|---|---|
NIC | ||||
SuperNIC (no Arm) | ||||
SuperNIC (no Arm) | ||||
DPU | ||||
DPU | ||||
SuperNIC (Arm inactive) | ||||
DPU / storage processor | ||||
Ethernet switch | ||||
Ethernet switch | ||||
InfiniBand switch | ||||
InfiniBand switch |
⚠ = not confirmed on a fetched primary source (hover for why). Facts as of DOCA 3.5.0 (Sep 2026). Selections are saved.
2Five rules that pick the part
The reference architectures do not publish a cable catalogue, so these rules are assembled from the interconnect docs and the RA hardware pages. Each one names its evidence.
- Same rack, switch-to-NIC 3 m or less, OSFP both ends → passive DAC in the MCP family. Cheapest, zero optics power, at the cost of bulk and cable mass.[1]
- Within a row, 3–50 m, multimode plant → MMA4Z00-NS twin-port with MFP7E10 straight fibre, or MFP7E20 where you need a 1:2 splitter into two 400G endpoints.[2]
- Cross-row, beyond 50 m, or a single-mode plant → a single-mode class part such as the MMS4X00 family. More money and more power per port.[1] (The exact single-mode reach figure in the research notes came from a search-result title rather than a fetched page — treat it as approximate and confirm the specific part before quoting.)
- Dual plane → twin-port transceiver at the NIC; single plane → single-port OSFP. The HGX RA states the choice in exactly those terms, which makes it a BOM fork rather than a field option.[6]
- Liquid-cooled or co-packaged-optics switches change the story entirely. SN6800-LD is listed with co-packaged optics, and Q3450 is available in a CPO variant — the pluggable disappears and with it the per-port optics line.[8][9]
The GB300 RA’s only cabling statement is a philosophy rather than a part list: “NVIDIA Ethernet interconnects make extensive use of pluggable optical transceivers and detachable optical fibers for easier installation, inspection, and debugging.”[7] The SuperPOD RA specifies the customer edge by type instead of part number — “DR1 single-mode” for uplinks and user storage.[10] Both are useful precisely because they tell you what the RA will not decide for you.
3The workhorse part, read line by line
Learn one part properly and the rest of the catalogue becomes readable. MMA4Z00-NS is “800Gb/s Twin-port OSFP, 2x400Gb/s Multimode 2xSR4, 50m”: the line rate is “400Gb/s for both 400GbE Ethernet and NDR InfiniBand based on the 100G-PAM4 modulation” per port, the twin port carries “Two, 4-channel MPO-12/APC optical connectors with two 4-channel fiber cables”, and the part is “only used in Quantum-2 and Spectrum-4 OSFP air-cooled switches”.[2]
Three consequences fall straight out of that description. It serves Ethernet and InfiniBand from the same part number, so the optics line is not a differentiator between the two fabrics at 400G.[2] It is multimode with a 50 m ceiling, so it fails silently as a design choice the moment a spine moves rows.[2] And it is air-cooled-switch only, so a liquid-cooled or CPO switch invalidates it entirely.[2][8]
The companion parts complete the link: straight multimode fibre MFP7E10-Nxxx, the 1:2 splitter MFP7E20-N0xx, the single-port OSFP equivalent MMA4Z00-NS400, and the QSFP112 equivalent MMA1Z00-NS400.[2] NVIDIA states it supplies multimode crossover and straight fibre “up to 100-meters straight and 50-meters for splitters” — a statement now carried on the product page itself as re-fetched on 2026-09-09, where the research notes had previously flagged it as a search snippet.[2]
Power is 15 W. The page says “Twin-port multimode OSFP transceivers remain at 15 Watts for all configurations linking OSFP switches, OSFP and QSFP112 adapters, and DPUs simultaneously”, confirmed again on 2026-09-09.[2] An older PDF search snippet of the same product quoted 8 W. Quote the fetched page, and say out loud that a conflicting figure exists, because this number is about to be multiplied by 64.
4Power and thermals nobody budgeted
Switch power is not a rounding error in an AI hall, and neither is switch geometry. The SN5600 series is 2U at 3.39 in x 17.2 in x 31 in — 86.2 x 438 x 788 mm — which is deeper than many legacy network racks, and all three models are listed with reverse airflow.[4] Airflow direction is a thermal fault when mixed in a shared rack, not a preference.[4]
The operating temperature splits the family: SN5610 tolerates “0–40ºC” while SN5600 and SN5600D are rated “0–35ºC”.[4] In a contained hot aisle that five degrees is the difference between a supported install and an unsupported one. Power input splits it again — SN5610 and SN5600 take 200–240 VAC while SN5600D is a DC bus bar part with a 40–60 VDC input range, built for the DC-busbar SuperPOD design.[4][11] Redundancy differs too: SN5610 has four PSUs in 2+2 and five fans at N+1, while SN5600 and SN5600D have two PSUs in 1+1 and four fans.[4]
On the InfiniBand side the swing factor is the cabling. QM9700 draws 747 W typical with passive cables and up to 1,720 W maximum with active cables; QM9790 is 640 W to 1,610 W and QM9701 is 720 W to 1,660 W.[5] Nearly a kilowatt of difference decided by a cable choice. Airflow is directional and temperature bounded on the same family: forward air flow 0–35 °C, reverse 0–40 °C.[5]
Then add the optics. At 15 W per twin-port module, a fully populated 64-cage SN5600 carries 64 x 15 W = 960 W of optics on top of the switch itself.[2][4] (Per-module figure published; the multiplication is mine — no NVIDIA page prints a fully-lit optics total.) Set that against the compute racks it has to live near: DGX B300 racks “exceed 50 kW” on the Ethernet RA and are about 56 kW on the XDR RA, while a GB300 NVL72 rack draws up to 142 kW from eight 33 kW power shelves and is liquid cooled.[11][12]
⚠ The fetched MMA4Z00-NS page states 15 W; a search snippet of the same product PDF says 8 W. Quote 15 W — the higher figure is the fetched one and the one that fails safe in a PDU budget — and flag the conflict. A fully lit 64-cage SN5600 is 64 x 15 W = 960 W of optics on top of the switch's own draw. All of these counts are derived arithmetic.
Single-port MMA4Z00-NS400 wattage is not published on the fetched page. ⚠
- Same rack, ≤3 m, OSFP both ends → passive DAC (MCP family). Cheapest, zero optics power, adds cable bulk.
- Within a row, 3–50 m, multimode plant → MMA4Z00-NS twin-port + MFP7E10 straight fiber; split to two 400G endpoints with MFP7E20 or an MCP7Y00 DAC splitter.
- Cross-row, >50 m or a single-mode plant → single-mode MMS4X00 / DR-class. More money and more power per port.
- Dual-plane B300/GB300 → twin-port transceiver at the NIC, split to two leaves. Single-plane → single-port OSFP. This is a BOM fork, not a field option.
- CPO switches (SN6800-LD, Q3450) remove the pluggable entirely — the optics line disappears from the BOM.
MCP passive copper · MCA active copper · MFS AOC · MMA/MMS transceivers · MFP fiber accessories.
SN5600: 2U, 788 mm deep, reverse airflow. Power 200–240 VAC (SN5600D is 40–60 VDC busbar). Operating 0–35 ºC.
FAE angle
Ask three questions before quoting optics: multimode or single-mode plant already installed? longest leaf-to-spine run in meters, measured not estimated? dual-plane or single-plane? Those three answers pick the part number.
MMA4Z00-NS overview · LinkX interconnect5Judging someone else's optics line
Evaluation is the skill here, so make it a fixed sequence. Given an optics line on a quote, five checks decide whether it is defensible.
One: quantity against cages, not logical ports. Transceivers go in cages. An SN5600 has 64 cages presenting 128 logical 400 GbE ports and a QM9700 has 32 cages carrying 64 NDR400 ports.[4][5] A transceiver count that matches the logical port count is double what fits.
Two: part class against measured reach. 50 m is the multimode ceiling on the workhorse part.[2] Ask for the longest run in metres and ask whether it was measured or estimated.
Three: transceiver type against plane count. Twin-port for dual plane, single-port OSFP for single plane.[6] A drawing showing two fabrics with single-port parts on the BOM is internally inconsistent, and the drawing is usually the correct half.
Four: optics power against the PDU budget. Add the fully-lit figure, name it as your arithmetic, and state which wattage source you used.[2]
Five: switch placement against depth airflow and temperature. 788 mm deep, reverse airflow, and a 0–35 °C ceiling on SN5600 versus 0–40 °C on SN5610.[4] If the rack is 800 mm and contained, say so before the order ships.
The cluster from the previous two lessons: 32 HGX B300 nodes, 256 GPUs, dual plane, 8 compute leaves and 4 spines on SN5600-class switches, 2 converged switches, 2 SN2201 for out-of-band. A vendor sends a quote with 512 single-port OSFP transceivers, 256 multimode fibres and no optics power line. Judge it.
- Server-facing links. Dual plane gives 256 GPUs x 2 = 512 links of 400 Gb/s, 256 per plane.[6] On the NIC side dual plane means twin-port transceivers, so 256 twin-port modules at the NIC and 512 400G endpoints at the leaves.[6] The quote’s 512 single-port parts are the wrong class — that is the single-plane part number.
- Reach class for those links. In-row, node-to-leaf. If the run is 3 m or less and both ends are OSFP, passive DAC in the MCP family is legitimate and removes 15 W per module from the power budget.[1][2] If it is 3–50 m in a multimode plant, MMA4Z00-NS twin-port with MFP7E10 straight fibre, or MFP7E20 where a 1:2 splitter into two 400G endpoints is wanted.[2]
- Leaf-to-spine links. 8 leaves x 64 uplinks = 512 links of 400 Gb/s across both planes at a non-blocking split. (The 64-uplink split is my assumption from the previous lesson, not an NVIDIA figure.)[4] Measure the run: same row keeps you on MMA4Z00-NS; cross-row past 50 m moves the whole class to single mode.[2]
- Out-of-band links. 1 Gb RJ45 copper to the SN2201s — no optics at all, plus four 100 GbE uplinks per SN2201 if they are aggregated.[13]
- Customer edge. At least 2x 100GbE DR1 single-mode.[10] DR1 is single mode, so this is the one link class that cannot be served by the multimode plant.
- Optics power. Take the leaf switches: 8 leaves fully lit at 64 cages x 15 W = 960 W each, so about 7.7 kW of optics across the compute leaves alone before the spines.[2][4] (Per-module figure published; both multiplications mine.) Add the switch draw and hand the total to facilities.
- Placement check. SN5600 at 788 mm deep, reverse airflow, 0–35 °C; SN5610 buys the extra five degrees.[4] If the network rack is contained and shares an aisle with 50 kW compute racks, recommend SN5610 and say why.[11]
- The verdict, in the form you would send it. “Three problems. The transceiver class is single-port and the design is dual plane, so the part number is wrong on 512 lines. There is no optics power line and the compute leaves alone add roughly 7.7 kW. And I need the longest leaf-to-spine run in metres before I can confirm the multimode parts at all.”
The customer decides to go single plane on the same 32 nodes to cut cost. Rebuild and judge.
- Server-facing links = 256 GPUs x ____ plane = ____ links of 400 Gb/s.[6]
- Transceiver class at the NIC = ____________ because the design is now ____________ .[6]
- Leaves per plane changed from ____ to ____ , so leaf-to-spine link count is now ____ .
- Optics power for the compute leaves = ____ leaves x ____ cages x ____ W = ____ W.[2]
- What did the customer give up in GPU bandwidth according to the RA, in one number?[6]
- Which single line of the optics BOM would have to be re-ordered if they later moved back to dual plane?
A Dell account sends you the optics section of a 64-node HGX B300 quote. It lists: 1,024 twin-port MMA4Z00-NS at the NICs; 1,024 MMA4Z00-NS at the leaves; MFP7E10 straight fibre for every link; 2x 100GbE multimode transceivers for the customer edge; no PDU line for optics. The site drawing shows spines in a separate row about 70 m away and network racks in a contained hot aisle beside 50 kW compute racks.
Produce a written judgement that:
- Names every line that is wrong and the published sentence that makes it wrong.
- Corrects the leaf-to-spine part class and explains what changes about cost and power.
- Corrects the customer-edge line.
- Adds the missing optics power line with the arithmetic shown and labelled as derived.
- Makes a switch-SKU recommendation for the contained hot aisle and defends it with the published temperature figures.
- Ends with the one question you still need answered before you would sign off.
Acceptance criteria: every correction cites a page rather than an opinion; every multiplication is marked as yours; the 15 W versus 8 W conflict is named where the power figure is used.
Episode 5 — Case closed
Two lines get added and one part number changes. Optics power goes in beside the chassis draw at roughly 960 W for a fully lit leaf,[2] and the leaf model moves to the SN5610, because its 0–40 °C ceiling is what a contained aisle beside racks exceeding 50 kW actually asks for.[4][11] What you say to the customer: “Nothing here changes your design. It changes two numbers that were missing from it, and both are far cheaper to correct now than during install week.” The hail model trains on the full fabric and the insurer files its rates on the date it picked. The SE closes the promise spreadsheet with every row green, and the operator’s labels still match the as-built map, port for port.
Lab
Read the real optics in the Dell lab and check them against the rules. Read-only throughout — ethtool -m reads the module EEPROM and changes nothing.
- Pre-flight inventory. List interfaces and identify which are optical:
Record which physical port each netdev maps to before reading any module, so a part number can be tied to a specific run.ip -br link ibdev2netdev - Read every populated module.
Expected: for each optical port,for i in $(ls /sys/class/net | grep -v lo); do echo "== $i"; sudo ethtool -m $i 2>/dev/null | head -25; doneVendor name,Vendor PN,Identifier,Transceiver type, plus length fields and current temperature and optical power readings. “If not”: a copper DAC or an unpopulated cage returns nothing or a minimal EEPROM — record that as a DAC link rather than treating it as a failure. - Map each Vendor PN to a LinkX prefix rule. MCP passive copper, MCA active copper, MFS AOC, MMA/MMS transceiver.[1] Write the category beside each port in your inventory.
- Check reach against the real run. Take the length fields from the EEPROM and compare them against the physical cable run you can see. A 50 m-class multimode part on a run that leaves the rack row is the failure this lesson exists to prevent.[2]
- Read temperature and optical power.
Expected: module temperature well inside the switch’s own envelope, and receive power within the vendor’s stated range. Record the numbers; a module running near the top of its temperature range in a lab is a preview of what it will do in a contained hot aisle.[4]sudo ethtool -m <iface> | grep -Ei 'temperature|power|bias' - Count what a fully lit switch would draw. Multiply the cage count of whatever switch is present by 15 W and write the total next to the switch’s own published draw.[2][5] Mark the multiplication as yours.
- Write the two-line field summary. One line for the part classes actually installed, one line for the reach mismatch you did or did not find. Those two lines are what you would say on a customer call.
- Optional, customer lab only. Put a twin-port and a single-port OSFP transceiver side by side and confirm by eye which one a dual-plane design needs.[6] Read-only; do not unseat a module from a live port.
Build the optics BOM for the 32-node dual-plane cluster, then price its power.
- List the link classes. Create a table with one row per class: node-to-leaf, leaf-to-spine, converged node-to-switch, OOB node-to-SN2201, customer edge. Columns: count, reach band, plane implication, part class, fibre accessory.
Expected:cd ~/projects/ra-sizing cat > optics.py <<'PY' LINKS = [ ("node-to-leaf", 512, "0-3 m or 3-50 m", "twin-port dual plane", "MCP DAC or MMA4Z00-NS + MFP7E10"), ("leaf-to-spine", 512, "measure it", "per plane", "MMA4Z00-NS if under 50 m else single-mode"), ("converged", 78, "in row", "n/a", "400G class"), ("oob", 95, "copper", "n/a", "RJ45 Cat6"), ("customer-edge", 2, "site dependent", "n/a", "DR1 single-mode 100GbE"), ] def optics_power(cages, watts=15.0): return cages*watts PY python3 -c "import optics; print(optics.optics_power(64)); print(optics.optics_power(64*8))"960.0W for one fully lit 64-cage switch and7680.0W across eight compute leaves.[2][4] - Apply the five rules to each row and write the part class you would quote, with the rule number beside it.[1][2][6]
- Mark the plane fork explicitly. Write one line stating that the node-to-leaf row changes part number - not quantity - if the design moves to single plane.[6]
- Add the optics power to the switch power and produce a single kW figure for the network racks. Label the multiplication as yours; no NVIDIA page prints a fully-lit optics total.[2]
- Record the wattage conflict. Write two sentences: what the fetched product page says (15 W, confirmed 2026-09-09), what the older PDF snippet said (8 W), which one you would quote and why.[2]
- Do the thermal sanity check on paper. Compare the SN5600 ceiling of 0–35 °C and the SN5610 ceiling of 0–40 °C against a contained hot aisle beside racks exceeding 50 kW, and write the SKU recommendation you would defend.[4][11]
- Check depth. 788 mm plus cable bend radius against the rack the customer actually has.[4] Write down what you would ask them to measure.
Retrieval check
10 questions from memory. Answer before looking anything up; misses become flashcards.
Explain it to a Dell SE
A Dell SE sends you an optics line on a quote and asks whether it is right. Explain in five sentences the three questions you ask before you can answer and what each answer eliminates.
Sources
Facts in this lesson were checked against MMA4Z00-NS overview page re-fetched 2026-09-09 - power confirmed at 15 Watts for all configurations and the 100 m straight / 50 m splitter fibre statement confirmed on the page itself; HGX AI Factory Networking Hardware and NVL72 AI Factory Networking Hardware as recorded 2026-09-07; NVIDIA Spectrum SN5600 series Dell-branded datasheet and QM97XX Specifications as recorded 2026-09-07; DGX SuperPOD B300 architecture and NVL72 AI Factory components as recorded in content/research/ra/part1.md 2026-09-07. Dates are when each page was fetched.
- Networking Interconnect — NVIDIA Networking Docs (LinkX part-number families) · fetched 2026-09-07
- Overview — MMA4Z00-NS 800Gb/s Twin-port OSFP 2x400Gb/s Multimode 2xSR4 50m · fetched 2026-09-09
- NVIDIA LinkX Interconnect — cables and transceivers · fetched 2026-09-07
- NVIDIA Spectrum SN5600 Series Switches Datasheet (Dell-branded) · fetched 2026-09-07
- Specifications — QM97XX 1U NDR 400Gbps InfiniBand Switch Systems User Manual · fetched 2026-09-07
- Networking Hardware — NVIDIA HGX AI Factory (B300) Enterprise RA · fetched 2026-09-07
- Networking Hardware — NVIDIA NVL72 AI Factory (GB300) Enterprise RA · fetched 2026-09-07
- NVIDIA Spectrum Ethernet Switches — product family port matrix · fetched 2026-09-07
- NVIDIA InfiniBand Switches — Quantum-X800 family and CPO variants · fetched 2026-09-07
- Network Fabrics — DGX B300 SuperPOD Spectrum-4 Ethernet and DC Busbar Power RA · fetched 2026-09-07
- DGX SuperPOD Architecture — DGX B300 Spectrum-4 Ethernet and DC Busbar Power RA · fetched 2026-09-07
- System Hardware and Components — NVIDIA NVL72 AI Factory (GB300) Enterprise RA · fetched 2026-09-07
- AI Networking Switches — Dell USA catalog · fetched 2026-09-07
The same idea elsewhere
Other lessons that cover this ground, sometimes from another course's angle.