The Ethernet vs InfiniBand decision framework
S4·E1Two quotes, one server · The site office of a converted mail-sorting hall, Tuesday morning
Builds on: Which reference architecture governs this deal, Who do I call: three support paths, one asset tag
Before you read: what do you already know?
3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.
After this lesson you can
- Judge which fabric a deal should use from ten cited criteria rather than from preference.
- Quote the published scale and feature claims for Spectrum-X and for Quantum InfiniBand without inflating either.
- Identify the criteria where no fetched page supports a claim and say so out loud in front of a customer.
- Defend a recommendation in under 200 words with a source for every load-bearing sentence.
Episode 1 — Two quotes, one server
Painted lanes from the hall’s mail-sorting days run under the racks; the Dell SE has parked his laptop on one and his cold coffee on another. Two quotes for the same 256 GPUs lie open side by side, differing by a row of racks. One was priced from the Spectrum-4 Ethernet reference architecture, where “Traffic per rail of the DGX B300 systems is always one hop away from the other 64 nodes in a SU” and rack power “exceeds 50 kW”.[1][2] The other was priced from the Quantum-X800 InfiniBand edition of the same build, where “Each group of 72 nodes is rail-aligned” and the rack is “~56kW”.[3][4] The network lead has both open on one screen and a notebook ruled into two columns. Her question: which of you miscounted.
Nobody did. That is what the documents are for. Reference architectures exist because a GPU cluster is a building-sized system, not a server purchase: somebody has to publish the ratios — nodes per scalable unit, rails per node, kilowatts per rack — so that a quote, a rack elevation and a cable order check against the same arithmetic, not against each other. NVIDIA publishes more than one: the HGX AI Factory Enterprise RA covers “32 nodes (256 GPUs) to 128 nodes (1,024 GPUs)” and is a different document again.[5]
The scalable unit is a property of the document, not of the server.
The board that funds this hall meets quarterly, and the quote is due the Friday after next. Before anyone argues fabrics, find out which document already answered it.
1The question behind the question
When a Dell account asks “Ethernet or InfiniBand?”, almost nobody is asking for a protocol comparison. They are asking whether the thing they are about to sign is buildable and supportable. So the first criterion is not performance, it is compliance: which published reference architecture is this customer bound to, and what does that document already decide for them?[5][6]
The clearest illustration is a single server sized two ways. NVIDIA publishes a DGX B300 SuperPOD in a Spectrum-4 Ethernet edition with DC busbar power, where “Traffic per rail of the DGX B300 systems is always one hop away from the other 64 nodes in a SU” and rack power “exceeds 50 kW”.[1][2] It also publishes a Quantum-X800 InfiniBand edition with AC power, where “Each group of 72 nodes is rail-aligned” and the rack is “~56kW”.[3][4] Same server, two documents, two scalable-unit sizes, two rack-power statements, two BOMs. If the customer’s NVIDIA account team quoted from one of them, the fabric question is already answered and your job is arithmetic, not persuasion.[2][4]
The Enterprise RAs sit under those. The HGX AI Factory RA documents deployments “ranging from 32 nodes (256 GPUs) to 128 nodes (1,024 GPUs)”, and the NVL72 AI Factory RA is “fully tested” to “8 SUs (Scalable Units)” and scoped to “multi-user, single tenant workloads”.[5][6] Those are the documents most Dell PowerEdge deals live inside, and both are Spectrum-X Ethernet designs.[6]
Not decidable yet — 3 decisive criteria still unanswered
Weighted points — Ethernet 0 · InfiniBand 0 · neutral 0 · unknown 13 (of 13). Criteria 1, 3 and 4 count double.
Does the reference architecture the customer must comply with already pick a fabric?
Your answer: Unknown
The DGX SuperPOD B300 RA in its Spectrum-4 Ethernet / DC-busbar edition sizes an SU at 64 nodes = 512 GPUs, four DGX B300 per rack, and states "The rack-level power consumption per rack exceeds 50 kW". The GB300 NVL72 Enterprise RA runs Spectrum-X at 800 Gb/s per GPU and is "fully tested" to 8 SUs.
The same DGX B300 server has an XDR InfiniBand / AC-PDU edition of the SuperPOD RA where an SU is 72 nodes = 576 GPUs, racks draw "~56kW" on three 2U PDUs, and "UFM 3.5 nodes are connected to four (4) FNM ports on the Q3400 switches".
Nothing on the fetched pages publishes a rule for mixing the two editions inside one SU, or a migration path from 64-node SUs to 72-node SUs.
Same server, two RAs, different SU arithmetic — and the SU is the unit of quoting. Ask which RA revision the customer's NVIDIA account team quoted from before you ask what they prefer.
- Answer these first: 1. Which RA binds them; 3. Reduction-heavy collectives; 4. Storage constraint.
- These ten criteria are an assembled framework, not a published NVIDIA scoring model. The evidence in each cell is published; the weighting is ours.
2Ten criteria, each with its evidence
The framework below is assembled from the published pages; every row cites where its evidence comes from, and the honest rows are the ones where the evidence is thin.
| # | Criterion | What the pages support |
|---|---|---|
| 1 | Which RA binds them | 64-node vs 72-node SU on the same DGX B300; Enterprise RAs are Spectrum-X[1][3][6] |
| 2 | Scale target | Spectrum-X multiplane “up to 128K GPUs in two tiers”; the NVIDIA Quantum InfiniBand family “Over 10,000 nodes in two-level fat tree” — a family-level benefit statement, not a per-generation ceiling[7][8] |
| 3 | Reduction-heavy collectives | SHARP runs reductions in the switch via sharp_am and libsharp; Quantum-X800 lists “SHARP v4”[10][9] |
| 4 | Storage constraint | PowerScale’s back end is Ethernet and Dell-qualified on Cumulus 5.9.1[11] |
| 5 | Operations skills on staff | Ethernet reuses the RFC 7938 eBGP Clos and ECMP; InfiniBand adds SM, PKeys, UFM[12][16] |
| 6 | Multi-tenancy model | NCP-AIN pairs “BGP-EVPN to isolate tenant workloads” with “partition keys (PKeys)”[13] |
| 7 | Adaptive routing at your speed | Spectrum-4 AR runs at 400G and 200G and “does not support adaptive routing on 800G links”[14] |
| 8 | Version-discipline burden | Spectrum-X is a validated stack of matched versions; IB’s analogue is UFM and SM alignment[15][16] |
| 9 | Optics and power at 400G | The same MMA4Z00-NS twin-port class serves air-cooled Quantum-2 and Spectrum-4[17] |
| 10 | Support path | Dell resells both families and offers support end to end; no page publishes a defect RACI[18][19] |
Three of those carry real decision weight. Criterion 1 usually settles it. Criterion 4 settles it whenever the customer has already bought Ethernet-only storage — Dell states PowerScale requires a private back-end network, isolated from the front end, and that “Dell does not support connecting any other devices to the back-end switches”.[11] Criterion 3 is the only one that can reverse an Ethernet-leaning answer, because in-switch reduction is a mechanism, not an adjective.[10]
The rest are close enough on published evidence that they should not carry a recommendation. Both scale ceilings are far above an enterprise buying under 1,024 GPUs.[7][8][5] The optics are the same class of part.[17] And NVIDIA itself publishes the strongest operational claim for the other side: InfiniBand “is actually simpler to deploy and maintain”.[20] Quote that sentence yourself before the customer finds it.
3What nobody publishes, and how to say it
An FAE’s credibility is spent fastest on claims that cannot be checked. Four gaps in this framework are worth naming explicitly.
No Spectrum-X in-switch reduction. The Spectrum-X platform page claims 1.6x performance over off-the-shelf Ethernet, the 128K-GPU multiplane figure, 1.9x cross-site NCCL with Spectrum-XGS, and photonics numbers — and it claims nothing about in-network reduction.[7] SHARP has an install guide, a manager process in UFM and a library that links into the application; it is an InfiniBand feature.[10] Say “correct, that is an InfiniBand capability” and move to whether the workload’s collectives are reduction-bound.
No defect-ownership RACI. Dell’s blog offers help from design consulting to post-deployment troubleshooting and Dell’s catalog sells both Spectrum Ethernet and Quantum-X800 switches, but no fetched page delineates who owns a switch defect in a Dell-sold AI Factory.[18][19] Teach it as a known unknown with two questions to ask instead: whose asset tag is on the switch, and which contract number covers it.
No head-to-head benchmark. NVIDIA’s InfiniBand operations blog is operational — UFM tooling, maintenance cadence, alert workflows — and contains no MLPerf reference and no comparison to Ethernet.[20] Do not attribute one to it.
Version pins that collide. The Spectrum-X validated solution stack publishes matched rows across switch NOS, NIC firmware, DOCA-Host, NetQ and NCCL; the PowerScale back-end document pins Cumulus Linux 5.9.1 on SN5600.[15][11] A single switch cannot satisfy both pins. That is a design constraint you raise in the review, not a paperwork detail you discover at cutover.
4Turning ten criteria into two hundred words
A recommendation that reads like a preference gets re-litigated by the next person who walks into the room. One that names its evidence and its gaps survives review. The pattern: state the binding constraint, state the two facts the customer already owns, state the one criterion that could reverse the call, and list what you could not verify.
Customer A — enterprise, 256 GPUs of HGX B300, PowerScale F710 already installed, ops team runs BGP everywhere.
Recommendation: Spectrum-X Ethernet.
Evidence walk: the deal sits inside the HGX AI Factory Enterprise RA, which documents 32 nodes (256 GPUs) to 128 nodes (1,024 GPUs) and is a Spectrum-X design — criterion 1 answered by compliance.[5] The storage is PowerScale, whose back end is Ethernet and must be a private isolated network with no other devices attached — criterion 4 answered by an asset they already own, and it also tells us the AI fabric and the storage back end will be separate switches because Dell qualified SN5600 there on Cumulus 5.9.1 while the Spectrum-X stack pins a different NOS row.[11][15] Criterion 5 favours Ethernet: their staff already run the RFC 7938 eBGP Clos and ECMP that the underlay is built from.[12] Criterion 3 is the reversal test: nothing in the ask says the training is reduction-bound, so the SHARP argument does not apply, and I say plainly that SHARP has no published Spectrum-X equivalent.[10]
Thin evidence, stated: no published defect RACI for a Dell-sold switch, so we will name the support contract explicitly in the design note.[18][19]
Customer B — research site, 2,000 GPUs, all-reduce-dominated training, no storage constraint stated.
Recommendation: InfiniBand is a legitimate candidate and I will not close it on the call.
Evidence walk: criterion 3 is live — SHARP performs reductions in the switch, with sharp_am in UFM and libsharp linked into the application, and Quantum-X800 lists SHARP v4.[10][9] Criterion 2 does not decide: the NVIDIA Quantum InfiniBand switching page publishes “Over 10,000 nodes in two-level fat tree” as a family-level highlight and Spectrum-X publishes 128K GPUs in two tiers, so both clear 2,000.[8][7] Criterion 5 counts against InfiniBand only if they have no fabric-management staff, since it introduces subnet manager, PKeys and UFM as new objects.[16][13]
Next step, not a recommendation: get one reduction-heavy job profiled. The decision turns on a measurement nobody in the room has.
Write the recommendation for a customer with 512 GPUs of HGX B300, an existing InfiniBand HPC cluster with UFM already staffed, and an Ethernet-only object store.
- Criterion 1 — the governing RA is ____ and it scopes ____ so the fabric it assumes is ____.[5]
- Criterion 4 — the storage forces ____ for at least the ____ network.
- Criterion 5 — this customer’s existing skills point toward ____ because they already run ____.[16]
- Criterion 3 — the question I must ask before recommending anything is ____.
- The criterion where I have no published evidence is ____ and the two questions I ask instead are ____ and ____.[18]
- My recommendation in one sentence: ____.
A Dell account brings you a customer with 384 GPUs of HGX B300, a mixed storage estate (PowerScale for home directories, a partner all-flash array on Ethernet for the training set), an ops team of three who run Ansible against Arista today, and a CTO who has read that “InfiniBand is simpler”.
Produce a written recommendation under 200 words plus a criterion table. Acceptance: every load-bearing sentence carries a source; the CTO’s sentence is quoted and answered rather than contradicted; at least two criteria are explicitly marked neutral or unknown rather than forced; the PowerScale NOS-pin consequence appears; and the note names one measurement that would change your answer.
Episode 1 — What went in the right-hand column
Twenty minutes, and nobody has opened a benchmark. The storage is PowerScale, which forces Ethernet for at least the back end and pins its own Dell-qualified NOS version, and no one has profiled a training job to see whether the collectives are reduction-bound.[11][10] Two of the three questions are answered by assets the customer already owns; the third is a measurement they will go and take. The recommendation goes in the network lead’s left column and the thin-evidence criteria in her right.
Then she closes the notebook and says the sentence she actually came to say: in 2019 they ran RoCE on somebody else’s switches and spent six months chasing PFC storms, and she is not doing that again.
Lab
Not applicable on the Dell-lab BlueField-3 and ConnectX hosts: this lesson decides between two fabrics and the lab has switches for neither. Nothing here is mutating; every step is a read.
Pre-flight inventory (5 minutes, useful anyway): on the Dell-lab host, record ibstat output, ibdev2netdev mapping and the negotiated link speed per port, so your written recommendations reference a machine you have actually touched.
Optional, in a customer lab or on borrowed InfiniBand time — book an hour on a live IB fabric and run the read-only set: ibstat, ibnodes, iblinkinfo and ibdiagnet. The objective is not the output, it is the operational model. Write down the three objects that have no Ethernet equivalent — the subnet manager, PKeys and UFM — and one sentence each on who on the customer’s staff would own them at 3 a.m.[16][13] Rollback: none required, all four commands are read-only diagnostics; do not run any ibportstate or SM configuration command on a fabric you do not own.
Fill the ten-criterion matrix for two invented but realistic Dell customers, citing a source for every cell. No hardware needed; the deliverable is two written recommendations.
- Build the blank matrix: ten rows, columns for “criterion”, “customer answer”, “evidence for Ethernet + URL”, “evidence for InfiniBand + URL”, “not claimed anywhere”. Expected: the last column is not empty. If it is, you are asserting where nobody published.
- Customer 1 — enterprise, 256 GPUs of HGX B300, PowerScale F710, one ops team that already runs eBGP. Fill all ten rows. Expected: criteria 1, 4 and 5 answered from documents the customer already owns.[5][11][12]
- Customer 2 — research site, 2,000 GPUs, reduction-heavy training. Fill all ten rows. Expected: criterion 3 is the only one that flips, and criterion 2 is neutral because both published ceilings clear 2,000.[7][8]
- Open the interactive in criteria view and set your answers for each customer. Compare its tally and caveat list to yours. Expected: disagreements are traceable to a cell where you weighted an unpublished claim. If the tool says “not claimed” for a cell you filled confidently, delete your cell.
- Write each recommendation in under 200 words, with a source per load-bearing sentence and an explicit list of thin-evidence criteria. Expected: at least two criteria per customer marked neutral or unknown.
- Read both aloud against a clock. Expected: under 90 seconds each. If you exceed it, you are arguing rather than citing.
Retrieval check
10 questions from memory. Answer before looking anything up; misses become flashcards.
Explain it to a Dell SE
Explain to a Dell account executive in four sentences how you decide Ethernet or InfiniBand for a customer without using the word better.
Sources
Facts in this lesson were checked against NVIDIA Spectrum-X platform page, Quantum-X800 page, NCP-AIN blueprint and Cumulus Linux 5.13 ECMP page re-fetched 2026-09-07. Dates are when each page was fetched.
- Network Fabrics — DGX SuperPOD B300 Spectrum-4 Ethernet and DC Busbar RA · fetched 2026-09-07
- DGX SuperPOD Architecture — B300 Spectrum-4 Ethernet and DC Busbar RA · fetched 2026-09-07
- Network Fabrics — DGX SuperPOD B300 Quantum-X800 InfiniBand and AC Power RA · fetched 2026-09-07
- DGX SuperPOD Architecture — B300 Quantum-X800 InfiniBand and AC Power RA · fetched 2026-09-07
- NVIDIA HGX AI Factory Enterprise Reference Architecture (index) · fetched 2026-09-07
- Overview — NVIDIA NVL72 AI Factory Enterprise RA · fetched 2026-09-07
- NVIDIA Spectrum-X Ethernet Platform · fetched 2026-09-07
- NVIDIA InfiniBand Switches · fetched 2026-09-07
- NVIDIA Quantum-X800 InfiniBand Platform · fetched 2026-09-07
- NVIDIA SHARP Installation · fetched 2026-09-07
- Dell PowerScale: Ethernet Back-End Network Overview H16346.8 · fetched 2026-09-07
- RFC 7938 — Use of BGP for Routing in Large-Scale Data Centers · fetched 2026-09-07
- NVIDIA-Certified Professional: AI Networking (NCP-AIN) blueprint · fetched 2026-09-07
- Equal Cost Multipath Load Sharing incl. Adaptive Routing | Cumulus Linux 5.13 · fetched 2026-09-07
- NVIDIA Spectrum-X Validated Solution Stack · fetched 2026-09-07
- NVIDIA UFM Unified Fabric Manager · fetched 2026-09-07
- MMA4Z00-NS 800Gb/s Twin-port OSFP transceiver overview · fetched 2026-09-07
- Dell-NVIDIA Partnership Powers High-Performance AI Fabric Solutions · fetched 2026-09-07
- AI Networking Switches | Dell USA · fetched 2026-09-07
- Simplifying Network Operations for AI with NVIDIA Quantum InfiniBand · fetched 2026-09-07
- Networking Physical Topologies — NVIDIA HGX AI Factory · fetched 2026-09-07
The same idea elsewhere
Other lessons that cover this ground, sometimes from another course's angle.
- What Spectrum-X is, and what it is notSpectrum-X course · Same ground: numbers, adaptive-routing and claims
- Scenario: sizing and positioningDOCA course · Same ground: positioning, storage and quote
- Who do I call: three support paths, one asset tagElsewhere in this course · Same ground: powerscale, versions and front