Skip to content

The Ethernet vs InfiniBand decision framework

S4·E1Two quotes, one server · The site office of a converted mail-sorting hall, Tuesday morning

S4·E1Evaluate~30 minsources checked todayverified against NVIDIA Spectrum-X platform page, Quantum-X800 page, NCP-AIN blueprint and Cumulus Linux 5.13 ECMP page re-fetched 2026-09-07

Builds on: Which reference architecture governs this deal, Who do I call: three support paths, one asset tag

Before you read: what do you already know?

3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.

After this lesson you can

  • Judge which fabric a deal should use from ten cited criteria rather than from preference.
  • Quote the published scale and feature claims for Spectrum-X and for Quantum InfiniBand without inflating either.
  • Identify the criteria where no fetched page supports a claim and say so out loud in front of a customer.
  • Defend a recommendation in under 200 words with a source for every load-bearing sentence.

Episode 1 — Two quotes, one server

The situation · The site office of a converted mail-sorting hall, Tuesday morning

Painted lanes from the hall’s mail-sorting days run under the racks; the Dell SE has parked his laptop on one and his cold coffee on another. Two quotes for the same 256 GPUs lie open side by side, differing by a row of racks. One was priced from the Spectrum-4 Ethernet reference architecture, where “Traffic per rail of the DGX B300 systems is always one hop away from the other 64 nodes in a SU” and rack power “exceeds 50 kW”.[1][2] The other was priced from the Quantum-X800 InfiniBand edition of the same build, where “Each group of 72 nodes is rail-aligned” and the rack is “~56kW”.[3][4] The network lead has both open on one screen and a notebook ruled into two columns. Her question: which of you miscounted.

Nobody did. That is what the documents are for. Reference architectures exist because a GPU cluster is a building-sized system, not a server purchase: somebody has to publish the ratios — nodes per scalable unit, rails per node, kilowatts per rack — so that a quote, a rack elevation and a cable order check against the same arithmetic, not against each other. NVIDIA publishes more than one: the HGX AI Factory Enterprise RA covers “32 nodes (256 GPUs) to 128 nodes (1,024 GPUs)” and is a different document again.[5]

The scalable unit is a property of the document, not of the server.

The board that funds this hall meets quarterly, and the quote is due the Friday after next. Before anyone argues fabrics, find out which document already answered it.

1The question behind the question

When a Dell account asks “Ethernet or InfiniBand?”, almost nobody is asking for a protocol comparison. They are asking whether the thing they are about to sign is buildable and supportable. So the first criterion is not performance, it is compliance: which published reference architecture is this customer bound to, and what does that document already decide for them?[5][6]

The clearest illustration is a single server sized two ways. NVIDIA publishes a DGX B300 SuperPOD in a Spectrum-4 Ethernet edition with DC busbar power, where “Traffic per rail of the DGX B300 systems is always one hop away from the other 64 nodes in a SU” and rack power “exceeds 50 kW”.[1][2] It also publishes a Quantum-X800 InfiniBand edition with AC power, where “Each group of 72 nodes is rail-aligned” and the rack is “~56kW”.[3][4] Same server, two documents, two scalable-unit sizes, two rack-power statements, two BOMs. If the customer’s NVIDIA account team quoted from one of them, the fabric question is already answered and your job is arithmetic, not persuasion.[2][4]

The Enterprise RAs sit under those. The HGX AI Factory RA documents deployments “ranging from 32 nodes (256 GPUs) to 128 nodes (1,024 GPUs)”, and the NVL72 AI Factory RA is “fully tested” to “8 SUs (Scalable Units)” and scoped to “multi-user, single tenant workloads”.[5][6] Those are the documents most Dell PowerEdge deals live inside, and both are Spectrum-X Ethernet designs.[6]

Unknown

Not decidable yet — 3 decisive criteria still unanswered

Weighted points — Ethernet 0 · InfiniBand 0 · neutral 0 · unknown 13 (of 13). Criteria 1, 3 and 4 count double.

Criterion 1 · decisive

Does the reference architecture the customer must comply with already pick a fabric?

Your answer: Unknown

Evidence — Spectrum-X Ethernet

The DGX SuperPOD B300 RA in its Spectrum-4 Ethernet / DC-busbar edition sizes an SU at 64 nodes = 512 GPUs, four DGX B300 per rack, and states "The rack-level power consumption per rack exceeds 50 kW". The GB300 NVL72 Enterprise RA runs Spectrum-X at 800 Gb/s per GPU and is "fully tested" to 8 SUs.

Evidence — Quantum InfiniBand

The same DGX B300 server has an XDR InfiniBand / AC-PDU edition of the SuperPOD RA where an SU is 72 nodes = 576 GPUs, racks draw "~56kW" on three 2U PDUs, and "UFM 3.5 nodes are connected to four (4) FNM ports on the Q3400 switches".

Not claimed anywhere we fetched

Nothing on the fetched pages publishes a rule for mixing the two editions inside one SU, or a migration path from 64-node SUs to 72-node SUs.

FAE angle

Same server, two RAs, different SU arithmetic — and the SU is the unit of quoting. Ask which RA revision the customer's NVIDIA account team quoted from before you ask what they prefer.

Recommendation caveats
  • Answer these first: 1. Which RA binds them; 3. Reduction-heavy collectives; 4. Storage constraint.
  • These ten criteria are an assembled framework, not a published NVIDIA scoring model. The evidence in each cell is published; the weighting is ours.
Start with every criterion unknown. Click a criterion to read the evidence for both sides and the line that says what is not claimed anywhere, then set the customer's answer and watch the tally move.

2Ten criteria, each with its evidence

The framework below is assembled from the published pages; every row cites where its evidence comes from, and the honest rows are the ones where the evidence is thin.

# Criterion What the pages support
1 Which RA binds them 64-node vs 72-node SU on the same DGX B300; Enterprise RAs are Spectrum-X[1][3][6]
2 Scale target Spectrum-X multiplane “up to 128K GPUs in two tiers”; the NVIDIA Quantum InfiniBand family “Over 10,000 nodes in two-level fat tree” — a family-level benefit statement, not a per-generation ceiling[7][8]
3 Reduction-heavy collectives SHARP runs reductions in the switch via sharp_am and libsharp; Quantum-X800 lists “SHARP v4”[10][9]
4 Storage constraint PowerScale’s back end is Ethernet and Dell-qualified on Cumulus 5.9.1[11]
5 Operations skills on staff Ethernet reuses the RFC 7938 eBGP Clos and ECMP; InfiniBand adds SM, PKeys, UFM[12][16]
6 Multi-tenancy model NCP-AIN pairs “BGP-EVPN to isolate tenant workloads” with “partition keys (PKeys)”[13]
7 Adaptive routing at your speed Spectrum-4 AR runs at 400G and 200G and “does not support adaptive routing on 800G links”[14]
8 Version-discipline burden Spectrum-X is a validated stack of matched versions; IB’s analogue is UFM and SM alignment[15][16]
9 Optics and power at 400G The same MMA4Z00-NS twin-port class serves air-cooled Quantum-2 and Spectrum-4[17]
10 Support path Dell resells both families and offers support end to end; no page publishes a defect RACI[18][19]

Three of those carry real decision weight. Criterion 1 usually settles it. Criterion 4 settles it whenever the customer has already bought Ethernet-only storage — Dell states PowerScale requires a private back-end network, isolated from the front end, and that “Dell does not support connecting any other devices to the back-end switches”.[11] Criterion 3 is the only one that can reverse an Ethernet-leaning answer, because in-switch reduction is a mechanism, not an adjective.[10]

The rest are close enough on published evidence that they should not carry a recommendation. Both scale ceilings are far above an enterprise buying under 1,024 GPUs.[7][8][5] The optics are the same class of part.[17] And NVIDIA itself publishes the strongest operational claim for the other side: InfiniBand “is actually simpler to deploy and maintain”.[20] Quote that sentence yourself before the customer finds it.

3What nobody publishes, and how to say it

An FAE’s credibility is spent fastest on claims that cannot be checked. Four gaps in this framework are worth naming explicitly.

No Spectrum-X in-switch reduction. The Spectrum-X platform page claims 1.6x performance over off-the-shelf Ethernet, the 128K-GPU multiplane figure, 1.9x cross-site NCCL with Spectrum-XGS, and photonics numbers — and it claims nothing about in-network reduction.[7] SHARP has an install guide, a manager process in UFM and a library that links into the application; it is an InfiniBand feature.[10] Say “correct, that is an InfiniBand capability” and move to whether the workload’s collectives are reduction-bound.

No defect-ownership RACI. Dell’s blog offers help from design consulting to post-deployment troubleshooting and Dell’s catalog sells both Spectrum Ethernet and Quantum-X800 switches, but no fetched page delineates who owns a switch defect in a Dell-sold AI Factory.[18][19] Teach it as a known unknown with two questions to ask instead: whose asset tag is on the switch, and which contract number covers it.

No head-to-head benchmark. NVIDIA’s InfiniBand operations blog is operational — UFM tooling, maintenance cadence, alert workflows — and contains no MLPerf reference and no comparison to Ethernet.[20] Do not attribute one to it.

Version pins that collide. The Spectrum-X validated solution stack publishes matched rows across switch NOS, NIC firmware, DOCA-Host, NetQ and NCCL; the PowerScale back-end document pins Cumulus Linux 5.9.1 on SN5600.[15][11] A single switch cannot satisfy both pins. That is a design constraint you raise in the review, not a paperwork detail you discover at cutover.

4Turning ten criteria into two hundred words

A recommendation that reads like a preference gets re-litigated by the next person who walks into the room. One that names its evidence and its gaps survives review. The pattern: state the binding constraint, state the two facts the customer already owns, state the one criterion that could reverse the call, and list what you could not verify.

Two recommendations, written to survive review

Customer A — enterprise, 256 GPUs of HGX B300, PowerScale F710 already installed, ops team runs BGP everywhere.

Recommendation: Spectrum-X Ethernet.

Evidence walk: the deal sits inside the HGX AI Factory Enterprise RA, which documents 32 nodes (256 GPUs) to 128 nodes (1,024 GPUs) and is a Spectrum-X design — criterion 1 answered by compliance.[5] The storage is PowerScale, whose back end is Ethernet and must be a private isolated network with no other devices attached — criterion 4 answered by an asset they already own, and it also tells us the AI fabric and the storage back end will be separate switches because Dell qualified SN5600 there on Cumulus 5.9.1 while the Spectrum-X stack pins a different NOS row.[11][15] Criterion 5 favours Ethernet: their staff already run the RFC 7938 eBGP Clos and ECMP that the underlay is built from.[12] Criterion 3 is the reversal test: nothing in the ask says the training is reduction-bound, so the SHARP argument does not apply, and I say plainly that SHARP has no published Spectrum-X equivalent.[10]

Thin evidence, stated: no published defect RACI for a Dell-sold switch, so we will name the support contract explicitly in the design note.[18][19]

Customer B — research site, 2,000 GPUs, all-reduce-dominated training, no storage constraint stated.

Recommendation: InfiniBand is a legitimate candidate and I will not close it on the call.

Evidence walk: criterion 3 is live — SHARP performs reductions in the switch, with sharp_am in UFM and libsharp linked into the application, and Quantum-X800 lists SHARP v4.[10][9] Criterion 2 does not decide: the NVIDIA Quantum InfiniBand switching page publishes “Over 10,000 nodes in two-level fat tree” as a family-level highlight and Spectrum-X publishes 128K GPUs in two tiers, so both clear 2,000.[8][7] Criterion 5 counts against InfiniBand only if they have no fabric-management staff, since it introduces subnet manager, PKeys and UFM as new objects.[16][13]

Next step, not a recommendation: get one reduction-heavy job profiled. The decision turns on a measurement nobody in the room has.

Episode 1 — What went in the right-hand column

How it ended

Twenty minutes, and nobody has opened a benchmark. The storage is PowerScale, which forces Ethernet for at least the back end and pins its own Dell-qualified NOS version, and no one has profiled a training job to see whether the collectives are reduction-bound.[11][10] Two of the three questions are answered by assets the customer already owns; the third is a measurement they will go and take. The recommendation goes in the network lead’s left column and the thin-evidence criteria in her right.

Then she closes the notebook and says the sentence she actually came to say: in 2019 they ran RoCE on somebody else’s switches and spent six months chasing PFC storms, and she is not doing that again.

Lab

Not applicable on the Dell-lab BlueField-3 and ConnectX hosts: this lesson decides between two fabrics and the lab has switches for neither. Nothing here is mutating; every step is a read.

Pre-flight inventory (5 minutes, useful anyway): on the Dell-lab host, record ibstat output, ibdev2netdev mapping and the negotiated link speed per port, so your written recommendations reference a machine you have actually touched.

Optional, in a customer lab or on borrowed InfiniBand time — book an hour on a live IB fabric and run the read-only set: ibstat, ibnodes, iblinkinfo and ibdiagnet. The objective is not the output, it is the operational model. Write down the three objects that have no Ethernet equivalent — the subnet manager, PKeys and UFM — and one sentence each on who on the customer’s staff would own them at 3 a.m.[16][13] Rollback: none required, all four commands are read-only diagnostics; do not run any ibportstate or SM configuration command on a fabric you do not own.

Retrieval check

10 questions from memory. Answer before looking anything up; misses become flashcards.

Explain it to a Dell SE

Explain to a Dell account executive in four sentences how you decide Ethernet or InfiniBand for a customer without using the word better.

14 flashcards for this lesson — 0 in deck. Spaced review lives at /review.

Sources

Facts in this lesson were checked against NVIDIA Spectrum-X platform page, Quantum-X800 page, NCP-AIN blueprint and Cumulus Linux 5.13 ECMP page re-fetched 2026-09-07. Dates are when each page was fetched.

  1. Network Fabrics — DGX SuperPOD B300 Spectrum-4 Ethernet and DC Busbar RA · fetched 2026-09-07
  2. DGX SuperPOD Architecture — B300 Spectrum-4 Ethernet and DC Busbar RA · fetched 2026-09-07
  3. Network Fabrics — DGX SuperPOD B300 Quantum-X800 InfiniBand and AC Power RA · fetched 2026-09-07
  4. DGX SuperPOD Architecture — B300 Quantum-X800 InfiniBand and AC Power RA · fetched 2026-09-07
  5. NVIDIA HGX AI Factory Enterprise Reference Architecture (index) · fetched 2026-09-07
  6. Overview — NVIDIA NVL72 AI Factory Enterprise RA · fetched 2026-09-07
  7. NVIDIA Spectrum-X Ethernet Platform · fetched 2026-09-07
  8. NVIDIA InfiniBand Switches · fetched 2026-09-07
  9. NVIDIA Quantum-X800 InfiniBand Platform · fetched 2026-09-07
  10. NVIDIA SHARP Installation · fetched 2026-09-07
  11. Dell PowerScale: Ethernet Back-End Network Overview H16346.8 · fetched 2026-09-07
  12. RFC 7938 — Use of BGP for Routing in Large-Scale Data Centers · fetched 2026-09-07
  13. NVIDIA-Certified Professional: AI Networking (NCP-AIN) blueprint · fetched 2026-09-07
  14. Equal Cost Multipath Load Sharing incl. Adaptive Routing | Cumulus Linux 5.13 · fetched 2026-09-07
  15. NVIDIA Spectrum-X Validated Solution Stack · fetched 2026-09-07
  16. NVIDIA UFM Unified Fabric Manager · fetched 2026-09-07
  17. MMA4Z00-NS 800Gb/s Twin-port OSFP transceiver overview · fetched 2026-09-07
  18. Dell-NVIDIA Partnership Powers High-Performance AI Fabric Solutions · fetched 2026-09-07
  19. AI Networking Switches | Dell USA · fetched 2026-09-07
  20. Simplifying Network Operations for AI with NVIDIA Quantum InfiniBand · fetched 2026-09-07
  21. Networking Physical Topologies — NVIDIA HGX AI Factory · fetched 2026-09-07

The same idea elsewhere

Other lessons that cover this ground, sometimes from another course's angle.