Skip to content

Seven objections and the published answers

S4·E2Six months of PFC storms · The customer's network operations room, day two of the Dell AI Factory workshop

S4·E2Evaluate~25 minsources checked todayverified against Cumulus Linux 5.13 RoCE and ECMP pages, NetQ 4.8 adaptive-routing monitoring page, ultraethernet.org and the Spectrum-X validated stack, re-fetched 2026-09-07

Builds on: The Ethernet vs InfiniBand decision framework

Before you read: what do you already know?

3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.

After this lesson you can

  • Answer each of the seven recurring fabric objections with a published sentence rather than an adjective.
  • Judge which objection must be conceded and state the concession in one sentence.
  • Produce the artifact that backs each answer - a command output or a version row - instead of a slide.
  • Distinguish an objection about mechanism from an objection about risk and answer each in its own currency.

Episode 2 — Six months of PFC storms

The situation · The customer's network operations room, day two of the Dell AI Factory workshop

The network lead has been quiet all morning. When the RoCE slide goes up she taps the notebook twice and repeats it — six months, 2019, somebody else’s switches — and then adds the three words she uses instead of a question: show me the counter. The Dell SE looks at you over a fresh coffee he will also forget to drink. Every instinct offers a reassuring adjective.

Do not reach for it. She is not attacking the design; she is asking whether some published thing exists that stops her six months happening twice. It does, and it exists because that experience was common enough to be engineered away: the Ethernet Storage Fabric guide states that “Cumulus Linux simplifies the RoCE configuration to a single command and automatically applies best practice settings for optimal performance”, covering buffer pools, priority mappings, PFC and ECN thresholds.[2] If PFC is a red line for her, nv set qos roce mode lossy is a documented ECN-only path rather than an argument.[1] And there is a price attached that you volunteer rather than hide, because NetQ states that “RoCE lossless mode must be enabled to display adaptive routing data”.[4]

Answer with a published sentence, name the artifact you will show, then say what the evidence does not cover.

Her objection has an answer either way, and both halves of it were written by somebody else. There are seven of these, and she has four more on the page.

1An objection is a request for evidence

Seven objections come up in nearly every AI-fabric conversation. They are not attacks; each one is a customer asking whether some published thing exists. The failure mode for an FAE is answering a request for evidence with an adjective — “robust”, “proven”, “industry-leading” — which the customer’s other vendor can match word for word and which does not survive their internal review.[10]

The discipline is one sentence long: answer with a published sentence or a version row, then name the artifact you would show. If you cannot find the sentence, say so and convert the objection into a question you will go and answer. That is a better meeting outcome than a confident claim that gets checked next week.

Two of the seven have real teeth. “Ethernet cannot do in-network reduction” is correct and must be conceded.[7] “We would have to learn a whole new protocol” has published support on both sides, because NVIDIA itself writes that InfiniBand “is actually simpler to deploy and maintain”.[9] The remaining five have clean, checkable answers.

Answer out loud in under 90 seconds first, then reveal the published counter and grade yourself: did you state a published fact, or an adjective?

Objection 1

RoCE is lossy and PFC storms will kill us.

Objections view. Select each of the seven, read the published counter and its source, then close the panel and say the answer out loud in under 90 seconds before you look again.

2Objections one to three: mechanism questions

1. “RoCE is lossy and PFC storms will kill us.” The published answer is that the configuration surface is one command: the Ethernet Storage Fabric guide states that “Cumulus Linux simplifies the RoCE configuration to a single command and automatically applies best practice settings for optimal performance”, covering “buffer pool settings, traffic classification and priority mappings, PFC and ECN thresholds, and queuing and scheduling settings”.[2] NVUE defaults to lossless, so nv set qos roce and nv set qos roce mode lossless are equivalent, and a customer who refuses PFC outright has a documented path: nv set qos roce mode lossy is ECN only, no PFC.[1] Carry the trap with you: “If you enable mode lossy, configuring nv set qos roce without a mode does not change the RoCE mode” — you must set mode lossless explicitly to go back.[1] The artifact is nv show qos roce and nv show interface <iface> qos roce counters.[1]

2. “ECMP hashing collides elephant flows.” Two published mechanisms, not one. Adaptive routing is documented on “Switches with the Spectrum-4 ASIC at 400G and 200G speeds”, enabled globally with nv set router adaptive-routing enable on and per interface on every port in the same ECMP route.[3] And the reference architectures already remove the single-hash exposure: each 800 Gb/s SuperNIC port is broken into “2x400 Gb/s interfaces”, each landing on “a different leaf switch”, each leaf “part of an independent fabric”.[5][6] Two caveats you volunteer rather than hide: adaptive routing “does not make use of resilient hashing” and is unsupported on layer-3 subinterfaces, SVIs, bonds and bond members; and enabling or disabling it “reloads the switchd service”, so it is a maintenance-window change.[3] The monitoring artifact carries its own prerequisites — NetQ documents that adaptive-routing monitoring is supported on Spectrum-4, “requires a switch fabric running Cumulus Linux 5.5.0 or above”, and that “RoCE lossless mode must be enabled to display adaptive routing data”, with lossy-mode switches appearing in the UI but showing none.[4] That matters directly: a customer who took the lossy path in objection 1 has traded away the AR telemetry in objection 2.

3. “Ethernet cannot do in-network reduction.” Correct. SHARP performs reductions inside the fabric, with sharp_am running in UFM and libsharp linked into the client application over IPoIB on TCP port 6126, and Quantum-X800 lists “SHARP v4” among its features.[7][14] The Spectrum-X platform page claims 1.6x over off-the-shelf Ethernet, 128,000 GPUs in two tiers with multiplane, and photonics efficiency numbers — and claims no in-network reduction.[12] Concede in one sentence, then ask the only question that matters: are these collectives reduction-bound enough to pay for a second operational model?

3Objections four to seven: risk questions

The last four are not about mechanism. They are about risk: staffing risk, maturity risk, budget risk and timing risk. Answering them with mechanism misses.

4. “We’d have to learn a whole new protocol.” The Ethernet answer is that the underlay is RFC 7938’s eBGP Clos, where “ECMP is the fundamental load-sharing mechanism used by a Clos topology” and a single ASN is allocated to all Tier 1 devices with a unique ASN per Tier 3 device — the same design the customer’s existing data centre already runs — plus RoCE QoS on top.[8][1] The InfiniBand answer, published by NVIDIA, is that it “is actually simpler to deploy and maintain”.[9] Present both and let the customer’s own staffing decide. This is a staffing question wearing a protocol costume.

5. “800G is unproven.” Point at the table. The Spectrum-X validated solution stack has published versioned rows from “April 2024 (v 1.0.1)” through “September 2026 (v2.3.1)”, validating three GPU platforms — GB300, B300 and H200 — with each row naming matched versions for switch NOS, ConnectX-8 and BlueField-3 firmware, DOCA-Host, NetQ, NCCL, HPC-X, Network Operator and DTS.[10] A published cadence of matched rows is evidence; “battle-tested” is not.

6. “Can we start small and grow?” Single plane is the documented cost option: “a single 1x 400 Gb/s connection” per GPU into “a leaf switch within a single compute fabric that scales to 1024 interfaces of 400 Gb/s”, explicitly at 50% of the bandwidth and “well-suited for workloads that do not warrant maximum throughput”.[5] Same switches, same NICs — and the RA is explicit that the cages are common: “The use of OSFP transceiver modules enables seamless migration between single plane and dual plane topologies, allowing flexibility in adapting to evolving performance needs.”[5] Quote that sentence first, then say what it does not cover: the RA also expresses the choice as “dual-plane using twin-port transceivers or single-plane using single-port OSFP transceivers”,[15] so “seamless” means the cage accepts either module, not that the upgrade is free — growing later still buys new optics on every node and a second fabric of leaves that does not exist yet.

7. “What about Ultra Ethernet — should we wait?” State facts and stop. The consortium’s mission is to “Deliver an Ethernet based open, interoperable, high performance, full-communications stack architecture to meet the growing network demands of AI & HPC at scale”; specification version 1.0.3 is released; NVIDIA and Dell are both listed General Members.[11] No page we fetched supports a product-timing claim, so any date in your answer is yours, not NVIDIA’s.

4Ninety seconds, four moves

An answer that runs long stops being evidence and starts being advocacy. Four moves, in order: restate the objection in their words; give the published sentence or version row; name the artifact you will show; state the residual risk the evidence does not cover.[10]

Positioning: three customers, one afternoon · decision 1/7A 0 · P 0 · S 0

Brief — Back-to-back calls: a VMware farm (200× R760), an AI training pod (64× XE9780 on Spectrum-X) and a storage-heavy multi-tenant inference cluster. Each has objections.

Dell account team + three end customers: VMware farm: "200 R760s on vSphere 8; NSX eats about 20% of our cores. Should we go DPU?"

Three customers in one afternoon, each with objections. Pick the answer that quotes a published fact rather than the one that sounds most confident, then read the scoring on accuracy, positioning and safety.
One objection, three drafts

Objection, verbatim from the customer: “We ran RoCE in 2019 on a different vendor’s switches and spent six months chasing PFC storms. Why is this different?”

Draft 1 — the adjective version (fails). “Spectrum-X is a purpose-built AI fabric with best-in-class congestion control, so PFC storms are not a concern.” Nothing here is checkable and none of it is a published sentence. The customer’s other vendor will say the same words about their own product this afternoon.

Draft 2 — the mechanism dump (fails differently). Twelve minutes on buffer pools, DSCP-to-switch-priority mappings, ECN thresholds and the CNP traffic class. All true, all cited, and it answers a question nobody asked: the objection was about six months of their engineers’ time, not about queueing.

Draft 3 — the one that works.

  1. Restate: “So the concern is not whether RoCE can be tuned, it is whether your team will spend another half-year tuning it.”
  2. Published sentence: Cumulus “simplifies the RoCE configuration to a single command and automatically applies best practice settings”, covering buffer pools, classification and priority mappings, PFC and ECN thresholds, and queuing.[2] And if PFC is a red line for you, nv set qos roce mode lossy is a documented ECN-only mode.[1]
  3. Artifact: “I will bring you nv show qos roce and nv show interface swp1 qos roce counters from a switch, so you can see the applied state rather than take my word for the defaults.”[1]
  4. Residual risk, volunteered: “Two things that evidence does not cover. First, if you choose lossy mode you lose adaptive-routing data in NetQ, which needs lossless plus Cumulus 5.5.0 or later on Spectrum-4.[4] Second, the RoCE defaults are per-ASIC — the docs say configuration generated for one Spectrum generation is not applicable to another — so the config from your 2019 fabric is not a starting point here.”[1]

Elapsed: about 75 seconds. The customer now owns a decision (lossless vs lossy) instead of a doubt.

Episode 2 — Into the minutes

How it ended

She gets no reassurance. She gets one command, the trap that rides along with it — set mode lossy and a bare nv set qos roce will not put it back — the telemetry she trades away if she takes it, and a promise of nv show qos roce output from a real switch on Thursday instead of a slide.[1][4] When she asks whether to wait for Ultra Ethernet, the NVIDIA PM on the bridge says “not announced”, so you give her what is published and stop: specification version 1.0.3 released, NVIDIA and Dell both listed General Members.[11]

At lunch, procurement forwards a competitor’s design guide with one sentence highlighted in yellow, and asks what it means.

Lab

Capture the NIC-side evidence behind objections 1 and 2 on the Dell-lab BlueField-3 or ConnectX host. Read-only throughout: no firmware writes, no mode change, no QoS change, so no rollback is required — but record the pre-flight state anyway so any later drift is attributable.

Pre-flight inventory: mlxfwmanager --query for board and firmware, ibdev2netdev for the device-to-interface map, and ip -br link for the interface list. Save all three to a file before you touch anything.

  1. Read the negotiated link speed and duplex on the fabric-facing interface: ethtool <iface> | head -20. Expected: the negotiated speed matches what the RA assumes for the plane design; if it negotiated lower, that is the artifact, not a footnote.
  2. Capture the congestion counters that back objection 1: ethtool -S <iface> | grep -i -E 'pause|ecn|cnp'. Expected: named counters for pause frames, ECN-marked packets and CNPs. If the grep returns nothing, record the exact driver and firmware version — the counter names differ by generation and “no counters” is a finding you must be able to explain.
  3. Repeat step 2 under load if you have a traffic generator available, and diff against the idle capture. Expected: you can point at which counter moved. This is the difference between telling a customer that congestion signalling exists and showing them it fired.
  4. Read the transceiver actually installed: ethtool -m <iface> | head -30. Expected: vendor, part number and type. Compare it to what a dual-plane design would require — a twin-port module — and note the mismatch if your lab card is single-port.[15]
  5. Write two sentences per captured artifact explaining which objection it answers and what it does not prove. Expected: at least one honest “this does not prove” line per artifact, because a host counter says nothing about switch buffer behaviour.
  6. Optional, in a customer lab only: repeat step 2 on a switch-attached host during a real all-reduce and keep the output with the design note. Nothing in this lab modifies the host; if you did change anything by accident, the rollback is to restore the interface state recorded in the pre-flight inventory.

Retrieval check

10 questions from memory. Answer before looking anything up; misses become flashcards.

Explain it to a Dell SE

Explain to a Dell SE why 'RoCE is lossy and PFC will storm' is answered with a command rather than a reassurance.

14 flashcards for this lesson — 0 in deck. Spaced review lives at /review.

Sources

Facts in this lesson were checked against Cumulus Linux 5.13 RoCE and ECMP pages, NetQ 4.8 adaptive-routing monitoring page, ultraethernet.org and the Spectrum-X validated stack, re-fetched 2026-09-07. Dates are when each page was fetched.

  1. RDMA over Converged Ethernet - RoCE | Cumulus Linux 5.13 · fetched 2026-09-07
  2. Cumulus Linux Configuration Guide for Ethernet Storage Fabrics · fetched 2026-09-07
  3. Equal Cost Multipath Load Sharing incl. Adaptive Routing | Cumulus Linux 5.13 · fetched 2026-09-07
  4. Adaptive Routing monitoring | Cumulus NetQ 4.8 · fetched 2026-09-07
  5. Networking Physical Topologies — NVIDIA HGX AI Factory · fetched 2026-09-07
  6. Networking Physical Topologies — NVIDIA NVL72 AI Factory · fetched 2026-09-07
  7. NVIDIA SHARP Installation · fetched 2026-09-07
  8. RFC 7938 — Use of BGP for Routing in Large-Scale Data Centers · fetched 2026-09-07
  9. Simplifying Network Operations for AI with NVIDIA Quantum InfiniBand · fetched 2026-09-07
  10. NVIDIA Spectrum-X Validated Solution Stack · fetched 2026-09-07
  11. Ultra Ethernet Consortium · fetched 2026-09-07
  12. NVIDIA Spectrum-X Ethernet Platform · fetched 2026-09-07
  13. NVIDIA DSX Air Platform for AI Factory Simulation · fetched 2026-09-07
  14. NVIDIA Quantum-X800 InfiniBand Platform · fetched 2026-09-07
  15. Networking Hardware — NVIDIA HGX AI Factory · fetched 2026-09-07
  16. NVIDIA Spectrum SN5600 Series Switches Datasheet (Dell-branded) · fetched 2026-09-07

The same idea elsewhere

Other lessons that cover this ground, sometimes from another course's angle.