Skip to content

Positioning multipath: what ships, what is a specification

S5·E5A quarter later, the RFP paragraph · NVIDIA briefing room, a Thursday, one quarter after the acceptance review

S5·E5Create~40 minsources checked todayverified against Re-fetched 2026-09-11: the NVIDIA MRC blog, the IBTA Release 2.1 overview deck, arXiv 2605.04333v1 and 2605.21187, the OCP MRC 1.0 URL (HTTP 403), the Broadcom MRC article, Ultra Ethernet Specification v1.0.3 and the UEC compliance page, the NVIDIA Spectrum-X product page and the 2026-08-24 giga-scale blog, the UEC members page (logos render in JavaScript, no member name extractable), the NVUE adaptive-routing set/unset and show references, the Cumulus Linux 5.14 and 5.18 Equal Cost Multipath Load Sharing pages, the rdma-core v65.0 mlx5dv_create_qp(3) man page, the Tech Preview Spectrum-X NIC Configuration page, the NCCL 2.31.2 environment variables page, the DOCA 3.5.0 out-of-order data placement page, the Dell-hosted SN5600 datasheet and the Dell PowerEdge XE9680 Technical Guide

Builds on: Multipath RC: a new transport in the annex, not in the API, NCP-AIN: what this course covers and what it does not

Before you read: what do you already know?

3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.

After this lesson you can

  • Separate the four different things a customer calls "multipath" and answer the one that was actually asked.
  • Tier every claim you repeat as specification, preprint, vendor blog, marketing or not-published, with its date.
  • Refuse the four sentences about MRC, UEC and adaptive routing that the public record does not support, and say the grounded replacement instead.
  • Compose a one-page, dated answer to an RFP networking paragraph in which every line carries its evidence tier and its re-verify date.

Episode 5 — A quarter later, the RFP paragraph

The situation · NVIDIA briefing room, a Thursday, one quarter after the acceptance review

The account manager has the one-page map from last quarter printed out, and she has kept the receipt. Nobody bought exam seats. What she is holding now is worse: a customer RFP whose entire networking section is one paragraph, and the paragraph asks for “a multipath RDMA transport, UEC-certified switching, and adaptive routing across all available paths.” Response due Tuesday. The Dell SE is already at row 63 of his spreadsheet of promises, his coffee is cold, and he wants one word: yes. Procurement, on speaker, wants to know the lead time on whatever the word turns out to mean. The NVIDIA PM, from the corner, says “not announced” to a question nobody has asked yet.

Multipath transports exist because a reliable connection has always been pinned to one path. A competitor states the problem better than our own collateral does: RoCE “requires in-order, lossless delivery of packets in a connection. The in-order requirement forces the network to use a single path per connection, and resulting collisions between different high-bandwidth flows inevitably creates congestion.”[5] The industry’s answer has a name, MRC, and it is genuinely standardised — the IBTA added Multipath Reliable Connection to InfiniBand Architecture Specification Volume 1 Release 2.1 as Annex 21 on July 31, 2026.[3] It is also, on every public page, a thing with no switch to flip: the one NVIDIA post that announces it names no queue pair, no verbs call, no configuration knob of any kind.[1]

A quarter is exactly long enough for a specification to become a requirement in somebody’s RFP.

Start with what that paragraph is actually asking for.

1Four things a customer calls multipath

The paragraph names three things and means four. Sort them before you answer, because answering the wrong one costs you the room.

MRC, the transport. NVIDIA’s definition: “MRC enables a single RDMA connection to distribute traffic across multiple network paths, improving throughput, load balancing and availability for large-scale AI training fabrics.”[1] The designers describe it as extending “the RoCEv2 Reliable Connection (RC) transport protocol to support multi-path operation borrowing several features of UET”, supporting the normal verbs interface and queue-pair abstraction “but only for write and write-with-immediate operations.”[2] One naming note and then never again: the term is MRC, not MPRC — no vendor, specification or repository uses MPRC — and inside NVIDIA’s own DOCA corpus “MRC” already means Multi-Rack Connectivity, a service, so define the acronym before you use it.[1][2]

Switch-side adaptive routing. A load balancing feature on the switch, enabled with switch-only NVUE commands, disabled by default, with both state toggles “Introduced in Cumulus Linux 5.1.0”.[11] This is the one you can configure this afternoon — carrying the support list of the release the customer runs, because from Cumulus Linux 5.16 the page adds a constraint that is absent in 5.14: “Adaptive routing is only supported on the NVIDIA Spectrum-X networking platform”.[19][20]

More queue pairs per connection. NCCL_IB_QPS_PER_CONNECTION, “a number between 1 and 128, default is 1”, whose stated rationale is routing entropy: “This can be useful on multi-level fabrics which need multiple queue pairs to have good routing entropy.”[13] More hash draws, not per-packet spreading of one connection.

UET packet spraying. The Ultra Ethernet Specification defines three entropy modes — single path, oblivious spraying, and path-aware spraying where the entropy value changes “based on feedback provided by the destination as to which EVs saw congestion.”[6] That is a different transport, not a mode of RoCE.

A fifth thing hides underneath and is not any of the four: out-of-order data placement, which NVIDIA documents for “ConnectX-5 adapters and above; RC and XRC QPs; DC transport” under its InfiniBand chapter.[14] It is a receiver capability, not a path-selection mechanism. Never let it be introduced as “our MRC”.

Positioning: three customers, one afternoon · decision 1/7A 0 · P 0 · S 0

Brief — Back-to-back calls: a VMware farm (200× R760), an AI training pod (64× XE9780 on Spectrum-X) and a storage-heavy multi-tenant inference cluster. Each has objections.

Dell account team + three end customers: VMware farm: "200 R760s on vSphere 8; NSX eats about 20% of our cores. Should we go DPU?"

Positioning drill: three customers, one afternoon. Only the fourth objection is about spraying and reordering; the transferable part is scoring yourself on accuracy, positioning and safety before you write the RFP page.

2What NVIDIA's own pages will and will not back you on

Exactly one NVIDIA page names MRC in the multipath sense, and it is a corporate blog post published 2026-05-06 and updated 2026-06-25.[1] Say “in the multipath sense” out loud, because a plain string search does not agree with you: searching the DOCA documentation set on 2026-09-11 returns one more occurrence of the letters MRC, a bug-fix row where MRC expands to Multi-Rack Connectivity, the DOCA service. That is a search result with a date, not a page you can send. Know precisely what the blog post carries, because in a meeting you will be quoting it from memory.

It carries the functional definition above, the collaboration list — NVIDIA “collaborated on MRC development with AMD, Broadcom, Intel, Microsoft and OpenAI” — and the provenance sentence: MRC was “proven first in production with performance optimized on NVIDIA Spectrum-X Ethernet hardware and now released as an open specification through the Open Compute Project”, where “released” hyperlinks to an OCP document.[1][4] That document could not be read: the OCP URL returns HTTP 403 to every automated fetch, so its contents, packet formats and field widths are unverified, and you must not upgrade “released as an open specification” into “an OCP standard” or “royalty-free”.[4]

What the page does not carry is the whole positioning argument. It names no QP type, no verbs API, no driver, no firmware setting and no configuration knob; the words “queue pair”, ibv_ and mlxconfig do not appear on it, and neither does “RoCE”.[1] Its only hardware phrase is the generic “NVIDIA ConnectX SuperNICs” — ConnectX-8 appears zero times.[1] ConnectX-8 enters the story only through the designers’ preprint, which names it exactly once in eighteen pages and reports no ConnectX-8 testbed, figure or measurement at all; every published experiment there runs on other vendors’ NICs.[2]

What NVIDIA documentation does back you on is narrower and more useful. Switch-side adaptive routing has NVUE commands and a version history.[11] Out-of-order data placement has a page and a published scope, and the DOCA page carries neither the flag nor the word “verbs” — the opt-in creation flag MLX5DV_QP_CREATE_OOO_DP and its capability query mlx5dv_ooo_recv_wrs_caps are documented in rdma-core, which is where you go for the query-then-set sequence, the both-ends requirement and the fact that there is no automatic fallback.[14][18] And the single NVIDIA documentation page that ties NIC settings to Spectrum-X is the Kubernetes NIC Configuration Operator page — marked Tech Preview in both its title and a warning, limited to “ConnectX-8 (device ID 1023) and BlueField-3 SuperNIC (device ID a2dc)”, and dependent on a DOCA package that is not public: “To access the package, contact your NVIDIA CPM.”[12]

3The numbers, and whose they are

Every number in this conversation belongs to a tier. Say the tier out loud with the number and you will never have to retract one.

Tier What it means Example in this topic
Specification A public normative document you can cite by clause Ultra Ethernet Specification v1.0.3, 574 pages, RFC 2119 keywords throughout[6]
Standards summary A body’s public overview of a members-only text IBTA’s Release 2.1 deck describing Annex 21[3]
Preprint Non-peer-reviewed, primary for design intent only arXiv:2605.04333v1, the MRC paper[2]
Vendor blog Authored by the vendor’s own architects, no method Broadcom’s MRC article; NVIDIA’s giga-scale post[5][8]
Marketing A product page or datasheet headline “Accelerate AI network performance by 1.6x over off-the-shelf (OTS) Ethernet”[10]
Not published The probe came back empty Any MRC configuration surface[1]

Three worked examples of the tiering, because this is where FAEs get caught. The 1.6x headline — “Accelerate AI network performance by 1.6x over off-the-shelf (OTS) Ethernet” — carries no footnote, no workload, no cluster size and no definition of the baseline it beats, and the product page carries no chart to support it either: marketing tier, quotable only as “NVIDIA’s product page claims”, and only with the metric named.[10] The 98% of theoretical line rate figure is better than its reputation but narrower than its usage: the blog summarises, and the paper behind it reports a 1st-percentile bandwidth of 377.23 Gbps on a 64-node subset of a Hopper cluster, under a worst-case allocation forcing every flow through a spine.[8][9] The 2.68 ms failover number is not measured against traditional Ethernet at all: the comparison is against a software NCCL load-balancing reference on the same Spectrum-X multiplane fabric, with a different benchmark, reaching 75% of line rate on three surviving planes of a 4-plane, 1152-GPU cluster — and in the single-plane case the RDMA connection crashes outright.[9]

The competitor’s numbers get the same treatment, not a worse one. Broadcom’s “sustain a 15:1 long-running many-to-one traffic pattern on every output port without building a trim queue” is an asserted capability from a vendor blog with no packet size, port speed, buffer configuration, metric or report behind it, and the figure appears on neither the Tomahawk 5 nor the Tomahawk 6 product page.[5] And the Ultra Ethernet Specification makes no performance claim comparing UET to anything — the percentages that are in its 574 pages are workload characterisations (“Up to 85%”, “20-40% overall BW”) and configuration defaults, not comparisons — so any “UEC says X percent faster” number came from a deck, not from the specification.[6]

The last one is the cross-check that ends most arguments: no primary source compares UET, MRC and SRD on the same benchmark. Each design describes itself and at best measures against TCP or against plain RoCEv2 with ECMP.[6][2][5] Any four-column comparative chart in front of you is synthesised.

Eight claims a customer, a competitor or your own deck will put in front of you. Predict first — the verdict, the probe and the evidence stay hidden until you commit.

Evidence tiers — what you may quote as fact, what you must attribute, what you must not use at allproduct documentationQuote it as fact. It carries commands, version scope and restrictions.standards bodyQuote it with its date and its access terms. Public summary is not normative text.preprintAttribute it to the paper, never to the vendor. Check whether the claim has a measurement behind it.vendor blogAnnouncement prose. Quote verbatim, attribute by name, and expect no methodology.marketing / datasheetA number here is a claim, not a measurement. Name the metric or do not use it.could not fetchSay "I could not read it". A 403 is a fact about the fetch, not about the document.no source at any tierThe probe came back empty. The negative finding IS the deliverable.
8 of 8 claims shown
  • On a slide in every deck, including yours. The network lead asks what the baseline was.

    Predict:
  • From procurement, after a competitor put "UEC certified" on a quote.

    Predict:
  • From a well-read customer engineer who found the DOCA page and assumed it applied.

    Predict:
  • Read straight off the Spectrum-X product page, usually while positioning a NIC refresh.

    Predict:
  • Nobody says this one. It is the thing you offer after refusing the other four.

    Predict:
  • The cheap alternative the customer will ask about once you tell them MRC has no knob.

    Predict:
  • From the competitor’s article, handed to you across the table by the network lead.

    Predict:
  • From the Dell SE, who needs a model number for the quote by Friday.

    Predict:
How the three verdicts are defined

Supported. A primary page — product documentation or a standards body — states it, and you can quote the sentence.

Unsupported. The primary pages do not state it. Either the probe came back empty, or the only thing asserting it is a blog, a deck or a datasheet with no method.

Unverifiable. The document that would settle it could not be read (403, members-only), or the record is silent in both directions.

Pick a claim on the left, commit a prediction, then read the probe that decided it.

A tier is not a quality judgement. A vendor blog can be the only public record of a real thing — it just cannot be quoted as documentation, and it never answers “what do I type”.

Eight claims a customer, a competitor or your own deck will put in front of you. Predict supported, unsupported or unverifiable for each, then read the probe that decided it.

4UEC-certified is not a thing

Procurement put “UEC-certified” in the RFP because a competitor put it on a quote. It describes something the Ultra Ethernet Consortium does not issue.

What UEC actually publishes is a specification. Version 1.0.3, dated July 16, 2026, is a public 574-page PDF with no registration, licensed CC BY-ND 4.0, and written in RFC 2119 conformance keywords from end to end.[6] Quote its conformance density only with the method attached — the customer has the same free PDF and a Ctrl-F key: text extraction on 2026-09-11 counted 969 occurrences of MUST, 98 of them MUST NOT, plus 49 SHALL.[6] Concede that cheerfully — it is a stronger public artifact than MRC has today, and saying so costs you nothing.

The specification defines four packet delivery modes, and says the set is deliberately closed: RUD, ROD, RUDI and UUD, with the note that “Unreliable ordered delivery is not specified. The four modes provided map to all common fielded use cases.”[6] ROD is the RC-shaped one and pays the RC-shaped price — packets “MUST be transmitted in order over the network interface using a single network path (i.e., using a single entropy value)”.[6] RUD is the design point MRC resembles: once-and-only-once delivery, packets “delivered to the semantic sublayer in the order they arrive from the network”, selective retransmission, direct data placement out of order.[6] It also defines three profiles, and is internally inconsistent about the middle one’s name — chapters 1 to 3 call it AI Full, chapter 7 calls it AI Extended, and the public compliance checklists say AI EXTENDED. Do not tell a customer one of the two names is stale.[6][7]

What no document defines is a certification. The compliance page links four documents — a Readme, a test-bed recommendations document, and a Transport and a PHY-LL matrix/checklist spreadsheet — and it is the Readme, not the page, that carries the sentence introducing the other three: they are “provided to implementors for the purpose of Ultra Ethernet [UE] Compliance self-attestation”.[7] The implementer tests “either by themselves or with the help of a third party” and then declares compliance in its own literature in a mandated format naming spec version, profile and checklist version.[7] There is no certification, no mark, no plugfest, no interop programme and no compliant-products list — and the page’s “UEC Member Declarations” link, which looks like one, is a Necessary-Claims patent register.[7] So the procurement-grade question is “show me your UE compliance statement — spec version, profile, checklist version”, and it is a better question than the one in the RFP.

Two more limits worth carrying. UEC does not forbid PFC: its position is that UET-CC “is designed for operation in a best-effort network and assumes that PFC is not enabled”, and “PFC SHOULD NOT be used anywhere in a best-effort network” — SHOULD-strength and conditional, not a prohibition.[6] And NVIDIA’s own posture on UEC has no page: the Spectrum-X product page contains zero occurrences of “Ultra Ethernet” or “UEC”.[10] Do not fill the gap with a membership tier either: the UEC members page renders its logos in JavaScript, so an automated fetch on 2026-09-11 extracted no vendor name at all, NVIDIA’s included — a fact about the page, not about the membership.[17] Say there is nothing official to point to rather than improvising a position in the room.

MRC and UET are relatives, not enemies, and the honest bridge is already published on both sides. The MRC paper says MRC “draws upon lessons from Ultra Ethernet Transport (UET)” and that “Like UET, MRC employs packet spraying, adaptive load balancing based on ECN, out-of-order memory placement of received data, selective retransmission, and uses packet trimming to mitigate incast.”[2] Broadcom’s article says MRC “shares many design principles with the Ultra Ethernet Transport (UET) specification, a key difference being that it appears as a minimal extension of the current Verbs specification, while UET is a clean slate redefinition with a larger scope.”[5] Say “appears as”, not “is” — the same page calls MRC’s design decisions “effectively … a completely new transport protocol with wide-ranging network implications”.[5]

5The hardware question underneath the architecture question

Every transport paragraph eventually becomes a slot question, and the slot question is where an FAE gets caught fastest.

Start with what NVIDIA publishes about the endpoints. Its marketing FAQ gives three SuperNIC generations: BlueField-3 for Hopper-era systems, ConnectX-8 for Blackwell at “800 Gb/s total throughput via 2×400 G, PCIe Gen6”, and ConnectX-9 for Vera Rubin NVL72 at “1,600 Gb/s per GPU via 4×200 G SerDes, PCIe Gen6”.[10] Quote those sentences and derive nothing from them: the ConnectX-8 figure is per NIC while the ConnectX-9 figure is per GPU, and four times 200 multiplies to 800, not 1,600.[10]

Then the Dell reality, which is what the customer can buy. The XE9680 Technical Guide carries a mixed-NIC rule that decides where the cards physically go: “the card type with the smaller quantity (≤2 cards) must be installed in Slot 31 and Slot 40 first”, after which “the remaining card type then fills the other available slots following the standard SPM order”, applying “to all high-speed NIC/DPU card types, including but not limited to BF3, CX-7, and CX-6”.[16] The same guide carries a flat exclusion that has ended PoCs after the order landed: “PowerEdge XE9680 system with MI300X GPUs does not support SmartNIC/DPUs.”[16] Say “SmartNIC/DPU”, not “Dell says no BlueField” — the guide never uses the word BlueField in 68 pages, and the MI300X configuration still supports several Mellanox NIC options.[16] Ask which GPU the box shipped with before you position a DPU.

And the switch side of the quote has an absence in it that you must state as an absence. The Dell-hosted SN5600 datasheet contains zero occurrences of “adaptive”: it describes Spectrum-X, RoCE “with extensions for AI cloud servers” and “256-way equal-cost multi-path (ECMP) routing for load balancing and redundancy”, and never names adaptive routing at all.[15] No Dell document reachable by any method publishes adaptive-routing support at model granularity for an SN-series switch. The defensible sentence is “no Dell document I can reach publishes it” — never “Dell does not support it”. Two halves, two verdicts. What Dell publishes is settled: the datasheet is in front of you and the word is not in it.[15] What Dell supports stays open, because the Dell SONiC release notes that would settle it could not be reached by any method I tried — record that as a dated probe of your own, not as a fact about Dell’s support matrix.

PowerEdgeBF-3 offeringDPU modeNIC modeAux powervSphere DSENotable KB
R660
16G
R760
16G
R760XA
16G (GPU)
XE9680
16G HGX H100/H200
XE9680L
16G liquid
R7725
17G (AMD)
R770
17G (Intel)
XE9780 / XE9785
17G HGX B300 — Dell AI Factory

✓ yes · ✗ no · ◐ conditional · ? unknown · — n/a. Click a cell for the evidence.

⚠ = not confirmed on a fetched primary source (hover for why). Facts as of DOCA 3.5.0 (Sep 2026). Selections are saved.

Which Dell PowerEdge platforms carry which silicon, and where the cell is a question mark because no Dell page states it. Click a cell for the document behind it.

6Writing the one page

The deliverable is not a meeting, it is one page the account manager can paste into the RFP response without editing. Five sections, in this order.

1. Restate the requirement. In your own words, name which of the four things each clause is asking for, and say so before you answer. “Multipath RDMA transport” is MRC; “UEC-certified switching” is a certification that does not exist; “adaptive routing across all available paths” is the switch feature, which is the one that is real and configurable.[11]

2. One row per claim, with tier and date. Every line carries its evidence tier and the date you fetched the page. MRC is standards-summary tier, IBTA Annex 21, 2026-07-31, plus a preprint and one vendor blog; its configuration surface is not-published tier.[3][1] Adaptive routing is documentation tier, NVUE, disabled by default.[11] The NIC-side Spectrum-X profile is documentation tier but Tech Preview, ConnectX-8 and BlueField-3 SuperNIC only, and needs a non-public package.[12]

3. What we can demonstrate in your lab this quarter. Two things, not three. Switch-side adaptive routing with its real scope attached — and the scope depends on the Cumulus release, so quote the one they run.[19][20] Say in the same paragraph what the demo does not prove: the show command prints configuration, not counters, and its documented sample output is a state toggle plus a link-utilization-threshold value under an applied column, with no AR packet, rerouted-flow or event count anywhere on the page.[21] And NCCL_IB_QPS_PER_CONNECTION for path diversity, with its ceiling stated in the same breath: each QP still rides one path at a time, so extra queue pairs buy more hash draws, not per-packet spreading of one connection, and NVIDIA publishes no recommended value.[13] Out-of-order data placement does not go in this section, however tempting it is. NVIDIA documents it under its InfiniBand chapter and says nothing about RoCE in either direction, so on an Ethernet fabric what you can demonstrate is reading the capability and its published scope — not showing it engage.[14][18] It belongs in the claim table as a documentation-tier row with that scope written out.

4. What we decline to assert, and why. Four sentences, each with its reason: that ConnectX-8 supports MRC as an NVIDIA statement; that MRC is an OCP standard or royalty-free; that adaptive routing requires a SuperNIC; and that any product is UEC-certified.[1][4][11][7]

5. What we need from you. The Cumulus release the switches will run — the NVUE syntax changed after 5.14[11] and the support statement changed again at 5.16, where “only supported on the NVIDIA Spectrum-X networking platform” appears for the first time[19][20]; the GPU model in the chassis, because it decides whether a SmartNIC or DPU is supported at all[16]; and their vendor’s UE compliance statement if the UEC clause stays in.[7]

Then date the page and set the re-verify interval. This lesson exists because the previous episode’s answer was re-checked one quarter later; the page you write today deserves the same treatment, and the date in its footer is what makes that possible.

Case closed — the paragraph, answered in one page

How it ended

The page goes out Monday. The transport clause gets a standards-summary row and a not-published row, side by side, dated.[3][1] The UEC clause comes back as a better question for their vendor: spec version, profile, checklist version.[7] The adaptive-routing clause becomes a lab booking, with the Cumulus release named as an open question.[11] The SE adds no row to his spreadsheet because there is nothing left to promise. The network lead writes one line in the notebook: tier, then date, then sentence. The operator labels the lab pair RE-VERIFY 2026-12-11.

What you say: “I will write down anything I can show you a command for, and nothing I can’t.”

The strongest thing you can say about a specification-stage technology is exactly what nobody publishes about it, and the date you checked.

Worked → faded → problem

The paragraph: “a multipath RDMA transport, UEC-certified switching, and adaptive routing across all available paths.”

Section 1 — What we read the requirement to mean. Three clauses, three different things. Clause 1 is the MRC transport. Clause 2 names a certification; we address it as a compliance question instead. Clause 3 is switch-side adaptive routing, which is a shipping, documented feature.

Section 2 — Claims and tiers, all pages fetched 2026-09-11.

Claim Tier Evidence
MRC is a real, standardised transport Standards summary IBTA Vol. 1 Rel. 2.1, Annex 21, published 2026-07-31
MRC was co-developed by six companies Vendor blog NVIDIA, with AMD, Broadcom, Intel, Microsoft and OpenAI
MRC extends RC for multipath, write and write-with-immediate at the transport level Preprint arXiv:2605.04333v1
MRC has a configuration surface Not published No vendor names an MRC QP-type constant, verbs symbol, mlxconfig parameter, env var, counter, firmware release note or perftest transport, and no vendor publishes a command that enables it
The furthest any vendor goes on an MRC API Vendor blog Broadcom: NIC vendors “can easily support the new MRC QP type” through libibverbs, “and manage falling back to RC when it is not available” — no symbol, header or command named
MRC runs on ConnectX-8 Preprint Named once, no ConnectX-8 measurement in the paper
The normative OCP MRC 1.0 text Not obtained The OCP URL returns HTTP 403 to every automated fetch
Switch-side adaptive routing Documentation NVUE set and unset commands, disabled by default, state toggles introduced in Cumulus Linux 5.1.0; support list is release-dependent (see section 3)
Out-of-order data placement Documentation ConnectX-5 and above, RC and XRC QPs, DC transport, documented under NVIDIA’s InfiniBand chapter; the opt-in flag and its capability query are in rdma-core; NVIDIA is silent about RoCE in both directions
Spectrum-X NIC-side AR and CC profile Documentation, Tech Preview ConnectX-8 and BlueField-3 SuperNIC only; requires a package obtained from an NVIDIA CPM
“UEC-certified” Not published UEC publishes self-attestation checklists; no certification, mark, plugfest or products list
1.6x over off-the-shelf Ethernet Marketing Product-page headline, no footnote, workload, cluster size or baseline

Section 3 — What we can demonstrate in your lab this quarter. Switch-side adaptive routing, configured and read back, with the scope stated for the release you run. On Cumulus 5.14 the documented scope is: Spectrum-4 at 400G and 200G, adaptive-routing-eligible RoCEv2 unicast, Layer 3 interfaces, next-hop router interfaces in the default VRF; not on 800G links, not on Layer 3 subinterfaces, SVIs, bonds or bond members; enabled on every port of the same ECMP route. On Cumulus 5.16 and later the same page adds a fourth constraint that is absent in 5.14 — “Adaptive routing is only supported on the NVIDIA Spectrum-X networking platform” — so the scope we quote here depends on the release the switches actually run, which is why it is also the first item in Section 5. Note that the show command prints configuration, not counters — it proves AR is configured, not that it is doing anything, and the sample output is release-dependent too (enable on on 5.14 and earlier, state enabled from 5.15). Alongside it: NCCL_IB_QPS_PER_CONNECTION, range 1 to 128, default 1, with its ceiling stated. Out-of-order data placement is not in this section: NVIDIA documents it under the InfiniBand chapter and says nothing about RoCE either way, so on your Ethernet fabric we can show you the capability and its scope, not the feature engaging. It is a row in the table above instead.

Section 4 — What we decline to assert. (a) That NVIDIA states ConnectX-8 supports MRC — NVIDIA’s page says “ConnectX SuperNICs” and names no generation. (b) That MRC is an OCP standard or royalty-free — the announcement says “released as an open specification through the Open Compute Project” and nothing more, and we could not read the document. (c) That adaptive routing requires a SuperNIC — it is enabled with switch-only commands. (d) That any product is UEC-certified — no such programme exists.

Section 5 — What we need from you. The Cumulus release the switches will run; the GPU model in each chassis; and, if the UEC clause stays, your vendor’s UE compliance statement naming spec version, profile and checklist version.

Footer: pages fetched 2026-09-11. Re-verify by 2026-12-11.

Lab

Read-only on the Dell lab pair and, if you have one, a Cumulus switch. No firmware, no mode, no QoS and no switch configuration is changed by any step; every command below only reads state. Capture the output of step 1 as the record you will attach to the RFP page.

  1. Pre-flight inventory on both hosts: ibv_devinfo, ibdev2netdev, ofed_info -s and show_gids. Record all four. These are the facts your page is allowed to assert about the hardware in front of you.
  2. Ask the card for a multipath knob: sudo mlxconfig -d <device> query and filter the output for ADAPTIVE, MULTIPATH, OOO, RETRANS, SLOW_RESTART, TX_WINDOW and MRC. Expected: no such parameter exists. This is the single most useful thing to have run before a customer asks you to enable MRC, because you can say you looked.
  3. Ask the benchmark tool which transports it knows: ib_write_bw --help and read the -c connection-type list. Expected: RC, UC, UD, XRC, DC and SRD. There is no MRC transport option. Note that SRD here is the AWS transport type, not anything to do with MRC.
  4. Check the counter surface for the same words: ethtool -S <interface> and the hw_counters directory for the device, filtered for mrc, multipath and spray. Expected: nothing matches. NVIDIA publishes no MRC telemetry counter and none appears on the card.
  5. Read the out-of-order capability rather than assuming it. Confirm the documented scope first — ConnectX-5 and above, RC and XRC QPs, DC transport, documented under the InfiniBand chapter — then note on your page that NVIDIA is silent about RoCE in both directions rather than writing either a yes or a no.
  6. If a Cumulus switch is available: nv show router adaptive-routing and nv show interface <interface-id> router adaptive-routing. Check the release first with nv show system, because the expected output depends on it: on 5.14 and earlier the configured state prints as enable on, and from 5.15 as state enabled. In both cases expect a link-utilization-threshold value under an applied column — and no packet count, rerouted-flow count or event count at all. If you see state enabled on a switch you assumed was 5.14, the switch is newer than your notes, not broken. Record that distinction — it is the honest answer to “show me the counter” for this feature.
  7. Diff every observation against the expectations above and against your own last run. Attach the dated diff to the RFP page as its evidence appendix, and put the next re-verify date in the footer.

Retrieval check

10 questions from memory. Answer before looking anything up; misses become flashcards.

Explain it to a Dell SE

Explain to the Dell SE, in four sentences, why you will not write "supports MRC" on the RFP response, and what you will write instead.

15 flashcards for this lesson — 0 in deck. Spaced review lives at /review.

Sources

Facts in this lesson were checked against Re-fetched 2026-09-11: the NVIDIA MRC blog, the IBTA Release 2.1 overview deck, arXiv 2605.04333v1 and 2605.21187, the OCP MRC 1.0 URL (HTTP 403), the Broadcom MRC article, Ultra Ethernet Specification v1.0.3 and the UEC compliance page, the NVIDIA Spectrum-X product page and the 2026-08-24 giga-scale blog, the UEC members page (logos render in JavaScript, no member name extractable), the NVUE adaptive-routing set/unset and show references, the Cumulus Linux 5.14 and 5.18 Equal Cost Multipath Load Sharing pages, the rdma-core v65.0 mlx5dv_create_qp(3) man page, the Tech Preview Spectrum-X NIC Configuration page, the NCCL 2.31.2 environment variables page, the DOCA 3.5.0 out-of-order data placement page, the Dell-hosted SN5600 datasheet and the Dell PowerEdge XE9680 Technical Guide. Dates are when each page was fetched.

  1. NVIDIA Spectrum-X — the Open, AI-Native Ethernet Fabric — Sets the Standard for Gigascale AI, Now With MRC (NVIDIA Blog) · fetched 2026-09-11
  2. Resilient AI Supercomputer Networking using MRC and SRv6 (arXiv:2605.04333v1) · fetched 2026-09-11
  3. What's new – Release 2.1, Vol. 1 and 2 — General Overview (IBTA slide deck) · fetched 2026-09-11
  4. OCP Multipath Reliable Connection (MRC) Specification 1.0 — NOT FETCHED (HTTP 403) · fetched 2026-09-11
  5. Enabling AI Networking @ Scale with Multi-path Reliable Connections (MRC) — Broadcom, 2026-05-06 · fetched 2026-09-11
  6. Ultra Ethernet Specification v1.0.3 (July 16, 2026) · fetched 2026-09-11
  7. UEC Compliance (Readme plus PHY-LL and Transport checklists) · fetched 2026-09-11
  8. Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules (2026-08-24) · fetched 2026-09-11
  9. High-speed Networking for Giga-Scale AI Factories (arXiv:2605.21187) · fetched 2026-09-11
  10. NVIDIA Spectrum-X Ethernet Platform for AI Networking (product page) · fetched 2026-09-11
  11. Adaptive Routing — NVUE 5.x Set and Unset Commands · fetched 2026-09-11
  12. [TECH PREVIEW] NVIDIA Spectrum-X NIC Configuration — Network Operator v25.10.0 · fetched 2026-09-11
  13. Environment Variables — NCCL 2.31.2 documentation · fetched 2026-09-11
  14. Out-of-order Data Placement — DOCA 3.5.0 · fetched 2026-09-11
  15. NVIDIA Spectrum SN5600 Series Switches (Datasheet, Dell-hosted, JUL25) · fetched 2026-09-11
  16. Dell PowerEdge XE9680 Technical Guide (Rev. A09, July 2025) · fetched 2026-09-11
  17. Members — Ultra Ethernet Consortium — MEMBER NAMES NOT EXTRACTABLE (logo block renders in JavaScript; static HTML carries only the tier headings) · fetched 2026-09-11
  18. mlx5dv_create_qp(3) — rdma-core v65.0 (MLX5DV_QP_CREATE_OOO_DP and mlx5dv_ooo_recv_wrs_caps) · fetched 2026-09-11
  19. Equal Cost Multipath Load Sharing — Cumulus Linux 5.18 (adaptive-routing support list and limitations) · fetched 2026-09-11
  20. Equal Cost Multipath Load Sharing — Cumulus Linux 5.14 (adaptive-routing support list and limitations) · fetched 2026-09-11
  21. Adaptive Routing — NVUE 5.x Show Commands (sample output) · fetched 2026-09-11