Who do I call: three support paths, one asset tag
S3·E5Dell says NVIDIA, NVIDIA says Dell · a bridge call at 02:10, day three of a flapping leaf, eleven days to the freeze
Builds on: Reading a real Dell AI Factory BOM, PowerScale back-end networking and the version-pin trap
Before you read: what do you already know?
3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.
After this lesson you can
- Distinguish the three published support paths for a switch in a Dell AI deployment by the evidence that identifies each.
- Judge whether a given claim about Dell and NVIDIA support responsibility is published or inferred.
- Route a real symptom to a front door and defend the routing in one sentence.
- Assemble the evidence bundle that lets a case survive a handoff between vendors.
Episode 5 - Dell says NVIDIA, NVIDIA says Dell
Three days, one leaf dropping links intermittently, two vendors on the bridge, each explaining that the other one owns it. The operator who called it in has stopped describing the fault. The network lead says “show me the counter” and nobody produces one. Then somebody offers the sentence that makes it worse: “NVIDIA owns switch RMAs in Dell AI Factories.” Nobody can say where that came from, and by tomorrow it will be the case title.
You ask two questions instead. Whose asset tag is on the box, and which support contract number covers it. Both are facts the customer can look up while the call runs, and between them they select a front door. Here the answer is Dell: the leaves were bought as Dell PowerSwitch line items, and Dell’s own announcement names them - SN5600 and SN2201 Ethernet switches and Quantum-X800 InfiniBand switches - as “backed by Dell ProSupport and Deployment Services”.[1] The solution sheet says it commercially too: “ProSupport+ or ProSupport One”, “ProDeploy”, “ProConsult”.[3]
Entitlement exists so that somebody is contractually obliged, which is why it attaches to a transaction and a role rather than to a symptom. What nobody publishes is which company owns the defect once a case is open. The page people expect to answer it, Dell’s partnership blog, promises “expert assistance every step of the way” and delineates nothing.[2] The gap is real and knowable, which makes saying so stronger than filling it.
The front door follows the asset, never the symptom. Three paths, one asset tag. Learn which is which.
1Three front doors
There are three published support routes for a switch in a Dell AI deployment, and they do not overlap.
Dell-sold PowerEdge with Dell PowerSwitch. Dell’s May 2025 announcement names the hardware and the service in the same breath: Dell PowerSwitch SN5600 and SN2201 Ethernet switches and NVIDIA Quantum-X800 InfiniBand switches, “backed by Dell ProSupport and Deployment Services to provide expert guidance at every stage of AI deployment”, with the Ethernet switches available in 2H 2025.[1] The AI Factory solution sheet carries the same picture at the deal level, listing “ProSupport+ or ProSupport One”, “ProDeploy” and “ProConsult” as the services attached to a solution that includes the switches.[3] Dell’s own AI switch catalog sells the NVIDIA Spectrum SN-series, the NVIDIA Quantum InfiniBand line and Broadcom-based PowerSwitch Z-series from the same page, so “bought from Dell” covers more silicon than people expect.[6]
DGX SuperPOD. The DGX SuperPOD B300 reference architecture scopes itself to “Single-tenant, multi-user enterprise environments with NVIDIA-led installation and professional support through NVIDIA Technical Account Managers”.[4] That is a different company leading both the install and the support relationship, and Dell is not in the path.[4]
PowerScale back end on SN5600. The PowerScale document routes this case somewhere else again: the switch “is supported via the Dell Technologies ETC program, please consult Dell Technologies account team for more details”, with manual configuration and a pinned Cumulus Linux 5.9.1.[5] Same switch model as path one, different program, different people, different version.[5]
Dell platform: POST, iDRAC, vSphere, firmware
2What is published, and what is not
The honest map has a hole in it, and knowing precisely where the hole is makes you more useful, not less.
What is published: the front doors above, each with a source.[1][4][5] What is not published, on any Dell or NVIDIA page fetched for this course, is a responsibility matrix saying which company owns a Spectrum switch defect once a case is open in a Dell-sold AI Factory.[1][2][3] The Dell partnership blog is the page people expect to answer it, and it does not: it offers “expert assistance every step of the way” from “design consulting to post-deployment troubleshooting” and contains no delineation of who handles what.[2]
One item from the research notes has now moved from rumour to fact. The claim that Dell ProSupport backs both the Spectrum Ethernet and the Quantum-X800 InfiniBand switches previously existed only as a search summary; it is stated on Dell’s own announcement page, which names the switch models and the services together.[1] Treat that as verified. Treat the defect-level split as still unpublished, and say so in those words.[1][2]
That leaves two questions that always have answers, and they are the ones to ask first: whose asset tag is on the affected device, and which support contract number covers it.[3][5] Both are facts the customer can look up while you are on the call, and between them they select the front door without anyone having to interpret a partnership.
3The evidence bundle that survives a handoff
Whichever door you knock on, the case is only as good as what you attach. Build the same bundle every time.
Identity first: the service tag of the affected system, and the asset owner of the affected switch.[3] Then the adapter inventory - part numbers, PSIDs and firmware from mlxfwmanager --query - because the first question on any fabric escalation is which versions are actually installed.[7] Then the target: the validated-stack row the cluster is meant to match, which as of v2.3.1 in September 2026 is switch NOS 5.18.1, ConnectX-8 firmware 40.50.1002, BlueField-3 firmware 32.50.1002, DOCA-Host 3.5.0-082 and NCCL 2.30.7.[7] Then the role of the device and the document that pins it, which for a PowerScale back-end switch is a different pin entirely at Cumulus 5.9.1.[5] Then symptom, timestamps and what changed.
A bundle in that shape converts an argument into a diff, and it keeps its value when the case is handed from Dell to NVIDIA or from a partner to an account team - which is the whole point, given that nobody publishes who performs that handoff.[1][2]
Brief — Customer's R760 with a BlueField-3 B3220 (DPU mode) lost the DPU after a BIOS update. Host sees nothing. They are about to reflash the DPU firmware "to be safe".
Dell L2 support engineer, bridge call with the customer: We are about to reflash the DPU firmware from the host — go ahead?
4Judging a routing call
Bloom level for this lesson is evaluation, so the skill being trained is judging someone else’s answer, including your own from last week.
A good routing call has four properties. It names the path and the evidence that selected it - asset tag and contract, not symptom.[3][5] It separates what is published from what is inferred, out loud.[1][2] It contains no commitment on behalf of a company that has not published one.[2] And it names the next action with an owner, even when the answer is “this goes to the account team”.[5]
A bad routing call usually fails in one of three recognisable ways. It routes by symptom, so an InfiniBand link fault goes to “the InfiniBand people” regardless of who sold the switch.[4] It routes by relationship, so it goes to whoever answered last time. Or it invents a policy - “NVIDIA owns switch RMAs” - that no page states, which is the failure that survives longest because it sounds like inside knowledge.[1][2]
The tell for all three is the same: the sentence does not contain an asset tag, a contract number or a citation.
Eleven days to spare
The case moves because it finally has a door and a bundle: service tag, the switch’s asset owner and role, the adapter inventory from mlxfwmanager --query, and the validated-stack row the cluster is meant to match.[3][7] The invented sentence about RMAs comes out of the case text and nothing replaces it. The leaf is swapped inside the entitlement that actually covers it, the cluster runs before the freeze, and row 41 of the promise spreadsheet finally closes.[1] The operator labels the surplus optics box “cages: 64” and goes home. The NVIDIA PM, asked what comes after this generation, says it is not announced.
Given. Three colleagues route the same symptom - a leaf switch dropping links intermittently - in three different clusters. Judge each.
Case A. Cluster: Dell-sold AI Factory, XE9680 nodes, SN5600 leaves, Dell asset tags, ProSupport+ contract on the solution.[3] Colleague routes to Dell ProSupport, attaches service tag, mlxfwmanager --query output, switch NOS version and the validated-stack row, and writes “front door is Dell ProSupport; these switches are backed by Dell ProSupport and Deployment Services”.[1][7] Verdict: sound. Path selected by asset and contract, claim carries a published source, evidence bundle complete.
Case B. Cluster: DGX SuperPOD, NVIDIA-led install, TAM assigned.[4] Colleague routes to Dell because the customer also buys PowerEdge servers from Dell. Verdict: wrong path, right instinct. The relationship is real but the architecture is scoped to NVIDIA-led installation and professional support through NVIDIA Technical Account Managers, so the TAM is the front door.[4] Fix: re-route, keep the evidence bundle, tell the Dell account team it exists so nobody is surprised.
Case C. Cluster: PowerScale back end on an SN5600 pair.[5] Colleague routes to Dell ProSupport as a standard switch case and adds “NVIDIA owns the Cumulus defect”. Verdict: two errors. The back-end SN5600 runs through the Dell Technologies ETC program with the account team, and it is pinned to Cumulus Linux 5.9.1 - so the first thing to check is whether the switch is even on the qualified version.[5] The ownership sentence is unpublished and must come out.[1][2]
Rewrite of case C in one sentence. “This is a PowerScale back-end switch supported through the Dell Technologies ETC program, so the case goes to the Dell account team; the switch reports NOS version X against a qualified pin of Cumulus Linux 5.9.1, and I have attached the service tag, the fabric listing and the version evidence.”[5]
Judge these two. Fill the blanks.
- Cluster: Dell-sold AI Factory with Quantum-X800 InfiniBand switches from Dell. Colleague routes to NVIDIA because “InfiniBand is NVIDIA’s”. Verdict: ____. Published sentence that decides it: ____.[1]
- Cluster: mixed - GPU fabric on SN5610 bought from Dell, PowerScale back end on a separate S5232F-ON pair. Symptom is on the storage pair. Front door: ____. Evidence that selects it: ____ and ____.[5][6]
- For each case, write the one sentence you would put at the top of the case, containing an asset identifier, the path and no unsourced commitment: ____.
- Which of the two cases involves a version pin, and which pin is it? ____.[5][7]
- Name one thing in your answers that you could not verify from a published page, and say how you would phrase it to the customer: ____.[2]
A Dell SE forwards a customer email: “Our AI cluster has been dropping links on one leaf for three days. Dell says talk to NVIDIA, NVIDIA says talk to Dell. Who actually owns this?” The cluster is a Dell-sold AI Factory: XE9780 nodes, SN5610 leaves purchased as Dell PowerSwitch, PowerScale F710 on a separate S5232F-ON back-end pair, ProSupport+ on the solution.
Write the reply. It must contain: (a) the front door with the published sentence that selects it; (b) the two identity facts you need from the customer before anything else; (c) the evidence bundle you are asking them to attach, item by item, with the command that produces each item you can name; (d) an explicit statement of what is not published about responsibility, phrased so it does not sound like an excuse; (e) the one question you are taking to the Dell account team and why it belongs there; and (f) a note on whether the PowerScale pair is in scope for this case and how you know.
Acceptance criteria: no sentence commits a company to a responsibility no page states; every routing claim carries a citation; the reply distinguishes the front door from defect ownership; and it names at least one thing you will confirm rather than assert.[1][2][5]
Lab
Goal: assemble a real evidence bundle from the Dell lab so the format is muscle memory before a customer needs it. Read-only throughout - no firmware updates, no configuration changes, no case actually opened.
- Pre-flight identity. Capture the service tag and entitlement view.
Expected: a service tag matching the asset label, plus model and firmware summary. Ifracadm getsysinfo racadm getsvctagracadmis unavailable, read the same fields from the iDRAC support page. This is line one of every case. - Adapter inventory with versions.
Expected: device list, part numbers, PSIDs and current firmware per adapter.[7] If not:sudo mst start && sudo mst status -v sudo mlxfwmanager --querymst status -vprints nothing when the mst driver is not loaded - runsudo mst startagain and checklsmod | grep mst. Ifmlxfwmanager --queryreports no devices, note it: an empty inventory is itself a case fact. Note any device where two update paths could apply - a Dell DUP and an NVIDIA bundle both targeting the same card is a finding in itself. - Host-side version context.
Expected: the DOCA-Host or driver version and the kernel. If not: if both commands come back empty the host is on an inbox driver - record “inbox, version fromofed_info -s 2>/dev/null || dpkg -l | grep -i doca | head uname -rmodinfo mlx5_core” rather than leaving the line blank, because “no DOCA installed” is a comparison result, not a missing measurement. Compare against the validated-stack row for September 2026 and write the diff, even if the diff is empty.[7] - Link and transceiver state for the affected port, if you are reproducing a link symptom.
Expected: negotiated speed and any error counters. If not:sudo ethtool ens1f0np0 | head -12 sudo ethtool -S ens1f0np0 | grep -iE "err|drop|fcs" | headethtoolreports “No such device” when the interface name differs - runip -br linkfirst and use the name it prints; if the counter grep returns nothing, say “no error counters incrementing at this time” with the timestamp rather than omitting the line. Counters with timestamps are worth more than adjectives in a case. - Assemble the bundle as a single text file in this order: service tag, switch asset owner and role, adapter inventory, host versions, target validated-stack row, symptom with timestamps, what changed and when.[5][7] Expected: one file you could paste into a case with no editing. If not: if any of the seven items is missing, keep its heading and write what is missing and why - a bundle with a named gap survives a handoff, and a bundle with a silent gap gets the case sent back.
- Practise the sentence. In one sentence, state which of the three paths this lab host would belong to and the fact that selects it.[1][4][5] Expected: a sentence containing an asset identifier and a path, and no claim about who owns a defect.
- Deliverable: the evidence-bundle file and the one-sentence routing statement.
Goal: build the routing tree and the two emails, then audit your own sourcing.
- Draw the three-path decision tree. Each branch is a question, not a category: whose asset tag is on the device; is this architecture NVIDIA-led with a TAM; does the device serve a PowerScale back end.[1][4][5] Expected: three leaves, each carrying a verbatim source sentence.
- Mark every claim in your tree that came from a search summary rather than a page you or the research fetched.[1][2] Expected: the Dell ProSupport statement covering Quantum-X800 now sits on the fetched announcement page and can be marked verified, while the defect-level responsibility split remains unpublished and must be marked as such.
- Write two customer emails from the same symptom - a leaf switch dropping links - routed to two different front doors, each naming the evidence attached and each stating what you are not claiming.[1][5] Expected: two emails whose first paragraphs differ only in the asset facts, not in the symptom description.
- Take the three bad routing patterns from segment 4 and write one line for each showing how it would look in a real case title.[2][4] Expected: three titles that a reviewer can reject on sight.
- Run the escalation scenario above and record your score by dimension, then write two sentences on where you lost points and what evidence would have prevented it.[7] Expected: an honest note, not a high score.
- Deliverable: the tree with sources, the two emails, and the sourcing audit marking verified against unpublished.
Retrieval check
10 questions from memory. Answer before looking anything up; misses become flashcards.
Explain it to a Dell SE
Explain to a Dell L2 engineer in four sentences why the same symptom on two different clusters goes to two different front doors.
Sources
Facts in this lesson were checked against Dell newsroom announcement of 2025-05-19 and the Dell-NVIDIA AI fabric blog re-checked 2026-09-09 for an explicit support RACI (none found); PowerScale H16346.8 ETC footnote and DGX SuperPOD B300 abstract, 2026-09-09 and 2026-09-07. Dates are when each page was fetched.
- Dell Technologies Unveils Next Generation Enterprise AI Solutions with NVIDIA (newsroom announcement) · fetched 2026-09-09
- Dell-NVIDIA Partnership Powers High-Performance AI Fabric Solutions · fetched 2026-09-07
- Dell AI Factory with NVIDIA Solution ID 19845005.1 · fetched 2026-09-07
- Abstract - NVIDIA DGX SuperPOD DGX B300 Reference Architecture · fetched 2026-09-07
- Dell PowerScale: Ethernet Back-End Network Overview (H16346.8, March 2025) · fetched 2026-09-09
- AI Networking Switches | Dell USA · fetched 2026-09-07
- NVIDIA Spectrum-X Validated Solution Stack · fetched 2026-09-07
The same idea elsewhere
Other lessons that cover this ground, sometimes from another course's angle.
- Scenario: escalationsDOCA course · Same ground: Evidence bundle, support and case
- Dell PowerSwitch SN-series: who supports whatSpectrum-X course · Same ground: support, triage and route
- Who actually ships SRv6: the Dell and NVIDIA support matrixSRv6 course · Same ground: evidence, support and published