Skip to content

Who do I call: three support paths, one asset tag

S3·E5Dell says NVIDIA, NVIDIA says Dell · a bridge call at 02:10, day three of a flapping leaf, eleven days to the freeze

S3·E5Evaluate~20 minsources checked todayverified against Dell newsroom announcement of 2025-05-19 and the Dell-NVIDIA AI fabric blog re-checked 2026-09-09 for an explicit support RACI (none found); PowerScale H16346.8 ETC footnote and DGX SuperPOD B300 abstract, 2026-09-09 and 2026-09-07

Builds on: Reading a real Dell AI Factory BOM, PowerScale back-end networking and the version-pin trap

Before you read: what do you already know?

3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.

After this lesson you can

  • Distinguish the three published support paths for a switch in a Dell AI deployment by the evidence that identifies each.
  • Judge whether a given claim about Dell and NVIDIA support responsibility is published or inferred.
  • Route a real symptom to a front door and defend the routing in one sentence.
  • Assemble the evidence bundle that lets a case survive a handoff between vendors.

Episode 5 - Dell says NVIDIA, NVIDIA says Dell

The situation · a bridge call at 02:10, day three of a flapping leaf, eleven days to the freeze

Three days, one leaf dropping links intermittently, two vendors on the bridge, each explaining that the other one owns it. The operator who called it in has stopped describing the fault. The network lead says “show me the counter” and nobody produces one. Then somebody offers the sentence that makes it worse: “NVIDIA owns switch RMAs in Dell AI Factories.” Nobody can say where that came from, and by tomorrow it will be the case title.

You ask two questions instead. Whose asset tag is on the box, and which support contract number covers it. Both are facts the customer can look up while the call runs, and between them they select a front door. Here the answer is Dell: the leaves were bought as Dell PowerSwitch line items, and Dell’s own announcement names them - SN5600 and SN2201 Ethernet switches and Quantum-X800 InfiniBand switches - as “backed by Dell ProSupport and Deployment Services”.[1] The solution sheet says it commercially too: “ProSupport+ or ProSupport One”, “ProDeploy”, “ProConsult”.[3]

Entitlement exists so that somebody is contractually obliged, which is why it attaches to a transaction and a role rather than to a symptom. What nobody publishes is which company owns the defect once a case is open. The page people expect to answer it, Dell’s partnership blog, promises “expert assistance every step of the way” and delineates nothing.[2] The gap is real and knowable, which makes saying so stronger than filling it.

The front door follows the asset, never the symptom. Three paths, one asset tag. Learn which is which.

1Three front doors

There are three published support routes for a switch in a Dell AI deployment, and they do not overlap.

Dell-sold PowerEdge with Dell PowerSwitch. Dell’s May 2025 announcement names the hardware and the service in the same breath: Dell PowerSwitch SN5600 and SN2201 Ethernet switches and NVIDIA Quantum-X800 InfiniBand switches, “backed by Dell ProSupport and Deployment Services to provide expert guidance at every stage of AI deployment”, with the Ethernet switches available in 2H 2025.[1] The AI Factory solution sheet carries the same picture at the deal level, listing “ProSupport+ or ProSupport One”, “ProDeploy” and “ProConsult” as the services attached to a solution that includes the switches.[3] Dell’s own AI switch catalog sells the NVIDIA Spectrum SN-series, the NVIDIA Quantum InfiniBand line and Broadcom-based PowerSwitch Z-series from the same page, so “bought from Dell” covers more silicon than people expect.[6]

DGX SuperPOD. The DGX SuperPOD B300 reference architecture scopes itself to “Single-tenant, multi-user enterprise environments with NVIDIA-led installation and professional support through NVIDIA Technical Account Managers”.[4] That is a different company leading both the install and the support relationship, and Dell is not in the path.[4]

PowerScale back end on SN5600. The PowerScale document routes this case somewhere else again: the switch “is supported via the Dell Technologies ETC program, please consult Dell Technologies account team for more details”, with manual configuration and a pinned Cumulus Linux 5.9.1.[5] Same switch model as path one, different program, different people, different version.[5]

SymptomBlueField-3 on a Dell Pow…SymptomDell platform: POST, iDRA…Symptom"No Memory Found" POST e…SymptomvSphere Distributed Serv…SymptomCard disappeared after a…SymptomFirmware: Dell DUP vs NV…
Symptom

Dell platform: POST, iDRAC, vSphere, firmware

The Dell-platform branch. Walk it once and notice how much of it is establishing which artefact belongs to whom - lifecycle log, service tag, firmware source - before anything is fixed. That is the shape of a routing decision.

2What is published, and what is not

The honest map has a hole in it, and knowing precisely where the hole is makes you more useful, not less.

What is published: the front doors above, each with a source.[1][4][5] What is not published, on any Dell or NVIDIA page fetched for this course, is a responsibility matrix saying which company owns a Spectrum switch defect once a case is open in a Dell-sold AI Factory.[1][2][3] The Dell partnership blog is the page people expect to answer it, and it does not: it offers “expert assistance every step of the way” from “design consulting to post-deployment troubleshooting” and contains no delineation of who handles what.[2]

One item from the research notes has now moved from rumour to fact. The claim that Dell ProSupport backs both the Spectrum Ethernet and the Quantum-X800 InfiniBand switches previously existed only as a search summary; it is stated on Dell’s own announcement page, which names the switch models and the services together.[1] Treat that as verified. Treat the defect-level split as still unpublished, and say so in those words.[1][2]

That leaves two questions that always have answers, and they are the ones to ask first: whose asset tag is on the affected device, and which support contract number covers it.[3][5] Both are facts the customer can look up while you are on the call, and between them they select the front door without anyone having to interpret a partnership.

3The evidence bundle that survives a handoff

Whichever door you knock on, the case is only as good as what you attach. Build the same bundle every time.

Identity first: the service tag of the affected system, and the asset owner of the affected switch.[3] Then the adapter inventory - part numbers, PSIDs and firmware from mlxfwmanager --query - because the first question on any fabric escalation is which versions are actually installed.[7] Then the target: the validated-stack row the cluster is meant to match, which as of v2.3.1 in September 2026 is switch NOS 5.18.1, ConnectX-8 firmware 40.50.1002, BlueField-3 firmware 32.50.1002, DOCA-Host 3.5.0-082 and NCCL 2.30.7.[7] Then the role of the device and the document that pins it, which for a PowerScale back-end switch is a different pin entirely at Cumulus 5.9.1.[5] Then symptom, timestamps and what changed.

A bundle in that shape converts an argument into a diff, and it keeps its value when the case is handed from Dell to NVIDIA or from a partner to an account team - which is the whole point, given that nobody publishes who performs that handoff.[1][2]

Escalation: DPU disappears after BIOS update · decision 1/6A 0 · P 0 · S 0

Brief — Customer's R760 with a BlueField-3 B3220 (DPU mode) lost the DPU after a BIOS update. Host sees nothing. They are about to reflash the DPU firmware "to be safe".

Dell L2 support engineer, bridge call with the customer: We are about to reflash the DPU firmware from the host — go ahead?

Run the escalation scenario. Watch how much of the score comes from triage order and from refusing to blame before evidence exists - the two habits that decide whether a cross-vendor case closes or stalls.

4Judging a routing call

Bloom level for this lesson is evaluation, so the skill being trained is judging someone else’s answer, including your own from last week.

A good routing call has four properties. It names the path and the evidence that selected it - asset tag and contract, not symptom.[3][5] It separates what is published from what is inferred, out loud.[1][2] It contains no commitment on behalf of a company that has not published one.[2] And it names the next action with an owner, even when the answer is “this goes to the account team”.[5]

A bad routing call usually fails in one of three recognisable ways. It routes by symptom, so an InfiniBand link fault goes to “the InfiniBand people” regardless of who sold the switch.[4] It routes by relationship, so it goes to whoever answered last time. Or it invents a policy - “NVIDIA owns switch RMAs” - that no page states, which is the failure that survives longest because it sounds like inside knowledge.[1][2]

The tell for all three is the same: the sentence does not contain an asset tag, a contract number or a citation.

Eleven days to spare

How it ended

The case moves because it finally has a door and a bundle: service tag, the switch’s asset owner and role, the adapter inventory from mlxfwmanager --query, and the validated-stack row the cluster is meant to match.[3][7] The invented sentence about RMAs comes out of the case text and nothing replaces it. The leaf is swapped inside the entitlement that actually covers it, the cluster runs before the freeze, and row 41 of the promise spreadsheet finally closes.[1] The operator labels the surplus optics box “cages: 64” and goes home. The NVIDIA PM, asked what comes after this generation, says it is not announced.

Evaluating three routing decisions

Given. Three colleagues route the same symptom - a leaf switch dropping links intermittently - in three different clusters. Judge each.

Case A. Cluster: Dell-sold AI Factory, XE9680 nodes, SN5600 leaves, Dell asset tags, ProSupport+ contract on the solution.[3] Colleague routes to Dell ProSupport, attaches service tag, mlxfwmanager --query output, switch NOS version and the validated-stack row, and writes “front door is Dell ProSupport; these switches are backed by Dell ProSupport and Deployment Services”.[1][7] Verdict: sound. Path selected by asset and contract, claim carries a published source, evidence bundle complete.

Case B. Cluster: DGX SuperPOD, NVIDIA-led install, TAM assigned.[4] Colleague routes to Dell because the customer also buys PowerEdge servers from Dell. Verdict: wrong path, right instinct. The relationship is real but the architecture is scoped to NVIDIA-led installation and professional support through NVIDIA Technical Account Managers, so the TAM is the front door.[4] Fix: re-route, keep the evidence bundle, tell the Dell account team it exists so nobody is surprised.

Case C. Cluster: PowerScale back end on an SN5600 pair.[5] Colleague routes to Dell ProSupport as a standard switch case and adds “NVIDIA owns the Cumulus defect”. Verdict: two errors. The back-end SN5600 runs through the Dell Technologies ETC program with the account team, and it is pinned to Cumulus Linux 5.9.1 - so the first thing to check is whether the switch is even on the qualified version.[5] The ownership sentence is unpublished and must come out.[1][2]

Rewrite of case C in one sentence. “This is a PowerScale back-end switch supported through the Dell Technologies ETC program, so the case goes to the Dell account team; the switch reports NOS version X against a qualified pin of Cumulus Linux 5.9.1, and I have attached the service tag, the fabric listing and the version evidence.”[5]

Lab

Goal: assemble a real evidence bundle from the Dell lab so the format is muscle memory before a customer needs it. Read-only throughout - no firmware updates, no configuration changes, no case actually opened.

  1. Pre-flight identity. Capture the service tag and entitlement view.
    racadm getsysinfo
    racadm getsvctag
    Expected: a service tag matching the asset label, plus model and firmware summary. If racadm is unavailable, read the same fields from the iDRAC support page. This is line one of every case.
  2. Adapter inventory with versions.
    sudo mst start && sudo mst status -v
    sudo mlxfwmanager --query
    Expected: device list, part numbers, PSIDs and current firmware per adapter.[7] If not: mst status -v prints nothing when the mst driver is not loaded - run sudo mst start again and check lsmod | grep mst. If mlxfwmanager --query reports no devices, note it: an empty inventory is itself a case fact. Note any device where two update paths could apply - a Dell DUP and an NVIDIA bundle both targeting the same card is a finding in itself.
  3. Host-side version context.
    ofed_info -s 2>/dev/null || dpkg -l | grep -i doca | head
    uname -r
    Expected: the DOCA-Host or driver version and the kernel. If not: if both commands come back empty the host is on an inbox driver - record “inbox, version from modinfo mlx5_core” rather than leaving the line blank, because “no DOCA installed” is a comparison result, not a missing measurement. Compare against the validated-stack row for September 2026 and write the diff, even if the diff is empty.[7]
  4. Link and transceiver state for the affected port, if you are reproducing a link symptom.
    sudo ethtool ens1f0np0 | head -12
    sudo ethtool -S ens1f0np0 | grep -iE "err|drop|fcs" | head
    Expected: negotiated speed and any error counters. If not: ethtool reports “No such device” when the interface name differs - run ip -br link first and use the name it prints; if the counter grep returns nothing, say “no error counters incrementing at this time” with the timestamp rather than omitting the line. Counters with timestamps are worth more than adjectives in a case.
  5. Assemble the bundle as a single text file in this order: service tag, switch asset owner and role, adapter inventory, host versions, target validated-stack row, symptom with timestamps, what changed and when.[5][7] Expected: one file you could paste into a case with no editing. If not: if any of the seven items is missing, keep its heading and write what is missing and why - a bundle with a named gap survives a handoff, and a bundle with a silent gap gets the case sent back.
  6. Practise the sentence. In one sentence, state which of the three paths this lab host would belong to and the fact that selects it.[1][4][5] Expected: a sentence containing an asset identifier and a path, and no claim about who owns a defect.
  7. Deliverable: the evidence-bundle file and the one-sentence routing statement.

Retrieval check

10 questions from memory. Answer before looking anything up; misses become flashcards.

Explain it to a Dell SE

Explain to a Dell L2 engineer in four sentences why the same symptom on two different clusters goes to two different front doors.

14 flashcards for this lesson — 0 in deck. Spaced review lives at /review.

Sources

Facts in this lesson were checked against Dell newsroom announcement of 2025-05-19 and the Dell-NVIDIA AI fabric blog re-checked 2026-09-09 for an explicit support RACI (none found); PowerScale H16346.8 ETC footnote and DGX SuperPOD B300 abstract, 2026-09-09 and 2026-09-07. Dates are when each page was fetched.

  1. Dell Technologies Unveils Next Generation Enterprise AI Solutions with NVIDIA (newsroom announcement) · fetched 2026-09-09
  2. Dell-NVIDIA Partnership Powers High-Performance AI Fabric Solutions · fetched 2026-09-07
  3. Dell AI Factory with NVIDIA Solution ID 19845005.1 · fetched 2026-09-07
  4. Abstract - NVIDIA DGX SuperPOD DGX B300 Reference Architecture · fetched 2026-09-07
  5. Dell PowerScale: Ethernet Back-End Network Overview (H16346.8, March 2025) · fetched 2026-09-09
  6. AI Networking Switches | Dell USA · fetched 2026-09-07
  7. NVIDIA Spectrum-X Validated Solution Stack · fetched 2026-09-07

The same idea elsewhere

Other lessons that cover this ground, sometimes from another course's angle.