Scenario: a qualification request
S9·E1One line of email and a five-week window · Hotel lobby in Round Rock, 06:40, before the Dell Platform Group meeting
Builds on: Writing a DOCA validation plan for an OEM server, The validated stack and version pinning, Dell-specific bring-up: iDRAC, BIOS, CPLD, aux power, DSE
Before you read: what do you already know?
3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.
After this lesson you can
- Scope a Dell qualification request for a BlueField-3 SKU into the questions that must be answered before any test runs.
- Produce a qualification plan with a test matrix, pass/fail evidence, risks mapped to Dell KB numbers, and named deliverables.
- Separate what NVIDIA publishes, what Dell publishes, and what neither confirms, and write each in the right voice.
- Adapt the plan to a different SKU and PowerEdge model without re-deriving it.
Episode 1 — One line of email and a five-week window
You are on a lobby couch with a laptop. The Dell SE lands beside you with a cup that went cold in the car, opens the spreadsheet where he logs every promise he makes, and reads out the overnight request: qualify the B3140H SuperNIC in the PowerEdge XE9780 for Spectrum-X. Behind that line is a medical-imaging AI provider whose proof-of-concept window closes in five weeks.
Nothing in it is safe. The B3140H is a SuperNIC SKU: one QSFP112 port up to 400GbE, 8 Arm cores, 16 GB DDR5, 75 W, shipping in NIC mode with the Arm inactive.[3][4] The XE9780 is Dell’s HGX B300 chassis, whose reference architecture puts eight ConnectX-8 SuperNICs east-west at 1:1 with the GPUs and one BlueField-3 DPU north-south in ECPF mode.[1] The validated stack agrees: B300 rows carry ConnectX-8 firmware, BlueField-3 sits in the H200 rows.[2] The requested card is not the card the architecture expects.
The validated stack exists for exactly this morning: switch OS, firmware, BFB, DOCA-Host and the collective library are developed against each other; DOCA-OFED is Level 1 against firmware and BFB, every other profile and every DOCA service Level 2, in lock-step.[2][5] Someone wrote down what was tested together, so a lab result means something.
Procurement joins to ask their only question, lead time, and this once they are right. NVIDIA’s own description of the role makes the first deliverable scoping, not testing.[20] A request is not a plan until every version in it is pinned to one validated-stack row. You open a list of questions, not a test plan.
1Read the request before you answer it
The request arrives as one line: Dell Platform Group asks you to qualify the B3140H SuperNIC in the PowerEdge XE9780 for Spectrum-X. NVIDIA’s own description of the FAE role is design-in, on-site bring-up, issue replication, and owning the product through its lifecycle, so the first deliverable is scoping, not testing.[20]
Three published facts reshape the ask before you write anything. First, the B3140H is a SuperNIC SKU: HHHL, one QSFP112 port up to 400GbE, 8 Arm cores, 16 GB DDR5, 75 W slot power, and it ships in NIC mode with the Arm cores inactive.[3][4] Second, the XE9780 is Dell’s HGX B300 chassis, and the HGX B300 reference architecture puts eight ConnectX-8 SuperNICs east-west at a 1:1 GPU ratio and one BlueField-3 DPU (B3240 recommended) north-south in ECPF/DPU mode.[1] Third, the Spectrum-X validated stack tracks that pairing: the H200 rows list BlueField-3 firmware, the B300 and GB300 rows list ConnectX-8 firmware.[2]
So a BlueField-3 SuperNIC east-west in an HGX B300 box is not the RA default. It was the Hopper-era design: NVIDIA’s November 2023 announcement used PowerEdge XE9680 servers with BlueField-3 SuperNICs and Spectrum-4 switches for the Israel-1 reference architecture; the release names the switch family, not a model number.[17] The ask can still be legitimate, for example a 400G design point, a north-south NIC-mode card, or a Dell-internal test, but you must learn which before you pick a firmware row.
There is also a Dell-side gap. Dell’s Partner DPU 352-BBFH, a dual-port 400GbE BlueField-3, lists the XE9780 as a compatible system.[10] The XE9780 manual’s DPU specifications page rendered without a DPU table when fetched, so no Dell page you can cite confirms the B3140H in that chassis.[13] That is not a blocker; it is the first question.
Brief — Email, Tuesday: "We want to offer the BlueField-3 B3140H SuperNIC in the XE9780 for Spectrum-X customers. Can NVIDIA qualify it? I need a plan by Friday."
Dell platform PM (17G AI servers): So — can you get it qualified? What do you need from us to start?
2The questions to ask first
Write the questions down and send them; do not ask them one at a time on calls. Each row below is a question, the reason it matters, and where the answer comes from.
| Axis | Question to Dell PG | Why it matters |
|---|---|---|
| Sourcing | Dell catalogue SuperNIC line, Dell Partner DPU, or NVIDIA-channel card? | Dell supports the Partner DPU when used with a Dell system; DSE is excluded for NVIDIA Channel cards.[10][9] |
| Role and fabric | East-west 400G on Spectrum-4 SN5600, or north-south? Cumulus or SONiC on the leaf? | Adaptive routing, DDP and congestion control are co-designed between the Spectrum-X switch and the SuperNIC.[16] |
| Mode | Stays in NIC mode or will be flipped to DPU mode? | Flip is INTERNAL_CPU_OFFLOAD_ENGINE=0 plus a power cycle; a B3140H in DPU mode is in scope for KB 000227031.[4][6] |
| BlueField firmware | Target 32.50.1002? Anything below 32.46.3048 in the lab? | 32.50.1002 is the DOCA 3.5.0 table and v2.3.1 value; below 32.46.3048 is the PCIe training bug of KB 000379421.[14][2][7] |
| DOCA-Host | 3.5.0-082 per v2.3.1? Which profile? | doca-ofed is Level 1 against firmware and BFB; doca-all, doca-networking, doca-roce and doca-host-basic are Level 2.[2][5][15] |
| BFB | Only if DPU mode: BFB 3.5.0 on Ubuntu 24.04 64k (bundle default), aligned with the firmware? If HBN is in scope, note its release notes name Ubuntu 22.04 and confirm the variant with NVIDIA. | Dell’s KB itself says keep the BFB aligned with the updated firmware.[14][7] |
| Slot and power | Which riser and slot (R7725: RC 2/3/4/5/8 slots 2 and 7; R770: RC 1/2 slots 31 and 36 or RC 6-2/11-2 slots 7 and 2)? Any 150 W DPU in the same design? | B3140H is 75 W with no connector; B3220/B3240 are 150 W and a missing 8-pin PCIe cable halts the card at boot.[3][11][12][24] |
| OS or hypervisor | Ubuntu 22.04/24.04, RHEL 9.x/10.x, or ESXi? | General Support lists the Linux matrix; ESXi brings the DSE rules: Dell-SKU only, reinstall on switch, vSphere 8.0 U3b or later.[14][9][19] |
| Server firmware | BIOS, iDRAC, CPLD levels in the lab? | The No Memory Found fault is version-specific: CPLD 1.1.5/1.1.7, iDRAC 7.10.50.00, BF firmware 32.40.1000.[6] |
| Update discipline | Will the lab power-cycle after every DPU firmware update? | KB 000300192: a warm reboot is not enough.[8] |
3The plan: matrix, evidence, timeline, deliverables
The matrix has the same dimensions every time: mode (NIC, DPU, zero-trust), DOCA-Host profile, firmware and BFB (the validated-stack row plus N-1 to exercise the Level-1 claim), server firmware (BIOS, iDRAC, CPLD), host OS, and riser/slot/power.[4][5][2] For a B3140H staying in NIC mode the BFB column collapses to “not applicable” and the DSE column disappears unless ESXi is in play.[4][19]
Each test case names its pass/fail evidence:
- Enumeration:
lspci -d 15b3:shows the PF;mlxconfig -d /dev/mst/<device> q INTERNAL_CPU_OFFLOAD_ENGINEreports the intended mode; fail is a Lifecycle Log “Device not detected” with PR8.[4][7] - Boot synchronisation: 20 or more cold, warm and graceful cycles; fail is “No Memory Found” or a fatal PCIe error in the LC log.[6][8]
- Provisioning (DPU mode only):
bfb-installreachesINFO[MISC]: DPU is readyin/dev/rshim0/miscandssh ubuntu@192.168.100.2overtmfifo_net0works.[21] - Datapath (DPU mode only): OVS-DOCA offload confirmed with
ovs-appctl dpctl/offload-stats-show, plus one deliberately unsupported case to confirm software fallback rather than a hang.[22] - Performance: DOCA Bench and the customer’s RDMA or NCCL tests recorded per mode and profile, with DTS (1.26.5 in v2.3.1) exporting to Prometheus as the evidence store.[28][29][2]
Regression triggers re-run the enumeration, boot-sync and offload subsets: any BIOS, iDRAC or CPLD change; any BlueField firmware change; any BFB or DOCA change, with October GA as the Level-2 boundary; any service version; any riser or power change.[5][6]
A timeline template you can propose: week 0 scoping call plus written questions; weeks 1 to 2 inventory, enumeration and alignment to one validated-stack row; weeks 3 to 4 boot-sync soak and datapath; week 5 performance and report. These durations are your proposal to Dell PG. Nothing in NVIDIA or Dell documentation fixes them, and you should say so.
Deliverables: the scoping email; the matrix; a results table with one row per server, riser, SKU, mode, profile, firmware, BFB and OS including log excerpts (LC log, dmesg for mlx5_core, rshim console); an issue list mapped to KB numbers; the validated-stack row each result maps to; and an open escalation log.[2][7]
4Risks, and how to write them down honestly
Every risk gets a source and a voice. Use “Dell’s KB says” for Dell facts, “NVIDIA has stated” for NVIDIA facts, and “not confirmed by a page I can cite” for everything else.
- PCIe training failure below firmware 32.46.3048: Dell’s KB path is power cycle, update firmware plus matching BFB, then a service request for a replacement card.[7]
- No Memory Found after BIOS updates: Dell’s KB says wait about a minute for the automatic reboot; the fix lands in a CPLD release; the B3140H is affected only if moved to DPU mode.[6]
- Firmware update without a power cycle leaves a fatal PCIe error in the LC log; plan the cycle into every update step.[8]
- Level-2 lock-step: a DOCA service or a non-OFED profile on a July release may not work with the next October firmware.[5]
- Aux power: NVIDIA’s specifications page requires the supplementary 8-pin ATX connector on the four DPU SKUs (B3240, B3220, B3210, B3210E) and lists 75 W minimum through the slot for every SKU; SuperNICs such as the B3140H have no connector. Dell lists B3220 and B3240 at 150 W, and a field report shows a B3210 halting with “ATX power not detected” until a PCIe 8-pin cable was fitted. Dell cable-kit part numbers per riser are not published on the pages fetched.[3][11][24]
- Availability: the R7725 manual marks the B3140H “may not be available at launch”; the XE9780 DPU page shows no table. Route dates to Dell product management.[11][13]
- Version horizon: DOCA 3.5.0 shipped 2 September 2026 and NVIDIA has announced DOCA 3.6.0 as the last release for ConnectX-4 Lx and ConnectX-5; a mixed fleet needs a branch plan.[26]
- Guarantees: every NVIDIA document states it is not a commitment to deliver any functionality and that specifications may change without notice; quote that rather than promising.[25]
5Worked scoping email, faded, then a new SKU and server
The email below is the artefact Dell PG expects within a day of the request. It states what is known, asks what is not, and proposes the plan.
Subject: Scoping the B3140H qualification in XE9780 for Spectrum-X
Thanks for the request. Before I commit a plan I need four answers, because they decide which validated-stack row we test against.
- Role. Is the B3140H the east-west SuperNIC at 400G, or a north-south card? The HGX B300 RA pairs eight ConnectX-8 SuperNICs east-west with one BlueField-3 DPU north-south, and the validated stack tracks ConnectX-8 firmware for B300 rows and BlueField-3 firmware for H200 rows.[1][2]
- Sourcing. Dell catalogue SuperNIC line, Partner DPU, or NVIDIA-channel? Dell’s Partner DPU 352-BBFH lists XE9780 as compatible, but that is a dual-port 400GbE part, and the XE9780 DPU specification page shows no table, so please confirm the B3140H line.[10][13]
- Mode. I assume NIC mode (factory default, Arm inactive). If DPU mode is required we add BFB 3.5.0 provisioning and KB 000227031 to scope.[4][6]
- Platform levels. BIOS, iDRAC, CPLD in the lab, host OS, and whether any ESXi hosts are involved (DSE is a Dell-SKU-only story per KB 000225111).[6][9]
Proposed baseline: Spectrum-X validated stack v2.3.1: BlueField-3 FW 32.50.1002, DOCA-Host 3.5.0-082, Cumulus 5.18.1, NetQ 5.1.0, DTS 1.26.5.[2] Host profile doca-ofed for the Level-1 matrix and doca-all for the Level-2 matrix.[5][15]
Plan: week 0 questions; weeks 1 to 2 inventory, enumeration, firmware alignment (any card below 32.46.3048 is updated first per KB 000379421); weeks 3 to 4 twenty-plus power cycles plus datapath; week 5 performance and report.[7] Durations are my proposal.
Risks: KB 000379421 (PCIe training), KB 000227031 (No Memory Found if DPU mode), KB 000300192 (power cycle after firmware), Level-2 October boundary.[7][6][8][5]
Deliverables: matrix, results table with log excerpts, KB-mapped issue list, validated-stack mapping, escalation log.
Subject: Scoping the ____ qualification in ____ for ____
- Role. East-west or north-south? The HGX B300 RA pairs eight ____ SuperNICs east-west with one ____ north-south; validated-stack ____ rows track ConnectX-8 firmware.[1][2]
- Sourcing. Dell catalogue, Partner DPU, or ____? Dell supports the Partner DPU “when used with a Dell system”; DSE is excluded for ____ cards.[10][9]
- Mode. Assume ____ mode (factory default for SuperNICs). If DPU mode: add BFB ____ and KB ____ to scope.[4][6]
- Platform levels. BIOS, ____, ____ levels; host OS; any ESXi.
Baseline: v2.3.1: BF-3 FW ____, DOCA-Host ____, Cumulus ____.[2] Profiles: ____ for Level 1, ____ for Level 2.[5]
Plan: ____ weeks; cards below FW ____ updated first per KB ____.[7]
Risks: KB ____ / ____ / ____ and the ____ GA Level-2 boundary.
Dell PG now asks you to qualify a B3240 DPU in the PowerEdge R7725 for HBN (BGP/EVPN on the DPU) with Dell PowerSwitch leaves.
Write the scoping email. Acceptance criteria:
- Mode is DPU (HBN requires DPU or zero-trust mode) and the memory floor is checked: HBN 3.5.0 does not support 8 GB SKUs; the B3240 has 32 GB.[18][3]
- Slot and power are named: R7725 riser configs RC 2/3/4/5/8, slots 2 and 7, 150 W, and the 8-pin PCIe auxiliary cable is listed as a bring-up prerequisite with the halt symptom.[11][24]
- The leaf requirement is stated: a BGP/EVPN-capable ToR; Dell SN-series run Cumulus Linux or SONiC.[18]
- HBN is a DOCA service, therefore Level 2: firmware, BFB 3.5.0 and HBN 3.5.0 move together within one cycle.[5][18]
- KB 000227031 and KB 000300192 appear as risks because the card runs in DPU mode and firmware will be updated.[6][8]
- At least one sentence uses “not confirmed by a page I can cite” (for example Dell aux-cable part numbers).
Nine o'clock, and a shorter meeting than expected
One page of questions, one page of matrix: sourcing, fabric role, mode, firmware, DOCA-Host profile, riser and power, and the two things you will not state. No availability date for a SuperNIC the R7725 manual itself marks as possibly unavailable at launch, and no claim that the XE9780 supports it, because the Dell page that would say so rendered without a DPU table.[11][13] The platform lead asks for one number and gets it: nothing below firmware 32.46.3048 goes into the lab.[7]
What you say: “Answer these four and I will pin the whole matrix to one validated-stack row by Friday.”[2]
The first cards ship to the customer’s cage that week. On day nine, at 01:40, one of them answers rshim and nothing else.
Lab
Goal: a five-minute narrated demo for a Dell audience showing a DPU-to-NIC mode flip and an OVS-DOCA offload check on the lab BlueField-3. Mode changes are firmware-parameter writes that require a power cycle; treat this as a maintenance window with a written rollback.[4][27]
- Pre-flight inventory (read-only, record everything on camera):
sudo mst start && sudo mst status -v;lspci -d 15b3: -nn;sudo mlxconfig -d /dev/mst/mt41692_pciconf0 q INTERNAL_CPU_OFFLOAD_ENGINE;sudo flint -d /dev/mst/mt41692_pciconf0 q;cat /dev/rshim0/misc. Expected:ENABLED(0)(DPU mode), firmware version andINFO[MISC]: DPU is ready.[4][21] If not: stop; the card is not in a known state and there is no rollback anchor.[27] - Offload check in DPU mode (on the Arm, read-only):
sudo ovs-appctl dpctl/offload-stats-showandsudo ovs-appctl coverage/show. Narrate what an offloaded flow count looks like for Dell.[22] If not (no OVS-DOCA bridge): say so on camera and skip to step 3; do not configure OVS during a mode-flip window. - Flip to NIC mode (mutating):
sudo mlxconfig -d /dev/mst/mt41692_pciconf0 s INTERNAL_CPU_OFFLOAD_ENGINE=1, then a full server power cycle, not a warm reboot. Rollback:s INTERNAL_CPU_OFFLOAD_ENGINE=0plus another power cycle.[4][8] - Verify NIC mode:
sudo mlxconfig -d /dev/mst/mt41692_pciconf0 q INTERNAL_CPU_OFFLOAD_ENGINEshowsDISABLED(1); the Arm console over rshim is silent because the Arm cores are inactive; the host sees a ConnectX-class adapter.[4][23] If not: the power cycle did not happen; repeat it before touching anything else. - Roll back to DPU mode with the inverse write and a power cycle, then confirm
INFO[MISC]: DPU is readyreturns in/dev/rshim0/misc.[4][21] If not: check/dev/rshim0/consolefor boot output and follow the escalation lesson. - Close the recording by mapping what you showed to the Dell KBs that make the power-cycle rule non-negotiable.[8]
- Open the FaeScenarioSim above (it starts on the qualification scenario) and play all six decisions, choosing the questions you would actually send. Expected: each choice is scored 0 to 2 on accuracy, positioning and safety, the customer replies in character, and the debrief shows the expert path. If not: replay and compare your picks with the ten axes in Segment 2.
- Write the scoping email for the B3140H request using the worked template, in your own words, with every version tied to the v2.3.1 row and every risk tied to a KB number.[2][7] Expected: 300 to 500 words, four numbered questions, a baseline, a plan, risks and deliverables. If not: you skipped an axis; check sourcing and mode first.
- Build the matrix as a CSV with columns server, riser, slot, SKU, mode, profile, firmware, BFB, OS, test, evidence, result. Fill the B3140H NIC-mode rows only. Expected: BFB and DSE columns read “n/a” for NIC mode.[4] If not: you are testing a DPU-mode feature on a NIC-mode card.
- Self-check: every unknown is phrased as a question; every Dell statement says “Dell’s KB says”; nothing promises a date. Expected: zero sentences you could not attribute. If not: rewrite the sentence or delete it.
Retrieval check
9 questions from memory. Answer before looking anything up; misses become flashcards.
Explain it to a Dell SE
Explain to a Dell platform engineer, in five sentences, why your first reply to a qualification request is a list of questions rather than a test plan.
Sources
Facts in this lesson were checked against DOCA 3.5.0 docs, Spectrum-X validated stack v2.3.1, Dell KBs 000227031/000379421/000300192/000225111, Dell R7725/R770/XE9780 manuals, 2026-09-06. Dates are when each page was fetched.
- NVIDIA HGX AI Factory Enterprise Reference Architecture: Components · fetched 2026-09-06
- NVIDIA Spectrum-X Validated Solution Stack · fetched 2026-09-06 · DOCA 3.5.0
- NVIDIA BlueField-3 Networking Platform User Guide: Specifications · fetched 2026-09-06
- BlueField Modes of Operation · fetched 2026-09-06 · DOCA 3.5.0
- DOCA Dependency Compatibility Policy · fetched 2026-09-06 · DOCA 3.5.0
- Dell KB 000227031: No Memory Found event on NVIDIA BlueField-3 enabled PowerEdge servers · fetched 2026-09-06
- Dell KB 000379421: BlueField-3 DPU PCIe Initialization Failure · fetched 2026-09-06
- Dell KB 000300192: XE9680L power cycle required after updating DPU firmware · fetched 2026-09-06
- Dell KB 000225111: NVIDIA Channel DPU VMware DSE Support · fetched 2026-09-06
- Dell Partner DPU, NVIDIA BlueField-3 Dual Port 400GbE (352-BBFH) · fetched 2026-09-06
- Dell PowerEdge R7725 Installation and Service Manual: DPU specifications · fetched 2026-09-06
- Dell PowerEdge R770 Installation and Service Manual: DPU specifications · fetched 2026-09-06
- Dell PowerEdge XE9780 Installation and Service Manual: DPU specifications (page rendered navigation only) · fetched 2026-09-06
- DOCA General Support (OS matrix and firmware table) · fetched 2026-09-06 · DOCA 3.5.0
- DOCA Profiles · fetched 2026-09-06 · DOCA 3.5.0
- NVIDIA Spectrum-X Network Platform Architecture whitepaper (July 2024, third-party mirror) · fetched 2026-09-06
- NVIDIA Spectrum-X available from Dell, HPE, Lenovo (Nov 2023) · fetched 2026-09-06
- HBN Service Release Notes (HBN 3.5.0) · fetched 2026-09-06 · DOCA 3.5.0
- Broadcom KB 379391: Configuring BlueField-3 into DPU mode for vSphere · fetched 2026-09-06
- NVIDIA: Senior Field Application Engineer, Networking System Hardware Design (posting) · fetched 2026-09-06
- BF-Bundle Installation and Upgrade · fetched 2026-09-06 · DOCA 3.5.0
- OVS-DOCA Hardware Acceleration · fetched 2026-09-06 · DOCA 3.5.0
- HowTo Configure NVIDIA BlueField-3 to NIC Mode on VMware vSphere 8.0 · fetched 2026-09-06
- NVIDIA Developer Forums: BlueField-3 (DPU mode) stuck in FW pre-initializing, no host netdevs (aux power) · fetched 2026-09-06
- DOCA Backward Compatibility Policy (PDF) · fetched 2026-09-06 · DOCA 3.5.0
- DOCA 3.5.0 Changes and New Features · fetched 2026-09-06 · DOCA 3.5.0
- NVIDIA/skills: doca-hardware-safety · fetched 2026-09-06
- DOCA Bench · fetched 2026-09-06 · DOCA 3.5.0
- DOCA Telemetry Service Guide · fetched 2026-09-06 · DOCA 3.5.0
The same idea elsewhere
Other lessons that cover this ground, sometimes from another course's angle.
- Writing a DOCA validation plan for an OEM serverElsewhere in this course · Same ground: vmware, plan and adapt
- PowerEdge support matrixElsewhere in this course · Same ground: vmware, consignment and partner
- Why infrastructure moved onto the NICElsewhere in this course · Same ground: sku, DOCA-Host and SuperNIC