Skip to content

HGX reference architecture roles

S7·E2The ninth card on the quote · Dell design review, Round Rock, Thursday morning

S7·E2Analyze~25 minsources checked todayverified against HGX AI Factory Enterprise RA components page, BlueField Modes of Operation (DOCA 3.5.0), BlueField-3 User Guide specifications, ConnectX-8 User Manual, Spectrum-X Validated Solution Stack v2.3.1, Dell press releases May 2025 and Mar 2026 — fetched 2026-09-06

Builds on: Spectrum-X architecture, Modes of operation: DPU, NIC, Zero-Trust

Before you read: what do you already know?

3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.

After this lesson you can

  • Assign each network element of an HGX B300 node (eight ConnectX-8, one BlueField-3, the switches) to its reference-architecture role and traffic class.
  • Explain why the north-south BlueField-3 runs in DPU (ECPF) mode while the east-west ConnectX-8 adapters carry GPU traffic with no DPU in the path.
  • Justify the 1:1 GPU-to-NIC ratio and the 2 x 400G breakout guidance from the published node specification.
  • Identify which Dell chassis and PowerSwitch parts map onto the reference architecture and where the public documents stop assigning roles.
  • Analyze a proposed bill of materials for role mismatches such as a SuperNIC SKU ordered for the north-south DPU slot.

Episode 2 — The ninth card on the quote

The situation · Dell design review, Round Rock, Thursday morning

Thirty-two HGX B300 nodes are on the projector and the room wants to sign today, because storm season does not move for a purchase order. You are at the end of the table with the Dell SE, the carrier’s infrastructure architect, the network lead and his notebook — three pages fuller than in cage 4 — and a procurement person who has asked twice about lead time and nothing else. Line nine of the quote reads one BlueField-3 B3140H per server. The SE explains it cheerfully: the customer asked for a SuperNIC, this is the cheaper BlueField-3, and it is already row 41 on the Promises tab. Everyone looks at you: saving, or return?

An AI node has two network jobs and they pull in opposite directions. GPU-to-GPU collectives want the shortest path from GPU memory to wire with nothing to manage, so the reference architecture solders eight ConnectX-8 SuperNICs onto the HGX B300 baseboard at up to 800 Gb/s per GPU, one per GPU.[1] Everything else — the customer network, storage, the infrastructure services the operations team wants isolated from the host — rides one BlueField-3 DPU per server in Embedded Function (ECPF) or DPU mode, where the Arm cores own the NIC and run those services.[1][2] Two products, two modes; a SuperNIC SKU ships in NIC mode with no Arm OS to host anything.[2]

In an AI node the mode is the role: NIC mode moves GPU traffic, DPU mode runs infrastructure. You ask the room to walk the node with you, card by card.

1Node anatomy: eight GPUs, eight SuperNICs, one DPU

The HGX AI Factory reference architecture describes a node with “8 Blackwell GPUs per node with 288GB HBM3e per GPU (2.30TB total per node)”.[1] For east-west traffic it specifies “Eight NVIDIA ConnectX-8 SuperNICs on the HGX baseboard for the East/West” at “up to 800 Gb/s (2 x 400Gb/s Ethernet) per GPU”, and states that the ConnectX-8 “is integrated onto the NVIDIA HGX B300 baseboard, maintaining a 1:1 GPU-to-NIC ratio”.[1] The fabric guidance is a minimum of “8x 400 Gb/s Ethernet NICs” per node and a recommended “16x 400 Gb/s Ethernet NICs using breakout”, meaning each 800G ConnectX-8 is split into two 400G fabric ports.[1]

The ninth adapter is different in kind: “One NVIDIA BlueField-3 DPU per server” for the north-south side, installed as a PCIe card rather than on the baseboard.[1] Some summaries of this RA describe the node as 2 CPUs, 8 GPUs, 9 NICs at 800 Gb/s; the components page fetched for this lesson does not use that label, so treat it as a mnemonic, not a specification name.

The host side matters for qualification: the RA asks for a minimum of “48 physical CPU cores per socket” (56 recommended), “Minimum of 2TB system memory”, “Minimum of 500GB/s memory bandwidth”, “Minimum 2 TB NVMe drive per CPU socket” for training, a “1 TB NVMe boot drive”, TPM 2.0 and GPU management over “SMBPBI over SMBus (OOB)”.[1] The control plane is “up to eight control plane nodes”, for example “two for Base Command Manager, two for Slurm head nodes, and three for Kubernetes control plane nodes”.[1]

2The north-south BlueField-3: DPU mode, converged traffic, own management plane

The RA’s north-south section is titled converged Ethernet networking and defines the traffic as “North/South (Customer and Storage Network) traffic: This involves traffic between NVIDIA HGX systems and any external resources”.[1] “NVIDIA recommends using the NVIDIA BlueField-3 B3240 DPU with NVIDIA HGX B300 platforms”; the B3220 is listed as the alternative.[1] The B3240 is the dual-port QSFP112 DPU with up to 400GbE per port, 16 Arm cores and 32GB DDR5; the B3220 is the dual-port 200GbE DPU.[3]

The operating mode is stated: “Embedded Function (ECPF) or DPU mode, which is commonly used as a default”.[1] In DPU mode “The Arm cores of BlueField are active, and the embedded Arm system runs services that manage the NIC resources and data path”, and “The NIC resources and functionality are owned and controlled by the embedded Arm subsystem”.[2] That is the mode in which the DOCA services of modules 4 and 5 (OVS-DOCA, HBN, storage emulation, telemetry) exist; a card in NIC mode has no Arm OS to run them.[2]

Power is a hard prerequisite. The RA warns that “Some BlueField-3 configurations for north-south connectivity can draw more than 75 W and therefore require both standard PCIe slot power and an additional PCIe power connector”.[1] The BlueField-3 hardware guide states that DPU SKUs need the supplementary 8-pin ATX power connector while SuperNIC SKUs run on 75 W slot power.[4] A field thread shows the failure mode: a DPU-mode BlueField-3 stuck in firmware pre-initializing with no host netdevs until the aux cable was attached.[14]

Management of the DPU itself rides separately: every BlueField-3 has a 1GbE RJ45 out-of-band port and an integrated BMC, and the host reaches the Arm side over RShim, “the SoC management interface in the BlueField SoC”.[4][15] So the node has three planes: GPU east-west on ConnectX-8, converged customer and storage north-south through the DPU, and DPU management out-of-band.

HGX B300 node (e.g. Dell PowerEdge XE9780)N-S onlyGPU0B300GPU1B300GPU2B300GPU3B300GPU4B300GPU5B300GPU6B300GPU7B300CX-8800GCX-8800GCX-8800GCX-8800GCX-8800GCX-8800GCX-8800GCX-8800GBlueField-3 DPU (B3240)DPU/ECPF mode · N-S · OVS-DOCA/HBN/DTSE-W leaves: SN5600Spectrum-X · RoCE · adaptive routingN-S leafstorage · mgmt · tenant

HGX B300 node in a Spectrum-X AI factory

Reference architecture per node: 8 Blackwell Ultra GPUs, "Eight NVIDIA ConnectX-8 SuperNICs per NVIDIA HGX B300 baseboard" at a "1:1 GPU-to-NIC ratio" for east-west, and "One NVIDIA BlueField-3 DPU per server" in "ECPF or DPU mode" for north-south.

Why it matters. Dell's HGX B300 chassis are PowerEdge XE9780/XE9785; Dell lists its Partner DPU (BlueField-3 dual-port 400GbE) as compatible with XE9780. Spectrum-X validated stack v2.3.1 (Sep 2026): Cumulus 5.18.1, BF-3 FW 32.50.1002, DOCA-Host 3.5.0-082, NCCL 2.30.7.

FAE note. Click a GPU/NIC pair, the DPU, or either fabric. The key question customers ask: why is the DPU not in the GPU traffic path?

Source: docs.nvidia.com

AI-factory view: eight ConnectX-8 ports go straight to the east-west leaf; the single BlueField-3 in DPU mode fronts the north-south customer and storage network and hosts the offloaded services.

3Why the DPU is not in the east-west path

Three facts rule the DPU out of the GPU fabric. First, bandwidth: the fastest BlueField-3 DPU, the B3240, offers up to 400GbE per port, while the RA gives every GPU its own 800 Gb/s ConnectX-8.[1][3] Both attach over the same link class in this design — the RA specifies “Eight Gen5 x16 links and one Gen4 x2 link per NVIDIA HGX B300 baseboard. One Gen5 x16 link per DPU, SuperNIC or adapter” — so the difference is the port speed the adapter drives, not the slot it sits in.[1] The ConnectX-8 silicon itself supports “PCIe Gen6 (64GT/s) through x16 edge connector”, which matters for future platforms, not for this baseboard.[5] One 400G card cannot serve eight GPUs at 800 Gb/s each.

Second, role: the Spectrum-X endpoint functions the fabric relies on (Direct Data Placement, telemetry-based congestion control, per-packet spraying eligibility) are SuperNIC functions, and the whitepaper describes the BlueField-3 SuperNIC as natively running in NIC mode with queue-pair state in on-card memory.[6] A card doing infrastructure work in DPU mode is a different product with a different job.

Third, lifecycle: a BlueField in DPU mode carries a full software bundle. DOCA’s own definition is that the BF-Bundle is “the software package installed on the BlueField Arm cores”, including the Arm OS, the runtime and the platform firmware.[12] A ConnectX-8 has only NIC firmware (the 40.x line, 40.50.1002 in DOCA 3.5.0) and host drivers to maintain.[13] Putting a DPU on every GPU port would multiply the provisioning surface by nine for no bandwidth gain.

The historical contrast makes the split obvious. In November 2023 the Israel-1 design on PowerEdge XE9680 used BlueField-3 SuperNICs for east-west with SN5600 switches, because 400G was the Hopper-era per-GPU speed.[7] Blackwell doubles the per-GPU port to 800G, which only ConnectX-8 provides, and the BlueField-3 moves to the north-south slot where its Arm cores do useful work.[1][5]

4Switches and the Dell mapping

The SN5600 is the Spectrum-4 switch: 51.2 Tb/s, “up to 64 ports of 800GbE or 128 ports of 400GbE within a single 2U switch”, and it was the switch in the Dell XE9680 Israel-1 design.[6][7] The validated stack pins Cumulus for Spectrum-4 at 5.18.1 in its September 2026 row for GB300, B300 and H200 alike, which tells you the SN5600 class is still the validated Spectrum-X leaf and spine for HGX B300.[11]

Dell’s own announcements name the switch pair: “Dell PowerSwitch SN5600, SN2201 Ethernet” for Spectrum-X in 2H 2025, and in March 2026 PowerSwitch SN5610 and SN2201 “supporting Cumulus Linux and SONiC”.[8][9] Be precise about what is not published: the RA components page fetched for this lesson names no switch model and assigns no role to an SN2201. Its usual placement as the low-speed management switch is a reasonable reading of a 1G-class part beside a 51.2T fabric switch, but confirm it against the customer’s specific RA or Dell design document rather than assert it.

The Dell chassis mapping is public. PowerEdge XE9780 and XE9785 are Dell’s HGX B300 air-cooled servers, with XE9780L and XE9785L as the liquid-cooled variants.[8] The Dell “Partner DPU, NVIDIA BlueField-3 Dual Port 400GbE” (Dell part 352-BBFH) lists PowerEdge XE8712 and XE9780 as compatible systems, which is consistent with the RA’s B3240 north-south role.[10] The validated stack’s B300 table pins ConnectX-8 firmware and a B300 FW/SW package but has no BlueField-3 row, while the H200 table pins BlueField FW Bundle 3.5.0 and BlueField-3 firmware 32.50.1002; lesson 7.3 shows how to combine them for an XE9780 with a north-south DPU.[11]

Map a BOM to RA roles: XE9780 + SN5600 + BlueField-3 + ConnectX-8

Customer BOM: 32 x PowerEdge XE9780 (HGX B300), each with 8 x ConnectX-8 on the baseboard and 1 x BlueField-3 dual-port 400GbE (Dell 352-BBFH); 8 x SN5600; 2 x SN2201.

  1. East-west count. 8 GPUs per node and 8 ConnectX-8 per node gives 1:1 as the RA requires; 32 nodes x 8 x 2 (breakout) = 512 fabric ports at 400G, or 256 at 800G. Role: east-west SuperNIC, firmware family 40.x, no BFB.[1][13]
  2. North-south count and SKU. One BlueField-3 per node as required. 352-BBFH is a dual-port 400GbE BlueField-3, which matches the B3240 class rather than a SuperNIC; Dell lists the XE9780 as compatible. Role: north-south DPU, provisioned with a BF-Bundle.[1][10][12]
  3. Power. A dual-port 400GbE DPU SKU needs the supplementary 8-pin PCIe power connector; add the cable to the checklist for the XE9780 slot.[4][1]
  4. Mode. DPU (ECPF) mode is the RA default for this card; confirm at bring-up with mlxconfig and do not leave a card that shipped in NIC mode unexamined.[1][2]
  5. Switches. SN5600 runs Cumulus 5.18.1 in the September 2026 validated stack; role leaf and spine. SN2201: role not assigned by the RA page; write to-be-confirmed and ask for the customer’s design document.[11][9]

Output: a five-row role table with part, role, mode, firmware family, provisioning path, and one open item (SN2201 placement).

Episode 2 closes — line nine, corrected before the order ships

How it ended

The quote leaves the room with a B3240 in the north-south slot, the auxiliary 8-pin PCIe power cable on the configuration, and the eight baseboard ConnectX-8 untouched because they were never a choice.[1][4] One open item stays on the page: the RA assigns no role to the SN2201, so its placement is marked to-be-confirmed. What you tell the architect: “The card on line nine is a SuperNIC — it ships in NIC mode with eight Arm cores and no bundle. That slot needs a DPU SKU in DPU mode, with the power cable.”[3] Then procurement, occasionally right, asks the lead time on a cable nobody costed. It lands the pod in the last change window before the storm-season freeze — which is when the architect asks which versions go on it.

Lab

Read-only. Goal: decide whether the lab BlueField-3 could fill the RA’s north-south role and record the evidence.

  1. sudo mst start && sudo mst status -v. Record the device path and PCI address. If not: lspci | grep -i mellanox.
  2. sudo mlxfwmanager --query. Record Part Number, Description and PSID. Match the part to the spec page families: DPU SKUs (B3240, B3220, B3210, B3210E: 16 Arm cores, 32GB, 8-pin aux power) or SuperNIC SKUs (B3140H, B3140L, B3220L, B3210L: 8 cores, 16GB, 75 W).[3][4]
  3. sudo mlxconfig -d /dev/mst/<dev> q INTERNAL_CPU_OFFLOAD_ENGINE. ENABLED(0) is DPU mode (the RA’s north-south default); DISABLED(1) is NIC mode. Record; do not change.[2][1]
  4. ls -l /dev/rshim* on the host. A present /dev/rshim0 shows the host-side management channel to the Arm side exists; if absent, note whether the BMC owns RShim or the driver is not loaded, and stop there (no changes).[15]
  5. Visually confirm whether an 8-pin PCIe power cable is attached to the card. Write the verdict: DPU SKU + DPU mode + aux power present means the card could take the north-south role; anything else, state which condition fails.[1][4]

Retrieval check

10 questions from memory. Answer before looking anything up; misses become flashcards.

Explain it to a Dell SE

Explain to a Dell SE, in four sentences, why an HGX B300 node has nine NVIDIA network cards and why only one of them is a DPU.

12 flashcards for this lesson — 0 in deck. Spaced review lives at /review.

Sources

Facts in this lesson were checked against HGX AI Factory Enterprise RA components page, BlueField Modes of Operation (DOCA 3.5.0), BlueField-3 User Guide specifications, ConnectX-8 User Manual, Spectrum-X Validated Solution Stack v2.3.1, Dell press releases May 2025 and Mar 2026 — fetched 2026-09-06. Dates are when each page was fetched.

  1. NVIDIA HGX AI Factory Enterprise Reference Architecture — Components · fetched 2026-09-06
  2. BlueField Modes of Operation · fetched 2026-09-06 · DOCA 3.5.0
  3. NVIDIA BlueField-3 Networking Platform User Guide — Specifications · fetched 2026-09-06
  4. NVIDIA BlueField-3 Networking Platform User Guide — Introduction · fetched 2026-09-06
  5. NVIDIA ConnectX-8 SuperNIC User Manual — Introduction · fetched 2026-09-06
  6. NVIDIA Spectrum-X Network Platform Architecture whitepaper (July 2024; third-party mirror of the NVIDIA PDF) · fetched 2026-09-06
  7. NVIDIA Spectrum-X available from Dell, HPE, Lenovo (Nov 20 2023) · fetched 2026-09-06
  8. Dell Technologies Unveils Next Generation Enterprise AI Solutions with NVIDIA (May 19 2025) · fetched 2026-09-06
  9. Dell AI Factory with NVIDIA Delivers Proven Path to Enterprise AI ROI (Mar 16 2026) · fetched 2026-09-06
  10. Dell Partner DPU, NVIDIA BlueField-3 Dual Port 400GbE (352-BBFH) · fetched 2026-09-06
  11. NVIDIA Spectrum-X Validated Solution Stack · fetched 2026-09-06 · DOCA 3.5.0
  12. DOCA Overview · fetched 2026-09-06 · DOCA 3.5.0
  13. DOCA General Support (firmware table) · fetched 2026-09-06 · DOCA 3.5.0
  14. Forum: BlueField-3 (DPU mode) stuck in FW pre-initializing — missing 8-pin PCIe aux power · fetched 2026-09-06
  15. BlueField Platform Software Troubleshooting Guide: SoC Management Interface (RShim) · fetched 2026-09-06

The same idea elsewhere

Other lessons that cover this ground, sometimes from another course's angle.