Skip to content

Reading node nomenclature: 2-8-9-400

S1·E2Nine adapters and a cable nobody ordered · Speakerphone in the Dell lab, Tuesday, three days before the BOM freezes

S1·E2Analyze~20 minsources checked todayverified against Enterprise RA program whitepaper key building blocks, HGX AI Factory components and Appendix B, NVL72 AI Factory components and physical topologies, all re-fetched 2026-09-07

Builds on: Which reference architecture governs this deal

Before you read: what do you already know?

3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.

After this lesson you can

  • Decompose an Enterprise RA node string into CPU count, GPU count, NIC count and per-GPU east-west speed.
  • Split the NIC digit into east-west SuperNICs and north-south DPUs for a given platform.
  • Produce the node string for a machine from its adapter inventory and defend each digit with a page.
  • State which NVIDIA-Certified Storage level a given node configuration requires.

Episode 2 — Nine adapters and a cable nobody ordered

The situation · Speakerphone in the Dell lab, Tuesday, three days before the BOM freezes

Procurement joins the call for eleven minutes and spends all of them on one line of the quote. She reads it out - “two eight nine four hundred” - then asks why eight GPUs need nine network cards, and whether the ninth is something she can cut. The SE says it is the standard config. Row 42 of his spreadsheet says “node line, Tuesday”.

She is right to ask. The Enterprise RA program whitepaper names node configurations with four numbers in a fixed order - CPUs, GPUs, NICs and east-west traffic per GPU - and the HGX row is 2-8-9-400.[1] Appendix B of the HGX AI Factory RA lists the hardware behind that nine: eight on-baseboard ConnectX-8 SuperNICs in a dual-port 400G configuration, plus one BlueField-3 B3240 DPU with 2x400G ports.[2] Eight plus one. The ninth adapter is the entire north-south side of the node.

The convention exists because clusters are quoted, certified and cabled per node type, and the industry needed one token that fixes socket count, GPU count, adapter count and fabric speed at once. It travels beyond NVIDIA: Cisco titles its own AI POD design document after the NVIDIA 2-8-9-400 Enterprise Reference Architecture.[8]

Two more things are missing from her line. The storage partner has to be NVIDIA-Certified at the Enterprise level, the level tied to 2-8-9-400.[1] And the BOM needs an 8-pin ATX 12V PCIe cable, because B3240 configurations can draw more than 75 W.[3]

Say the string out loud as a sentence before you agree to anything on that line. Start by decomposing the four numbers.

1Four numbers, one sentence

The Enterprise RA program whitepaper names node configurations with four numbers in a fixed order: CPUs, GPUs, NICs, bandwidth.[1] It spells the pattern out in prose for the smallest one: 2 CPUs balanced with up to 4 GPUs plus 3 NICs with east-west traffic of 200 GbE per GPU.[1]

The whitepaper defines four configurations, each with its own node-count range:[1]

String Shape Node range
PCIe-Optimized 2-4-3-200 2 CPU, 4 GPU, 3 NIC, 200 GbE per GPU 8 to 32 nodes
PCIe-Optimized 2-8-5-200 2 CPU, 8 GPU, 5 NIC, 200 GbE per GPU 4 to 32 nodes
HGX 2-8-9-400 2 CPU, 8 GPU, 9 NIC, 400 GbE per GPU 4 to 32 nodes
Scale-out Grace 2-2-3-400 2 Grace CPU, 2 GPU, 3 NIC, 400 GbE per GPU 4 to 32 nodes

Two properties of the string are worth naming before anything else. First, the bandwidth digit is per GPU, not per node - it is the number the fabric sees on each GPU’s east-west path.[1] Second, the NIC digit counts every network adapter in the node, which is why it is always larger than the GPU count on the HGX and PCIe-optimized rows.

The nomenclature is public shorthand rather than NVIDIA-internal jargon: Cisco titles its own AI POD design document as an implementation of the NVIDIA 2-8-9-400 Enterprise Reference Architecture.[8] If a competitor is using the string in a document title, a Dell FAE should be able to decompose it on a whiteboard.

Which RA governs this deal?

Four customer facts → one named document, its scope sentence and its SU arithmetic.

GPU count256 GPUsband 4 / 9
Platform
Who installs
Rack power
Enterprise Reference ArchitecturesDGX SuperPOD Reference ArchitecturesProgram Whitepaper32–256 GPUs · defines the node names4/4HGX AI Factory (B300)32–128 nodes · 256–1,024 GPUs4/4NVL72 AI Factory (GB300)up to 8 SUs · 18 trays per SU2/4HGX AI Factory (H100/H200/B200)the Dell installed base3/4SuperPOD B300 — Spectrum-4 / DC busbarSU = 64 nodes = 512 GPUs0/4SuperPOD B300 — Quantum-X800 / ACSU = 72 nodes = 576 GPUs1/4SuperPOD GB200 NVL72SU = 8 NVL72 racks · NDR-era0/4
Shortlist:
4/4 customer factsLast updated May 18, 2026

NVIDIA HGX AI Factory — Enterprise Reference Architecture (B300)

Scope, in the document's own words

ranging from 32 nodes (256 GPUs) to 128 nodes (1,024 GPUs)

HGX AI Factory — index
Scalable Unit

No SU in this RA — it sizes in nodes. One node = "Eight NVIDIA B300 GPUs on an HGX B300 baseboard", so nodes = GPUs / 8.

  • Nodes at 8 GPUs each: 256 ÷ 8 = 32 nodes derived
  • SUs: This RA does not define a Scalable Unit — do not invent one
Rack power — Not published in this RA. It publishes node minimums instead: "Minimum of 48 physical CPU cores per socket", "Minimum of 2TB system memory", "Minimum of 500GB/s memory bandwidth".
Support path — Dell-built HGX in a PowerEdge chassis: Dell-branded, Dell-installed, Dell-supported. No NVIDIA TAM is implied by this document.
Node nomenclature — HGX 2-8-9-400. The "9" is 8 east-west SuperNICs + 1 north-south DPU: the appendix states "8 on baseboard NVIDIA ConnectX-8 SuperNIC with a dual-port 400G configuration" and "1 NVIDIA BlueField-3 B3240 DPU with 2x400G ports". The appendix itself never prints the string 2-8-9-400 — that definition lives in the program whitepaper. HGX AI Factory — Appendix B node configurations
Scale-up (NVLink, inside the accelerator) — HGX B300 baseboard: "Total Aggregate Bandwidth 14.4TB/s", "GPU-to-GPU Bandwidth 1800GB/s" — one NVLink domain of 8 GPUs.
Scale-out (the fabric, outside) — "800 Gb/s (2 x 400Gb/s Ethernet) per GPU" through eight ConnectX-8 SuperNICs — a 1:1 GPU-to-SuperNIC ratio.
⚠ Derived, not published: 1,800 GB/s scale-up = 14,400 Gb/s, against 800 Gb/s scale-out per GPU on B300 — about 18x. This ratio is arithmetic over the NVLink page and the HGX B300 components page; neither prints it.
What this RA does not publish
  • No BGP/EVPN configuration, no IP addressing plan, no MTU.
  • No RoCE QoS values — no DSCP, PFC or ECN settings.
  • No adaptive-routing or congestion-control parameters. "VLAN isolation would be used to provide logical separation of the networks above over the single physical fabric" is as far as it goes.
  • The certified-storage page names no protocols (NFS over RDMA, GPUDirect Storage), no tiers and no partner names.

FAE angle: this is the Dell PowerEdge case, and the plane decision is the BOM. Dual plane splits 800 Gb/s into "2x400 Gb/s interfaces", each landing on "a different leaf switch" in "an independent fabric that scales to 1024 interfaces of 400 Gb/s". Single plane is one 400 Gb/s link per GPU and the RA calls it a "50%" bandwidth reduction. Dual plane is recommended for 64+ node deployments.

Select the HGX B300 card and read its node nomenclature line, then switch the platform toggle to the H100/H200/B200 row and to GB300 NVL72 to see the same four-number idea applied to different adapter layouts.

2Splitting the NIC digit

The digit that carries the most information is the third one, because it is a sum of two different jobs.

Appendix B of the HGX AI Factory RA lists the B300 node hardware directly. Eight NVIDIA B300 SXM GPUs on an NVIDIA B300 baseboard, with GPU memory of up to 2304 GB. Eight on-baseboard NVIDIA ConnectX-8 SuperNICs with a dual-port 400G configuration. One NVIDIA BlueField-3 B3240 DPU with 2x400G ports and a 1 Gb RJ45 management port.[2] Eight plus one is nine.

The components page states the same split in fabric terms: 8x ConnectX-8 SuperNICs on the HGX baseboard delivering 800 Gb/s (2 x 400 Gb/s Ethernet) per GPU.[3] Read that carefully - the adapter is an 800 Gb/s device, and the string says 400, because the RA breaks each 800 Gb/s port into two 400 Gb/s interfaces that land on two different leaf switches.[3] The string carries the number the leaf port sees.

The ninth adapter is a different animal in the same box. The recommended north-south device is the NVIDIA BlueField-3 B3240 P-Series FHHL DPU at 400 GbE, in a single-slot or dual-slot form factor, and it requires 8-pin ATX 12V PCIe auxiliary power because configurations can draw more than 75 W.[3] That cable is a BOM line and it is the most common thing missing from a first delivery.

Here is the discipline point of this lesson. Appendix B never prints the string 2-8-9-400 next to the node it describes; the string is defined only in the program whitepaper.[2][1] Quote the appendix for the hardware and the whitepaper for the name, and never claim one page says both.

3The same idea on a rack-scale tray

The GB300 NVL72 is not a node in the HGX sense, so the arithmetic moves from the chassis to the tray.

One NVL72 rack carries 72 Blackwell Ultra GPUs and 36 Grace CPUs across 18 compute trays.[11] That is 4 GPUs and 2 CPUs per tray. Each compute tray carries 4 ConnectX-8 Host Channel Adapters delivering 800 Gb/s, plus one dual-port NVIDIA BlueField-3 B3240 DPU with an aggregate bandwidth of approximately 480 Gb/s.[4] Four plus one is five adapters.

So the tray reads as 2 CPUs, 4 GPUs, 5 NICs at 800 Gb/s per GPU. The NVL72 AI Factory overview page names the tray directly: the NVIDIA Enterprise RA using 2-4-5-800 (dual plane) node architecture with NVIDIA GB300 NVL72.[11] Note where it lives: the components and physical-topologies pages describe the same tray without printing the string, exactly as Appendix B does for 2-8-9-400.[4][5]

The GB300 fabric side is published clearly and is worth keeping with the tray count. Dual plane breaks the 800 Gb/s interface into 2x400 Gb/s interfaces; every such interface connects to a different leaf switch, and every such leaf switch is part of an independent fabric that scales to 1024 interfaces of 400 Gb/s.[5] Tracking of each plane, load balancing and failure handling is handled by the ConnectX-8 SuperNIC on the hardware level.[5]

4Storage certification rides on the string

The node string is not only a naming convention; it is the key into the storage program. NVIDIA-Certified Storage has two levels, and the whitepaper ties each level to specific node configurations: Foundation certifies storage partners for the 2-4-3-200 and 2-8-5-200 reference configurations, and Enterprise certifies storage partners for 2-8-9-400.[1]

That has a direct consequence in a Dell deal. If the customer is buying HGX nodes, a storage vendor certified only at Foundation level has not been certified against the configuration being sold. The string is therefore the shortest way to ask a storage partner the right question: not “are you NVIDIA certified” but “are you certified at Enterprise level for 2-8-9-400”.

The same logic runs backwards into the installed base. Dell’s published AI Factory solution lists NVIDIA Bluefield-3 (1x3220/8x3140H) per node against two PowerEdge XE9680 with 16 H200 SXM GPUs.[7] That is one B3220 DPU north-south and eight B3140H SuperNICs east-west - the same eight-plus-one shape as 2-8-9-400, one generation earlier.[7] On that generation the RA connects each compute and management node with two 200 GbE ports to two separate switches, so the bandwidth digit is 200 rather than 400.[6]

From an inventory to a node string

Request. “Here is a spec for a node we are being quoted. Two sockets. Eight GPUs on one baseboard. Eight ConnectX-8 adapters described as dual-port 400G on the baseboard, plus one BlueField-3 B3240 with two 400G ports and a 1 Gb RJ45. What is this in NVIDIA’s nomenclature, and what does it commit us to?”

1. CPUs. Two populated sockets, so the first digit is 2.[1]

2. GPUs. Eight SXM GPUs on the baseboard, so the second digit is 8. Appendix B describes exactly this node as eight NVIDIA B300 SXM GPUs on an NVIDIA B300 baseboard with GPU memory of up to 2304 GB.[2]

3. NICs. Count every network adapter, then classify. Eight ConnectX-8 SuperNICs face the compute fabric, one per GPU; one BlueField-3 B3240 DPU faces north-south. Eight plus one gives 9.[2]

4. Bandwidth. Do not read this off the adapter. The RA states the east-west number as 800 Gb/s (2 x 400 Gb/s Ethernet) per GPU, and the nomenclature carries the per-interface figure, so the digit is 400.[3]

Result: HGX 2-8-9-400, defined in the program whitepaper.[1]

What it commits them to. Storage partners must be NVIDIA-Certified at the Enterprise level, which is the level that certifies 2-8-9-400; Foundation-level certification covers only 2-4-3-200 and 2-8-5-200.[1] And the BOM needs an 8-pin ATX 12V PCIe cable for the DPU, because B3240 configurations can draw more than 75 W.[3]

One caveat to say out loud. The appendix that lists this hardware never prints the string itself - the name comes from the whitepaper.[2][1]

Eleven minutes well spent

How it ended

You read the string back as a sentence - two sockets, eight GPUs, nine adapters of which eight are east-west and one is the DPU, four hundred gigabit per GPU per interface - and procurement checks it against her line.[1] Then you name the two pages separately, because the appendix that lists this hardware never prints the string itself.[2][1] What you actually say: “The ninth adapter is the BlueField-3 B3240, and it needs an 8-pin auxiliary power cable that is not on this quote.”[3] She adds the cable and asks for its lead time. That evening the customer’s ML platform lead emails a second quote: rack-scale NVL72, and a note that the fabric is now a commodity.

Lab

Goal: write the node string for a real machine in the Dell lab. All steps read state only.

  1. Pre-flight. Record hostname, PowerEdge model and how many CPU sockets are populated: lscpu | grep -E 'Socket|Model name'. Expected: a socket count of 1 or 2. That is your first digit.
  2. GPU count: nvidia-smi -L on a GPU host. Expected: one line per GPU. If the host has no GPU, write 0 and continue - the adapter arithmetic is the part that matters here.
  3. Adapter inventory: lspci -nn | grep -i -E 'mellanox|nvidia' then mst status -v. Record every adapter with its PCI address. Expected: the two lists agree in count.
  4. Port inventory: ibdev2netdev maps each RDMA device to a netdev. Expected: one line per port, for example mlx5_0 port 1 ==> ens1f0np0 (Up). A dual-port card produces two lines and still counts as one adapter in the NIC digit.
  5. Classify each adapter as east-west SuperNIC or north-south DPU. A BlueField in DPU mode exposes an Arm side over rshim - ls /dev/rshim* and, if present, cat /dev/rshim0/misc. Expected: rshim present on a DPU and absent on a SuperNIC or plain ConnectX. Do not change any mode; this is an inventory step only.
  6. Write the machine’s string. Take the bandwidth digit from the RA page for its generation, not from the adapter label.[6][3]
  7. Cross-check against iDRAC. Open the iDRAC PCIe slot inventory and compare its adapter list to step 3. Explain any difference in one sentence - an OCP 3.0 card in the OCP slot is the usual cause and it is easy to miss on both sides.
  8. Acceptance: your string, the iDRAC slot list and mst status -v agree on the adapter count, or you can name why they do not.

Retrieval check

10 questions from memory. Answer before looking anything up; misses become flashcards.

Explain it to a Dell SE

Explain to a Dell platform engineer in four sentences what 2-8-9-400 means and why the ninth adapter is not a typo.

13 flashcards for this lesson — 0 in deck. Spaced review lives at /review.

Sources

Facts in this lesson were checked against Enterprise RA program whitepaper key building blocks, HGX AI Factory components and Appendix B, NVL72 AI Factory components and physical topologies, all re-fetched 2026-09-07. Dates are when each page was fetched.

  1. Key building blocks of Enterprise Reference Architectures · fetched 2026-09-07
  2. Appendix B - Node Configurations - NVIDIA HGX AI Factory · fetched 2026-09-07
  3. Components - NVIDIA HGX AI Factory · fetched 2026-09-07
  4. System Hardware and Components - NVIDIA NVL72 AI Factory · fetched 2026-09-07
  5. Networking Physical Topologies - NVIDIA NVL72 AI Factory · fetched 2026-09-07
  6. Networking Physical Topologies - NVIDIA HGX AI Factory (H100 H200 B200) · fetched 2026-09-07
  7. Dell AI Factory with NVIDIA Solution ID 19845005.1 · fetched 2026-09-07
  8. Cisco AI POD Infrastructure for the NVIDIA 2-8-9-400 Enterprise Reference Architecture · fetched 2026-09-07
  9. Specifications - NVIDIA ConnectX-8 SuperNIC User Manual · fetched 2026-09-07
  10. Networking Hardware - NVIDIA NVL72 AI Factory · fetched 2026-09-07
  11. Overview - NVIDIA NVL72 AI Factory · fetched 2026-09-07

The same idea elsewhere

Other lessons that cover this ground, sometimes from another course's angle.