Which reference architecture governs this deal
S1·E1Two leaf counts for the same server · Dell briefing room, Round Rock, Monday, four days before the BOM freezes
Before you read: what do you already know?
3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.
After this lesson you can
- Name the three families of NVIDIA network reference architecture and the customer question each one answers.
- Read a reference architecture's own scope statement and state the GPU range, the tenancy model and who installs the cluster.
- Decide which RA governs a deal from four customer facts: GPU count, GPU platform, who installs, and whether the hall has a DC busbar.
- Explain why the same DGX B300 server appears under two RAs with different scalable-unit sizes and different rack-power statements.
Episode 1 — Two leaf counts for the same server
The Dell SE keeps a spreadsheet of every promise he has made to this customer, and row 41 reads “leaf count, Monday”. It is Monday, his coffee went cold an hour ago, and row 41 is in trouble. He has costed the utility’s forecasting cluster around a repeating block of 512 GPUs. The customer’s network lead has the same server written in his graph-paper notebook and a block of 576. Neither can explain the other’s arithmetic, and both of them turn to you.
Neither is wrong. The Spectrum-4 Ethernet and DC busbar edition of the DGX SuperPOD B300 RA makes a scalable unit 64 DGX B300 nodes for 512 GPUs, four systems to a rack, with rack-level power consumption above 50 kW.[6] The Quantum-X800 InfiniBand and AC power edition makes the same server’s scalable unit 72 nodes for 576 GPUs, four to a rack with three 2U rack PDUs, at around 56 kW.[7] Same server, two documents, two SU arithmetics - and therefore two leaf counts, two cable counts and two rack counts.
This is why the shelf looks the way it does. NVIDIA does not publish one architecture per server; it publishes a tested design per fabric, per power delivery and per install model, and the scalable unit is the arithmetic unit each design is quoted in. The SuperPOD documents assume NVIDIA-led installation and premium support through an NVIDIA Technical Account Manager, while the Enterprise RAs assume a partner-built factory.[5][1]
A switch count means nothing until somebody names the governing document. Start with the three families of document.
1Three families of document, one question per family
NVIDIA does not publish one reference architecture. It publishes at least three families of them, and the first skill in this course is knowing which family a customer question lands in.[1][5] The DGX SuperPOD RAs describe NVIDIA-branded DGX systems installed by NVIDIA.[5] The Enterprise Reference Architectures describe partner-built AI factories.[1] The per-platform validated stacks, such as the Spectrum-X validated solution stack, publish matched version rows rather than topology.[13]
The Enterprise RA index lists these documents by exact title: Program Whitepaper, NVIDIA RTX PRO AI Factory, NVIDIA HGX AI Factory, NVIDIA NVL72 AI Factory, NVIDIA HGX AI Factory (HGX H100, H200, B200), NVIDIA Enterprise AI Factory Design Guide White Paper, and NVIDIA AI Enterprise: Software Reference Architecture, followed by eleven further deployment and scaling guides.[1] Read that list once and the vocabulary stops being fog. “The Enterprise RA” is not a document, it is a shelf.
The program whitepaper puts the shelf’s own boundary in writing: recommendations for building AI Factories for enterprise-class deployments ranging from 32 to 256 GPUs.[1][2] Individual platform RAs then publish wider ranges of their own, which is why quoting the program range as if it capped every document is a mistake a competitor will correct in front of you.[3]
Which RA governs this deal?
Four customer facts → one named document, its scope sentence and its SU arithmetic.
SuperPOD B300 — Spectrum-4 / DC busbar vs SuperPOD B300 — Quantum-X800 / AC
| SuperPOD B300 — Spectrum-4 / DC busbar | SuperPOD B300 — Quantum-X800 / AC | |
|---|---|---|
| Family | DGX SuperPOD Reference Architecture | DGX SuperPOD Reference Architecture |
| Scope, verbatim | “Single-tenant, multi-user enterprise environments with NVIDIA-led installation and professional support through NVIDIA Technical Account Managers” | “Each group of 72 nodes is rail-aligned. Traffic per rail of the DGX B300 systems is always one hop away from the other 72 nodes in a SU” |
| Scalable Unit | SU = 64 DGX B300 nodes = 512 GPUs. 4 SU = 256 nodes = 2,048 GPUs. 64 SU = "2000+ DGX B300 nodes". | SU = 72 DGX B300 nodes = 576 GPUs. 8 SU = 576 nodes = 4,608 GPUs. Maximum listed is "72+ SU with 2000+ DGX B300 nodes". |
| Rack power | "Four DGX B300s are within a single rack." "The rack-level power consumption per rack exceeds 50 kW." The design uses "the MGX-based, DC busbar powered design for the best datacenter density and efficiency". | "Four DGX B300 fit within a single rack" with "three (3) 2U rack PDUs for maximum redundancy". "The rack-level power consumption per rack is ~56kW". |
| Who installs / supports | NVIDIA-led installation and professional support through NVIDIA Technical Account Managers — that sentence is the support path, in the abstract, verbatim. | Same SuperPOD support model: NVIDIA-led install with a Technical Account Manager. |
| Power scheme | DC busbar (MGX) | AC rack PDU |
| Last updated | Last updated November 19, 2025 | variant of the same November 2025 SuperPOD family; the research note records no separate date |
| Node string | DGX B300 (NVIDIA-branded appliance) | DGX B300 PS (72 per SU) |
FAE angle: the SU is the unit of quoting. When a Dell account asks "how many switches", the honest answer is "how many SUs, and which RA" — 64-node vs 72-node SUs change leaf count, cable count and rack count. Check which RA revision the customer's NVIDIA account team quoted from, and whether the datacenter is busbar-capable.
2The scope sentence is the deciding sentence
Every one of these documents scopes itself in a sentence near the front, and that sentence is the fastest disqualifier you have.
The HGX AI Factory RA covers deployments from 32 nodes (256 GPUs) to 128 nodes (1,024 B300 HGX GPUs); its chapters are Abstract, Overview, Components, Networking Hardware, Networking Physical Topologies, Networking Logical Architecture, Software, NVIDIA-Certified Storage, Services, Summary, Appendix: Node Configurations and Notices; it was last updated May 18, 2026.[3]
The NVL72 AI Factory RA states that a fully tested system scales up to 8 SUs, where each scalable unit is a single NVL72 rack integrating 18 compute trays, each rack carrying 72 Blackwell Ultra GPUs and 36 Grace CPUs.[4] It also scopes its tenancy: the design is tailored to a configuration where users are all part of the same enterprise so that accounting and access control can be consolidated.[4] That is a single-tenant, multi-user document. A service provider selling isolated tenants is outside it.
The DGX SuperPOD B300 RA is designed for a single-tenant, multi-user enterprise environment, with NVIDIA-led installation and white-glove services for initial bring-up and commissioning, and premium support from an NVIDIA Technical Account Manager.[5] It covers InfiniBand, the NVLink network, Ethernet fabric topologies, storage system specifications, recommended rack layouts and wiring guides.[5]
Three sentences, three different customers. Note what none of them promise: a configuration guide. No Enterprise RA publishes BGP or EVPN configuration, IP addressing, MTU, or DSCP/PFC/ECN values - lesson 4 of this module makes that boundary explicit.[3][4]
3Who installs is the fork, and it forks three things at once
Scale narrows the shortlist. Who installs and commissions picks the branch. A DGX SuperPOD arrives with NVIDIA-led installation and an NVIDIA Technical Account Manager behind it.[5] An HGX baseboard inside a Dell PowerEdge is Dell-branded, Dell-installed and Dell-supported: Dell’s own AI Factory solution sheet lists ProSupport Plus or ProSupport One, ProDeploy and ProConsult against the same bill of materials that carries the servers, the SN5600 switches and the PowerScale array.[10]
That single fork moves three things together:
| What forks | Enterprise RA path | DGX SuperPOD path |
|---|---|---|
| Governing document | HGX AI Factory or NVL72 AI Factory[3][4] | DGX SuperPOD B300 or GB200 RA[5][9] |
| Install and support | Partner-installed; Dell ProSupport in a Dell deal[10][11] | NVIDIA-led install, NVIDIA TAM[5] |
| Switch SKU line | SN5610 and SN2201 named in the Enterprise RAs[3] | SN5600D and DC busbar on the Spectrum-4 edition[6] |
There is a third support path that catches Dell accounts by surprise. If an SN5600 is used as a PowerScale back-end switch it runs through the Dell Technologies ETC program and PowerScale pins it to Cumulus Linux 5.9.1 - a different version from the one the Spectrum-X validated stack pins for the AI fabric.[12][13] Module 3 develops that conflict; here it is enough to know a switch can belong to a support path that neither RA describes.
| Product | Speed | PCIe | Role | GPU generation |
|---|---|---|---|---|
NIC | ||||
SuperNIC (no Arm) | ||||
SuperNIC (no Arm) | ||||
DPU | ||||
DPU | ||||
SuperNIC (Arm inactive) | ||||
DPU / storage processor | ||||
Ethernet switch | ||||
Ethernet switch | ||||
InfiniBand switch | ||||
InfiniBand switch |
⚠ = not confirmed on a fetched primary source (hover for why). Facts as of DOCA 3.5.0 (Sep 2026). Selections are saved.
4Same server, two RAs: the 64 versus 72 node scalable unit
The most useful thing in this lesson is a discrepancy. The DGX B300 server appears in two SuperPOD RAs, and they do not agree on the scalable unit.
On the Spectrum-4 Ethernet and DC busbar edition, an SU is 64 DGX B300 nodes for 512 GPUs, four DGX B300 fit within a single rack, and rack-level power consumption per rack exceeds 50 kW.[6] The design uses an MGX-based DC busbar powered design for datacenter density, while noting that for a legacy datacenter without the possibility for a DC busbar it is still possible to build a SuperPOD with traditional PDU and AC powered EIA racks.[6]
On the Quantum-X800 InfiniBand and AC power edition, an SU is 72 nodes for 576 GPUs, four DGX B300 fit within a single rack with three 2U rack PDUs for maximum redundancy, and rack-level power is around 56 kW; the published scaling table runs to 18 SU (1,296 nodes, 9,216 GPUs), while the prose says the design can scale up to and beyond 72 SU with 2000+ DGX B300 nodes and directs you to contact NVIDIA for solutions of four scalable units or more.[7]
Same server. Different document. Different SU arithmetic, and therefore different leaf counts, cable counts and rack counts for the same GPU total. The SU is the unit of quoting, so “how many switches” is unanswerable until you know how many SUs and which RA.
Dates move too, independently. Re-fetched on 2026-09-07: both Enterprise RAs read Last updated on May 18, 2026[3][4], the Spectrum-4 SuperPOD edition read November 19, 2025[5], and the Quantum-X800 edition had moved to September 02, 2026[7]. A number you memorised from a document six months ago is a number you have to re-read.
Row 41, closed
You do not pick a number. You put both editions on the screen - 64-node SUs above 50 kW per rack against 72-node SUs at around 56 kW[6][7] - and then ask who installs and commissions the cluster, because that fork also decides the support path and the switch SKU line.[5][10] Dell installs, so an Enterprise RA governs and the DGX arithmetic leaves the room. What you actually say: “Before I give you a switch count, tell me which document your NVIDIA account team quoted from - the SU is 64 nodes in one and 72 in the other.” The SE closes row 41. Four minutes later procurement is pointing at the node line on the Dell quote, asking why eight GPUs need nine network cards.
Lab
Goal: place the hardware you actually have on the RA map. All steps are read-only; nothing here changes firmware or configuration.
- Pre-flight inventory. On each BlueField-3 or ConnectX host in the Dell lab record hostname, PowerEdge model and iDRAC firmware. Keep the list - later lessons in this course reuse it.
- Run
mst status -vand record the device list and PCI addresses. Expected: one line per adapter with anmtdevice name and a PCI B:D.F. If the command is missing, the DOCA-Host or MLNX_OFED tools are not installed - stop and note it rather than installing anything. - Run
mlxfwmanager --queryand record, per adapter, the Part Number, PSID, Description and current firmware version. Expected: a Description string that names the card family, for example a BlueField-3 or ConnectX part. - Run
lspci -nn | grep -i -E 'mellanox|nvidia'and confirm the adapter count matches step 2. A mismatch usually means an OCP 3.0 card thatmstenumerates differently - note it rather than assuming a fault. - Decide the RA row. If the card is a BlueField-3 B3140H or B3220, the matching Enterprise RA row is the H100/H200/B200 generation, which names B3140H SuperNICs east-west and B3220 DPUs north-south at two 200 GbE ports per node.[8] If it is a ConnectX-8, the row is the B300 generation instead.
- Write one sentence you could say to a customer, in this shape: “This node carries N adapters of family X, which puts it on the Y Enterprise RA row, so its converged network is Z GbE per port.” Check the last clause against the RA page rather than against memory.[8]
- Acceptance: the sentence names a specific RA document, not “the NVIDIA reference architecture”.
Goal: build the one-page comparison table you will keep open in every RA conversation. Nothing here needs hardware.
- Open the Enterprise RA index at
docs.nvidia.com/enterprise-reference-architectures/and copy the document titles verbatim into a list.[1] Expected: seven core documents plus eleven deployment and scaling guides. If your list is shorter, you are looking at a filtered view - clear the search box. - Open the program whitepaper’s key building blocks page and record its GPU range sentence verbatim.[2] Expected: a range of 32 to 256 GPUs. If not, the page moved - record what it now says and note the date you read it.
- Open the HGX AI Factory index and record the node and GPU range plus the Last updated line.[3] Expected on 2026-09-07: 32 nodes (256 GPUs) to 128 nodes (1,024 GPUs), last updated May 18, 2026.
- Open the NVL72 AI Factory overview and record the tested SU count, the trays per SU, the GPUs and CPUs per rack, and the tenancy sentence.[4] Expected: up to 8 SUs, 18 compute trays per SU, 72 GPUs and 36 Grace CPUs per rack.
- Open both DGX SuperPOD B300 architecture pages - the Spectrum-4/DC busbar one and the Quantum-X800/AC one - and record SU size, GPUs per SU, systems per rack, rack power and the Last updated line for each.[6][7] Expected: 64 nodes / 512 GPUs / exceeds 50 kW against 72 nodes / 576 GPUs / around 56 kW.
- Fill a six-column table: GPU-count range, SU size in nodes, who installs, who supports, rack power statement, page last-updated date. Mark with an asterisk every cell that has moved since 2026-09-07, when the two Enterprise RAs read May 18, 2026, the Spectrum-4 SuperPOD RA read November 19, 2025 and the Quantum-X800 SuperPOD RA read September 02, 2026.[3][5][7]
- Acceptance: you can point at one cell and say which document it came from without reopening the browser. If you cannot, the table is too long - cut it to the six columns.
Retrieval check
10 questions from memory. Answer before looking anything up; misses become flashcards.
Explain it to a Dell SE
Explain to a Dell SE in four sentences how you decide which NVIDIA reference architecture governs a deal, and why the answer changes the bill of materials and not just the paperwork.
Sources
Facts in this lesson were checked against NVIDIA Enterprise Reference Architectures index and program whitepaper, HGX AI Factory, NVL72 AI Factory, DGX SuperPOD B300 (Spectrum-4/DC busbar and Quantum-X800/AC), all re-fetched 2026-09-07. Dates are when each page was fetched.
- NVIDIA Enterprise Reference Architectures (index) · fetched 2026-09-07
- Key building blocks of Enterprise Reference Architectures · fetched 2026-09-07
- NVIDIA HGX AI Factory (index) · fetched 2026-09-07
- Overview - NVIDIA NVL72 AI Factory · fetched 2026-09-07
- Abstract - DGX SuperPOD B300 Spectrum-4 Ethernet and DC Busbar Power RA · fetched 2026-09-07
- DGX SuperPOD Architecture - DGX B300 Spectrum-4 Ethernet and DC Busbar Power RA · fetched 2026-09-07
- DGX SuperPOD Architecture - DGX B300 Quantum-X800 InfiniBand and AC Power RA · fetched 2026-09-07
- Networking Physical Topologies - NVIDIA HGX AI Factory (H100 H200 B200) · fetched 2026-09-07
- Key Components of the DGX SuperPOD - DGX GB200 Reference Architecture · fetched 2026-09-07
- Dell AI Factory with NVIDIA Solution ID 19845005.1 · fetched 2026-09-07
- Dell-NVIDIA Partnership Powers High-Performance AI Fabric Solutions · fetched 2026-09-07
- Dell PowerScale Ethernet Back-End Network Overview (H16346.8) · fetched 2026-09-07
- NVIDIA Spectrum-X Validated Solution Stack · fetched 2026-09-07
The same idea elsewhere
Other lessons that cover this ground, sometimes from another course's angle.
- HGX reference architecture rolesDOCA course · Same ground: hgx, power and reference
- Design review: an eight-rail IB pod on Dell XE9680InfiniBand course · Same ground: superpod, Scalable Unit and dgx
- Who actually ships SRv6: the Dell and NVIDIA support matrixSRv6 course · Same ground: scope, architecture and support