Skip to content

Scenario: a 256-GPU Dell AI Factory design review

S4·E4Friday's quote, Tuesday's review · A Dell account team call, four days before the quote goes to the board

S4·E4Create~35 minsources checked todayverified against HGX AI Factory Enterprise RA, Dell AI switch catalog, PowerScale back-end overview H16346.8 and the MMA4Z00-NS transceiver page, fetched 2026-09-07

Builds on: The sizing chain: GPUs, NICs, planes, leaves, spines, Reading a real Dell AI Factory BOM, The Ethernet vs InfiniBand decision framework

Before you read: what do you already know?

3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.

After this lesson you can

  • Produce a two-page design note for a 32-node HGX B300 cluster that names its governing reference architecture and cites every ratio.
  • Derive the full sizing chain from GPU count to optics count and label each derived number as derived.
  • Detect the two constraints planted in the brief - the PowerScale NOS pin and the legacy AC racks - before they reach the quote.
  • Compose the three questions that must be answered before any optics part number is quoted.

Episode 4 — Friday's quote, Tuesday's review

The situation · A Dell account team call, four days before the quote goes to the board

The SE has his spreadsheet on the shared screen, one row per promise made since the site walk, and a BOM pasted in from a DGX SuperPOD deck because the node count looked close enough. The brief under it is a single paragraph: 32 nodes of HGX B300 — 256 GPUs — PowerScale storage, the sorting hall’s existing multimode fiber plant, legacy AC racks, and an ops team that runs one IP fabric per NIC and automates it with Ansible.

Three sentences in that paragraph have already decided parts of the design, and two of them are quietly fatal to the pasted BOM. The governing document is the HGX AI Factory Enterprise RA, which scopes “ranging from 32 nodes (256 GPUs) to 128 nodes (1,024 GPUs)” — this customer sits exactly on its floor.[1] “Legacy AC racks” removes a whole SKU family before any arithmetic starts, because SN5600D takes DC bus bar input at “40-60 VDC” while SN5610 and SN5600 take “200–240 VAC”.[9] And “PowerScale” means a private back-end network where “Dell does not support connecting any other devices to the back-end switches”, so the spare leaf ports somebody is about to reuse are not available at all.[13]

A design review exists for exactly this: to catch the sentences in a brief that cost money on Friday and cost switches in March.

Say the chain out loud, and label every number the reference architecture does not print as derived.

Start with what the brief already decided, then do the arithmetic where everyone can hear it.

1The brief, and what it has already decided

The ask. A Dell account has a customer who wants 32 nodes of HGX B300 — 256 GPUs — with PowerScale storage. The hall has an existing multimode fiber plant and legacy AC racks. The ops team runs one IP fabric per NIC today and automates it with Ansible. They want a design review before the quote goes out.

Three things are decided before you open a spreadsheet.

The governing document. The HGX AI Factory Enterprise RA documents deployments “ranging from 32 nodes (256 GPUs) to 128 nodes (1,024 GPUs)” and is the Dell-installed HGX case; the DGX SuperPOD RAs cover NVIDIA-led installations with support “through NVIDIA Technical Account Managers”.[1][16] This customer sits at the exact floor of the Enterprise RA.

The node shape. Each HGX B300 baseboard carries “Eight NVIDIA ConnectX-8 SuperNICs”, “Up to 800 Gbps per adapter”, stated as “800 Gb/s (2 x 400Gb/s Ethernet) per GPU” — a 1:1 GPU-to-SuperNIC ratio — plus one BlueField-3 B3240 DPU north-south with two 400 GbE ports, which needs an “8-pin ATX 12V PCIe” connector because it can draw more than 75 W.[2][6] In the Enterprise RA program whitepaper’s nomenclature the 2-8-9 shape is “2 CPUs balanced with 8 GPUs plus 9 NICs”, but its “HGX 2-8-9-400” reference configuration is scoped to “Eligible GPU: NVIDIA HGX H100, H200, or B200” and to 4 to 32 nodes.[8] For B300 the governing document uses its own string: the HGX AI Factory RA labels this deployment “32 NVIDIA HGX B300 Servers (256 GPUs) in 2-8-9-800 configuration with NVIDIA Spectrum-X compute fabric”, counting the full 800 Gb/s port rather than the 400 GbE per-GPU east-west rate.[4] Write the string from the document you are sizing against, and say which document it came from.

The power envelope. Legacy AC racks rule out the DC-busbar SKUs immediately: SN5600D takes DC bus bar input at “40-60 VDC” while SN5610 and SN5600 take “200–240 VAC”.[9] Any BOM lifted from a DC-busbar SuperPOD design is unbuildable here.[16]

Which RA governs this deal?

Four customer facts → one named document, its scope sentence and its SU arithmetic.

GPU count256 GPUsband 4 / 9
Platform
Who installs
Rack power
Enterprise Reference ArchitecturesDGX SuperPOD Reference ArchitecturesProgram Whitepaper32–256 GPUs · defines the node names4/4HGX AI Factory (B300)32–128 nodes · 256–1,024 GPUs4/4NVL72 AI Factory (GB300)up to 8 SUs · 18 trays per SU2/4HGX AI Factory (H100/H200/B200)the Dell installed base3/4SuperPOD B300 — Spectrum-4 / DC busbarSU = 64 nodes = 512 GPUs0/4SuperPOD B300 — Quantum-X800 / ACSU = 72 nodes = 576 GPUs1/4SuperPOD GB200 NVL72SU = 8 NVL72 racks · NDR-era0/4
Shortlist:
4/4 customer factsLast updated May 18, 2026

NVIDIA HGX AI Factory — Enterprise Reference Architecture (B300)

Scope, in the document's own words

ranging from 32 nodes (256 GPUs) to 128 nodes (1,024 GPUs)

HGX AI Factory — index
Scalable Unit

No SU in this RA — it sizes in nodes. One node = "Eight NVIDIA B300 GPUs on an HGX B300 baseboard", so nodes = GPUs / 8.

  • Nodes at 8 GPUs each: 256 ÷ 8 = 32 nodes derived
  • SUs: This RA does not define a Scalable Unit — do not invent one
Rack power — Not published in this RA. It publishes node minimums instead: "Minimum of 48 physical CPU cores per socket", "Minimum of 2TB system memory", "Minimum of 500GB/s memory bandwidth".
Support path — Dell-built HGX in a PowerEdge chassis: Dell-branded, Dell-installed, Dell-supported. No NVIDIA TAM is implied by this document.
Node nomenclature — HGX 2-8-9-400. The "9" is 8 east-west SuperNICs + 1 north-south DPU: the appendix states "8 on baseboard NVIDIA ConnectX-8 SuperNIC with a dual-port 400G configuration" and "1 NVIDIA BlueField-3 B3240 DPU with 2x400G ports". The appendix itself never prints the string 2-8-9-400 — that definition lives in the program whitepaper. HGX AI Factory — Appendix B node configurations
Scale-up (NVLink, inside the accelerator) — HGX B300 baseboard: "Total Aggregate Bandwidth 14.4TB/s", "GPU-to-GPU Bandwidth 1800GB/s" — one NVLink domain of 8 GPUs.
Scale-out (the fabric, outside) — "800 Gb/s (2 x 400Gb/s Ethernet) per GPU" through eight ConnectX-8 SuperNICs — a 1:1 GPU-to-SuperNIC ratio.
⚠ Derived, not published: 1,800 GB/s scale-up = 14,400 Gb/s, against 800 Gb/s scale-out per GPU on B300 — about 18x. This ratio is arithmetic over the NVLink page and the HGX B300 components page; neither prints it.
What this RA does not publish
  • No BGP/EVPN configuration, no IP addressing plan, no MTU.
  • No RoCE QoS values — no DSCP, PFC or ECN settings.
  • No adaptive-routing or congestion-control parameters. "VLAN isolation would be used to provide logical separation of the networks above over the single physical fabric" is as far as it goes.
  • The certified-storage page names no protocols (NFS over RDMA, GPUDirect Storage), no tiers and no partner names.

FAE angle: this is the Dell PowerEdge case, and the plane decision is the BOM. Dual plane splits 800 Gb/s into "2x400 Gb/s interfaces", each landing on "a different leaf switch" in "an independent fabric that scales to 1024 interfaces of 400 Gb/s". Single plane is one 400 Gb/s link per GPU and the RA calls it a "50%" bandwidth reduction. Dual plane is recommended for 64+ node deployments.

The brief's four facts entered as toggles. Read the matching RA's scope statement and its 'what this RA does not publish' list before you size anything.

2The chain, said out loud

Say the chain in front of the customer. It converts a brand argument into arithmetic they can check.

Step 1 — GPUs to NICs. 256 GPUs at 1 ConnectX-8 SuperNIC per GPU = 256 SuperNICs, plus 32 BlueField-3 B3240 DPUs, one per node.[2][6]

Step 2 — NICs to fabric interfaces. Dual plane: each 800 Gb/s port breaks out to “2x400 Gb/s interfaces”, each landing on “a different leaf switch”, each leaf “part of an independent fabric that scales to 1024 interfaces of 400 Gb/s”.[3] So 256 × 2 = 512 server-facing 400G interfaces, 256 per planederived: the ratio and the breakout are published, the multiplication is mine.[3]

Step 3 — interfaces to leaves. SN5610 presents “64 total ports of 800 Gbps”; the SN5600-series datasheet gives the same silicon as 64 OSFP cages that become 128x400G logical ports.[5][9] A non-blocking leaf therefore has 64 downlinks and 64 uplinks at 400G — derived: the RA never prints a per-leaf split, so this is an assumption the customer’s port map must confirm.[9] On that assumption, 256 interfaces per plane ÷ 64 downlinks = 4 leaves per plane, 8 leaves total, plus spines at the chosen uplink ratio.[9]

Step 4 — the plane decision has an operational price. Two independent fabrics is exactly what this customer’s ops team does not do today.[3] Single plane is the documented alternative: “a single 1x 400 Gb/s connection” per GPU, which “reduces the total GPU bandwidth by 50%” and is “well-suited for workloads that do not warrant maximum throughput”.[3] Quote the RA’s own recommendation before the customer does: “In this Enterprise RA we recommend using a Dual Plane topology” — stated with no node-count qualifier — and the logical architecture puts the 32-, 64- and 128-node B300 builds on a dual-plane Spectrum-X compute fabric alike.[3][4] Single plane here is therefore a deliberate, documented departure from the RA’s recommendation on cost and ops-model grounds, not a size-based exemption. Put both options in the note with their BOM deltas, label the single-plane case as a departure, and let the customer choose; do not choose for them.

GPUs
256
Planes
Uplink ratio
256 GPUs32 nodes256 NICs800 Gb/s512 x400G2 planes8 leaves64 down / 64 up4 spines1:1HGX B300 · SN5610 · 2-planeserver-facing 400G: 512 · leaf-spine links: 512 · twin-port optics: 1,024converged ports: 78 · OOB ports: 78 · storage floor: 3,200 Gb/sRA documents 32 nodes (256 GPUs) to 128 nodes (1,024 GPUs)
link 1 / 6
HGX B300 · published ratio

GPUs → NICs

256 x ConnectX-8 SuperNIC

"Eight NVIDIA ConnectX-8 SuperNICs per HGX B300 baseboard", "800 Gb/s (2 x 400Gb/s Ethernet) per GPU" — a 1:1 GPU-to-SuperNIC ratio. The "9" in 2-8-9-400 is 8 east-west SuperNICs + 1 north-south BlueField-3 B3240. The DPU is the north-south NIC and is counted separately from the east-west SuperNICs — that is where the "9" in HGX 2-8-9-400 comes from.

FAE angle

The nomenclature string is the fastest audit of a quote: CPUs-GPUs-NICs-speed. If the NIC line does not read 2-8-9-400 on a Hopper or B300 node, something is missing — usually the north-south DPU.

256 GPUs x 1 NIC per GPU = 256 NICs, plus 1 DPU per node (32 DPUs)

HGX B300 Components

Node shape and SU arithmetic for this preset: HGX B300 RA · nomenclature (HGX 2-8-9-400) · B300 node configurations · SuperPOD SU and rack power · NVL72 networking hardware

Run the chain with show-math on, then set planes to 1 and record every line that changes. The delta between the two runs is the single-plane conversation.

3Converged, storage, out-of-band, and the pin nobody sees

Converged north-south. “Each compute and management node is connected with two 400 GbE ports to two separate switches” for redundancy, achieving “up to 40 GB/s per node”.[3][4] For 32 compute nodes plus the control plane — the RA supports “up to eight control plane nodes” — that is 64 compute-side 400G converged ports plus two per management node.[2] Derived: the port totals; published: the two-ports-per-node rule.

Storage. The RA’s rule is “approximately 12.5 Gb/s per GPU”, scaling linearly.[7] For 256 GPUs that is ≈3.2 Tb/s aggregate — derived.[7] Storage platforms must be in the NVIDIA-Certified Storage program, and that page names no protocols, tiers or partners, so make no claim about NFS over RDMA or GPUDirect Storage from it.[7]

Out-of-band. “1 Gb RJ45 connectivity for all the Nodes” into “NVIDIA SN2201 48-port 1Gb switches”, reaching BMC ports, DPU and SuperNIC management ports and switch OOB ports.[4] Count at least two per node — system BMC and BlueField-3 BMC — before switches and PDUs are added.[4] Dell also sells the S3248T-ON as an OOB alternative from the same catalog.[10]

The planted constraint. The customer has PowerScale. Dell requires PowerScale to have a “private back-end” network, isolated from the front end, with redundant switches, and states “Dell does not support connecting any other devices to the back-end switches”.[13] Where SN5600 is used there, support runs “via the Dell Technologies ETC program” and the document says to “ensure the Spectrum 4 switches are running the version Cumulus Linux 5.9.1. This is the version that Dell Technologies tested and qualified”.[13] Meanwhile the Spectrum-X validated solution stack pins its own NOS row for the AI fabric.[14] One switch cannot satisfy both pins, and the back end cannot host other devices anyway. The design note needs a separate storage back-end switch pair, and it needs to say why in one sentence with both citations. Dell’s own reference solution shows the same shape: SN5600 for the fabric with a separate S5232F-ON pair carrying the PowerScale storage cluster network.[18]

4Optics, power, and three questions

The brief says multimode plant, which narrows the family immediately — but only after three questions are answered: is the plant multimode or single-mode; what is the longest leaf-to-spine run in measured meters; and is the design dual-plane or single-plane?[11][5]

For a multimode plant at 32 nodes, the workhorse is MMA4Z00-NS: “400Gb/s for both 400GbE Ethernet and NDR InfiniBand” per port in a twin-port OSFP totalling 800 Gb/s, “Two, 4-channel MPO-12/APC optical connectors”, multimode, maximum reach 50 m, and used “only … in Quantum-2 and Spectrum-4 OSFP air-cooled switches”.[11] Its companions: straight fiber “MFP7E10-Nxxx” up to 50 m, the 1:2 splitter “MFP7E20-N0xx”, the single-port equivalent “MMA4Z00-NS400” and the QSFP112 equivalent “MMA1Z00-NS400”.[11] Anything inside a rack at ≤3 m can be passive DAC from the MCP family; the prefixes to know are MCP copper DAC, MCA active copper, MFS fiber AOC, MMA/MMS transceivers, MFP fiber accessories.[12] Runs beyond 50 m or on a single-mode plant move to the MMS/DR classes.[12]

Two numbers that end up in someone else’s budget. Optics power: the fetched MMA4Z00-NS overview states 15 Watts, while a search snippet of the same product’s PDF says 8 W — use the fetched page and flag the conflict in writing, because a fully populated 64-cage switch multiplies that difference by 64.[11] Switch depth and thermals: the SN5600 series is 2U and 788 mm deep with “Reverse” airflow, SN5610 tolerates “0–40ºC” while SN5600 and SN5600D tolerate “0–35ºC”, and PSU redundancy differs — SN5610 is 2+2 with five fans, SN5600 and SN5600D are 1+1 with four.[9] In a legacy AC hall with contained hot aisles, that 5 °C and that rail depth decide whether the design is installable.[9]

5The design note, and how it is graded

The deliverable is two pages a Dell account executive can hand to a customer. It is graded on whether it cites a page rather than asserts a number.

Positioning: three customers, one afternoon · decision 1/7A 0 · P 0 · S 0

Brief — Back-to-back calls: a VMware farm (200× R760), an AI training pod (64× XE9780 on Spectrum-X) and a storage-heavy multi-tenant inference cluster. Each has objections.

Dell account team + three end customers: VMware farm: "200 R760s on vSphere 8; NSX eats about 20% of our cores. Should we go DPU?"

Run the positioning scenario against the clock before you write. The scoring separates technical accuracy from positioning and safety - most first runs score high on accuracy and low on positioning.
The 256-GPU design note

1. Governing RA. HGX AI Factory Enterprise RA — it scopes “32 nodes (256 GPUs) to 128 nodes (1,024 GPUs)” and this is a Dell-installed HGX deployment, not an NVIDIA-led DGX SuperPOD.[1][16] Node shape is the 2-8-9-800 HGX B300 node in the governing RA’s own nomenclature; the whitepaper’s 2-8-9-400 configuration covers H100/H200/B200 at 4 to 32 nodes and is not this node.[4][8]

2. Plane decision. Recommend single plane with a documented upgrade path, and label it in the note as a departure from the RA: the RA says “In this Enterprise RA we recommend using a Dual Plane topology” with no node-count qualifier and shows dual plane at 32, 64 and 128 nodes alike.[3][4] The grounds are cost and the customer’s one-IP-fabric-per-NIC ops model, not cluster size. BOM consequence stated plainly: single plane is exactly 50% of dual-plane bandwidth, uses single-port OSFP transceivers rather than twin-port, and halves the leaf count.[3][5] The RA calls migration between the topologies “seamless” because the OSFP cage takes either module, but the modules and the second fabric of leaves still have to be bought.[3]

3. Sizing chain. 256 GPUs × 1 SuperNIC per GPU = 256 SuperNICs (published ratio).[2] Single plane → 256 × 400G server-facing interfaces (derived).[3] Leaf radix 64 cages → 128x400G logical, assumed 64 down / 64 up (derived, confirm against port map).[9] → 4 leaves (derived), spines per chosen uplink ratio. Dual-plane variant: 512 interfaces, 8 leaves (derived).[3]

4. Converged / storage / OOB. Converged: 2×400 GbE per compute and management node to two separate switches, up to 40 GB/s per node → 64 compute-side ports plus control plane (derived).[3][2] Storage floor: ≈12.5 Gb/s per GPU → ≈3.2 Tb/s (derived).[7] OOB: ≥2 × 1 GbE RJ45 per node into SN2201 48-port switches → ≥64 node ports plus switch and PDU ports (derived).[4]

5. Optics BOM. Multimode plant, all runs measured ≤50 m: MMA4Z00-NS400 single-port modules at the NIC for the single-plane design, MMA4Z00-NS twin-port where a switch cage feeds two 400G endpoints, MFP7E10-Nxxx straight fiber, MFP7E20-N0xx splitters where used, MCP-family passive DAC for intra-rack ≤3 m.[11][12] Optics power budgeted at 15 W per twin-port module with the 8 W PDF conflict flagged.[11]

6. Version-pin catch. PowerScale back end is a separate private switch pair, no other devices attached, pinned to Cumulus Linux 5.9.1 per Dell’s ETC qualification, and therefore not the same switches as the Spectrum-X-stack AI fabric.[13][14] Dell’s own reference solution uses a separate S5232F-ON pair for the PowerScale cluster network.[18]

7. Power and thermal. AC hall → SN5610 or SN5600 at 200–240 VAC, not SN5600D at 40–60 VDC.[9] Note SN5610’s 0–40 ºC envelope versus 0–35 ºC on SN5600, the 788 mm depth against the customer’s rail depth, and consistent “Reverse” airflow.[9] Server side: the XE9780 is a 10U air-cooled platform with twelve 3200 W PSUs and 8x CX-8 OSFP with the B300 GPU configuration.[17]

8. Support path. Dell-sold PowerEdge plus Dell PowerSwitch → Dell ProSupport is the front door, with ProDeploy and ProConsult in the solution.[15][18] State honestly that no fetched page publishes a defect-ownership split, and record instead the asset owner and the contract number for every switch.[15]

9. Three questions before optics are quoted. Multimode or single-mode plant; longest run in measured meters; dual-plane or single-plane. Answers dated and initialled in the note.[11]

10. Derived-number register. A short table listing every derived figure in the note — interface totals, leaf counts, converged and OOB port counts, storage aggregate, optics power — each with the published inputs it came from.

Episode 4 — Two pages and a register

How it ended

What goes back is two pages. It names its governing RA with the scope quote, walks the chain from 256 GPUs to a cable count, and marks every figure the RA does not print as derived — interface totals, leaf counts, the storage aggregate derived from ≈12.5 Gb/s per GPU.[1][7] The plane recommendation is stated as a departure from the RA’s published dual-plane recommendation, on cost and ops-model grounds.[3] The PowerScale back end becomes its own switch pair, and the DC-busbar line items are gone.[13][9] The night-shift operator walks the trays with a tape and comes back with a measured 38 m where the floor plan claimed 25.[11]

Then, label maker in hand, he asks the question nobody has costed: at 02:10, who exactly is qualified to pick up?

Lab

Ground one column of the BOM in measurement instead of assumption, using the Dell-lab BlueField-3 or ConnectX host as the concrete node. Everything here is a read; no firmware, mode or configuration change, so no rollback is required. Record the pre-flight state anyway.

Pre-flight inventory: mlxfwmanager --query (board, PSID, firmware), lspci | grep -i mellanox (card count and slots), ibdev2netdev (device-to-interface map) and ip -br link — saved to a file with a timestamp before anything else.

  1. Derive the node’s real nomenclature string in the Enterprise RA format CPUs-GPUs-NICs-speed from what is physically installed: count sockets with lscpu, count accelerators, count NICs, read the port speed. Expected: a string you can defend, and a note where the lab node differs from the 2-8-9-800 shape the governing RA gives the brief.[8]
  2. Count the actual out-of-band ports on the chassis: the system BMC/iDRAC port plus any DPU management port. Expected: the count matches or contradicts the “at least two per node” assumption in your OOB section — record which.[4]
  3. Read the installed transceivers: ethtool -m <iface> | head -30 on each fabric-facing interface. Expected: vendor, part number, type and nominal reach. Compare with the MMA4Z00-NS family and note whether the lab part is twin-port or single-port.[11]
  4. Read the negotiated speed with ethtool <iface> | grep -i speed and compare it to the 400G-per-interface assumption in the sizing chain. Expected: a number, not an expectation. Any mismatch goes into the design note as a measured fact.
  5. Substitute every measured value into the corresponding BOM column so at least one column is measured rather than assumed, and mark the substituted cells. Expected: your derived-number register shrinks by exactly the number of cells you measured.
  6. Optional, in a customer lab: measure one real cable run end to end with a tape or an OTDR rather than reading it off a floor plan, and put the measured metres and the date into the optics section.[11]

Retrieval check

10 questions from memory. Answer before looking anything up; misses become flashcards.

Explain it to a Dell SE

Explain to a Dell SE how you get from 256 GPUs to a cable count without asserting anything the reference architecture does not publish.

14 flashcards for this lesson — 0 in deck. Spaced review lives at /review.

Sources

Facts in this lesson were checked against HGX AI Factory Enterprise RA, Dell AI switch catalog, PowerScale back-end overview H16346.8 and the MMA4Z00-NS transceiver page, fetched 2026-09-07. Dates are when each page was fetched.

  1. NVIDIA HGX AI Factory Enterprise Reference Architecture (index) · fetched 2026-09-07
  2. Components — NVIDIA HGX AI Factory · fetched 2026-09-07
  3. Networking Physical Topologies — NVIDIA HGX AI Factory · fetched 2026-09-07
  4. Networking Logical Architecture — NVIDIA HGX AI Factory · fetched 2026-09-07
  5. Networking Hardware — NVIDIA HGX AI Factory · fetched 2026-09-07
  6. Appendix B Node Configurations — NVIDIA HGX AI Factory · fetched 2026-09-07
  7. NVIDIA-Certified Storage — NVIDIA HGX AI Factory · fetched 2026-09-07
  8. Key building blocks of Enterprise Reference Architectures · fetched 2026-09-07
  9. NVIDIA Spectrum SN5600 Series Switches Datasheet (Dell-branded) · fetched 2026-09-07
  10. AI Networking Switches | Dell USA · fetched 2026-09-07
  11. MMA4Z00-NS 800Gb/s Twin-port OSFP 2x400Gb/s Multimode 50m · fetched 2026-09-07
  12. Networking Interconnect (LinkX families and part-number prefixes) · fetched 2026-09-07
  13. Dell PowerScale: Ethernet Back-End Network Overview H16346.8 · fetched 2026-09-07
  14. NVIDIA Spectrum-X Validated Solution Stack · fetched 2026-09-07
  15. Dell-NVIDIA Partnership Powers High-Performance AI Fabric Solutions · fetched 2026-09-07
  16. Abstract — DGX SuperPOD B300 Spectrum-4 Ethernet and DC Busbar RA · fetched 2026-09-07
  17. Dell PowerEdge XE9780 Technical Guide · fetched 2026-09-07
  18. Dell AI Factory with NVIDIA Solution ID 19845005.1 component list · fetched 2026-09-07

The same idea elsewhere

Other lessons that cover this ground, sometimes from another course's angle.