Reading a real Dell AI Factory BOM
S3·E3An architecture with the reasoning deleted · a conference room, ninety minutes before the customer call
Builds on: The PowerEdge AI server line, Dell's switch catalog: SN, Q and Z in one price list
Before you read: what do you already know?
3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.
After this lesson you can
- Decompose a published Dell AI Factory bill of materials into compute, fabric, storage, rack and software roles.
- Infer the reference-architecture node shape from a single NIC line on a quote.
- Determine which BOM lines are fixed by an RA statement and which are Dell packaging choices.
- Re-size the same solution for a different GPU count and name every line that changes.
Episode 3 - An architecture with the reasoning deleted
The sheet procurement left on the pallet is one page long. Dell Solution 19845005.1: three PowerEdge R660, two XE9680, “16x NVIDIA H200SXM GPUs (Total 2256GB)”, one adapter line, three switch lines, three PowerScale F710, a rack and four PDUs.[1] It is the carrier’s existing pilot, and the Dell SE wants to tell them this afternoon that it is “the NVIDIA reference architecture”. Nobody in the room can say whether it is, because the sheet never uses the phrase and never cites a document.
You read one line out loud: “NVIDIA Bluefield-3 (1x3220/8x3140H)”.[1] Nine adapters per server, eight east-west and one north-south, in a two-socket eight-GPU box. The Enterprise RA whitepaper names node types as CPUs-GPUs-NICs-speed and defines “HGX 2-8-9-400” as “2 CPUs balanced with 8 GPUs plus 9 NICs with East-West traffic of 400 GbE per GPU”.[2] The pilot is an HGX 2-8-9-400 implementation and it never said so.
That is what a bill of materials is: an architecture with the reasoning deleted. The nomenclature exists so the ratios survive the deletion - strip every design sentence off a quote and the balance between CPUs, GPUs, adapters and speed is still legible to somebody who was not in the design meeting. It also means the sheet cannot defend itself. The same whitepaper scopes that node type to 4 to 32 nodes, and this deployment has two.[2][1]
Read the NIC line first: it is the fingerprint the configurator could not delete. Then read the rest backwards, line by line, and both of those facts come off the same page.
1The document, line by line
Dell publishes AI Factory solutions as short component lists with a solution ID. Solution 19845005.1 is one page and every line on it is a design decision.
Compute is “3x Dell PowerEdge R660” and “2x Dell PowerEdge XE9680” with “Dual Intel Platinum CPU/Total 64c/2.112TB H200Srvr RAM” and “16x NVIDIA H200SXM GPUs (Total 2256GB)”.[1] Adapters are one line: “NVIDIA Bluefield-3 (1x3220/8x3140H)”.[1] Fabric is three lines: “3x NVIDIA Spectrum-4 SN5600 400G BE + 200G FE”, “2x PowerSwitch S5232F-ON Storage Cluster 100G BE” and “1x NVIDIA SN2201 ToR”.[1] Storage is “3x PowerScale F710 (Total Storage 230 TB)”.[1] Rack and power are “1x APC 750x1200 42U Rack (A6921473)” and “4x PDU (AC021024)”, under the caveat “Actual rack configurations will vary based on power per rack. This solution is presented in a rack that can accommodate approximately 48/2kWs.”[1] Software is “5x Ubuntu OS”, “Upstream Kubernetes” and “19x NVAIE” with “(NVIDIA NeMo Microservices for Generative AI, NIMs, RAG Systems, BCMe)”.[1] Services are “ProSupport+ or ProSupport One”, “ProDeploy” and “ProConsult”.[1]
Five Ubuntu licences for five physical hosts - two XE9680 and three R660 - is the first small consistency check that passes.[1] The “19x NVAIE” line against sixteen GPUs is the first one that does not reconcile on the page, and the honest move is to say so rather than to invent a mapping.[1]
| Product | Speed | PCIe | Role | GPU generation |
|---|---|---|---|---|
NIC | ||||
SuperNIC (no Arm) | ||||
SuperNIC (no Arm) | ||||
DPU | ||||
DPU | ||||
SuperNIC (Arm inactive) | ||||
DPU / storage processor | ||||
Ethernet switch | ||||
Ethernet switch | ||||
InfiniBand switch | ||||
InfiniBand switch |
⚠ = not confirmed on a fetched primary source (hover for why). Facts as of DOCA 3.5.0 (Sep 2026). Selections are saved.
2Reading the NIC line backwards
The Enterprise RA whitepaper names node types with the pattern CPUs-GPUs-NICs-speed. It defines “PCIe-Optimized 2-4-3-200” as “2 CPUs balanced with up to 4 GPUs plus 3 NICs with East-West traffic of 200 GbE per GPU” for 8 to 32 nodes, “PCIe-Optimized 2-8-5-200” for 4 to 32 nodes, “HGX 2-8-9-400” as “2 CPUs balanced with 8 GPUs plus 9 NICs with East-West traffic of 400 GbE per GPU” for 4 to 32 nodes, and “Scale-out Grace 2-2-3-400”.[2]
Now apply it. Two Intel Platinum CPUs per XE9680, eight H200 GPUs per XE9680, and per server one B3220 plus eight B3140H equals nine adapters.[1][9] Two, eight, nine - the node is an HGX 2-8-9-400.[2] The nine splits by role exactly as the H100/H200/B200 RA describes: “BlueField-3 B3140H SuperNICs” connected to GPUs for the compute fabric, and “BlueField-3 B3220 DPUs” in compute nodes for the converged network at “two 200 GbE ports to two separate switches”.[4]
Two consequences follow immediately, and both are the kind of thing a customer will thank you for.
First, storage certification. The whitepaper states two certification levels, “Foundation” certifying 2-4-3-200 and 2-8-5-200 and “Enterprise” certifying 2-8-9-400.[2] A 2-8-9-400 deployment therefore points at the Enterprise level, and the certified-storage page defines the programme as “a comprehensive validation framework that ensures storage platforms meet stringent performance, quality, and interoperability standards for AI workloads” without naming protocols or partners.[5] So the right sentence is “this node type maps to the Enterprise certification level; ask which certification the storage line carries”, not “PowerScale is or is not certified”.
Second, scale. The whitepaper scopes HGX 2-8-9-400 to 4 to 32 nodes and this solution has two.[2][1] The node shape is right and the deployment is smaller than the published range. Say both.
3Three switch lines, three fabrics
Each switch line answers a different question, and mixing them up is the most common misreading of an AI quote.
“3x NVIDIA Spectrum-4 SN5600 400G BE + 200G FE” is the AI fabric, and one switch model is carrying both the back end at 400G and the front end at 200G.[1] That is possible because the SN5600 presents 64 OSFP cages that Dell prints as “64x 800 GbE OSFP” and NVIDIA counts as 256 ports of 200GbE, with a 160 MB fully shared packet buffer behind them.[6][7] The RA calls these roles the compute (east-west) fabric and the converged (north-south) fabric, and on this generation the converged side is 2x200 GbE per node - which is exactly what the 200G FE half of the line is terminating.[4]
“2x PowerSwitch S5232F-ON Storage Cluster 100G BE” is not the GPU-to-storage path. PowerScale requires a private back-end network “configured with redundant switches for high availability” that acts as “the backplane for the PowerScale cluster”, and “Dell does not support connecting any other devices to the back-end switches”.[8] Two switches, redundant, dedicated - the BOM line is that requirement, priced.
“1x NVIDIA SN2201 ToR” is the out-of-band network, matching the RA pattern of a network that “provides bulk management 1Gb RJ45 connectivity for all the Nodes” and is served by the “NVIDIA SN2201 48-port 1 Gb plus 4-port 100 GbE” switch.[11] The physical topology says what it reaches: it “connects all the base management controller (BMC) ports” plus “the 1 GbE switch management ports and the BlueField-3 SuperNIC and DPU management ports”.[4] One caution the RA itself implies: a single out-of-band switch is a single point of failure for management access, which is a design conversation the sheet does not have.
4Re-sizing the same solution
The value of reading a BOM backwards is that you can then run it forwards at a different size. Use the sizing chain from module 2 with this solution’s ratios: eight GPUs per node, one east-west adapter per GPU, one north-south DPU per node.[1][4]
GPUs → NICs
16 x BlueField-3 B3140H SuperNIC
This is the generation in the Dell AI Factory BOM (Solution ID 19845005.1): "NVIDIA Bluefield-3 (1x3220/8x3140H)" against 2x XE9680 with 16x H200 SXM. Converged is "two 200 GbE ports to two separate switches" — half the B300 figure. The DPU is the north-south NIC and is counted separately from the east-west SuperNICs — that is where the "9" in HGX 2-8-9-400 (Hopper shape: 1x B3220 + 8x B3140H) comes from.
The nomenclature string is the fastest audit of a quote: CPUs-GPUs-NICs-speed. If the NIC line does not read 2-8-9-400 on a Hopper or B300 node, something is missing — usually the north-south DPU.
16 GPUs x 1 NIC per GPU = 16 NICs, plus 1 DPU per node (2 DPUs)
HGX B300 ComponentsNode shape and SU arithmetic for this preset: H100/H200/B200 RA · nomenclature (HGX 2-8-9-400 (Hopper shape: 1x B3220 + 8x B3140H)) · B300 node configurations · SuperPOD SU and rack power · NVL72 networking hardware
At 16 GPUs the fabric is small enough that switch count is driven by roles rather than by radix: three fabrics need at least three switch groups regardless of port count.[1] As GPU count rises, the east-west interface count is the first thing to move, then leaves per plane, then spines. What does not move is the per-node shape - nine adapters - because that is fixed by the node type.[2]
Ninety minutes later
The SE walks into the call with sentences instead of a hope: the node shape is HGX 2-8-9-400 exactly, and two nodes sits below the 4 to 32 node range the whitepaper prints for that node type - both true, both cited.[2][1] What you actually say: “It is the published node shape, under the published node count. Say both and nobody gets to correct you later.”
Procurement has been reading the same page for a different reason, and she has found her saving. The sheet buys two switches purely for the PowerScale storage cluster network.[1] She wants to know why those cannot ride the new SN5600 pair. The storage architect is dialled in tomorrow.
Given. Dell Solution 19845005.1 as published.[1]
Step 1 - roles. Rewrite the sheet as a role table. Compute: 2x XE9680, 16 GPUs. Control plane: 3x R660. East-west: 8x B3140H per node. North-south: 1x B3220 per node. AI fabric: 3x SN5600 (400G back end, 200G front end). Storage cluster back end: 2x S5232F-ON at 100G. Out-of-band: 1x SN2201. Storage: 3x PowerScale F710, 230 TB. Rack and power: 1 APC 42U, 4 PDUs, approximately 48/2kWs. Software: 5x Ubuntu, upstream Kubernetes, 19x NVAIE. Services: ProSupport+ or ProSupport One, ProDeploy, ProConsult.[1]
Step 2 - node type. 2 CPUs, 8 GPUs, 9 NICs per XE9680 gives HGX 2-8-9-400.[2] Generation confirmed by the parts themselves: B3140H east-west and B3220 north-south is the H100/H200/B200 RA pairing.[4]
Step 3 - annotate each line with the RA statement it satisfies. East-west adapters: one per GPU, 400 GbE per GPU per the node type.[2] North-south: “two 200 GbE ports to two separate switches” per node.[4] Out-of-band: 1 Gb RJ45 for all nodes on SN2201.[11] Storage cluster back end: PowerScale’s private redundant back end with no other devices attached - a Dell requirement, not an NVIDIA RA statement.[8] Rack, PDUs, Ubuntu, Kubernetes, NVAIE quantity and the services lines: no RA statement covers them; they are Dell packaging and commercial choices.
Step 4 - findings worth raising. (a) Two nodes is below the 4 to 32 node range published for 2-8-9-400.[2] (b) 19 NVAIE against 16 GPUs is unexplained on the sheet.[1] (c) One SN2201 means a single out-of-band path.[1] (d) The 2-8-9-400 node type maps to the Enterprise storage certification level, so ask what the F710 line carries.[2][5] (e) Storage bandwidth at 16 GPUs is about 200 Gb/s, and that one is quoted rather than derived - the certified-storage page prints the example “a 16 GPU cluster would require around 200 Gb/s of aggregate storage bandwidth”.[5] Any other cluster size is your multiplication of the per-GPU rule, and is labelled as such.
Same method, new sheet. A Dell SE sends you a quote with 4x XE9680, 32 H200 GPUs, “NVIDIA Bluefield-3 (1x3220/8x3140H)” per server, 4x SN5600, 2x S5232F-ON, 1x SN2201, 4x PowerScale F710 and the same rack note. Fill the blanks.
- GPUs per node: ____. Adapters per node: ____ east-west plus ____ north-south. Node type: ____.[2]
- Is this deployment inside the published node-count range for that node type? ____ Evidence: ____.[2]
- East-west interfaces at one 400G interface per GPU, single plane: ____ interfaces. Show the multiplication and mark it derived.[2]
- Aggregate storage bandwidth by the published per-GPU rule: ____ Gb/s. Source: ____.[5]
- Which line would you question first, and what exactly would you ask? ____.
- Which lines have no RA statement behind them at all? ____.[8]
A Dell SE writes: “Customer liked 19845005.1 but wants 32 GPUs, and their facilities team says 30 kW per rack maximum. Can we just double it?”
Produce a written answer containing: (a) the re-sized role table with every quantity that changes and every quantity that does not; (b) the node type and whether 32 GPUs moves the deployment into or out of the published node-count range, with the citation; (c) the derived east-west interface count and the derived storage bandwidth, both labelled derived; (d) an explicit statement about the rack, comparing the sheet’s “approximately 48/2kWs” note against the customer’s 30 kW ceiling and saying what that means for the number of racks; (e) the two lines you would not double and why; and (f) the three questions you would send back to the account before anyone quotes.
Acceptance criteria: every published ratio carries a source; every multiplication is marked as your arithmetic; no claim that a specific storage product is or is not NVIDIA-certified; and the answer names at least one thing the solution sheet does not publish at all.[1][2][5][8]
Lab
Goal: produce the BOM row for the Dell lab’s own kit, using the same discovery you would run at a customer site. Read-only - no firmware updates, no configuration changes.
- Pre-flight inventory per host: model, service tag, CPU and memory.
Expected: model and service tag that match the asset label. If not:racadm getsysinfo sudo dmidecode -t system | head -12 free -gracadmis not installed on the host - read the same fields from the iDRAC web UI under Dashboard, or runsudo dmidecode -s system-serial-numberandsudo dmidecode -s system-product-name. Record them - the same values open a support case in lesson 5. - Enumerate accelerators if the host has any:
Expected: one line per GPU. If the command is absent the host is a control-plane-class node in BOM terms, not a compute node.nvidia-smi -L - Enumerate the adapters and their exact part identity:
Expected: part number, PSID, description and firmware per device. If not:sudo mst start && sudo mst status -v sudo mlxfwmanager --querymst status -vprints nothing when the mst driver is not loaded - runsudo mst startagain and checklsmod | grep mst; ifmlxfwmanageris missing, install themftpackage or read the part number fromsudo lspci -vv. Write the adapters in BOM notation - for example “NVIDIA Bluefield-3 (1x3220)” - so the row is directly comparable to the published sheet.[1] - Record link speed and cage type per port:
Expected: negotiated speed and the transceiver identifier. If not:sudo ethtool ens1f0np0 | head -12 sudo ethtool -m ens1f0np0 | head -12ethtool -mreturns “Cannot get module EEPROM information” on a port with a DAC or an empty cage - record “no module” or the copper cable identifier, which is itself a BOM fact. - Measure the power envelope. Read the per-host power draw from iDRAC (System, Power, Power Monitoring, or
racadm getsysinfooutput) and multiply by the number of hosts in your rack. Compare the total against the solution sheet’s “approximately 48/2kWs” rack note and state in one sentence whether that rack could hold the quoted solution as-is.[1] Expected: a number and a yes or no with the assumption named. If not: if iDRAC shows no power monitoring on this platform, use the PSU nameplate rating instead and say in the row that you used nameplate rather than measured draw - the two differ by a lot and the difference is the finding. This multiplication is yours, not Dell’s. - Deliverable: one BOM-format row for the lab - server model, accelerator count, adapter parts, switch, storage, rack power - plus the sentence from step 5.
Goal: turn the published solution into a role table you can defend, then re-size it.
- Rebuild the BOM as a table with columns: line as printed, quantity, role, RA statement it satisfies or “none”.[1][2][4][11] Expected: every network line has an RA statement; rack, PDU, OS, Kubernetes, NVAIE and services lines have none. That asymmetry is the point. If not: if you find an RA statement for the rack or the services lines, check that you are not quoting a general design principle as if it governed that line - “none” is the correct entry and it is what makes the network rows meaningful.
- Write two sentences explaining why one SN5600 model can serve both a 400G back end and a 200G front end, referring to cages and breakout rather than to throughput.[6][7] Expected: 64 OSFP cages, breakout to many 400G and 200G interfaces, one SKU and one spare - with the shared-maintenance caveat named. If not: if your two sentences reach for the 51.2 Tb/s figure, rewrite them - throughput never explains how many interfaces a switch can present.
- Open the FabricSizer above at this solution’s numbers and record the six computed step cards. Then set GPUs to 32 and record the same six.[2] Expected: a list of exactly which values changed. Mark each changed value as published ratio or derived arithmetic. If not: if the per-node adapter count changes when you raise the GPU count, you have altered a ratio rather than the size - reset the sizer to 8 GPUs per node and one NIC per GPU.
- Compute the storage bandwidth implied at 16 and at 32 GPUs using the published per-GPU rule, and write the sentence you would use with a customer that makes clear which part is published and which is your multiplication.[5] Expected: about 200 Gb/s at 16 GPUs - quoted, because the page prints that example - and about 400 Gb/s at 32 GPUs, labelled derived. If not: if you cannot find the 16-GPU example on the page, do not fall back to hedging the number - re-read the sizing paragraph, because quoting a published figure as your own arithmetic understates what you can defend.
- List every finding you would raise on the original sheet - node count below range, unexplained NVAIE quantity, single out-of-band switch, storage certification level to confirm.[1][2][5] Expected: four to six findings, each phrased as a question to the account team rather than as a criticism. If not: if a finding asserts that the BOM is wrong, rewrite it - every item on this list is a thing the sheet does not say, which is a question, not a defect.
- Deliverable: the annotated role table plus a one-paragraph answer to “can we just double it”. If not: if your paragraph doubles every line, check the rack and the out-of-band switch first - those are the two that do not scale by the same factor.
Retrieval check
10 questions from memory. Answer before looking anything up; misses become flashcards.
Explain it to a Dell SE
Explain to a Dell SE in five sentences how you can tell which reference architecture a quote implements by looking only at its network adapter line.
Sources
Facts in this lesson were checked against Dell AI Factory with NVIDIA Solution ID 19845005.1 re-fetched and read in full 2026-09-07; NVIDIA Enterprise RA whitepaper, HGX AI Factory components and H100/H200/B200 topologies, 2026-09-07. Dates are when each page was fetched.
- Dell AI Factory with NVIDIA Solution ID 19845005.1 · fetched 2026-09-07
- Key building blocks of Enterprise Reference Architectures · fetched 2026-09-07
- Components - NVIDIA HGX AI Factory Enterprise RA · fetched 2026-09-07
- Networking Physical Topologies - NVIDIA HGX AI Factory (H100 H200 B200) · fetched 2026-09-07
- NVIDIA-Certified Storage - NVIDIA HGX AI Factory · fetched 2026-09-07
- AI Networking Switches | Dell USA · fetched 2026-09-07
- NVIDIA Spectrum SN5600 Series Switches Datasheet (Dell-hosted) · fetched 2026-09-07
- Dell PowerScale: Ethernet Back-End Network Overview (H16346.8) · fetched 2026-09-07
- PowerEdge XE9680 | Dell USA · fetched 2026-09-07
- Networking Hardware - NVIDIA HGX AI Factory Enterprise RA · fetched 2026-09-07
- Networking Logical Architecture - NVIDIA HGX AI Factory Enterprise RA · fetched 2026-09-09
The same idea elsewhere
Other lessons that cover this ground, sometimes from another course's angle.
- Scenario: a 256-GPU Dell AI Factory design reviewElsewhere in this course · Same ground: powerscale, converged and derived
- Reading node nomenclature: 2-8-9-400Elsewhere in this course · Same ground: nomenclature, storage and bom
- Optics, cables and the power nobody budgetedElsewhere in this course · Same ground: bom, derived and power