The sizing chain: GPUs, NICs, planes, leaves, spines
S2·E3The quote that fit in half a chassis · A Dell deal desk review, the morning before it goes to the customer
Builds on: Dual plane: two fabrics, not a bond
Before you read: what do you already know?
3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.
After this lesson you can
- Compute server-facing interfaces leaves per plane and spines from a GPU count using the published ratios.
- Separate published ratios from your own arithmetic at every step and say which is which out loud.
- Convert switch cage counts into logical port counts for SN5600 SN5610 and QM9700 before quoting a leaf count.
- Check your own leaf and spine numbers against the one published spine-scaling table the RAs give you.
Episode 3 — The quote that fit in half a chassis
Phase two comes back as a 512-GPU quote with a switch list and a price procurement calls very competitive, which is usually the first warning. Procurement wants two numbers, price and lead time, and both improve when the switch count is small. The SE has the draft open, the coffee is cold again, and the promise spreadsheet has a fresh row in it.
Two lines are wrong and both happen to flatter the price. “QM9700, 64 transceivers”: a QM9700 presents 32 OSFP cages carrying 64 ports of NDR 400 Gb/s, so 64 transceivers is twice what physically fits in the chassis.[4] And the leaf count is half of what the design needs, because the interfaces were counted once per GPU instead of once per plane.[2]
Both errors come from one habit — reading logical ports as if they were cages. An SN5600 presents 64 cages that become 128 logical interfaces of 400 GbE.[3] That is why the sizing chain exists as a chain rather than an answer. A switch count cannot be checked by the person receiving it, but every multiplication can: GPUs to SuperNICs at the published 1:1 ratio,[1] NIC ports to server-facing interfaces through the plane choice,[2] then interfaces to leaves through a divisor you have to admit is your own.
Transceivers go in cages, leaf arithmetic runs on logical ports, and a quote that mixes the two is wrong in both directions at once.
Segment 1 says the chain in the order you would speak it.
1The chain, said out loud
Customers ask for a switch count. The better answer is the chain, spoken in front of them with their own GPU number in it, because every line is checkable:
GPUs x NICs-per-GPU x planes = server-facing interfaces. Server-facing interfaces per plane ÷ leaf downlink radix = leaves per plane. Leaf uplinks ÷ spine radix, at the chosen oversubscription = spines per plane.
Each of the first three factors is published. The HGX B300 RA fixes NICs per GPU at 1:1 — “Eight NVIDIA ConnectX-8 SuperNICs per NVIDIA HGX B300 baseboard”, “Up to 800 Gbps per adapter”, “800 Gb/s (2 x 400Gb/s Ethernet) per GPU”.[1] The plane factor is fixed by the topology choice: dual plane breaks each interface “to 2x400 Gb/s interfaces” that land on different leaves in independent fabrics, and single plane does not.[2]
The divisor is where you have to be careful, and it is the subject of the next segment. Everything after that divisor is arithmetic you own.
GPUs → NICs
256 x ConnectX-8 SuperNIC
"Eight NVIDIA ConnectX-8 SuperNICs per HGX B300 baseboard", "800 Gb/s (2 x 400Gb/s Ethernet) per GPU" — a 1:1 GPU-to-SuperNIC ratio. The "9" in 2-8-9-400 is 8 east-west SuperNICs + 1 north-south BlueField-3 B3240. The DPU is the north-south NIC and is counted separately from the east-west SuperNICs — that is where the "9" in HGX 2-8-9-400 comes from.
The nomenclature string is the fastest audit of a quote: CPUs-GPUs-NICs-speed. If the NIC line does not read 2-8-9-400 on a Hopper or B300 node, something is missing — usually the north-south DPU.
256 GPUs x 1 NIC per GPU = 256 NICs, plus 1 DPU per node (32 DPUs)
HGX B300 ComponentsNode shape and SU arithmetic for this preset: HGX B300 RA · nomenclature (HGX 2-8-9-400) · B300 node configurations · SuperPOD SU and rack power · NVL72 networking hardware
2Cages are not ports
The number that decides a leaf BOM is cages, not terabits. An SN5600 or SN5610 has 64 OSFP cages that present 128 logical interfaces of 400 GbE, 51.2 Tb/s in 2U, with a 160 MB fully shared packet buffer.[3][7] A QM9700 has 32 OSFP cages carrying 64 ports of NDR 400 Gb/s at 25.6 Tbps.[4] An SN5400 is 64 x 400 GbE or 128 x 200 GbE at 25.6 Tb/s.[7]
Transceivers go in cages. So a quote line reading “64 transceivers, QM9700” has ordered twice what physically fits, and a quote reading “128 transceivers, SN5600” has done the same.[4][3] Dell’s own catalog reinforces the distinction by printing both numbers together: “Q3401-RD and Q3400-RA: 144 x 800Gb/s over 72 OSFP cages” and “QM9790, QM9701, QM9700: 64x 400 Gb/s over 32 OSFP cages”.[10]
The second trap is the one nobody prints. The RAs do not publish how many of a leaf’s logical ports face servers and how many face spines. The familiar non-blocking split — half down, half up, so 64 downlinks and 64 uplinks on a 128-port SN5600 leaf — is arithmetic on the datasheet radix, not an NVIDIA statement.[3] Use it, but label it every time.
3Checking your own numbers against the one published table
There is exactly one place where an RA prints the output of this chain rather than its inputs: the GB200 NVL72 SuperPOD InfiniBand spine-scaling table. It has four rows and it starts at two SUs — 1,152 GPUs is 2 SU with 6 core groups, 32 IB leaves and 24 IB spines per SU, so 64 leaves and 48 spines in total; 2,304 GPUs is 4 SU and 96 spines; 4,608 GPUs is 8 SU and 192; 9,216 GPUs is 16 SU with 6 core groups, 24 switches per core group, 32 IB leaves and 24 IB spines per SU, so 512 leaves and 384 spines in total.[5] Every row of the table uses 6 core groups — what scales with SU count is switches per core group, not the number of core groups. Every row also repeats the same per-SU figures, 32 leaves and 24 spines, and that is where the single-SU number comes from: the RA prints no 576-GPU row, but it does say each SU contains 4 SLGs of 8 leaf and 6 spine switches, so one SU is 24 spines by arithmetic on published text rather than by a printed row.[5] Four rows plus one derived case is the whole table. The ratio looks linear, which is exactly why inventing a fifth row is dangerous.
The same RA also explains the shape of the fabric it is counting. Each SU “contains 4 SLGs to match with the number of IB rails (which equals the number of GPUs per compute tray)”, and each spine-leaf group has 8 leaf switches and 6 spine switches for a fully non-blocking fat tree.[5] Put that next to the QM9700 radix and the reason becomes visible: 32 cages give 64 NDR400 ports, so a leaf that spends all 64 on one rail’s downlinks has zero cages left for uplinks.[4] The SLG exists to solve that. (The zero-uplink observation is arithmetic on the published radix; the RA does not print it.)
Scale anchors from the Ethernet side complete the picture. The Spectrum-4 and DC-busbar DGX B300 RA sizes an SU at 64 nodes and 512 GPUs, four SUs at 256 nodes and 2,048 GPUs, and 64 SUs at over 2,000 nodes.[8] The Quantum-X800 edition of the same server sizes an SU at 72 nodes and 576 GPUs.[9] When someone says “per the reference architecture”, the SU size tells you which document they are holding.
4The two things that quietly break the arithmetic
Both traps have already appeared, and both are worth naming as a pair because they are the two questions to ask on any quote before you compute anything.
Cages versus logical ports. Ask which unit each switch line is counted in. The tell is a transceiver quantity that matches the logical port count instead of the cage count.[4][3]
Dual versus single plane. The plane factor is a x2 on server-facing interfaces and therefore on leaves. It is visible in the optics line as twin-port versus single-port OSFP, and the RA fixes single plane at exactly 50 percent of the GPU bandwidth.[2] A quote that assumes single plane and a drawing that shows dual plane differ by a factor of two in leaf count.
Use the sizer below to feel the size of each error. Change planes from 2 to 1 and watch the leaf count halve; change the leaf SKU from sn5600 to qm9700 and watch it change again because the radix moved from 128 logical ports to 64.[3][4]
⚠ QM9700 (InfiniBand NDR) is InfiniBand; the HGX B300 Enterprise RA specifies Spectrum Ethernet for the compute fabric. Quote it only if the customer chose InfiniBand deliberately.
GPUs → NICs
512 x ConnectX-8 SuperNIC
"Eight NVIDIA ConnectX-8 SuperNICs per HGX B300 baseboard", "800 Gb/s (2 x 400Gb/s Ethernet) per GPU" — a 1:1 GPU-to-SuperNIC ratio. The "9" in 2-8-9-400 is 8 east-west SuperNICs + 1 north-south BlueField-3 B3240. The DPU is the north-south NIC and is counted separately from the east-west SuperNICs — that is where the "9" in HGX 2-8-9-400 comes from.
The nomenclature string is the fastest audit of a quote: CPUs-GPUs-NICs-speed. If the NIC line does not read 2-8-9-400 on a Hopper or B300 node, something is missing — usually the north-south DPU.
512 GPUs x 1 NIC per GPU = 512 NICs, plus 1 DPU per node (64 DPUs)
HGX B300 ComponentsNode shape and SU arithmetic for this preset: HGX B300 RA · nomenclature (HGX 2-8-9-400) · B300 node configurations · SuperPOD SU and rack power · NVL72 networking hardware
Customer: 32 HGX B300 nodes, dual plane, SN5600-class leaves and spines, non-blocking. Say the chain out loud.
- GPUs. 32 nodes x 8 GPUs per node = 256 GPUs. The node shape is published: eight B300 GPUs on an HGX B300 baseboard.[1]
- NICs. 1:1 with GPUs — “Eight NVIDIA ConnectX-8 SuperNICs per NVIDIA HGX B300 baseboard”, up to 800 Gbps per adapter. So 256 SuperNIC ports at 800 Gb/s.[1]
- Server-facing interfaces. Dual plane breaks each 800 Gb/s port to 2x400 Gb/s onto two leaves in two independent fabrics: 256 x 2 = 512 x 400 Gb/s, i.e. 256 per plane.[2] (Ratios published; multiplication mine.)
- Sanity check against the published cap. Each plane “scales to 1024 interfaces of 400 Gb/s”; 256 is well inside it, so one plane pair covers this cluster with room to grow.[2]
- Leaves. SN5600 radix is 64 cages / 128 logical 400 GbE.[3] Assume a non-blocking split of 64 downlinks + 64 uplinks — this is my arithmetic, not an NVIDIA figure. 256 ÷ 64 = 4 leaves per plane, 8 leaves total.
- Spines. Uplinks per plane = 4 leaves x 64 = 256 x 400 Gb/s. On SN5600-class spines with 128 x 400 GbE each, at 1:1 non-blocking: 256 ÷ 128 = 2 spines per plane, 4 spines total, each leaf carrying 32 links to each spine.[3] (Derived.)
- State the answer with its labels. “Eight leaves and four spines, on the assumption of a half-and-half leaf split. NVIDIA publishes the 1:1 GPU-to-NIC ratio, the 2x400 breakout, and the 128-port leaf radix. The split and everything downstream of it are mine — change the split and the numbers move.”
- Cross-check the shape. Compare against the GB200 single SU (576 GPUs, 24 IB spines across 4 SLGs, a figure derived from the per-SU column rather than printed as a row) — different platform and different radix, so expect different absolute numbers, but a leaf-to-spine ratio in the same neighbourhood.[5]
Same customer decides to save money and go single plane on the same 32 nodes.
- GPUs = 32 x ____ = ____ .[1]
- SuperNIC ports = ____ at a ____ ratio to GPUs.[1]
- Single plane means each GPU takes ____ x 400 Gb/s into ____ fabric, so server-facing interfaces = ________ .[2]
- Compare to the per-fabric cap of ____ interfaces of 400 Gb/s.[2]
- Leaves = ____ ÷ ____ = ____ , where the divisor is ______________ and is labelled ____________ .[3]
- Spines at 1:1 = ____ .
- In one sentence: what did the customer save in switches, and what did the RA say they gave up in bandwidth?[2]
A Dell account is quoting a 72-node cluster of the same HGX B300 nodes and wants a 3:1 oversubscribed fabric to save on spines, still dual plane, still SN5600-class.
Produce, showing every step:
- Server-facing 400 Gb/s interfaces in total and per plane.
- Leaves per plane under a 3:1 oversubscription — state explicitly what you assumed about the downlink/uplink split and why 3:1 changes it.
- Spines per plane and the total switch count.
- One sentence naming which numbers came from a published page and which are yours.
- One risk sentence about what 3:1 does to a rail-optimized collective.
Acceptance criteria: every published figure carries its source; every derived figure is marked; the oversubscription assumption is stated as a number before it is used.
Episode 3 — Case note: two lines, one label
The transceiver quantity comes down to what 32 cages can physically hold,[4] and the leaf count doubles once the plane factor is applied to the server-facing interfaces instead of being assumed away.[2] The half-and-half leaf split stays in the document, now written as an assumption with its number stated before it is used.[3] What you say to the deal desk: “Every figure in here is either quotable from a page or labelled as my arithmetic, and the customer is welcome to argue with the second kind.” Procurement gets its lead time. Nobody remarks that the quote still contains exactly one network — and the pod ships.
Lab
Count physical things in the Dell lab and prove the cage-to-port rule to yourself. Read-only.
- Pre-flight inventory. List what switches and hosts you actually have and their optics:
Record the count of physical cages on the front panel of whatever switch is present — count them by eye, and photograph the panel.ip -br link ls /sys/class/net - Prove one cage carries two logical ports on the host side. On a host with an 800G-class port:
Expected on a broken-out cage: two netdevs at 400 Gb/s each rather than one at 800. If you see one netdev at the full rate, the cage is not broken out and the host will present a single interface — which is the configuration that loses adaptive routing on Spectrum-4.ibdev2netdev ethtool <iface> | grep -E 'Speed|Supported link modes' - Read the optics identity in the same breath.
Expected: vendor part number and media type. Note whether the part is a twin-port or single-port class part — that decides whether a break-out is even physically possible on this link.sudo ethtool -m <iface> | head -20 - Do the arithmetic against what you counted. Take the cage count you photographed, double it for a 2x400 break-out, and compare against the datasheet logical-port figure for that model.[3][4] Any mismatch is a model identification error, not a maths error.
- Write the two-line takeaway. One line for the cage-to-logical-port ratio you measured, one line for the transceiver class you found. These are the two lines you will say on a customer call.
- Optional, customer lab only. Count cages on a real SN5600 or QM9700 and reconcile them against the quote’s transceiver quantities. Read-only; change nothing on a production switch.
Build the model, then try to break it against the one published table.
- Write the model. A small script or spreadsheet taking GPU count, GPUs per node, NICs per GPU, planes, leaf logical-port radix and downlink fraction; emitting NIC ports, server interfaces total and per plane, leaves per plane, total leaves and spines.
Expected: 512 interfaces, 256 per plane, 4 leaves per plane, 8 leaves, 2 spines per plane, 4 spines.mkdir -p ~/projects/ra-sizing && cd ~/projects/ra-sizing cat > size.py <<'PY' def size(gpus, gpus_per_node, nics_per_gpu, planes, leaf_logical, downlink_frac, spine_logical, oversub=1.0): nodes = gpus / gpus_per_node nic_ports = gpus * nics_per_gpu ifaces = nic_ports * planes # 400G server-facing interfaces per_plane = ifaces / planes downlinks = leaf_logical * downlink_frac leaves_per_plane = -(-per_plane // downlinks) # ceiling uplinks = leaves_per_plane * (leaf_logical - downlinks) / oversub spines_per_plane = -(-uplinks // spine_logical) return dict(nodes=nodes, nic_ports=nic_ports, ifaces=ifaces, per_plane=per_plane, leaves_per_plane=leaves_per_plane, leaves=leaves_per_plane*planes, spines_per_plane=spines_per_plane, spines=spines_per_plane*planes) PY python3 -c "import size; print(size.size(256,8,1,2,128,0.5,128))" - Validate against the only published table. Feed the GB200 shape and compare to the RA’s rows of 1,152 GPUs to 48 total IB spines, 2,304 to 96, 4,608 to 192 and 9,216 to 384, and to the per-SU figures every row repeats - 32 leaves and 24 spines.[5] Record where your model disagrees rather than tuning it silently until it matches.
- Explain the disagreement. The GB200 fabric is built as spine-leaf groups — 8 leaf plus 6 spine per SLG, 4 SLGs per SU — not as a flat leaf-and-spine tier, so a flat model should be expected to differ.[5] Write one sentence naming the structural reason.
- Run the 32-node HGX B300 worked example through the script and confirm you reproduce 512 server-facing interfaces and 256 per plane.[1][2]
- Mark every derived number. In the script’s output, tag
leaves_per_plane,spines_per_planeand anything computed fromdownlink_fracas derived. The RAs do not print a per-leaf split.[3] - Cross-check the radix constants you hard-coded: SN5600 128 x 400 GbE from 64 cages, QM9700 64 x NDR400 from 32 cages.[3][4] A wrong constant here is invisible and poisons every answer.
Retrieval check
10 questions from memory. Answer before looking anything up; misses become flashcards.
Explain it to a Dell SE
A customer asks how many switches they need for 512 GPUs. Explain in five sentences the chain you would say out loud and which single number in it is yours rather than NVIDIA's.
Sources
Facts in this lesson were checked against HGX AI Factory RA Components and Networking Physical Topologies re-fetched 2026-09-07; QM97XX Specifications re-fetched 2026-09-07 (32 OSFP cages 25.6 Tbps); Cumulus Linux 5.13 ECMP/adaptive routing re-fetched 2026-09-07; GB200 SuperPOD Network Fabrics and DGX SuperPOD B300 architecture pages as recorded in content/research/ra/part1.md 2026-09-07. Dates are when each page was fetched.
- Components — NVIDIA HGX AI Factory (B300) Enterprise RA · fetched 2026-09-07
- Networking Physical Topologies — NVIDIA HGX AI Factory (B300) Enterprise RA · fetched 2026-09-07
- NVIDIA Spectrum SN5600 Series Switches Datasheet (Dell-branded) · fetched 2026-09-07
- Specifications — QM97XX 1U NDR 400Gbps InfiniBand Switch Systems User Manual · fetched 2026-09-07
- Network Fabrics — DGX GB200 NVL72 SuperPOD Reference Architecture · fetched 2026-09-07
- Networking Physical Topologies — NVIDIA NVL72 AI Factory (GB300) Enterprise RA · fetched 2026-09-07
- NVIDIA Spectrum Ethernet Switches — product family port matrix · fetched 2026-09-07
- DGX SuperPOD Architecture — DGX B300 Spectrum-4 Ethernet and DC Busbar Power RA · fetched 2026-09-07
- DGX SuperPOD Architecture — DGX B300 with Quantum-X800 InfiniBand and AC Power RA · fetched 2026-09-07
- AI Networking Switches — Dell USA catalog · fetched 2026-09-07
The same idea elsewhere
Other lessons that cover this ground, sometimes from another course's angle.
- Topologies: fat-tree, rail-optimized, twin-planeInfiniBand course · Same ground: slg, gb200 and sizing
- Design review: an eight-rail IB pod on Dell XE9680InfiniBand course · Same ground: radix, superpod and sizing
- Dell's switch catalog: SN, Q and Z in one price listElsewhere in this course · Same ground: Cages versus logical ports, Scale-out and quoting