Skip to content

Dell AI Factory scenarios and the NCA-AIIO positioning

S4·E5The numbers you are allowed to say out loud · A Dell account meeting the week after acceptance, a deck due Friday

S4·E5Create~35 minsources checked todayverified against Dell AI Factory with NVIDIA 2-8-9-400 solution brief (PDF text extracted 2026-09-09); Dell Open Ethernet for AI blog re-fetched 2026-09-09; NVIDIA Enterprise Reference Architectures index re-fetched 2026-09-09; NCA-AIIO certification page re-fetched 2026-09-09; NVIDIA Network Operator v26.7.0 and GPU Operator platform support re-fetched 2026-09-09

Builds on: Spectrum-X rails inside Kubernetes, BlueField in Kubernetes: DPF versus Network Operator, The triage ladder: sos-report, error strings and the checklists

Before you read: what do you already know?

3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.

After this lesson you can

  • Compose a qualification answer to 'which CNI do you support?' that is two-layer and entirely made of published facts.
  • Compose an escalation intake form that collects the four versions and the fabric evidence before any theory is offered.
  • Compose a positioning answer that separates Network Operator from DPF and converged from split fabric without asserting an uncited design.
  • Calibrate the NCA-AIIO claim honestly against the published blueprint instead of marketing this course as exam preparation.

Episode 5 — The numbers you are allowed to say out loud

The situation · A Dell account meeting the week after acceptance, a deck due Friday

Acceptance is signed and the account team wants the next twelve nodes on a slide. Procurement is here for two questions, price and lead time on switches. The Dell SE has the promise spreadsheet open at the summary tab and is on his second cold coffee. An NVIDIA PM dialled in for the roadmap half and has already used his phrase twice: not announced.

Half of what this room asks can be quoted exactly and half cannot, and knowing which half is which is the job. Dell publishes a solution brief for the XE9680-based AI Factory endorsed by NVIDIA Enterprise Reference Architectures and aligned with the 2-8-9-400 configuration: four PowerEdge R670 management servers, up to twelve XE9680 workers, two Spectrum-4 SN5610 carrying the converged network, one Spectrum SN2201 for out-of-band, PowerScale F710 storage — identical across the upstream Kubernetes and OpenShift columns except the cluster manager.[1] That is quotable line by line. What the digits mean is not: NVIDIA’s Enterprise RA index lists the pattern beside 2-8-10-400 and 2-8-9-800 and decodes none of them.[3] Procurement’s question has a printed answer too, in a footnote — larger scale clusters are supported and require additional network switching.[1]

Quote the rows you can point at, and turn everything else into a question.

Three deliverables come out of this lesson, and each one is built only from statements you can point at.

1The bill of materials you are allowed to quote

There is exactly one Dell AI Factory configuration in this course that can be quoted line by line, and it is worth memorising because it ends arguments. Dell publishes a solution brief for the PowerEdge XE9680 based Dell AI Factory with NVIDIA, “endorsed by NVIDIA Enterprise Reference Architectures (Enterprise RAs) aligned with 2-8-9-400 configuration”.[1]

Component Upstream Kubernetes Red Hat OpenShift Container Platform
Management Server 4 x PowerEdge R670 4 x PowerEdge R670
Worker nodes Up to 12 PowerEdge XE9680 Up to 12 PowerEdge XE9680
Networking 2 x Spectrum-4 SN5610 converged, 1 x Spectrum SN2201 OOB same
Storage PowerScale F710 PowerScale F710
Cluster Manager Base Command Manager Dell Automation Platform
AI software NVIDIA AI Enterprise NVIDIA AI Enterprise plus OpenShift AI

Each XE9680 carries “two high-performance processors and eight NVIDIA H200 SXM GPUs”, and the four R670 servers “provide dedicated resources for orchestration and system administration”.[1] The table carries one footnote that matters more than its size suggests: “Larger scale clusters supported and require additional network switching.”[1] Two switches for the converged network is a starting point, not a ceiling.

The label itself is the trap. NVIDIA’s Enterprise Reference Architecture index lists patterns including 2-8-9-400, 2-8-10-400, 2-8-9-800 and 2-8-5-200, and explains none of the digits anywhere on the page.[3] The obvious reading is not stated by either vendor, so treat 2-8-9-400 as an undecoded pattern name. What the index does state is scale: Enterprise RAs cover AI Factories “ranging from 32 to 256 GPUs”.[3]

PowerEdgeBF-3 offeringDPU modeNIC modeAux powervSphere DSENotable KB
R660
16G
R760
16G
R760XA
16G (GPU)
XE9680
16G HGX H100/H200
XE9680L
16G liquid
R7725
17G (AMD)
R770
17G (Intel)
XE9780 / XE9785
17G HGX B300 — Dell AI Factory

✓ yes · ✗ no · ◐ conditional · ? unknown · — n/a. Click a cell for the evidence.

⚠ = not confirmed on a fetched primary source (hover for why). Facts as of DOCA 3.5.0 (Sep 2026). Selections are saved.

Open the Dell tab and read every row as a sentence you could say on a call. Anything you cannot source to a row is a question, not a claim.

2Two fabrics, two design points, and the question about which NOS they bought

The brief separates two networks with two different design points, and conflating them is a common source of confused escalations. The GPU network: “Spectrum-X provides a non-blocking, rail-optimized spine-leaf fabric that ensures deterministic, low-diameter GPU paths. With RDMA/RoCEv2, PFC, ECN, adaptive routing, and congestion control, the fabric delivers low-latency, predictable, and highly scalable performance for LLM training workloads.”[1] The enterprise network: “The platform uses an EVPN-VXLAN fabric that supports scalable north-south traffic, workload mobility, and tenant segmentation through a modern BGP-based control plane. EVPN-Multihoming (EVPN-MH) provides active/active server connectivity without MLAG.”[1]

Those are different problems solved by different mechanisms on the same purchase order. Rail optimisation and congestion control decide whether a collective finishes on time; EVPN-MH decides whether a server keeps both uplinks live without a chassis pairing protocol. When a customer says “the fabric”, ask which one.

The deployment layer is also published: “Using Kubernetes operators—including the NVIDIA GPU Operator, NVIDIA Network Operator, NVIDIA NIM Operator and Dell CSI Operator—the environment can be provisioned, configured, and managed using declarative, cloud-native patterns.”[1] Dell also claims the design “aligns directly with the NVIDIA AI Enterprise infrastructure support matrix and the Spectrum-X validated solution stack”.[1] That claim is the reason the version arithmetic in this course belongs in a Dell conversation at all: alignment to a support matrix is only meaningful if the four versions on the cluster are inside their windows.[6][7]

Then there is the network operating system. Dell announced that it is “delivering NVIDIA Spectrum-X Ethernet with Enterprise SONiC Distribution by Dell Technologies—a validated, open networking solution engineered for AI”, naming RoCEv2 for direct GPU memory access over Ethernet, adaptive routing, congestion management using ECN marking and PFC, and enhanced telemetry.[2] The announcement names no switch model numbers and no SONiC release numbers.[2] That absence is load-bearing for anyone practising on Dell Enterprise SONiC 4.5.1 in containerlab: the lab image can teach the PFC, ECN and DSCP-trust half of the RoCE story, and it cannot be claimed to be the validated release.

3Deliverable one and two: the qualification answer and the escalation intake

An FAE produces a small number of repeatable artifacts. The first is the qualification answer to “which CNI do you support?” — a question that assumes one layer where there are two. The answer has a fixed shape: NVIDIA does not replace the primary CNI, whatever the platform ships; it adds Multus plus a device plugin for the fast path, and what it does constrain is the secondary path and the driver.[8] Then the one check that earns the answer: confirm the primary CNI has not claimed the RDMA-capable PFs as pod uplinks, because a node where the same interface is both the Calico or Cilium uplink and the device plugin’s selector gives a working cluster that silently loses RDMA whenever the CNI reconfigures the interface.[8]

The second is the escalation intake form, and its whole purpose is to make you collect evidence before offering theory. Four versions, in this order, then two commands:

# the four versions, in dependency order
kubectl version                                   # 1. Kubernetes
helm list -A                                      # 2+3. GPU Operator, Network Operator
kubectl get nicclusterpolicy nic-cluster-policy -o yaml \
  | grep -A3 ofedDriver                           # 4. DOCA-OFED driver tag

# the two pieces of evidence no theory survives
kubectl describe node <node> | grep -A20 Allocatable
kubectl netop-sosreport --verbose --log-lines 10000

Kubernetes comes first because it bounds both operators: Network Operator 26.7.0 states its Kubernetes range as >=1.32 and <=1.36 while GPU Operator 26.7.x states 1.33—1.37, so only 1.33 through 1.36 satisfies both published windows.[6][7] The driver tag comes last because it is what decides whether the RDMA and GPUDirect path can bind at all; 26.7.0 ships GA doca3.5.0-26.07-0.7.7.0-0 and LTS doca3.2.2-25.10-2.4.1.0-4.[6] Collect the sos-report before changing anything, because it captures previous-container logs you are about to destroy.[10]

Qualification request · decision 1/6A 0 · P 0 · S 0

Brief — Email, Tuesday: "We want to offer the BlueField-3 B3140H SuperNIC in the XE9780 for Spectrum-X customers. Can NVIDIA qualify it? I need a plan by Friday."

Dell platform PM (17G AI servers): So — can you get it qualified? What do you need from us to start?

Run the qualification scenario twice: once answering with a product name, once with the two-layer answer plus the uplink check. Compare where each conversation ends up.

4Deliverable three: the positioning answer, and the escalation that is really a NUMA problem

The positioning answer has three forks, and each one is settled by a published line rather than an opinion.

Network Operator versus DPF is settled by the platform matrix: the Network Operator supports BlueField-3 and BlueField-3 SuperNIC in NIC mode only.[6] A customer who wants DPU-mode offload with services managed on the card is asking for DPF, which treats the card as a second computer with its own cluster. A customer who wants pods on the host to reach the fabric fast is asking for the Network Operator. That single line resolves most of the scoping conversation without any architecture debate.

Converged versus split fabric is settled by a question, not by an assertion. The published 2-8-9-400 brief is converged: two SN5610 carry the converged network with one SN2201 for out-of-band.[1] Older Dell XE9680 material described a three-network split of frontend, backend and out-of-band, but that material could not be fetched for this course, so it is not cited here and must not be presented as a second design you have read. Ask which generation the customer is quoting.

Which NOS they bought is settled by segment 2. Ask it early.

The third fork is where positioning turns into an escalation. “Training is slow” on a dual-socket XE9680 with eight GPUs and multiple NICs is very often a NIC on one socket feeding a GPU on another. Topology Manager’s default policy is none, which attempts no alignment at all; single-numa-node requires every requested resource to come from one NUMA node or the pod is rejected, and restricted admits only when all resources align to a common set.[9] The Kubernetes-side fix is single-numa-node with pod scope plus CPU Manager on static, and the pod requesting the GPU and the NIC resource in the same container so both hint providers vote.[9][8] Then warn the customer about the inversion before it happens: pods start getting rejected with TopologyAffinityError because no single NUMA node has both a free GPU and a free NIC.[9] That rejection is the system working. The answers are restricted with prefer-closest-numa-nodes, or fixing the physical NIC placement — not switching Topology Manager back off.[9]

NUMA node 0 · socket 0free: 4 GPU · 2 NIC · 24 CPUnvidia.com/gpuGPU 0GPU 1GPU 2GPU 3nvidia.com/hostdev (SR-IOV PF/VF)ens2f0ens2f1NUMA node 1 · socket 1free: 4 GPU · 2 NIC · 24 CPUnvidia.com/gpuGPU 4GPU 5GPU 6GPU 7nvidia.com/hostdev (SR-IOV PF/VF)ens3f0ens3f1inter-socket link (UPI / xGMI) — a GPU on one node reading a NIC on the other pays this hop on every transfer
One container asks for
nvidia.com/gpu1
nvidia.com/hostdev1
cpu8

The Network Operator GPUDirect example requests nvidia.com/hostdev: 1 and nvidia.com/gpu: 1 in the same container, with securityContext.capabilities.add: ["IPC_LOCK"]. That single co-request is what makes NUMA alignment possible at all.

Click a GPU or a NIC

Selecting a device explains it and lets you move it to the other socket — the physical change an FAE actually recommends. Devices are keyboard reachable: Tab, then Enter.

Set the policy, the placement and the request, then submit the pod to see the admission decision.

topologyManagerPolicy: single-numa-node

All resources must come from one NUMA node or the pod is rejected.

The strongest guarantee and the one that produces TopologyAffinityError most often. Pair it with topologyManagerScope: pod and a pod that requests GPU and NIC in the same container.

All containers in the pod land on a common set of NUMA nodes. Pod-scope resource math uses the effective-request formula: the max of the sum of the app containers, or the max of the init containers.

Exclusive CPUs for Guaranteed pods with integer CPU requests; BestEffort and Burstable pods stay in the shared pool. Only then does CPU become a Hint Provider and vote with the GPU and the NIC.

Both options sit behind the TopologyManagerPolicyOptions feature gate. prefer-closest-numa-nodes defaults to false and applies to best-effort and restricted only. ⚠ max-allowable-numa-nodes is recorded in the course notes as “default unlimited, caps how many NUMA nodes a pod's resources may span”; the upstream page could not be re-read to confirm whether it caps the pod span or the machine's NUMA count — verify before quoting it to a customer.

CPU Manager static options
  • full-pcpus-only=trueGA 1.33+
  • distribute-cpus-across-numa=trueBeta 1.33+
  • align-by-socket=trueAlpha 1.25+
  • distribute-cpus-across-cores=trueAlpha 1.31+
  • strict-cpu-reservation=trueGA 1.35+
  • prefer-align-cpus-by-uncorecache=trueGA 1.36+

Static-policy reservation precedence, first match wins: --reserved-cpus (explicit list, highest), then --kube-reserved, then --system-reserved. The reservation must be greater than zero.

Switching CPU Manager from none to static: drain the node, stop kubelet, rm /var/lib/kubelet/cpu_manager_state, update the config, start kubelet. Skipping the state-file delete is a classic silent failure — the new policy simply does not take effect.

What you would actually apply
# /var/lib/kubelet/config.yaml
topologyManagerPolicy: "single-numa-node"
topologyManagerScope: "pod"
cpuManagerPolicy: "static"
spec:
  containers:
    - name: trainer
      securityContext:
        capabilities:
          add: ["IPC_LOCK"]
      resources:
        limits:      # requests must match limits → Guaranteed QoS
          nvidia.com/gpu: "1"
          nvidia.com/hostdev: "1"
          cpu: "8"

On a Spectrum-X rail deployment the NIC side is requested as nvidia.com/rail0 / nvidia.com/rail1 instead of nvidia.com/hostdev; on the RDMA shared plugin it is rdma/rdma_shared_device_a. Same alignment question, different resource name.

FAE angle

“The job runs at half speed” on a dual-socket PowerEdge is usually a NIC on socket 0 feeding a GPU on socket 1. The Kubernetes-side fix is single-numa-node + scope: pod + CPU Manager static, with nvidia.com/gpu and nvidia.com/hostdev in the same container so both hint providers vote. Then the failure inverts into TopologyAffinityError — and that rejection is the system working.

Map the box before you argue about policy: lstopo, numactl --hardware, nvidia-smi topo -m, cat /sys/class/net/<pf>/device/numa_node.

Documented limitation: a hard limit of 8 NUMA nodes per system; systems with more are unsupported, with no workaround. Check it before recommending an NPS4 BIOS setting on a high-core-count AMD node — dual socket × NPS4 is already 8.

Dell publishes the 2-8-9-400 configuration on PowerEdge XE9680: up to 12 worker nodes, "each equipped with two high-performance processors and eight NVIDIA H200 SXM GPUs". Two sockets, eight GPUs — that is the board this planner draws.

Topology Manager aligns pods of all QoS classes, but only for resources whose Hint Providers actually supply topology hints. Topology Manager requires Kubernetes v1.18 or later.

GPUDirect RDMA moves data NIC↔GPU without touching host memory, which is exactly why the NIC and the GPU want to hang off the same NUMA node. Verify the module path with lsmod | grep nvidia and the transport with NCCL_DEBUG=INFO.

⚠ The per-node NIC count and the 24 allocatable CPUs per NUMA node in this diagram are illustrative: the Dell brief confirms two sockets and eight H200 SXM GPUs, but the “9 NICs” reading of 2-8-9-400 is unverified.

Place the NICs on the wrong socket first and watch the pod get admitted under policy none, then tighten to single-numa-node and read the rejection. Predict the verdict before you press.

5Calibrating the certification claim

The NCA-AIIO is worth holding and it is not what this course is. The published facts: 50 questions, “One hour”, $125, online and remotely proctored, and “This certification is valid for two years from issuance. Recertification may be achieved by retaking the exam.”[4] The recommended preparation is the self-paced “AI Infrastructure and Operations Fundamentals” course, typical completion time seven hours.[4]

The blueprint is an associate-level document. It defines three domains — Essential AI knowledge at 38%, AI Infrastructure at 40%, AI Operations at 22% — across 22 numbered objectives, and describes the audience as “IT professionals new to AI operations and infrastructure” across roles “from technical pre-sales to data center operations”.[5] Nothing in those 22 objectives names CNI, Multus, SR-IOV, RDMA or the Network Operator.[5]

So the honest sentence is: this course over-serves a subset of objectives rather than covering the exam. It contributes to 2.5 “Identify key components and considerations of a cluster of an accelerated infrastructure”, 2.7 “Determine networking requirements for AI workloads”, 2.8 “Identify and describe data center networking protocols and key concepts”, 2.9 “Identify high-speed data center network options and their use cases”, 2.10 “Explain the purpose and benefits of a DPU in a data center”, 3.1, 3.2 and 3.4.[5] It contributes almost nothing to Domain 1 and to the monitoring objective 3.3, and Domains 1 and 3 together carry 60% of the exam.[5] Someone who studies only this course will be over-prepared on rails and under-prepared on training-versus-inference architecture, power and cooling, and DCGM.

Say that plainly in an interview. “This material goes deeper than the exam on Kubernetes networking and does not cover most of Domain 1” is a stronger statement about your judgement than any claim of exam readiness.

Composing a one-page positioning brief for a Dell OEM opportunity

The ask: a Dell account team is quoting an AI Factory to a customer who currently runs OpenShift, has ConnectX-7 in existing nodes, mentions “DPUs” once in the RFP, and asks in writing “do you support Cilium?”. Produce a one-page brief.

  1. Open with the bill of materials you can quote. Four R670 management nodes, up to twelve XE9680 with eight H200 SXM each, two SN5610 converged plus one SN2201 out-of-band, PowerScale F710, and on OpenShift the cluster manager is Dell Automation Platform with OpenShift AI added to NVIDIA AI Enterprise.[1] Every one of those is a printed row.
  2. Answer the CNI question two-layer. OpenShift ships its own primary CNI and the Network Operator does not replace it; it adds Multus plus a device plugin for the fast path, and constrains the secondary path and the driver.[8] Add the OpenShift-specific consequence: on OpenShift, SR-IOV resources use the openshift.io prefix, so any runbook written against a vanilla cluster will name the wrong resource.[8]
  3. Scope the DPU mention rather than answering it. State the platform line: BlueField-3 SuperNIC is supported by the Network Operator in NIC mode only.[6] Then ask the disambiguating question in writing: do you want the card to be a fast NIC for host pods, or do you want to run services on the card? The first is Network Operator, the second is DPF. Do not price either until they answer.
  4. State the version window as a requirement, not a recommendation. Network Operator 26.7.0 supports Kubernetes 1.32 through 1.36 and GPU Operator 26.7.x supports 1.33 through 1.37, so build inside 1.33 through 1.36, and record the DOCA-OFED tag the design assumes.[6][7]
  5. Name one risk with its evidence command. On dual-socket XE9680 nodes, NIC-to-GPU NUMA alignment is a real performance risk; propose single-numa-node with pod scope and warn in the same paragraph that pods may then be rejected with TopologyAffinityError, which is the intended behaviour.[9]
  6. Close with what you did not assert. The digits in 2-8-9-400 are an undecoded pattern label,[3] and no claim is made that any specific Dell Enterprise SONiC release carries the validated Spectrum-X feature set, because the announcement names none.[2]

The shape to keep: quoted rows, then a two-layer answer, then a question instead of a guess, then one number that constrains the build, then one named risk, then an explicit list of the claims you declined to make.

Case closed

How it ended

The deck goes out with a quoted bill of materials, one intersection — build on Kubernetes 1.33 through 1.36, because Network Operator 26.7.0 states >=1.32 and <=1.36 and GPU Operator 26.7.x states 1.33—1.37 — and two open questions instead of two guesses.[1][6][7] The PM’s roadmap stays not announced, and nothing in the deck depended on it.

The case ends where it started. The network lead’s notebook holds four commands and one command prefix; the operator’s labels are finally correct; row 14 is green and rows 19 and 22 became quarters of their own. The last sentence you say: every number in here is a line from a published page, and where I could not find a page I wrote a question instead.[2]

Lab

Read-only throughout. The Dell lab BF-3 and ConnectX host is used as the evidence source that turns the intake form from a template into a filled artifact. No step changes configuration, so no rollback is required; if you want to also exercise a mutating step, do it in the previous lesson’s lab where the rollback is specified.

  1. Pre-flight inventory. Record the host name, uname -r, lspci | grep -i mellanox, and whether a Kubernetes cluster is running on this host at all. Expected result: a short block you paste at the top of the filled form. If there is no cluster, steps 2 and 3 become “not applicable on this host” — write that rather than skipping them silently.
  2. Fill section A with real output. Run kubectl version, helm list -A, and kubectl get nicclusterpolicy nic-cluster-policy -o yaml and copy the ofedDriver.version value verbatim. Expected result: four concrete strings. If nic-cluster-policy returns NotFound, record that: the operator only acts on a policy with exactly that name, so a NotFound here is a finding and not a missing step.
  3. Fill section B. Run kubectl describe node <node> | grep -A20 Allocatable and record which NIC or RDMA resources are advertised, if any. Then run the host ladder read-only — lsmod, ibstat, ibv_devinfo, mst status, dmesg | tail — and note which rungs pass. Expected result: you can say in one sentence whether this host’s fabric is visible to Kubernetes at all. If ibstat shows a port that is not Active, stop and record it; the Kubernetes layer is not the problem yet.
  4. Answer the DPU-mode question with evidence. Run mlxconfig -d <dev> q | grep -i internal_cpu read-only and state which of the two operator stories applies to this card as configured. Expected result: one sentence naming NIC mode or DPU mode plus the line of output that proves it. Do not change the mode here.
  5. Rehearse the escalation answer against your own output. Five minutes, timed, using only the strings you collected. Expected result: no sentence in the rehearsal contains a number you did not read off this host or off a cited page.
  6. Optional at a customer site. Repeat steps 2 and 3 against a live Spectrum-X or Dell SONiC fabric and write down what the fabric added to the answer that the single host could not: rail resources in Allocatable, per-rail sync status, and a switch-side view of PFC and ECN counters. Read-only, and ask before running anything on a production cluster.

Retrieval check

10 questions from memory. Answer before looking anything up; misses become flashcards.

Explain it to a Dell SE

Explain to a Dell SE, in five sentences, what you can state as fact about a Dell AI Factory network from published material, and where you have to stop and ask a question instead.

14 flashcards for this lesson — 0 in deck. Spaced review lives at /review.

Sources

Facts in this lesson were checked against Dell AI Factory with NVIDIA 2-8-9-400 solution brief (PDF text extracted 2026-09-09); Dell Open Ethernet for AI blog re-fetched 2026-09-09; NVIDIA Enterprise Reference Architectures index re-fetched 2026-09-09; NCA-AIIO certification page re-fetched 2026-09-09; NVIDIA Network Operator v26.7.0 and GPU Operator platform support re-fetched 2026-09-09. Dates are when each page was fetched.

  1. Dell AI Factory with NVIDIA — NVIDIA 2-8-9-400 configuration ERA endorsed (Solution Brief) · fetched 2026-09-09
  2. Dell — Open Ethernet for AI: NVIDIA Spectrum-X with Dell SONiC · fetched 2026-09-09
  3. NVIDIA Enterprise Reference Architectures — index · fetched 2026-09-09
  4. NVIDIA-Certified Associate: AI Infrastructure and Operations (NCA-AIIO) — certification page · fetched 2026-09-09
  5. NCA-AIIO Exam Study Guide (document 4694224 dated Jan26) · fetched 2026-09-07
  6. NVIDIA Network Operator v26.7.0 — Platform Support · fetched 2026-09-09
  7. NVIDIA GPU Operator — Platform Support · fetched 2026-09-09
  8. NVIDIA Network Operator v26.7.0 — Deployment Guide with Kubernetes · fetched 2026-09-07
  9. Kubernetes — Control Topology Management Policies on a node · fetched 2026-09-07
  10. NVIDIA Network Operator v26.7.0 — SOS-Report Collection Script · fetched 2026-09-07
  11. NVIDIA Network Operator v26.7.0 — Spectrum-X Quick Start · fetched 2026-09-07

The same idea elsewhere

Other lessons that cover this ground, sometimes from another course's angle.