Dell AI Factory scenarios and the NCA-AIIO positioning
S4·E5The numbers you are allowed to say out loud · A Dell account meeting the week after acceptance, a deck due Friday
Builds on: Spectrum-X rails inside Kubernetes, BlueField in Kubernetes: DPF versus Network Operator, The triage ladder: sos-report, error strings and the checklists
Before you read: what do you already know?
3 quick questions. Wrong answers are fine and expected; trying first makes the lesson stick.
After this lesson you can
- Compose a qualification answer to 'which CNI do you support?' that is two-layer and entirely made of published facts.
- Compose an escalation intake form that collects the four versions and the fabric evidence before any theory is offered.
- Compose a positioning answer that separates Network Operator from DPF and converged from split fabric without asserting an uncited design.
- Calibrate the NCA-AIIO claim honestly against the published blueprint instead of marketing this course as exam preparation.
Episode 5 — The numbers you are allowed to say out loud
Acceptance is signed and the account team wants the next twelve nodes on a slide. Procurement is here for two questions, price and lead time on switches. The Dell SE has the promise spreadsheet open at the summary tab and is on his second cold coffee. An NVIDIA PM dialled in for the roadmap half and has already used his phrase twice: not announced.
Half of what this room asks can be quoted exactly and half cannot, and knowing which half is which is the job. Dell publishes a solution brief for the XE9680-based AI Factory endorsed by NVIDIA Enterprise Reference Architectures and aligned with the 2-8-9-400 configuration: four PowerEdge R670 management servers, up to twelve XE9680 workers, two Spectrum-4 SN5610 carrying the converged network, one Spectrum SN2201 for out-of-band, PowerScale F710 storage — identical across the upstream Kubernetes and OpenShift columns except the cluster manager.[1] That is quotable line by line. What the digits mean is not: NVIDIA’s Enterprise RA index lists the pattern beside 2-8-10-400 and 2-8-9-800 and decodes none of them.[3] Procurement’s question has a printed answer too, in a footnote — larger scale clusters are supported and require additional network switching.[1]
Quote the rows you can point at, and turn everything else into a question.
Three deliverables come out of this lesson, and each one is built only from statements you can point at.
1The bill of materials you are allowed to quote
There is exactly one Dell AI Factory configuration in this course that can be quoted line by line, and it is worth memorising because it ends arguments. Dell publishes a solution brief for the PowerEdge XE9680 based Dell AI Factory with NVIDIA, “endorsed by NVIDIA Enterprise Reference Architectures (Enterprise RAs) aligned with 2-8-9-400 configuration”.[1]
| Component | Upstream Kubernetes | Red Hat OpenShift Container Platform |
|---|---|---|
| Management Server | 4 x PowerEdge R670 | 4 x PowerEdge R670 |
| Worker nodes | Up to 12 PowerEdge XE9680 | Up to 12 PowerEdge XE9680 |
| Networking | 2 x Spectrum-4 SN5610 converged, 1 x Spectrum SN2201 OOB | same |
| Storage | PowerScale F710 | PowerScale F710 |
| Cluster Manager | Base Command Manager | Dell Automation Platform |
| AI software | NVIDIA AI Enterprise | NVIDIA AI Enterprise plus OpenShift AI |
Each XE9680 carries “two high-performance processors and eight NVIDIA H200 SXM GPUs”, and the four R670 servers “provide dedicated resources for orchestration and system administration”.[1] The table carries one footnote that matters more than its size suggests: “Larger scale clusters supported and require additional network switching.”[1] Two switches for the converged network is a starting point, not a ceiling.
The label itself is the trap. NVIDIA’s Enterprise Reference Architecture index lists patterns including 2-8-9-400, 2-8-10-400, 2-8-9-800 and 2-8-5-200, and explains none of the digits anywhere on the page.[3] The obvious reading is not stated by either vendor, so treat 2-8-9-400 as an undecoded pattern name. What the index does state is scale: Enterprise RAs cover AI Factories “ranging from 32 to 256 GPUs”.[3]
| PowerEdge | BF-3 offering | DPU mode | NIC mode | Aux power | vSphere DSE | Notable KB |
|---|---|---|---|---|---|---|
R660 16G | ||||||
R760 16G | ||||||
R760XA 16G (GPU) | ||||||
XE9680 16G HGX H100/H200 | ||||||
XE9680L 16G liquid | ||||||
R7725 17G (AMD) | ||||||
R770 17G (Intel) | ||||||
XE9780 / XE9785 17G HGX B300 — Dell AI Factory |
✓ yes · ✗ no · ◐ conditional · ? unknown · — n/a. Click a cell for the evidence.
⚠ = not confirmed on a fetched primary source (hover for why). Facts as of DOCA 3.5.0 (Sep 2026). Selections are saved.
2Two fabrics, two design points, and the question about which NOS they bought
The brief separates two networks with two different design points, and conflating them is a common source of confused escalations. The GPU network: “Spectrum-X provides a non-blocking, rail-optimized spine-leaf fabric that ensures deterministic, low-diameter GPU paths. With RDMA/RoCEv2, PFC, ECN, adaptive routing, and congestion control, the fabric delivers low-latency, predictable, and highly scalable performance for LLM training workloads.”[1] The enterprise network: “The platform uses an EVPN-VXLAN fabric that supports scalable north-south traffic, workload mobility, and tenant segmentation through a modern BGP-based control plane. EVPN-Multihoming (EVPN-MH) provides active/active server connectivity without MLAG.”[1]
Those are different problems solved by different mechanisms on the same purchase order. Rail optimisation and congestion control decide whether a collective finishes on time; EVPN-MH decides whether a server keeps both uplinks live without a chassis pairing protocol. When a customer says “the fabric”, ask which one.
The deployment layer is also published: “Using Kubernetes operators—including the NVIDIA GPU Operator, NVIDIA Network Operator, NVIDIA NIM Operator and Dell CSI Operator—the environment can be provisioned, configured, and managed using declarative, cloud-native patterns.”[1] Dell also claims the design “aligns directly with the NVIDIA AI Enterprise infrastructure support matrix and the Spectrum-X validated solution stack”.[1] That claim is the reason the version arithmetic in this course belongs in a Dell conversation at all: alignment to a support matrix is only meaningful if the four versions on the cluster are inside their windows.[6][7]
Then there is the network operating system. Dell announced that it is “delivering NVIDIA Spectrum-X Ethernet with Enterprise SONiC Distribution by Dell Technologies—a validated, open networking solution engineered for AI”, naming RoCEv2 for direct GPU memory access over Ethernet, adaptive routing, congestion management using ECN marking and PFC, and enhanced telemetry.[2] The announcement names no switch model numbers and no SONiC release numbers.[2] That absence is load-bearing for anyone practising on Dell Enterprise SONiC 4.5.1 in containerlab: the lab image can teach the PFC, ECN and DSCP-trust half of the RoCE story, and it cannot be claimed to be the validated release.
3Deliverable one and two: the qualification answer and the escalation intake
An FAE produces a small number of repeatable artifacts. The first is the qualification answer to “which CNI do you support?” — a question that assumes one layer where there are two. The answer has a fixed shape: NVIDIA does not replace the primary CNI, whatever the platform ships; it adds Multus plus a device plugin for the fast path, and what it does constrain is the secondary path and the driver.[8] Then the one check that earns the answer: confirm the primary CNI has not claimed the RDMA-capable PFs as pod uplinks, because a node where the same interface is both the Calico or Cilium uplink and the device plugin’s selector gives a working cluster that silently loses RDMA whenever the CNI reconfigures the interface.[8]
The second is the escalation intake form, and its whole purpose is to make you collect evidence before offering theory. Four versions, in this order, then two commands:
# the four versions, in dependency order
kubectl version # 1. Kubernetes
helm list -A # 2+3. GPU Operator, Network Operator
kubectl get nicclusterpolicy nic-cluster-policy -o yaml \
| grep -A3 ofedDriver # 4. DOCA-OFED driver tag
# the two pieces of evidence no theory survives
kubectl describe node <node> | grep -A20 Allocatable
kubectl netop-sosreport --verbose --log-lines 10000Kubernetes comes first because it bounds both operators: Network Operator 26.7.0 states its Kubernetes range as >=1.32 and <=1.36 while GPU Operator 26.7.x states 1.33—1.37, so only 1.33 through 1.36 satisfies both published windows.[6][7] The driver tag comes last because it is what decides whether the RDMA and GPUDirect path can bind at all; 26.7.0 ships GA doca3.5.0-26.07-0.7.7.0-0 and LTS doca3.2.2-25.10-2.4.1.0-4.[6] Collect the sos-report before changing anything, because it captures previous-container logs you are about to destroy.[10]
Brief — Email, Tuesday: "We want to offer the BlueField-3 B3140H SuperNIC in the XE9780 for Spectrum-X customers. Can NVIDIA qualify it? I need a plan by Friday."
Dell platform PM (17G AI servers): So — can you get it qualified? What do you need from us to start?
4Deliverable three: the positioning answer, and the escalation that is really a NUMA problem
The positioning answer has three forks, and each one is settled by a published line rather than an opinion.
Network Operator versus DPF is settled by the platform matrix: the Network Operator supports BlueField-3 and BlueField-3 SuperNIC in NIC mode only.[6] A customer who wants DPU-mode offload with services managed on the card is asking for DPF, which treats the card as a second computer with its own cluster. A customer who wants pods on the host to reach the fabric fast is asking for the Network Operator. That single line resolves most of the scoping conversation without any architecture debate.
Converged versus split fabric is settled by a question, not by an assertion. The published 2-8-9-400 brief is converged: two SN5610 carry the converged network with one SN2201 for out-of-band.[1] Older Dell XE9680 material described a three-network split of frontend, backend and out-of-band, but that material could not be fetched for this course, so it is not cited here and must not be presented as a second design you have read. Ask which generation the customer is quoting.
Which NOS they bought is settled by segment 2. Ask it early.
The third fork is where positioning turns into an escalation. “Training is slow” on a dual-socket XE9680 with eight GPUs and multiple NICs is very often a NIC on one socket feeding a GPU on another. Topology Manager’s default policy is none, which attempts no alignment at all; single-numa-node requires every requested resource to come from one NUMA node or the pod is rejected, and restricted admits only when all resources align to a common set.[9] The Kubernetes-side fix is single-numa-node with pod scope plus CPU Manager on static, and the pod requesting the GPU and the NIC resource in the same container so both hint providers vote.[9][8] Then warn the customer about the inversion before it happens: pods start getting rejected with TopologyAffinityError because no single NUMA node has both a free GPU and a free NIC.[9] That rejection is the system working. The answers are restricted with prefer-closest-numa-nodes, or fixing the physical NIC placement — not switching Topology Manager back off.[9]
The Network Operator GPUDirect example requests nvidia.com/hostdev: 1 and nvidia.com/gpu: 1 in the same container, with securityContext.capabilities.add: ["IPC_LOCK"]. That single co-request is what makes NUMA alignment possible at all.
Selecting a device explains it and lets you move it to the other socket — the physical change an FAE actually recommends. Devices are keyboard reachable: Tab, then Enter.
Set the policy, the placement and the request, then submit the pod to see the admission decision.
All resources must come from one NUMA node or the pod is rejected.
The strongest guarantee and the one that produces TopologyAffinityError most often. Pair it with topologyManagerScope: pod and a pod that requests GPU and NIC in the same container.
All containers in the pod land on a common set of NUMA nodes. Pod-scope resource math uses the effective-request formula: the max of the sum of the app containers, or the max of the init containers.
Exclusive CPUs for Guaranteed pods with integer CPU requests; BestEffort and Burstable pods stay in the shared pool. Only then does CPU become a Hint Provider and vote with the GPU and the NIC.
Both options sit behind the TopologyManagerPolicyOptions feature gate. prefer-closest-numa-nodes defaults to false and applies to best-effort and restricted only. ⚠ max-allowable-numa-nodes is recorded in the course notes as “default unlimited, caps how many NUMA nodes a pod's resources may span”; the upstream page could not be re-read to confirm whether it caps the pod span or the machine's NUMA count — verify before quoting it to a customer.
full-pcpus-only=true— GA 1.33+distribute-cpus-across-numa=true— Beta 1.33+align-by-socket=true— Alpha 1.25+distribute-cpus-across-cores=true— Alpha 1.31+strict-cpu-reservation=true— GA 1.35+prefer-align-cpus-by-uncorecache=true— GA 1.36+
Static-policy reservation precedence, first match wins: --reserved-cpus (explicit list, highest), then --kube-reserved, then --system-reserved. The reservation must be greater than zero.
Switching CPU Manager from none to static: drain the node, stop kubelet, rm /var/lib/kubelet/cpu_manager_state, update the config, start kubelet. Skipping the state-file delete is a classic silent failure — the new policy simply does not take effect.
# /var/lib/kubelet/config.yaml topologyManagerPolicy: "single-numa-node" topologyManagerScope: "pod" cpuManagerPolicy: "static"
spec:
containers:
- name: trainer
securityContext:
capabilities:
add: ["IPC_LOCK"]
resources:
limits: # requests must match limits → Guaranteed QoS
nvidia.com/gpu: "1"
nvidia.com/hostdev: "1"
cpu: "8"On a Spectrum-X rail deployment the NIC side is requested as nvidia.com/rail0 / nvidia.com/rail1 instead of nvidia.com/hostdev; on the RDMA shared plugin it is rdma/rdma_shared_device_a. Same alignment question, different resource name.
“The job runs at half speed” on a dual-socket PowerEdge is usually a NIC on socket 0 feeding a GPU on socket 1. The Kubernetes-side fix is single-numa-node + scope: pod + CPU Manager static, with nvidia.com/gpu and nvidia.com/hostdev in the same container so both hint providers vote. Then the failure inverts into TopologyAffinityError — and that rejection is the system working.
Map the box before you argue about policy: lstopo, numactl --hardware, nvidia-smi topo -m, cat /sys/class/net/<pf>/device/numa_node.
Documented limitation: a hard limit of 8 NUMA nodes per system; systems with more are unsupported, with no workaround. Check it before recommending an NPS4 BIOS setting on a high-core-count AMD node — dual socket × NPS4 is already 8.
Dell publishes the 2-8-9-400 configuration on PowerEdge XE9680: up to 12 worker nodes, "each equipped with two high-performance processors and eight NVIDIA H200 SXM GPUs". Two sockets, eight GPUs — that is the board this planner draws.
Topology Manager aligns pods of all QoS classes, but only for resources whose Hint Providers actually supply topology hints. Topology Manager requires Kubernetes v1.18 or later.
GPUDirect RDMA moves data NIC↔GPU without touching host memory, which is exactly why the NIC and the GPU want to hang off the same NUMA node. Verify the module path with lsmod | grep nvidia and the transport with NCCL_DEBUG=INFO.
⚠ The per-node NIC count and the 24 allocatable CPUs per NUMA node in this diagram are illustrative: the Dell brief confirms two sockets and eight H200 SXM GPUs, but the “9 NICs” reading of 2-8-9-400 is unverified.
5Calibrating the certification claim
The NCA-AIIO is worth holding and it is not what this course is. The published facts: 50 questions, “One hour”, $125, online and remotely proctored, and “This certification is valid for two years from issuance. Recertification may be achieved by retaking the exam.”[4] The recommended preparation is the self-paced “AI Infrastructure and Operations Fundamentals” course, typical completion time seven hours.[4]
The blueprint is an associate-level document. It defines three domains — Essential AI knowledge at 38%, AI Infrastructure at 40%, AI Operations at 22% — across 22 numbered objectives, and describes the audience as “IT professionals new to AI operations and infrastructure” across roles “from technical pre-sales to data center operations”.[5] Nothing in those 22 objectives names CNI, Multus, SR-IOV, RDMA or the Network Operator.[5]
So the honest sentence is: this course over-serves a subset of objectives rather than covering the exam. It contributes to 2.5 “Identify key components and considerations of a cluster of an accelerated infrastructure”, 2.7 “Determine networking requirements for AI workloads”, 2.8 “Identify and describe data center networking protocols and key concepts”, 2.9 “Identify high-speed data center network options and their use cases”, 2.10 “Explain the purpose and benefits of a DPU in a data center”, 3.1, 3.2 and 3.4.[5] It contributes almost nothing to Domain 1 and to the monitoring objective 3.3, and Domains 1 and 3 together carry 60% of the exam.[5] Someone who studies only this course will be over-prepared on rails and under-prepared on training-versus-inference architecture, power and cooling, and DCGM.
Say that plainly in an interview. “This material goes deeper than the exam on Kubernetes networking and does not cover most of Domain 1” is a stronger statement about your judgement than any claim of exam readiness.
The ask: a Dell account team is quoting an AI Factory to a customer who currently runs OpenShift, has ConnectX-7 in existing nodes, mentions “DPUs” once in the RFP, and asks in writing “do you support Cilium?”. Produce a one-page brief.
- Open with the bill of materials you can quote. Four R670 management nodes, up to twelve XE9680 with eight H200 SXM each, two SN5610 converged plus one SN2201 out-of-band, PowerScale F710, and on OpenShift the cluster manager is Dell Automation Platform with OpenShift AI added to NVIDIA AI Enterprise.[1] Every one of those is a printed row.
- Answer the CNI question two-layer. OpenShift ships its own primary CNI and the Network Operator does not replace it; it adds Multus plus a device plugin for the fast path, and constrains the secondary path and the driver.[8] Add the OpenShift-specific consequence: on OpenShift, SR-IOV resources use the
openshift.ioprefix, so any runbook written against a vanilla cluster will name the wrong resource.[8] - Scope the DPU mention rather than answering it. State the platform line: BlueField-3 SuperNIC is supported by the Network Operator in NIC mode only.[6] Then ask the disambiguating question in writing: do you want the card to be a fast NIC for host pods, or do you want to run services on the card? The first is Network Operator, the second is DPF. Do not price either until they answer.
- State the version window as a requirement, not a recommendation. Network Operator 26.7.0 supports Kubernetes 1.32 through 1.36 and GPU Operator 26.7.x supports 1.33 through 1.37, so build inside 1.33 through 1.36, and record the DOCA-OFED tag the design assumes.[6][7]
- Name one risk with its evidence command. On dual-socket XE9680 nodes, NIC-to-GPU NUMA alignment is a real performance risk; propose
single-numa-nodewith pod scope and warn in the same paragraph that pods may then be rejected withTopologyAffinityError, which is the intended behaviour.[9] - Close with what you did not assert. The digits in 2-8-9-400 are an undecoded pattern label,[3] and no claim is made that any specific Dell Enterprise SONiC release carries the validated Spectrum-X feature set, because the announcement names none.[2]
The shape to keep: quoted rows, then a two-layer answer, then a question instead of a guess, then one number that constrains the build, then one named risk, then an explicit list of the claims you declined to make.
Now compose the same brief for a different opportunity: upstream Kubernetes rather than OpenShift, existing BlueField-3 cards the customer says are “in DPU mode”, a stated intent to run Run:ai, and a question about whether their existing Dell SONiC leaves are enough.
- Quote the Upstream Kubernetes column: the cluster manager is ____ and the AI software row is ____.
- Answer the CNI question two-layer and name the resource prefix that applies here, which is ____ rather than
openshift.io. - The cards are said to be in DPU mode. Cite the platform line — Network Operator supports BlueField-3 SuperNIC in ____ mode only — and state which product line that puts them in.
- State the Kubernetes window that satisfies both operators: ____ through ____.
- Answer the SONiC question with the one fact and the one absence: Dell announced ____ and named no ____.
- Name the version contradiction you must quote rather than resolve when Run:ai enters the design, and say why you quote both numbers.
Compose, from the course notes alone, a two-column comparison of the published converged 2-8-9-400 design against a split frontend and backend design, plus the questions that disambiguate which one a customer is quoting.
Acceptance criteria: every cell in the converged column carries a citation to the solution brief; the split column is explicitly labelled as not sourced to a page fetched for this course; the disambiguating questions are answerable by the customer without cluster access; and the page contains at least one sentence naming a claim you deliberately did not make and why.
Case closed
The deck goes out with a quoted bill of materials, one intersection — build on Kubernetes 1.33 through 1.36, because Network Operator 26.7.0 states >=1.32 and <=1.36 and GPU Operator 26.7.x states 1.33—1.37 — and two open questions instead of two guesses.[1][6][7] The PM’s roadmap stays not announced, and nothing in the deck depended on it.
The case ends where it started. The network lead’s notebook holds four commands and one command prefix; the operator’s labels are finally correct; row 14 is green and rows 19 and 22 became quarters of their own. The last sentence you say: every number in here is a line from a published page, and where I could not find a page I wrote a question instead.[2]
Lab
Read-only throughout. The Dell lab BF-3 and ConnectX host is used as the evidence source that turns the intake form from a template into a filled artifact. No step changes configuration, so no rollback is required; if you want to also exercise a mutating step, do it in the previous lesson’s lab where the rollback is specified.
- Pre-flight inventory. Record the host name,
uname -r,lspci | grep -i mellanox, and whether a Kubernetes cluster is running on this host at all. Expected result: a short block you paste at the top of the filled form. If there is no cluster, steps 2 and 3 become “not applicable on this host” — write that rather than skipping them silently. - Fill section A with real output. Run
kubectl version,helm list -A, andkubectl get nicclusterpolicy nic-cluster-policy -o yamland copy theofedDriver.versionvalue verbatim. Expected result: four concrete strings. Ifnic-cluster-policyreturns NotFound, record that: the operator only acts on a policy with exactly that name, so a NotFound here is a finding and not a missing step. - Fill section B. Run
kubectl describe node <node> | grep -A20 Allocatableand record which NIC or RDMA resources are advertised, if any. Then run the host ladder read-only —lsmod,ibstat,ibv_devinfo,mst status,dmesg | tail— and note which rungs pass. Expected result: you can say in one sentence whether this host’s fabric is visible to Kubernetes at all. Ifibstatshows a port that is not Active, stop and record it; the Kubernetes layer is not the problem yet. - Answer the DPU-mode question with evidence. Run
mlxconfig -d <dev> q | grep -i internal_cpuread-only and state which of the two operator stories applies to this card as configured. Expected result: one sentence naming NIC mode or DPU mode plus the line of output that proves it. Do not change the mode here. - Rehearse the escalation answer against your own output. Five minutes, timed, using only the strings you collected. Expected result: no sentence in the rehearsal contains a number you did not read off this host or off a cited page.
- Optional at a customer site. Repeat steps 2 and 3 against a live Spectrum-X or Dell SONiC fabric and write down what the fabric added to the answer that the single host could not: rail resources in Allocatable, per-rail sync status, and a switch-side view of PFC and ECN counters. Read-only, and ask before running anything on a production cluster.
Three written artifacts and one rehearsal. Nothing here touches hardware; the deliverables are documents you keep.
- Build the escalation intake form. One page. Section A: the four versions in dependency order with the exact command beside each —
kubectl version,helm list -A,kubectl get nicclusterpolicy nic-cluster-policy -o yaml, and theofedDriver.versiontag read from that output. Section B: the two evidence items,kubectl describe node <node> | grep -A20 Allocatableandkubectl netop-sosreport --verbose --log-lines 10000. Section C: three questions you ask before any theory — which NOS, which generation of the design, and which of NIC mode or DPU mode the BlueField cards are in. Expected result: the form fits on one page and contains no diagnosis. If it contains a hypothesis, delete the hypothesis. - Build the converged-versus-split comparison. Two columns. Fill the converged column only from the solution brief and mark each cell with the row it came from: management servers, worker nodes, networking, storage, cluster manager, orchestration, AI software. Fill the split column with the word “unsourced” wherever you cannot cite a page you fetched, and write the three questions that tell you which one a customer is quoting. Expected result: at least one cell in the split column reads “unsourced”. If none does, you invented a design. If not, re-read what you actually have citations for.
- Write and take a 20-question NCA-AIIO self-test. Allocate items by the published weights: roughly 8 from Domain 1, 8 from Domain 2, 4 from Domain 3. Write them from the objective texts, not from this course. Expected result: you will struggle to write the Domain 1 items and the 3.3 monitoring items from course material alone. That difficulty is the finding — record which objectives you could not source and where you would go for them. If you found Domain 1 easy to write, check whether you drifted back into networking content.
- Rehearse one scenario aloud for five minutes. Pick the qualification, the escalation or the positioning case, set a timer, and record yourself. Expected result: every number you say is traceable to a row or a page. If not, replay and mark each unsourced number; those are the sentences to rewrite before you say them to a customer.
Retrieval check
10 questions from memory. Answer before looking anything up; misses become flashcards.
Explain it to a Dell SE
Explain to a Dell SE, in five sentences, what you can state as fact about a Dell AI Factory network from published material, and where you have to stop and ask a question instead.
Sources
Facts in this lesson were checked against Dell AI Factory with NVIDIA 2-8-9-400 solution brief (PDF text extracted 2026-09-09); Dell Open Ethernet for AI blog re-fetched 2026-09-09; NVIDIA Enterprise Reference Architectures index re-fetched 2026-09-09; NCA-AIIO certification page re-fetched 2026-09-09; NVIDIA Network Operator v26.7.0 and GPU Operator platform support re-fetched 2026-09-09. Dates are when each page was fetched.
- Dell AI Factory with NVIDIA — NVIDIA 2-8-9-400 configuration ERA endorsed (Solution Brief) · fetched 2026-09-09
- Dell — Open Ethernet for AI: NVIDIA Spectrum-X with Dell SONiC · fetched 2026-09-09
- NVIDIA Enterprise Reference Architectures — index · fetched 2026-09-09
- NVIDIA-Certified Associate: AI Infrastructure and Operations (NCA-AIIO) — certification page · fetched 2026-09-09
- NCA-AIIO Exam Study Guide (document 4694224 dated Jan26) · fetched 2026-09-07
- NVIDIA Network Operator v26.7.0 — Platform Support · fetched 2026-09-09
- NVIDIA GPU Operator — Platform Support · fetched 2026-09-09
- NVIDIA Network Operator v26.7.0 — Deployment Guide with Kubernetes · fetched 2026-09-07
- Kubernetes — Control Topology Management Policies on a node · fetched 2026-09-07
- NVIDIA Network Operator v26.7.0 — SOS-Report Collection Script · fetched 2026-09-07
- NVIDIA Network Operator v26.7.0 — Spectrum-X Quick Start · fetched 2026-09-07
The same idea elsewhere
Other lessons that cover this ground, sometimes from another course's angle.
- NCP-AIN: what this course covers and what it does notRoCE course · Same ground: NCA-AIIO, positioning and blueprint
- Dell PowerSwitch SN-series: who supports whatSpectrum-X course · Same ground: sonic, positioning and network
- Mock interview and 90-day planDOCA course · Same ground: calibration, answer and escalation