Cluster Test and Verification

33% of the NCP-AII exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

System and Server Bring-up · Physical Layer Management · Control Plane Installation and Configuration · Cluster Test and Verification · Troubleshoot and Optimize

4.1 Single-node stress test

Official objective: “Perform a single-node stress test.”

Running the NVSM pre-flight stress test and DCGM diagnostics on one node.

Key points

  1. DCGM (Data Center GPU Manager) diagnostics come in run levels. A higher level runs longer and includes every test below it. Level 1 is the quick readiness check. Level 4 is the extended, longer-running hardware diagnostic.

    What NVIDIA says (3)

    “higher numbered tests include all beneath.”

    — DCGM User Guide: DCGM Diagnostics

    “4 - Extended (Longer-running System HW Diagnostics)”

    — DCGM User Guide: DCGM Diagnostics

    “Level 1 tests to use as a readiness metric”

    — DCGM User Guide: DCGM Diagnostics

  2. A single-node stress test checks one server alone, before testing the fabric. On DGX, NVIDIA runs it with NVSM.

    What NVIDIA says (2)

    “To run the tests, use NVSM.”

    — DGX H100/H200 Service Manual: Introduction

    “NVIDIA recommends running the pre-flight stress test before putting a system into a production environment or after servicing.”

    — DGX H100/H200 Service Manual: Introduction

Key terms: NVIDIA System Management DCGM diagnostics

Practice 4.1 (2 questions) Objective page

4.2 Running HPL

Official objective: “Execute HPL (High-Performance Linpack).”

How NVIDIA HPL maps MPI processes to GPUs and where its sample inputs live.

Key points

  1. HPL (High-Performance Linpack) solves a large dense linear system to stress GPUs and the network. MPI (Message Passing Interface) starts one process per rank. NVIDIA's HPL expects one GPU per MPI process, so the process count equals the GPU count.

    What NVIDIA says (1)

    “The NVIDIA HPL benchmark expects one GPU per MPI process. As such, set the number of MPI processes to match the number of available GPUs in the cluster.”

    — NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

  2. An MPI process, or rank, is one copy of the program. NVIDIA HPL maps one GPU to each rank.

    What NVIDIA says (1)

    “The NVIDIA HPL benchmark expects one GPU per MPI process.”

    — NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

Key terms: High-Performance Linpack

Practice 4.2 (2 questions) Objective page

4.3 Single-node NCCL and NVLink Switch

Official objective: “Perform single-node NCCL (including verifying NVLink Switch).”

Running nccl-tests on one node, Fabric Manager and checking NVLink peer-to-peer.

Key points

  1. nccl-tests check both the performance and the correctness of NCCL operations. All-reduce is the collective most training jobs use. '-g 8' uses all 8 GPUs. '-b' and '-e' set the smallest and largest message size. '-f 2' doubles the size each step.

    What NVIDIA says (2)

    “These tests check both the performance and the correctness of NCCL operations.”

    — NVIDIA nccl-tests

    “Run on single node with 8 GPUs ( -g 8 ), scanning from 8 Bytes to 128MiB (Mebibytes), doubling between each test ( -f 2 ) : $ ./build/all_reduce_perf -b 8 -e 128M -f 2 -g 8”

    — NVIDIA nccl-tests

  2. NVLink Switch systems (HGX/DGX) need Fabric Manager (FM) to set up the NVSwitch fabric between GPUs. FM runs as the nvidia-fabricmanager service. NVIDIA says FM must be running to restore NVLink peer-to-peer after MIG mode is disabled.

    What NVIDIA says (2)

    “registers the daemon as the nvidia-fabricmanager system service.”

    — NVIDIA Fabric Manager User Guide

    “To successfully restore GPU NVLink peer-to-peer capability after the MIG mode is disabled on these systems, the FM service must be running.”

    — NVIDIA Fabric Manager User Guide

  3. Peer-to-peer (P2P) means one GPU reads or writes another GPU's memory directly. nvidia-smi topo -p2p prints a matrix of P2P status. The capability is p for PCIe and n for NVLink.

    What NVIDIA says (2)

    “You can use nvidia-smi topo -p2p <capability> to print a matrix of P2P status between GPU pairs.”

    — NCCL Troubleshooting: GPU troubleshooting

    “The <capability> value is p for PCIe and n for NVLink.”

    — NCCL Troubleshooting: GPU troubleshooting

Key terms: nccl-tests Fabric Manager GPU peer-to-peer

Try it: NVLink Topology

Practice 4.3 (3 questions) Objective page

4.4 Signal quality on cables

Official objective: “Validate cables by verifying signal quality.”

Counters, BER and eye opening with mlxlink, and fabric-wide error checks.

Key points

  1. BER (bit error rate) is the share of bits received wrong. The eye opening measures how clean the signal is. A wider eye means a cleaner signal. mlxlink checks and debugs link status. -c shows counters and BER. -e shows the eye.

    What NVIDIA says (3)

    “The mlxlink tool is used to check and debug link status and related issues.”

    — NVIDIA Firmware Tools: mlxlink Utility

    “-c |--show_counters Show Physical Counters and BER Info”

    — NVIDIA Firmware Tools: mlxlink Utility

    “-e |--show_eye Show Eye Opening Info”

    — NVIDIA Firmware Tools: mlxlink Utility

  2. Error counters count events such as symbol errors on each port. Rising counters point to a bad cable or module. ibqueryerrors checks every port in the fabric.

    What NVIDIA says (1)

    “The default behavior is to report the port error counters which exceed a threshold for each port in the fabric. The default threshold is zero (0).”

    — MLNX_OFED: InfiniBand Fabric Utilities

  3. ibdiagnet scans the fabric with directed route packets. Directed route means the packet follows a path given hop by hop, so it works even before routing is set up. Its stages include error counter checks and link width and speed checks. A link at lower width or speed than planned is a cabling problem.

    What NVIDIA says (2)

    “Scans the fabric using directed route packets and extracts all the available information regarding its connectivity and devices.”

    — MLNX_OFED: InfiniBand Fabric Utilities

    “Link width and speed checks”

    — MLNX_OFED: InfiniBand Fabric Utilities

Key terms: mlxlink Bit error rate ibdiagnet

Try it: InfiniBand Fabric

Practice 4.4 (3 questions) Objective page

4.5 Confirming cabling

Official objective: “Confirm cabling is correct.”

Comparing real cabling with the planned topology using cable validation and IB discovery tools.

Key points

  1. Topology is the planned map of which port connects to which. The cable validation tool runs as a UFM plugin. It uses managed switches over the management network, so it works without a running subnet manager.

    What NVIDIA says (3)

    “Cable validation is the process of validating the actual cable deployment, against the expected topology (from the planning).”

    — InfiniBand Cluster Bring-up Procedure: Cable Validation

    “The tool is not dependent on a working SM, or any working IB communication at all, but utilizes the management interfaces of all the connected network devices.”

    — InfiniBand Cluster Bring-up Procedure: Cable Validation

    “The validation can utilize managed switches only.”

    — InfiniBand Cluster Bring-up Procedure: Cable Validation

  2. A GUID (globally unique identifier) names each InfiniBand device, like a MAC address. ibnetdiscover walks the subnet and prints every node and link.

    What NVIDIA says (1)

    “Performs InfiniBand subnet discovery and outputs a human readable topology file. GUIDs, node types, and port numbers are displayed as well as port LIDs and node descriptions.”

    — MLNX_OFED: InfiniBand Fabric Utilities

  3. iblinkinfo shows each port's link, so you can spot links down or at the wrong width or speed.

    What NVIDIA says (1)

    “Reports link info for each port in an InfiniBand fabric, node by node.”

    — MLNX_OFED: InfiniBand Fabric Utilities

  4. Each managed switch knows its neighbors on every port. The tool's agents collect this and compare it with the planned topology file. It repeats the check, so you can validate the cluster as you build it.

    What NVIDIA says (2)

    “Perform IB neighbors searchCompare the result to the expected topologyReport it to the main tool”

    — InfiniBand Cluster Bring-up Procedure: Cable Validation

    “This allows to bring up the cluster gradually / incrementally, and to validate the deployment in smaller pieces, rather than all at once.”

    — InfiniBand Cluster Bring-up Procedure: Cable Validation

Key terms: Cable validation ibdiagnet

Try it: InfiniBand Fabric

Practice 4.5 (4 questions) Objective page

4.6 Switch firmware and software

Official objective: “Confirm FW/SW on switches.”

Checking switch versions are aligned with UFM or flint.

Key points

  1. Version drift means devices run different releases. That can cause odd behavior. The guide checks switch ASICs, transceivers and HCAs, and recommends UFM for the check. An HCA (host channel adapter) is the InfiniBand NIC in a server.

    What NVIDIA says (2)

    “The recommended guideline is to confirm that the versions among the cluster are aligned, or differ with up to 2 versions.”

    — InfiniBand Cluster Bring-up Procedure: Confirm Components' Firmware and Software Versions

    “The process can be done using UFM GUI ( which is recommended), or through MOFED commands.”

    — InfiniBand Cluster Bring-up Procedure: Confirm Components' Firmware and Software Versions

  2. An unmanaged switch has no switch OS of its own. You reach it in-band through its LID. A LID (local identifier) is the switch's InfiniBand address. On managed switches, the ASIC firmware is bundled in the switch OS.

    What NVIDIA says (2)

    “This section is applicable only to externally managed (unmanaged) switches (the ASIC firmware is bundled in NOS in managed systems).”

    — InfiniBand Cluster Bring-up Procedure: Confirm Components' Firmware and Software Versions

    “Check the firmware version, run flint -d lid-X -qq q”

    — InfiniBand Cluster Bring-up Procedure: Confirm Components' Firmware and Software Versions

  3. UFM is NVIDIA's fabric manager for InfiniBand. It monitors the fabric and pushes switch software. The SuperPOD guide upgrades InfiniBand switches through the UFM appliance.

    What NVIDIA says (2)

    “The high-speed InfiniBand fabrics are managed with NVIDIA Unified Fabric Manager (UFM).”

    — DGX SuperPOD Administration Guide: Managing High-Speed Fabrics

    “Check out the instructional guide on upgrading InfiniBand switches using the UFM appliance.”

    — DGX SuperPOD Deployment Guide: Upgrade InfiniBand Switches

  4. The UFM fabric health report runs a series of checks and lists errors and warnings. The bring-up guide expects every section green before moving on.

    What NVIDIA says (2)

    “UFM fabric health report contains the results of a series of checks that run on the fabric.”

    — InfiniBand Cluster Bring-up Procedure: UFM Fabric Health

    “Confirm that all fields are indicating green status”

    — InfiniBand Cluster Bring-up Procedure: UFM Fabric Health

Key terms: Unified Fabric Manager

Practice 4.6 (4 questions) Objective page

4.7 BlueField-3 firmware and software

Official objective: “Confirm FW/SW on BlueField-3.”

Listing BF-Bundle versions with bf-info and activating new NIC firmware.

Key points

  1. bf-info runs on the BlueField itself. It lists firmware such as ATF, UEFI, NIC and BMC firmware, plus drivers and tools. ATF (Arm Trusted Firmware) is the first-stage boot firmware.

    What NVIDIA says (2)

    “Retrieve installed packages and their versions as part of BF-Bundle installation:”

    — DOCA: BF-Bundle Installation and Upgrade

    “bf# sudo bf-info”

    — DOCA: BF-Bundle Installation and Upgrade

  2. New NIC firmware is written to flash but runs only after a firmware reset. If only software changed, an Arm reset is enough.

    What NVIDIA says (2)

    “If the NIC firmware was updated, perform a firmware reset ( mlxfwreset ) or a graceful shutdown and power cycle.”

    — DOCA: BF-Bundle Installation and Upgrade

    “If the NIC firmware was not updated, perform a BlueField Arm reset (software reset/reboot).”

    — DOCA: BF-Bundle Installation and Upgrade

Key terms: BlueField DPU RShim

Practice 4.7 (2 questions) Objective page

4.8 Transceiver firmware

Official objective: “Confirm FW on transceivers.”

Querying module firmware and why each side updates only its near-end modules.

Key points

  1. MFT (NVIDIA Firmware Tools) includes flint, the firmware burning tool. With --linkx, flint works on the cable or transceiver attached to a device port instead of the device itself.

    What NVIDIA says (2)

    “Query the Transceiver firmware information.”

    — InfiniBand Cluster Bring-up Procedure: Transceiver Firmware Installation

    “flint -d lid-1 --linkx --downstream_device_ids 1 q”

    — InfiniBand Cluster Bring-up Procedure: Transceiver Firmware Installation

  2. A cable has a module at each end. NVIDIA says each NIC or switch can update only its own attached modules.

    What NVIDIA says (2)

    “Each device (NIC, Switch) can update only the modules connected directly to it, not the far end.”

    — InfiniBand Cluster Bring-up Procedure: Transceiver Firmware Installation

    “Updating the far-end transceiver requires the same operation to be done at the far-end switch(es).”

    — InfiniBand Cluster Bring-up Procedure: Transceiver Firmware Installation

Key terms: Transceiver

Practice 4.8 (2 questions) Objective page

4.9 ClusterKit node assessment

Official objective: “Run ClusterKit to perform a multifaceted node assessment.”

What ClusterKit tests, how it flags bad results and how scopes and stress runs work.

Key points

  1. ClusterKit is part of HPC-X, NVIDIA's HPC communication toolkit. It tests latency, bandwidth, GPU communication, collectives and CPU/GPU stress. That is why NVIDIA calls it multifaceted.

    What NVIDIA says (3)

    “ClusterKit is a multipurpose node assessment tool for high-performance clusters”

    — HPC-X: ClusterKit

    “SLURM or passwordless ssh connectivity across the hosts.”

    — HPC-X: ClusterKit

    “ClusterKit is the recommended tool for InfiniBand end-to-end performance validation.”

    — InfiniBand Cluster Bring-up Procedure: Performance Testing

  2. A pairwise test measures pairs of nodes. ClusterKit judges each pair against the best result, not a fixed number. Latency over 2.1 times the minimum is also bad by default.

    What NVIDIA says (2)

    “Message bandwidths (BWs) less than 93% (by default) of the maximum”

    — HPC-X: Running ClusterKit

    “Message latencies that are 2.1 times (by default) above the minimum are considered 'bad,'”

    — HPC-X: Running ClusterKit

  3. Oversubscription means a switch has less uplink bandwidth than downlink bandwidth. Cross-rack links then measure lower by design. With --topo-file, ClusterKit divides the expected maximum by the oversubscription ratio.

    What NVIDIA says (2)

    “ClusterKit can account for oversubscription using a topology information file.”

    — HPC-X: Running ClusterKit

    “The ratio of total downlink bandwidth to total uplink bandwidth is referred to as the oversubscription ratio.”

    — HPC-X: Running ClusterKit

  4. A scope is a set of nodes, such as one rack. The scope_info file defines them. Comparing scopes shows whether one rack behaves differently from the others.

    What NVIDIA says (1)

    “The purpose of scoped tests is to analyze how similar sets of nodes (such as all nodes in a single rack) behave and whether there are differences among them.”

    — HPC-X: Running ClusterKit

  5. Testing under stress shows how the fabric behaves when nodes are hot and busy. --with-stress loads CPU, GPU or both. -Y sets the duration in minutes.

    What NVIDIA says (2)

    “To run stress tests, use the --with-stress[=<STRESS_TYPES>] option with the ClusterKit binary.”

    — HPC-X: Running ClusterKit

    “If the tests finish earlier than the specified <TEST_TIME> , they will re-run.”

    — HPC-X: Running ClusterKit

  6. A per-GPU pair test checks each GPU's network path on its own, so one bad GPU-NIC path stands out. -z makes the GPU-to-GPU tests pair GPUs by index.

    What NVIDIA says (1)

    “test corresponding GPU pairs: GPU0-to-GPU0, GPU1-to-GPU1, etc.”

    — HPC-X: ClusterKit Test Descriptions and Options

Key terms: ClusterKit

Practice 4.9 (6 questions) Objective page

4.10 NCCL east-west bandwidth

Official objective: “Run NCCL to verify E/W fabric bandwidth.”

Multi-node nccl-tests with MPI and reading busbw.

Key points

  1. East-west (E/W) traffic flows between compute nodes over the fabric. This run starts 64 MPI processes, one per GPU, across 8 nodes. MPI (Message Passing Interface) is the standard for launching multi-node jobs. The binaries must be built with MPI=1.

    What NVIDIA says (3)

    “Run 64 MPI processes on nodes with 8 GPUs each, for a total of 64 GPUs spread across 8 nodes.”

    — NVIDIA nccl-tests

    “mpirun -np 64 -N 8 ./build/all_reduce_perf -b 8 -e 8G -f 2 -g 1”

    — NVIDIA nccl-tests

    “(NB: The nccl-tests binaries must be compiled with MPI=1 for this case)”

    — NVIDIA nccl-tests

  2. algbw (algorithm bandwidth) is data size divided by time. busbw (bus bandwidth) corrects for the collective's traffic pattern, so it can be compared with hardware link speed. The nccl-tests README points to the busbw column.

    What NVIDIA says (1)

    “See the Performance page for explanation about numbers, and in particular the "busbw" column.”

    — NVIDIA nccl-tests

  3. Multi-node runs start one process per GPU with MPI. The test binaries must be built with MPI support.

    What NVIDIA says (1)

    “(NB: The nccl-tests binaries must be compiled with MPI=1 for this case)”

    — NVIDIA nccl-tests

Key terms: nccl-tests Bus bandwidth

Try it: AllReduce Deep Dive

Practice 4.10 (3 questions) Objective page

4.11 NCCL burn-in

Official objective: “Perform NCCL burn-in.”

Running nccl-tests in long loops with correctness checks, and triaging differences.

Key points

  1. A burn-in runs a load for a long time to catch parts that fail under sustained stress. -N sets run cycles, and 0 means run forever. -c checks results, which is slower on many GPUs.

    What NVIDIA says (2)

    “-N,--run_cycles <cycle count> run & print each cycle. Default : 1; 0=infinite.”

    — NVIDIA nccl-tests

    “-c,--check <check iteration count> perform count iterations, checking correctness of results on each iteration.”

    — NVIDIA nccl-tests

  2. NCCL is NVIDIA's GPU communication library. NVIDIA names it the tool for validating the fabric. A repeatable gap on the same hardware points to a component. Suspect nodes leave the batch partition for triage.

    What NVIDIA says (2)

    “However, over multiple runs on the same sets of hardware a difference is found, it can indicate an issue with some component of that system.”

    — DGX SuperPOD Administration Guide: System Health Checks and Debugging

    “If an issue is found, it should be removed from the batch partition for initial triage.”

    — DGX SuperPOD Administration Guide: System Health Checks and Debugging

Key terms: nccl-tests Burn-in

Practice 4.11 (2 questions) Objective page

4.12 HPL burn-in

Official objective: “Perform HPL burn-in.”

Using HPL as a math-heavy load for burn-in with the sample Slurm scripts.

Key points

  1. HPL (High-Performance Linpack) solves a large dense linear system. It loads the GPUs hard and also uses the network. NVIDIA's SuperPOD admin guide lists NCCL for the fabric and HPL for math-intensive applications with network communication.

    What NVIDIA says (2)

    “Math intensive applications with network communications”

    — DGX SuperPOD Administration Guide: System Health Checks and Debugging

    “The NVIDIA HPL benchmark expects one GPU per MPI process.”

    — NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

  2. HPL.dat is HPL's input file. It sets the problem size and process grid. NVIDIA ships samples of both input files and Slurm batch-job scripts with the package.

    What NVIDIA says (2)

    “Samples of Slurm batch-job scripts in sample-slurm directory”

    — NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

    “Samples of input files in sample-dat directory”

    — NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

Key terms: High-Performance Linpack Burn-in

Practice 4.12 (2 questions) Objective page

4.13 NeMo burn-in

Official objective: “Perform NeMo burn-in.”

Using NeMo training runs and NVIDIA recipes as an end-to-end load test.

Key points

  1. NeMo Framework is NVIDIA's framework for training and deploying large language models. A real training job is a good burn-in, because it loads GPUs, memory, network and storage together. NVIDIA's example uses synthetic data, so the first run does not depend on preparing a dataset.

    What NVIDIA says (2)

    “The document walks through using a DGX Cloud Slurm cluster as a user to launch a simple pretraining job, targeting synthetic data to minimize dependencies for initial use.”

    — DGX Cloud Slurm: Workload Examples

    “In order to pull the NeMo FW training container from NGC, the previously noted NGC API key needs to be added to a configuration file in the DGX Cloud Slurm cluster.”

    — DGX Cloud Slurm: Workload Examples

  2. A performance recipe is a packaged, repeatable benchmark run, such as NeMo pretraining. Each one is measured on NVIDIA reference architectures to set a baseline. Run the small system info recipe first. It catches setup problems before you spend hours of GPU time.

    What NVIDIA says (2)

    “we recommend running the system info recipe to collect basic system information and check for common cluster configuration issues”

    — NVIDIA Exemplar Performance: Performance Recipes README

    “These workloads are tested against NVIDIA Reference Architectures to establish baselines for comparison.”

    — NVIDIA Exemplar Performance: Performance Recipes README

Key terms: Burn-in

Practice 4.13 (2 questions) Objective page

4.14 Testing storage

Official objective: “Test storage.”

gdscheck and gdsio: verifying the GDS platform and measuring the data path.

Key points

  1. gdsio is the storage benchmark that ships with GDS. -x picks the transfer type. 0 is GDS. 1 is storage to CPU. 2 goes through the CPU to the GPU. Comparing 0 and 2 shows what GDS gains.

    What NVIDIA says (2)

    “xfer_type : 0 - Storage -> GPU ( GDS ) 1 - Storage -> CPU 2 - Storage -> CPU -> GPU”

    — GPUDirect Storage Configuration and Benchmarking Guide

    “-x 0 , the IO data path, in this case GDS.”

    — GPUDirect Storage Configuration and Benchmarking Guide

  2. A storage test needs enough parallel load and enough time to reach steady state. -w sets the number of IO threads. -T sets the duration in seconds. -i sets the IO size, which matters for throughput.

    What NVIDIA says (3)

    “-w 8 , 8 workers (8 IO threads)”

    — GPUDirect Storage Configuration and Benchmarking Guide

    “-T 120 , runs for 120 seconds.”

    — GPUDirect Storage Configuration and Benchmarking Guide

    “-i 1M , IO size (important for assessing throughput).”

    — GPUDirect Storage Configuration and Benchmarking Guide

  3. gdscheck is the GDS platform check tool. With -p it reports items such as IOMMU state and ACS on PCI switches.

    What NVIDIA says (2)

    “using gdscheck you should see the following output if the IOMMU is disabled on the system:”

    — GPUDirect Storage Best Practices Guide

    “Platform verification succeeded”

    — GPUDirect Storage Best Practices Guide

Key terms: GPUDirect Storage gdsio

Try it: GPUDirect Storage

Practice 4.14 (3 questions) Objective page