Troubleshoot and Optimize
12% of the NCP-AII exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
System and Server Bring-up · Physical Layer Management · Control Plane Installation and Configuration · Cluster Test and Verification · Troubleshoot and Optimize
5.1 Hardware fault triage
Reading Xid errors, checking fans and pulling a suspect node out of production.
Key points
An Xid is an error report from the NVIDIA driver, written to the kernel log with a number. PCIe (PCI Express) is the bus that connects the GPU to the CPU. Xid 79 means the GPU vanished from that bus. Hardware faults on the PCIe link often cause it.
What NVIDIA says (2)
“This event is logged when the GPU driver attempts to access the GPU over its PCI Express connection and finds that the GPU is not accessible.”
“This event is often caused by hardware failures on the PCI Express link”
ECC (error-correcting code) memory can fix single-bit errors. A double bit error cannot be fixed. Xid 48 reports that uncorrectable error. NVIDIA says a GPU reset or node reboot is needed. nvidia-smi can summarize ECC errors.
What NVIDIA says (3)
“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”
“A GPU reset or node reboot is needed to clear this error.”
“The tool nvidia-smi can provide a summary of ECC errors.”
A fan module holds two fans. If either fails, you replace the whole module. NVIDIA lists three ways to find it: the LED, nvsm show fans, and BMC sensor data.
What NVIDIA says (3)
“Run the nvsm show fans command and view the command output.”
“If a fan is running at an abnormal speed, then that fan needs to be replaced.”
“If either fan fails, then the entire module must be replaced.”
Triage means sorting a problem to decide what to do. A batch partition is the Slurm queue that runs normal jobs. Pulling the node out stops it from hurting more jobs while you test it.
What NVIDIA says (2)
“If an issue is found, it should be removed from the batch partition for initial triage.”
“For health checks, this is the NVIDIA System Management tool (nvsm).”
Key terms: NVIDIA System Management Xid error
5.2 Identifying faulty parts
Finding a failed PSU, GPU or card from LEDs, NVSM and run-to-run differences.
Key points
A PSU (power supply unit) feeds the system. The DGX H100 names them PSU0 to PSU5. A failed one shows an amber LED. nvsm show psus reports each PSU's health.
What NVIDIA says (2)
“Identify the broken power supply either by the amber color LED or by the power supply number”
“Look for any that do not report Status_Health=OK .”
A baseline is a known-good result to compare against. NVIDIA expects the same test on different parts of the system to perform similarly. A repeatable gap on the same hardware points to a component in that hardware.
What NVIDIA says (2)
“When running the following tests, you should expect that performance between runs of the same configuration on distinct parts of the system should run in a similar time or at a similar performance level.”
“However, over multiple runs on the same sets of hardware a difference is found, it can indicate an issue with some component of that system.”
5.3 Replacing parts
Hot-swapping PSUs and fans safely and getting an RMA for other parts.
Key points
Hot-swapping means replacing a part while the system runs. The DGX H100 runs at full capacity on four PSUs. All PSUs must come from the same manufacturer. An empty PSU bay disturbs airflow, so the swap must be fast.
What NVIDIA says (3)
“If the system is on, make sure at least 4 other power supplies are working by confirming the IN and OUT LEDs are lit green:”
“All PSUs in the system must be from the same manufacturer.”
“Once the power supply is out of the chassis, replace it with the new power supply in less than 30 seconds to avoid airflow disruptions in the system”
Verification after a repair means proving the new part works. The fan swap must take less than 30 seconds to avoid overheating. Then confirm with the BMC, the LED and nvsm.
What NVIDIA says (2)
“Replace the old fan with the new one within 30 seconds to avoid overheating of the system components.”
“Verifying that the amber LED on the fan module is extinguished”
An RMA (return merchandise authorization) is the number that tracks a return. NVIDIA says to get one from Enterprise Support and to use only NVIDIA-supplied replacements.
What NVIDIA says (2)
“Contact NVIDIA Enterprise Support to obtain an RMA number for any system or component that needs to be returned for repair or replacement.”
“When replacing a component, use only the replacement supplied to you by NVIDIA.”
Key terms: Customer-replaceable unit Return merchandise authorization
5.4 AMD and Intel server tuning
IOMMU, ACS and NUMA affinity for GPU and network traffic.
Key points
The IOMMU (I/O memory management unit) translates device memory addresses. It can force PCIe traffic through the CPU. GDS (GPUDirect Storage) moves data straight between storage and GPU memory, so NVIDIA turns the IOMMU off. The option name depends on the CPU vendor. You add it to GRUB_CMDLINE_LINUX_DEFAULT.
What NVIDIA says (2)
“For AMD CPUs, add amd_iommu=off . For Intel CPUs, add intel_iommu=off .”
“$ journalctl -k -b | grep -i iommu”
ACS (Access Control Services) is a PCIe feature that can redirect device-to-device traffic up to the CPU root complex. That breaks GPU Direct and can cause slowdowns or hangs. On bare metal, NVIDIA says to disable it on PCI switches. Virtual machines need ACS, so leave it on there.
What NVIDIA says (3)
“IO virtualization (also known as VT-d or IOMMU) can interfere with GPU Direct by redirecting all PCI point-to-point traffic to the CPU root complex, causing a significant performance reduction or even a hang.”
“You can check whether ACS is enabled on PCI bridges by running: sudo lspci -vvv | grep ACSCtl”
“Virtual machines require ACS to function, hence disabling ACS is not an option.”
NUMA (non-uniform memory access) means each CPU socket has its own nearby memory and devices. Reaching the far socket is slower. NVIDIA says each rank should use cores and memory close to its GPU and NIC. A rank is one process in the job.
What NVIDIA says (2)
“On NUMA systems, each rank should generally use CPU cores and host memory close to its GPU and, for multi-node jobs, its NIC.”
“Use nvidia-smi topo -m and lscpu --extended=CPU,NODE,SOCKET,CORE to inspect GPU, NIC, CPU, and NUMA locality.”
Affinity means tying a process to specific cores or memory. NVIDIA HPL takes --cpu-affinity and --mem-affinity. The docs give example values for DGX H100 and DGX A100, because the two CPU layouts differ.
What NVIDIA says (2)
“where --cpu-affinity is mapping to cores on the local node and --mem-affinity is mapping to NUMA-nodes on the local node.”
“The NVIDIA HPL benchmark expects one GPU per MPI process.”
Core association tells Slurm which CPU cores are closest to each GPU. Jobs then get cores near their GPU. The guide does this by hand for DGX A100. On DGX H100 it uses autodetect.
What NVIDIA says (1)
“For DGX A100 systems, clear the Type value and set the correct core association with each GPU entry for maximum performance.”
Key terms: Access Control Services IOMMU NUMA
5.5 Storage optimization
ACS, PCIe placement and IO size settings for GPUDirect Storage.
Key points
ACS forces peer-to-peer PCIe traffic up through the root complex. GDS then cannot bypass the CPU. NVIDIA says to disable ACS for best GDS performance. gdscheck -p lists the switches that have it on.
What NVIDIA says (2)
“For optimal GDS performance, disable ACS.”
“To list all of the PCI switches that have ACS enabled, issue /usr/local/cuda/gds/tools/gdscheck -p .”
P2P DMA (peer-to-peer direct memory access) lets a NIC write straight into GPU memory. It works best when the NIC and GPU share a PCIe switch. Crossing sockets adds a slow hop.
What NVIDIA says (2)
“For the P2P DMA to function efficiently, NICs, NVMes and GPUs should be under a PCIe switch when possible.”
“ensure at least one NIC is in the same CPU socket as the GPU.”
cuFile is the GDS library API. Its settings live in /etc/cufile.json. GDS splits a large read or write into chunks of max_direct_io_size. Bigger chunks mean fewer calls.
What NVIDIA says (2)
“For the requested IO size, GDS issues IO requests sequentially in chunks of reads/writes based on the max_direct_io_size parameter.”
“Larger values of max_direct_io_size will result in a reduced number of calls to the IO stack”
Key terms: GPUDirect Storage Access Control Services NUMA