Troubleshoot and Optimize

12% of the NCP-AII exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

System and Server Bring-up · Physical Layer Management · Control Plane Installation and Configuration · Cluster Test and Verification · Troubleshoot and Optimize

5.1 Hardware fault triage

Official objective: “Identify and troubleshoot hardware faults (e.g., GPU, fan, network card).”

Reading Xid errors, checking fans and pulling a suspect node out of production.

Key points

  1. An Xid is an error report from the NVIDIA driver, written to the kernel log with a number. PCIe (PCI Express) is the bus that connects the GPU to the CPU. Xid 79 means the GPU vanished from that bus. Hardware faults on the PCIe link often cause it.

    What NVIDIA says (2)

    “This event is logged when the GPU driver attempts to access the GPU over its PCI Express connection and finds that the GPU is not accessible.”

    — Xid Errors: Analyzing the Xid Catalog

    “This event is often caused by hardware failures on the PCI Express link”

    — Xid Errors: Analyzing the Xid Catalog

  2. ECC (error-correcting code) memory can fix single-bit errors. A double bit error cannot be fixed. Xid 48 reports that uncorrectable error. NVIDIA says a GPU reset or node reboot is needed. nvidia-smi can summarize ECC errors.

    What NVIDIA says (3)

    “This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”

    — Xid Errors: Analyzing the Xid Catalog

    “A GPU reset or node reboot is needed to clear this error.”

    — Xid Errors: Analyzing the Xid Catalog

    “The tool nvidia-smi can provide a summary of ECC errors.”

    — Xid Errors: Analyzing the Xid Catalog

  3. A fan module holds two fans. If either fails, you replace the whole module. NVIDIA lists three ways to find it: the LED, nvsm show fans, and BMC sensor data.

    What NVIDIA says (3)

    “Run the nvsm show fans command and view the command output.”

    — DGX H100/H200 Service Manual: Front Fan Module Replacement

    “If a fan is running at an abnormal speed, then that fan needs to be replaced.”

    — DGX H100/H200 Service Manual: Front Fan Module Replacement

    “If either fan fails, then the entire module must be replaced.”

    — DGX H100/H200 Service Manual: Front Fan Module Replacement

  4. Triage means sorting a problem to decide what to do. A batch partition is the Slurm queue that runs normal jobs. Pulling the node out stops it from hurting more jobs while you test it.

    What NVIDIA says (2)

    “If an issue is found, it should be removed from the batch partition for initial triage.”

    — DGX SuperPOD Administration Guide: System Health Checks and Debugging

    “For health checks, this is the NVIDIA System Management tool (nvsm).”

    — DGX SuperPOD Administration Guide: System Health Checks and Debugging

Key terms: NVIDIA System Management Xid error

Try it: ECC Error Lifecycle XID Fault Drill Slurm Scheduler

Practice 5.1 (4 questions) Objective page

5.2 Identifying faulty parts

Official objective: “Identify faulty cards, GPUs, and power supplies.”

Finding a failed PSU, GPU or card from LEDs, NVSM and run-to-run differences.

Key points

  1. A PSU (power supply unit) feeds the system. The DGX H100 names them PSU0 to PSU5. A failed one shows an amber LED. nvsm show psus reports each PSU's health.

    What NVIDIA says (2)

    “Identify the broken power supply either by the amber color LED or by the power supply number”

    — DGX H100/H200 Service Manual: Power Supply Replacement

    “Look for any that do not report Status_Health=OK .”

    — DGX H100/H200 Service Manual: Power Supply Replacement

  2. A baseline is a known-good result to compare against. NVIDIA expects the same test on different parts of the system to perform similarly. A repeatable gap on the same hardware points to a component in that hardware.

    What NVIDIA says (2)

    “When running the following tests, you should expect that performance between runs of the same configuration on distinct parts of the system should run in a similar time or at a similar performance level.”

    — DGX SuperPOD Administration Guide: System Health Checks and Debugging

    “However, over multiple runs on the same sets of hardware a difference is found, it can indicate an issue with some component of that system.”

    — DGX SuperPOD Administration Guide: System Health Checks and Debugging

Practice 5.2 (2 questions) Objective page

5.3 Replacing parts

Official objective: “Replace faulty cards, GPUs, and power supplies.”

Hot-swapping PSUs and fans safely and getting an RMA for other parts.

Key points

  1. Hot-swapping means replacing a part while the system runs. The DGX H100 runs at full capacity on four PSUs. All PSUs must come from the same manufacturer. An empty PSU bay disturbs airflow, so the swap must be fast.

    What NVIDIA says (3)

    “If the system is on, make sure at least 4 other power supplies are working by confirming the IN and OUT LEDs are lit green:”

    — DGX H100/H200 Service Manual: Power Supply Replacement

    “All PSUs in the system must be from the same manufacturer.”

    — DGX H100/H200 Service Manual: Power Supply Replacement

    “Once the power supply is out of the chassis, replace it with the new power supply in less than 30 seconds to avoid airflow disruptions in the system”

    — DGX H100/H200 Service Manual: Power Supply Replacement

  2. Verification after a repair means proving the new part works. The fan swap must take less than 30 seconds to avoid overheating. Then confirm with the BMC, the LED and nvsm.

    What NVIDIA says (2)

    “Replace the old fan with the new one within 30 seconds to avoid overheating of the system components.”

    — DGX H100/H200 Service Manual: Front Fan Module Replacement

    “Verifying that the amber LED on the fan module is extinguished”

    — DGX H100/H200 Service Manual: Front Fan Module Replacement

  3. An RMA (return merchandise authorization) is the number that tracks a return. NVIDIA says to get one from Enterprise Support and to use only NVIDIA-supplied replacements.

    What NVIDIA says (2)

    “Contact NVIDIA Enterprise Support to obtain an RMA number for any system or component that needs to be returned for repair or replacement.”

    — DGX H100/H200 Service Manual: Introduction

    “When replacing a component, use only the replacement supplied to you by NVIDIA.”

    — DGX H100/H200 Service Manual: Introduction

Key terms: Customer-replaceable unit Return merchandise authorization

Practice 5.3 (3 questions) Objective page

5.4 AMD and Intel server tuning

Official objective: “Execute performance optimization for AMD and Intel servers.”

IOMMU, ACS and NUMA affinity for GPU and network traffic.

Key points

  1. The IOMMU (I/O memory management unit) translates device memory addresses. It can force PCIe traffic through the CPU. GDS (GPUDirect Storage) moves data straight between storage and GPU memory, so NVIDIA turns the IOMMU off. The option name depends on the CPU vendor. You add it to GRUB_CMDLINE_LINUX_DEFAULT.

    What NVIDIA says (2)

    “For AMD CPUs, add amd_iommu=off . For Intel CPUs, add intel_iommu=off .”

    — NVIDIA GPUDirect Storage Installation and Troubleshooting Guide

    “$ journalctl -k -b | grep -i iommu”

    — NVIDIA GPUDirect Storage Installation and Troubleshooting Guide

  2. ACS (Access Control Services) is a PCIe feature that can redirect device-to-device traffic up to the CPU root complex. That breaks GPU Direct and can cause slowdowns or hangs. On bare metal, NVIDIA says to disable it on PCI switches. Virtual machines need ACS, so leave it on there.

    What NVIDIA says (3)

    “IO virtualization (also known as VT-d or IOMMU) can interfere with GPU Direct by redirecting all PCI point-to-point traffic to the CPU root complex, causing a significant performance reduction or even a hang.”

    — NCCL Troubleshooting: GPU troubleshooting

    “You can check whether ACS is enabled on PCI bridges by running: sudo lspci -vvv | grep ACSCtl”

    — NCCL Troubleshooting: GPU troubleshooting

    “Virtual machines require ACS to function, hence disabling ACS is not an option.”

    — NCCL Troubleshooting: GPU troubleshooting

  3. NUMA (non-uniform memory access) means each CPU socket has its own nearby memory and devices. Reaching the far socket is slower. NVIDIA says each rank should use cores and memory close to its GPU and NIC. A rank is one process in the job.

    What NVIDIA says (2)

    “On NUMA systems, each rank should generally use CPU cores and host memory close to its GPU and, for multi-node jobs, its NIC.”

    — NCCL Troubleshooting: Performance and tuning

    “Use nvidia-smi topo -m and lscpu --extended=CPU,NODE,SOCKET,CORE to inspect GPU, NIC, CPU, and NUMA locality.”

    — NCCL Troubleshooting: Performance and tuning

  4. Affinity means tying a process to specific cores or memory. NVIDIA HPL takes --cpu-affinity and --mem-affinity. The docs give example values for DGX H100 and DGX A100, because the two CPU layouts differ.

    What NVIDIA says (2)

    “where --cpu-affinity is mapping to cores on the local node and --mem-affinity is mapping to NUMA-nodes on the local node.”

    — NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

    “The NVIDIA HPL benchmark expects one GPU per MPI process.”

    — NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

  5. Core association tells Slurm which CPU cores are closest to each GPU. Jobs then get cores near their GPU. The guide does this by hand for DGX A100. On DGX H100 it uses autodetect.

    What NVIDIA says (1)

    “For DGX A100 systems, clear the Type value and set the correct core association with each GPU entry for maximum performance.”

    — DGX SuperPOD Deployment Guide: Slurm Setup

Key terms: Access Control Services IOMMU NUMA

Practice 5.4 (5 questions) Objective page

5.5 Storage optimization

Official objective: “Optimize storage.”

ACS, PCIe placement and IO size settings for GPUDirect Storage.

Key points

  1. ACS forces peer-to-peer PCIe traffic up through the root complex. GDS then cannot bypass the CPU. NVIDIA says to disable ACS for best GDS performance. gdscheck -p lists the switches that have it on.

    What NVIDIA says (2)

    “For optimal GDS performance, disable ACS.”

    — GPUDirect Storage Best Practices Guide

    “To list all of the PCI switches that have ACS enabled, issue /usr/local/cuda/gds/tools/gdscheck -p .”

    — GPUDirect Storage Best Practices Guide

  2. P2P DMA (peer-to-peer direct memory access) lets a NIC write straight into GPU memory. It works best when the NIC and GPU share a PCIe switch. Crossing sockets adds a slow hop.

    What NVIDIA says (2)

    “For the P2P DMA to function efficiently, NICs, NVMes and GPUs should be under a PCIe switch when possible.”

    — GPUDirect Storage Best Practices Guide

    “ensure at least one NIC is in the same CPU socket as the GPU.”

    — GPUDirect Storage Best Practices Guide

  3. cuFile is the GDS library API. Its settings live in /etc/cufile.json. GDS splits a large read or write into chunks of max_direct_io_size. Bigger chunks mean fewer calls.

    What NVIDIA says (2)

    “For the requested IO size, GDS issues IO requests sequentially in chunks of reads/writes based on the max_direct_io_size parameter.”

    — GPUDirect Storage Best Practices Guide

    “Larger values of max_direct_io_size will result in a reduced number of calls to the IO stack”

    — GPUDirect Storage Best Practices Guide

Key terms: GPUDirect Storage Access Control Services NUMA

Practice 5.5 (3 questions) Objective page