Troubleshooting and Optimization
23% of the NCP-AIO exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Installation and Deployment · Administration · Workload Management · Troubleshooting and Optimization
4.1 Troubleshooting Docker with GPUs
Selecting GPUs, runtime debug logs and lost GPU access.
Key points
The device list is quoted twice: single quotes outside, double quotes inside.
What NVIDIA says (1)
“For example, '"device=2,3"' enumerates GPUs 2 and 3 to the container.”
NVIDIA_VISIBLE_DEVICES is the environment variable the NVIDIA Container Runtime reads to choose GPUs.
What NVIDIA says (1)
“none No GPU is accessible, but driver capabilities are enabled.”
Debug logs show what the runtime hook did, which usually reveals the root cause.
What NVIDIA says (1)
“Edit your runtime configuration under /etc/nvidia-container-runtime/config.toml and uncomment the debug=... line.”
NVML (NVIDIA Management Library) is the API nvidia-smi uses. With systemd managing cgroups, a reload can strip the container's device access.
What NVIDIA says (2)
“For Docker, use cgroupfs as the cgroup driver for containers.”
“On systems where systemd is used to manage the cgroups of the container, reloading the systemd unit files ( systemctl daemon-reload ) is sufficient to trigger container updates and cause a loss of GPU access.”
SELinux (Security-Enhanced Linux) labels files and processes. This option stops labeling for the container.
What NVIDIA says (2)
“you might have to specify --security-opt=label=disable on the Docker or Podman command line”
“However, using this option disables SELinux separation in the container, and the container runs in an unconfined type.”
Key terms: NVIDIA Container Toolkit NVML
Try it: NGC Container Flow
4.2 Troubleshooting Fabric Manager
Service status, version match, logs and the not-ready error.
Key points
Fabric Manager (FM) configures the NVSwitch fabric so GPUs can talk over NVLink. It runs as a systemd service.
What NVIDIA says (1)
“To check FM, for Linux based OS distributions, run the following command: sudo systemctl status”
Without FM, the NVSwitch fabric is not set up, so new CUDA jobs are refused.
What NVIDIA says (1)
“A new CUDA job launch will fail with a cudaErrorSystemNotReady error.”
A mismatch between FM and the driver is a common cause of FM start failures.
What NVIDIA says (1)
“NVIDIA Fabric Manager Package (same version as the Driver package).”
Check this log first when FM fails to start or reports NVSwitch errors.
What NVIDIA says (1)
“Default Value LOG_FILE_NAME=/var/log/fabricmanager.log”
Key terms: Fabric Manager
4.3 Troubleshooting BCM
CMDaemon logs and debug mode, provisioning states and status.
Key points
CMDaemon is the BCM service on every node that carries out management commands.
What NVIDIA says (1)
“CMDaemon generates log messages in /var/log/cmdaemon from specific internal subsystems”
Debug output is off by default because it grows the log quickly. Turn it off when done.
What NVIDIA says (1)
“A global debug mode can be enabled in CMDaemon using cmdaemonctl”
The node-installer is the small system that provisions a node's image at boot.
What NVIDIA says (1)
“This state is entered from the INSTALLING state when the head node CMDaemon can no longer ping the node. It could indicate the node has crashed while running the node-installer.”
Provisioning copies software images to nodes. This command shows its queue and state.
What NVIDIA says (1)
“provisioningstatus Provisioning subsystem status Pending request: node001, node002”
Key terms: Base Command Manager CMDaemon
4.4 Troubleshooting Magnum IO
NCCL debug and interface selection, peer-to-peer tests and ACS for GDS.
Key points
Magnum IO is NVIDIA's set of data-movement libraries, including NCCL and GPUDirect Storage. NCCL_DEBUG controls NCCL's debug output.
What NVIDIA says (1)
“The NCCL_DEBUG variable controls the debug information that is displayed from NCCL. This variable is commonly used for debugging.”
It takes a list of name prefixes. A ^ excludes interfaces.
What NVIDIA says (1)
“The NCCL_SOCKET_IFNAME variable specifies which IP interfaces to use for communication.”
Peer-to-peer (P2P) is a direct GPU-to-GPU path. Turning it off is a test, not a fix.
What NVIDIA says (1)
“The NCCL_P2P_DISABLE variable disables the peer to peer (P2P) transport, which uses CUDA direct access between GPUs, using NVLink or PCI.”
GDS moves data straight between storage and GPU memory. ACS forces that traffic up through the CPU.
What NVIDIA says (1)
“For optimal GDS performance, disable ACS. Note To list all of the PCI switches that have ACS enabled, issue /usr/local/cuda/gds/tools/gdscheck -p .”
Key terms: Magnum IO NCCL GPUDirect Storage Access Control Services
Try it: AllReduce Deep Dive GPUDirect Storage
4.5 Storage performance
GDS compatibility mode, IOMMU, PCIe topology and benchmarking.
Key points
cuFile is the GDS API. Compatibility mode works without the direct path but is slower.
What NVIDIA says (1)
“In the /etc/cufile.json file, verify that allow_compat_mode is set to true . gdscheck -p displays whether the allow_compat_mode property is set to true .”
IOMMU (input-output memory management unit) translates device addresses. Here it forces a slower path.
What NVIDIA says (2)
“When the IOMMU setting is enabled, PCIe traffic will be routed through the CPU root ports.”
“Before you install GDS, you must disable IOMMU.”
Fewer hops between GPU and NIC means higher throughput for GDS.
What NVIDIA says (1)
“The optimal path between a GPU and NIC will be one PCIe switch path designated as PIX. The least optimal path is designated as SYS”
Benchmark with several thread counts. The number of I/O threads strongly affects results.
What NVIDIA says (1)
“the number of processes/threads generating IO is critical to determining maximum performance. The gdsio tool provides for specifying random reads or random writes”
Key terms: High-speed storage GPUDirect Storage IOMMU gdsio
Try it: GPUDirect Storage Storage Bottleneck
4.6 Troubleshooting NGC container deployment
GPU Operator pod failures, lost GPU access and readiness checks.
Key points
Operand pods wait for the driver and toolkit pods. Fix the driver pod first.
What NVIDIA says (2)
“Note that the operand pods will only come up when the driver daemonset and toolkit pods come up successfully.”
“kubectl logs -n gpu-operator nvidia-driver-daemonset-p97x5 -c nvidia-driver-ctr”
The container's device access was removed. A fresh container gets it back. Then apply a lasting fix.
What NVIDIA says (1)
“You need to delete the container after the issue occurs. When it is restarted, either manually or automatically depending on whether you use a container orchestration platform, it regains access to the GPU.”
Triton is NVIDIA's inference server. Its readiness endpoint returns 200 when ready.
What NVIDIA says (1)
“The HTTP request returns status 200 if Triton is ready and non-200 if it is not ready.”
CrashLoopBackOff means Kubernetes keeps restarting a failing container. A hook error points to the container toolkit.
What NVIDIA says (2)
“State: Waiting Reason: CrashLoopBackOff”
“When using the NVIDIA Container Runtime Hook (that is, the Docker --gpus flag or the NVIDIA Container Runtime in legacy mode) to inject requested GPUs and driver libraries into a container”
Key terms: NVIDIA GPU Operator NGC NVIDIA Container Toolkit NVML Triton Inference Server
Try it: Kubernetes GPU Ops