5.4 AMD and Intel server tuning

NCP-AII · Troubleshoot and Optimize (12% of the exam) · Official objective: “Execute performance optimization for AMD and Intel servers.”

IOMMU, ACS and NUMA affinity for GPU and network traffic.

Key points

  1. The IOMMU (I/O memory management unit) translates device memory addresses. It can force PCIe traffic through the CPU. GDS (GPUDirect Storage) moves data straight between storage and GPU memory, so NVIDIA turns the IOMMU off. The option name depends on the CPU vendor. You add it to GRUB_CMDLINE_LINUX_DEFAULT.

    What NVIDIA says (2)

    “For AMD CPUs, add amd_iommu=off . For Intel CPUs, add intel_iommu=off .”

    — NVIDIA GPUDirect Storage Installation and Troubleshooting Guide

    “$ journalctl -k -b | grep -i iommu”

    — NVIDIA GPUDirect Storage Installation and Troubleshooting Guide

  2. ACS (Access Control Services) is a PCIe feature that can redirect device-to-device traffic up to the CPU root complex. That breaks GPU Direct and can cause slowdowns or hangs. On bare metal, NVIDIA says to disable it on PCI switches. Virtual machines need ACS, so leave it on there.

    What NVIDIA says (3)

    “IO virtualization (also known as VT-d or IOMMU) can interfere with GPU Direct by redirecting all PCI point-to-point traffic to the CPU root complex, causing a significant performance reduction or even a hang.”

    — NCCL Troubleshooting: GPU troubleshooting

    “You can check whether ACS is enabled on PCI bridges by running: sudo lspci -vvv | grep ACSCtl”

    — NCCL Troubleshooting: GPU troubleshooting

    “Virtual machines require ACS to function, hence disabling ACS is not an option.”

    — NCCL Troubleshooting: GPU troubleshooting

  3. NUMA (non-uniform memory access) means each CPU socket has its own nearby memory and devices. Reaching the far socket is slower. NVIDIA says each rank should use cores and memory close to its GPU and NIC. A rank is one process in the job.

    What NVIDIA says (2)

    “On NUMA systems, each rank should generally use CPU cores and host memory close to its GPU and, for multi-node jobs, its NIC.”

    — NCCL Troubleshooting: Performance and tuning

    “Use nvidia-smi topo -m and lscpu --extended=CPU,NODE,SOCKET,CORE to inspect GPU, NIC, CPU, and NUMA locality.”

    — NCCL Troubleshooting: Performance and tuning

  4. Affinity means tying a process to specific cores or memory. NVIDIA HPL takes --cpu-affinity and --mem-affinity. The docs give example values for DGX H100 and DGX A100, because the two CPU layouts differ.

    What NVIDIA says (2)

    “where --cpu-affinity is mapping to cores on the local node and --mem-affinity is mapping to NUMA-nodes on the local node.”

    — NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

    “The NVIDIA HPL benchmark expects one GPU per MPI process.”

    — NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

  5. Core association tells Slurm which CPU cores are closest to each GPU. Jobs then get cores near their GPU. The guide does this by hand for DGX A100. On DGX H100 it uses autodetect.

    What NVIDIA says (1)

    “For DGX A100 systems, clear the Type value and set the correct core association with each GPU entry for maximum performance.”

    — DGX SuperPOD Deployment Guide: Slurm Setup

Key terms

Sample question

GDS needs the IOMMU off on a bare-metal x86 server. Which kernel option does NVIDIA give for AMD and for Intel CPUs?

Show the answer

Answer: AMD: amd_iommu=off. Intel: intel_iommu=off.

The IOMMU (I/O memory management unit) translates device memory addresses. It can force PCIe traffic through the CPU. GDS (GPUDirect Storage) moves data straight between storage and GPU memory, so NVIDIA turns the IOMMU off. The option name depends on the CPU vendor. You add it to GRUB_CMDLINE_LINUX_DEFAULT.

What NVIDIA says (2)

“For AMD CPUs, add amd_iommu=off . For Intel CPUs, add intel_iommu=off .”

— NVIDIA GPUDirect Storage Installation and Troubleshooting Guide

“$ journalctl -k -b | grep -i iommu”

— NVIDIA GPUDirect Storage Installation and Troubleshooting Guide

Practice 5.4 (5 questions) Full Troubleshoot and Optimize guide

← 5.3 Replacing parts · 5.5 Storage optimization →