System and Server Bring-up
31% of the NCP-AII exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
System and Server Bring-up · Physical Layer Management · Control Plane Installation and Configuration · Cluster Test and Verification · Troubleshoot and Optimize
1.1 Deployment and validation sequence
The order of work from first power-on to a validated cluster, and the checks that close each stage.
Key points
Bring-up follows an order: first boot setup, then a health check, then workload tests. NVSM (NVIDIA System Management) is the DGX tool for system health. NVIDIA's quick health check runs 'nvsm show health' and expects every check, and the overall status, to be Healthy.
What NVIDIA says (2)
“NVIDIA provides customers a diagnostics and management tool called NVIDIA System Management, or NVSM.”
“Verify that the output summary shows that all checks are Healthy and that the overall system status is Healthy.”
Bring-up is the process of taking a new cluster from boxes to production. Each stage in NVIDIA's checklist has entry criteria, so the order matters. Performance testing comes last and needs cluster verification done first.
What NVIDIA says (3)
“Cluster verification stage successfully completed”
“HPC-X package successfully installed and ClusterKit is working”
“All versions are aligned and confirmed”
Validation means checking the result against the spec. The SuperPOD guide loads the bcm-post-install module and runs bcm-validate-pod. It checks items such as the Slurm and Ubuntu versions on the head node.
What NVIDIA says (2)
“After the installation is completed, using the bcm-validate-pod command, one can verify whether all the settings are applied correctly according to the specification.”
“This module validates the completed pod setup with the expected configuration.”
Key terms: NVIDIA System Management
1.2 Network topologies for AI factories
The four SuperPOD fabrics, scalable units and the rail-optimized compute fabric.
Key points
A rail-optimized fabric connects GPU number N in every node to the same leaf switch, called a rail. A scalable unit (SU) is a building block of 32 nodes. Within an SU, same-rail traffic is one hop away. Traffic between rails goes through the spine.
What NVIDIA says (3)
“Traffic per rail of the DGX H100 systems is always one hop away from the other 31 nodes in a SU.”
“Traffic between nodes, or between rails, traverses the spine layer.”
“The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems”
A fabric is a network of switches, cables and adapters. SuperPOD keeps traffic types apart. Compute carries GPU-to-GPU traffic. Storage carries file data. In-band management carries cluster services. Out-of-band (OOB) management reaches the BMCs and device management ports.
What NVIDIA says (2)
“DGX SuperPOD configurations utilize four network fabrics:”
“-Compute Fabric | -Storage Fabric | -In-Band Management Network | -Out-of-Band Management Network”
In-band means the normal OS network of each node, as opposed to the separate BMC network. In SuperPOD it carries cluster services such as BCM and Slurm, and reaches outside services such as the NGC registry.
What NVIDIA says (2)
“Provides connectivity for the in-cluster services such as Base Command Manager, Slurm and to other services outside of the cluster such as the NGC registry, code repositories, and data sources.”
“Enables access to the home filesystem and storage pool.”
OOB ports include BMCs, PDUs and switch management ports. A PDU (power distribution unit) feeds power to a rack. Nobody runs jobs over these ports, and they control the hardware. So they sit on a separate, secured network.
What NVIDIA says (2)
“It connects the management ports of all devices including DGX and management servers, storage, networking gear, rack PDUs, and all other devices.”
“These are separate onto their own fabric because there is no use-case where users need access to these ports and are secured using logical network separation.”
Key terms: Out-of-band management network In-band management network Scalable unit Rail-optimized fabric
1.3 BMC, OOB and TPM setup
Reaching the BMC safely, giving it an address, the ports it uses, Redfish and the TPM in the SBIOS.
Key points
The BMC (baseboard management controller) is a small management computer in the server that works even when the OS is down. Redfish is a standard REST API for server management. The DGX H100 guide lists Redfish features including user accounts, BMC and SBIOS configuration, and power capping.
What NVIDIA says (2)
“Manage user accounts, privileges, and roles Manager sessions BMC configuration SBIOS configuration”
“To manage the maximum power consumption on a system through power capping using Redfish API”
The BMC (baseboard management controller) is a small management computer in the server. It works even when the OS is down. When the network has no DHCP, you give the BMC a static IP address. ipmitool is the command-line tool for the BMC. NVIDIA's guide first prints the LAN settings, then sets the address source to static, then sets the address, netmask and gateway.
What NVIDIA says (2)
“You will need to do this if your network does not support DHCP.”
“Set the IP address source to static. $ sudo ipmitool lan set 1 ipsrc static”
OOB (out-of-band) management means managing a server over a separate network, not the one that carries user traffic. The BMC is the OOB entry point, so it must be protected. NVIDIA recommends a dedicated management network with firewall protection. If remote access is needed, use an isolated path such as a VPN.
What NVIDIA says (3)
“NVIDIA recommends that you connect the BMC port in the DGX H100/H200 system to a dedicated management network with firewall protection.”
“it should be accessed through a secure method that provides isolation from the internet, such as through a VPN server.”
“DGX OS Server software installs Docker Engine which uses the 172.17.xx.xx subnet by default”
The TPM (Trusted Platform Module) is a security chip that stores keys and measures the boot process. The SBIOS (system BIOS) is the firmware setup program you reach with Del or F2 at boot. NVIDIA says not to change SBIOS settings beyond its documents. Enabling the TPM is one of the documented changes.
What NVIDIA says (3)
“Here are some occasions where it might be necessary to reconfigure settings in the SBIOS:”
“Enabling the TPM and Preventing the BIOS from Sending Block SID Requests”
“press the Del or F2 key to enter the BIOS Setup Utility.”
IPMI (Intelligent Platform Management Interface) is the standard protocol that ipmitool uses to talk to a BMC. Over the network it uses RMCP+ on port 623. The web UI and Redfish use HTTPS on 443.
What NVIDIA says (2)
“443 Redfish Redfish https with auth 623 RMCP+ IPMI”
“443 | HTTPS | Web User Interface”
A REST API is a web interface driven by HTTP methods such as GET and PATCH. DMTF is the standards body behind Redfish. Redfish can read inventory and health, change BMC and SBIOS settings, and cap power.
What NVIDIA says (2)
“Redfish is DMTF’s standard set of APIs for managing and monitoring a platform.”
“By default, Redfish support is enabled in the DGX H100/H200 BMC and the SBIOS.”
Key terms: Baseboard management controller Out-of-band management network Intelligent Platform Management Interface Redfish System BIOS Trusted Platform Module
1.4 Firmware upgrades and fault detection
Secure Flash, NIC firmware updates with mlxfwmanager, and confirming versions after a power cycle.
Key points
Firmware is the low-level software stored on the board, such as the BMC and SBIOS. Secure Flash checks the image signature before writing it, so only signed, verified firmware is installed.
What NVIDIA says (1)
“Secure Flash is implemented for the DGX H100/H200 to prevent unsigned and unverified firmware images from being flashed onto the system.”
OPN (ordering part number) identifies the card model. PSID (Parameter-Set Identification) is a 16-character ID in the firmware that marks the card configuration. The image must match both. After flashing, an AC power cycle makes the new firmware take effect. An AC power cycle means fully removing and restoring power.
What NVIDIA says (3)
“If MLNX_OFED is installed, use the mlxfwmanager tool to update the firmware.”
“PSID (Parameter-Set Identification) is a 16-ascii character string embedded in the firmware image”
“Perform an AC power cycle on the system for the firmware update to take effect.”
Each ConnectX port appears as an mlx5 device under /sys/class/infiniband. Its fw_ver file shows the running firmware. NVIDIA's step is to confirm all versions are the same.
What NVIDIA says (2)
“After the system starts, log in and confirm the firmware versions are all the same:”
“cat /sys/class/infiniband/mlx5_*/fw_ver”
mlxfwmanager is NVIDIA's firmware update and query tool for network adapters. --query lists each device with its current and available firmware, and its status.
What NVIDIA says (2)
“The mlxfwmanager is a firmware update and query utility which scans the system for available NVIDIA devices (only mst PCI devices) and performs the necessary firmware updates.”
“To query all the devices on the machine, use the following command line: # mlxfwmanager --query”
Key terms: Secure Flash Parameter-Set Identification mlxfwmanager
1.5 Power and cooling
PSU redundancy, N+1 feeds, power capping and matching cooling to the full heat load.
Key points
A PSU (power supply unit) converts rack power for the server. The DGX H100 has six. Four must be energized for the system to operate. That is why NVIDIA requires at least N+1 power with two power sources, each sized for 50% of peak load.
What NVIDIA says (2)
“Four of the six power supplies must be energized for the system to operate.”
“the data center must minimally provide N+1 power, where N equals two power sources. Each power source must be sized to support 50% of the total peak load.”
Hot aisle / cold aisle means racks face each other so cool air enters the fronts and hot air leaves the backs into a separate aisle. Containment walls off one aisle. Blanking panels cover empty rack spaces so hot air cannot leak back. NVIDIA says to optimize airflow first, before more drastic cooling mitigations.
What NVIDIA says (3)
“Before considering more drastic cooling mitigations, it is important to make sure that the airflow in the space is optimized and well managed.”
“unoccupied RU spaces in the rack should be covered with blanking panels”
“Rack densities greater than 4 DGX H100 Systems are not recommended, due to thermodynamic considerations”
TGP (Total Graphics Power) is the GPU's power budget. SMBPBI is an out-of-band channel that the BMC can use to set GPU power. The GPU's PMU (Performance Monitoring Unit) applies the most conservative limit.
What NVIDIA says (2)
“The GPU has three sources of power limits:”
“The GPU Performance Monitoring Unit (PMU) selects the most conservative policy to cap power”
Each power circuit is sized for 50% of the load. With N+1, N is two circuits. So the real heat load equals two circuits. Planning cooling for one circuit leaves the rack short.
What NVIDIA says (2)
“However, with the specified N+1 power provisioning, N equals two circuits.”
“Therefore, it is critical to align the cooling capacity with N, and not simply the capacity of a single power circuit.”
Three-phase power carries more power per circuit than single-phase. N+1 adds a spare circuit so one feed can fail. NVIDIA prefers 415 VAC, 32 A, three-phase, N+1 for dense racks.
What NVIDIA says (1)
“The preferred power for high-density deployment patterns is 415 VAC, 32A, three-phase, N+1.”
Key terms: Redfish N+1 power Rack power distribution unit Power capping Blanking panel
1.6 Installing GPU servers
Who installs a DGX, how it sits in the rack, first boot and reading the GPU-NIC topology.
Key points
nvidia-smi (System Management Interface) is the standard command-line tool for NVIDIA GPUs. 'nvidia-smi topo -m' prints a matrix of the links between all GPUs and NICs, plus the CPU and memory affinity of each GPU. Affinity tells you which CPU socket and memory are closest to a GPU.
What NVIDIA says (1)
“Topology connections and affinities matrix between the GPUs and NICs in the system nvidia-smi topo -m”
A warranty covers repair costs. NVIDIA ties it to an approved installer. The installer also performs first boot setup. The steps are documented so you can re-image the system later.
What NVIDIA says (2)
“Your DGX H100/H200 system must be installed by NVIDIA partner network personnel or NVIDIA field service engineers.”
“If not performed accordingly, your hardware warranty will be voided.”
A rack mount kit fixes a server into the rack posts. The DGX H100 kit is a shelf, not sliding rails. The system stays put. You service parts from the front or rear.
What NVIDIA says (1)
“The rack mount kit acts as a shelf in the rack, it does not allow the system to be moved once installed. All components are serviceable from the front or rear.”
First boot setup runs the first time the system powers on after delivery or re-imaging. It creates one admin account for the OS, the BMC and the GRUB boot loader, and sets up the primary network interface.
What NVIDIA says (2)
“Create an administrative user account for the system, BMC, and Grub boot loader.”
“Configure the primary network interface.”
Key terms: First boot setup
1.7 Validating installed hardware
Checking health with NVSM and the BMC after install, including the expected RAID re-sync.
Key points
Validating installed hardware means checking that every part is present and reports sane values. The BMC works out-of-band. Its GPU Information page lists each GPU's GUID, VBIOS version, InfoROM version and retired pages. Retired pages are memory pages the GPU has stopped using because of errors.
What NVIDIA says (1)
“Provides basic information on all the GPUs in the systems, including GUID, VBIOS version, InfoROM version, and number of retired pages for each GPU.”
RAID 1 keeps two identical copies of a drive. After first boot it rebuilds the mirror. This is expected. nvsm reports a warning until the rebuild ends.
What NVIDIA says (2)
“During this time, running the nvsm show health command reports a warning that the RAID volume is re-syncing.”
“The process can take an hour to complete.”
The BMC web UI lets you check hardware without logging in to the OS. The Sensor page shows readings. The GPU Information page shows each GPU's VBIOS and InfoROM version and its retired pages.
What NVIDIA says (2)
“Provides status and readings for system sensors, such as SSD, PSUs, voltages, CPU temperatures, DIMM temperatures, and fan speeds.”
“Provides basic information on all the GPUs in the systems, including GUID, VBIOS version, InfoROM version, and number of retired pages for each GPU.”
Key terms: Baseboard management controller NVIDIA System Management RAID 1
1.8 Cable types and transceivers
Telling optical from copper modules and reading module details with mlxlink.
Key points
A transceiver converts electrical signals to light for optical cables. DDM (Digital Diagnostic Monitoring) is the transceiver's built-in readings, such as temperature and optical power. mlxlink checks and debugs link status for passive, active and transceiver cables. Its --cable --ddm option reads the DDM data.
What NVIDIA says (2)
“The mlxlink tool is used to check and debug link status and related issues. The tool can be used on different links and cables (passive, active, transceiver and backplane).”
“--ddm Get cable Digital Diagnostic Monitoring information”
A transceiver converts electrical signals to light, or drives a copper link, at each cable end. Optical and copper types use different firmware images. You can tell them apart by vendor part number.
What NVIDIA says (2)
“Note that we can have both optical & copper type transceivers - these can be identified by Vendor PN and running show explicitly.”
“each transceiver type has different firmware image.”
Before trusting a link, confirm the cable or module is the type you planned. mlxlink --show_module prints its identity, such as OSFP, optical or copper, the part number and FW version.
What NVIDIA says (2)
“-m |--show_module Show Module Info”
“Cable Type : Optical Module (separated)”
Key terms: mlxlink Transceiver Digital Diagnostic Monitoring
1.9 Installing physical GPUs
Safe handling, which parts are customer-replaceable and what to check before driver install.
Key points
ESD (electrostatic discharge) is a static shock that can kill electronic parts. An ESD strap grounds you. Labels let you put every cable back where it was. NVIDIA also warns not to hold the tray by its ejection handles.
What NVIDIA says (3)
“Wear an ESD strap during any procedure that involves touching electronic components.”
“To avoid misconfigurations, label all the cables before unplugging them.”
“Do not hold the motherboard tray by the ejection handles.”
A CRU (customer-replaceable unit) is a part you may replace yourself. The DGX H100 list includes PSUs, fans, drives, DIMMs and some NICs. It does not include GPUs. For parts not covered, NVIDIA says to contact Enterprise Support.
What NVIDIA says (2)
“You can obtain the following components for replacement in your data center.”
“Contact NVIDIA Enterprise Support for replacement instructions and guidance for specific components if those instructions are not included in this document.”
The NVIDIA driver builds kernel modules against the running kernel. NVIDIA's pre-installation actions are: check the distribution, install matching kernel headers, and pick a driver version.
What NVIDIA says (2)
“Verify the system is running a supported version of Linux.”
“Verify the system has the correct kernel headers and development packages installed.”
Key terms: Customer-replaceable unit Electrostatic discharge
1.10 Validating hardware for workloads
DCGM diagnostic levels, soak loops and the pre-flight stress test.
Key points
Validating hardware for workloads means proving a node can run real jobs. DCGM diagnostics have levels with different run times. Level 1 is the quick readiness metric. Level 2 suits a job epilogue after a failure. Levels 3 and 4 are for an administrator's post-mortem.
What NVIDIA says (1)
“Level 1 tests to use as a readiness metric Level 2 tests to use as an epilogue on failure Level 3 and Level 4 tests to be run by an administrator as post-mortem”
A stress test runs the hardware at high load to expose weak parts. NVSM's stress-test can target GPUs, CPU, memory and storage, for a set duration.
What NVIDIA says (2)
“NVIDIA recommends running the pre-flight stress test before putting a system into a production environment or after servicing.”
“You can specify running the test on the GPUs, CPU, memory, and storage, and also specify the duration of the tests.”
A soak test keeps hardware under load for longer to expose intermittent faults. DCGM's --iterations runs the same suite in a loop.
What NVIDIA says (2)
“DCGM also supports running tests suites in loops using the --iterations option. Using this option allows for increasing the runtime duration of the tests.”
“dcgmi diag -r pcie --iterations 3”
Key terms: DCGM diagnostics
1.11 Third-party storage parameters
The storage fabric, NFS export settings and the GDS settings for local and network storage.
Key points
GDS (GPUDirect Storage) lets storage move data straight into GPU memory, skipping a copy through CPU memory. NVIDIA's configuration guide covers local (DAS) and network (NAS) storage and benchmarks them with gdsio, the tool installed with GDS.
What NVIDIA says (2)
“Local drive configurations (Direct Attached Storage - DAS) and Network storage (Network Attached Storage - NAS) are covered.”
“The benchmarking tool included when GDS is installed, gdsio, is covered and its use demonstrated.”
RDMA (remote direct memory access) moves data between machines without the CPU copying it. GDS uses the listed client addresses to choose RDMA devices for an NFS mount.
What NVIDIA says (2)
“fs:nfs:rdma_dev_addr_list [] Provides the list of IPv4 addresses for all the interfaces a single NFS mount can use.”
“This property is used by the cuFile dynamic routing feature to infer preferred RDMA devices.”
The storage fabric carries file data between DGX nodes and storage systems. NVIDIA wants high bandwidth per node and fabric features such as adaptive routing (AR). AR spreads traffic across paths.
What NVIDIA says (2)
“This is because the I/O per-node for the DGX SuperPOD must exceed 40 GBps.”
“The storage devices are connected at a 1:1 port to uplink ratio.”
Key terms: GPUDirect Storage
Try it: GPUDirect Storage