AI Operations
Blueprint + study plan Verified blueprint, study plan generator, and curated official resources. Practice content coming soon.
Start here
Practice questions for this exam are coming. Start with a study plan.
Build a study planToday's study plan
Enter your exam date and weekly hours for a day-by-day plan weighted by the official domains.
Build my planReadiness by domain
Tracks your accuracy and coverage in each official domain as you practise.
Mock exam
Unlocks when the question bank reaches 30 questions.
Exam details and more tools
Everything below is folded away to keep this page calm. Open a section, or switch on Advanced mode (in the More menu) to have them open by default.
More study tools
- Ask Aegis: the tutor button at the bottom right answers from the course content with citations.
- Xid error reference: every in-use NVIDIA Xid code with NVIDIA's recommended actions.
- HGX hardware guide: a beginner's tour of an 8-GPU server.
Exam facts
- Level
- Professional
- Duration
- 120 minutes
- Questions
- 30
- Hands-on lab
- Yes
- Price
- $500
- Delivery
- Online, remotely proctored
- Language
- English
- Validity
- Two years from issuance
- Scoring
- Pass/fail (no numeric score is reported)
- Prerequisites
- Two to three years of operational experience working in a data center with NVIDIA hardware solutions. The candidate should be able to monitor and manage all the parts of a data center infrastructure in support of AI workloads.
Exam facts verified against NVIDIA's NCP-AIOL page on 2026-10-08. NVIDIA's certification index card shows this exam as "NCP-AIO"; the exam page heading and the registration link use "NCP-AIOL". Aegis uses the registration code. Format: 30 multiple-choice questions plus 3 hands-on lab exercises inside the 120-minute session (per the official page). Always confirm on nvidia.com before booking.
Official blueprint
Installation and Deployment — 31%
- 1.1 Describe the Mission Control toolkit
- 1.2 Use BCM’s Base View interface to monitor cluster performance, resource utilization, and node health in real time.
- 1.3 Manage job scheduling and resource allocation using BCM’s workload manager (e.g., SLURM or Kubernetes)
- 1.4 Apply patches, update firmware, and synchronize software images across cluster nodes using BCM
- 1.5 Administer user accounts, roles, and permissions to ensure secure access to the cluster using BCM
- 1.6 Configure and monitor network settings for cluster nodes, DPUs, and switches using BCM
- 1.7 Diagnose and resolve cluster issues, such as job failures, node outages, or resource bottlenecks, using BCM.
- 1.8 Use BCM to organize and configure compute nodes into categories based on hardware or workload requirements.
- 1.9 Using BCM, maintain documentation and generate reports on cluster usage, performance, and issues.
- 1.10 Install and initialize Kubernetes on NVIDIA hosts using BCM
- 1.11 Deploy DOCA Services on DPU Arm
- 1.12 Install Run:ai
- 1.13 Install Slurm
Administration — 23%
- 2.1 Administer Slurm cluster.
- 2.2 Describe data center architecture for AI Workloads
- 2.3 Administer Run:ai
- 2.4 Administer Kubernetes
- 2.5 Configure MIG
Workload Management — 23%
- 3.1 Deploy inference workloads with Kubernetes
- 3.2 Deploy inference workloads with Run:ai
- 3.3 Deploy training workloads with Slurm
- 3.4 Deploy training workloads with Run:ai
- 3.5 Use system management tools to troubleshoot issues
- 3.6 Allocate resources between teams with Run:ai, Slurm and Kubernetes
- 3.7 Deploy containers from NGC
Troubleshooting and Optimization — 23%
- 4.1 Troubleshoot Docker
- 4.2 Troubleshoot the fabric manager service for NVLink and NVSwitch systems
- 4.3 Troubleshoot Base Command Manager
- 4.4 Troubleshoot Magnum IO components
- 4.5 Troubleshoot storage performance
- 4.6 Troubleshoot the deployment of a container from NGC
Study roadmap
- Baseline. Read the official exam page and study guide, then skim every domain below to find gaps.
- Installation and Deployment (31%). Study the NVIDIA training and docs linked below for this domain; keep notes per objective.
- Administration (23%). Study the NVIDIA training and docs linked below for this domain; keep notes per objective.
- Workload Management (23%). Study the NVIDIA training and docs linked below for this domain; keep notes per objective.
- Troubleshooting and Optimization (23%). Study the NVIDIA training and docs linked below for this domain; keep notes per objective.
- Exam rehearsal. Re-read your weakest domain notes and the study guide.