Introductory MLOps Practices
10% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management
6.1 Monitoring and optimizing ML pipelines
Why models degrade, operational metrics and reliable orchestration.
Key points
Monitoring means tracking a deployed model's behavior to analyze its performance. Software logic is fixed in code; a model's quality depends on the data it sees.
What NVIDIA says (2)
“Monitoring a machine learning model after deployment is vital, as models can break and degrade in production.”
“Monitoring is not a one-time action that you do and forget about.”
Operational metrics show whether the service is healthy and fast. They complement model-quality metrics. I/O means input/output.
What NVIDIA says (1)
“Examples include: Latency IO/memory/disk use System reliability (uptime) Auditability”
Reliable pipelines need retries, run history and clear failure reports. An orchestrator provides them. MLOps means machine learning operations.
What NVIDIA says (1)
“Orchestration: Prefect flows chain the stages together, track runs, and handle failures”
Key terms: Prefect MLOps Model monitoring
6.2 Tracking experiments
MLflow, Weights & Biases and why every run must be traceable.
Key points
Experiment tracking records what was run, with which settings, and how it scored. That lets you compare runs and reproduce the best one. ML means machine learning. XGBoost means eXtreme Gradient Boosting.
What NVIDIA says (1)
“Experiment tracking: MLflow logs hyperparameters, model artifacts, and evaluation metrics for every run”
Weights & Biases (W&B) is an experiment tracking platform. Like MLflow, it records runs so you can debug and reproduce models. ML means machine learning.
What NVIDIA says (1)
“Weights & Biases : W&B’s tools help many ML users build better models faster, debugging and reproducing their models with just a few lines of code”
MLOps means machine learning operations: practices for running AI reliably in production. Tracking ties each model to its data, code and settings.
What NVIDIA says (1)
“AI models require careful tracking through cycles of experiments, tuning and retraining.”
Key terms: MLflow Weights & Biases MLOps Experiment tracking Model artifact
6.3 Saving, loading and predicting
pickle and joblib, the pickle security risk, and ONNX export.
Key points
Serialization turns a model into bytes that can be stored and loaded later. All single-GPU cuML estimators support the standard Python methods.
What NVIDIA says (2)
“Single GPU Model Serialization # All single-GPU cuML estimators support serialization using standard Python libraries.”
“This notebook demonstrates how to save and load cuML models using various serialization methods, including pickle, joblib, and cross-platform deployment strategies.”
Loading a pickle can execute code embedded in the file. Treat model files like programs.
What NVIDIA says (2)
“Security Warning # Only unpickle or deserialize models from trusted sources.”
“Malicious pickle data can execute arbitrary code during deserialization, potentially compromising your entire system.”
ONNX is a portable model format. Exporting to it removes the cuML dependency at inference time.
What NVIDIA says (2)
“Use as_sklearn() to convert the cuML model to a scikit-learn estimator, then pass it to skl2onnx.convert_sklearn() .”
“The resulting .onnx file can be loaded with ONNX Runtime for inference on both CPU and GPU, with no cuML dependency at inference time.”
Key terms: Model serialization ONNX
6.4 Detecting drift
Data drift, prediction drift and monitoring without ground truth.
Key points
Drift means the world the model sees has changed. Statistical tests on feature distributions flag it early.
What NVIDIA says (2)
“Data drift : Changes in distribution between the training data and production data can be monitored to check for drift: this is done by detecting changes in the statistical properties of feature values over time.”
“Statistical tests should be used to detect drift, and predictive performance should be monitored to evaluate the model’s performance over time.”
Ground truth means the correct label. When it is delayed, the distribution of the model's own outputs is a useful proxy.
What NVIDIA says (2)
“Prediction drift: When it is not possible to acquire ground truth labels, predictions must be monitored.”
“If there is a drastic change in the distribution of predictions, something has potentially gone wrong.”
Key terms: Model monitoring Data drift Prediction drift
6.5 Artifacts and configurations for reproducibility
Model registries, champion aliases and configs in version control.
Key points
A model registry stores model versions with names and aliases. An alias like 'champion' is a stable pointer that moves when a better model wins. MLOps means machine learning operations.
What NVIDIA says (1)
“The champion alias always points to the current production model: In a typical production cycle, start with a full pipeline run to establish a baseline champion.”
Version control records every change to files over time. Configs in version control can rebuild an environment exactly. MLOps means machine learning operations.
What NVIDIA says (1)
“Storing infrastructure configurations in version control so production environments can be quickly replicated, whether for fault-tolerance or to create a replica environment for development.”
Key terms: MLflow Model registry Model artifact
6.6 Benchmarking and choosing hardware
Timing CPU and GPU runs fairly, and Multi-Instance GPU for small jobs.
Key points
Benchmarking means timing the same workload on different hardware. Wall-clock time is the real elapsed time. HPO means hyperparameter optimization.
What NVIDIA says (1)
“In the notebook demo below, we compare benchmarking results to show how GPU can accelerate HPO tuning jobs relative to CPU.”
MIG partitions one GPU into isolated instances, each with its own resources. It suits many small workloads; trade-offs include no NVLink between instances. NVLink means NVIDIA's high-speed GPU interconnect.
What NVIDIA says (1)
“Multi-Instance GPU is a technology that allows partitioning a single GPU into multiple instances, making each one seem as a completely independent GPU.”
Key terms: Multi-Instance GPU Benchmarking