Introductory MLOps Practices

10% of the NCA-ADS exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Data Manipulation and Preparation · Machine Learning With RAPIDS · Data Science Pipelines and Workflow Automation · Descriptive Analysis and Visualization · Foundations of Accelerated Data Science · Introductory MLOps Practices · Advance Data Structures · Software and Environment Management

6.1 Monitoring and optimizing ML pipelines

Official objective: “Monitoring and optimizing machine learning (ML) pipelines for performance and reliability”

Why models degrade, operational metrics and reliable orchestration.

Key points

  1. Monitoring means tracking a deployed model's behavior to analyze its performance. Software logic is fixed in code; a model's quality depends on the data it sees.

    What NVIDIA says (2)

    “Monitoring a machine learning model after deployment is vital, as models can break and degrade in production.”

    — A Guide to Monitoring Machine Learning Models in Production

    “Monitoring is not a one-time action that you do and forget about.”

    — A Guide to Monitoring Machine Learning Models in Production

  2. Operational metrics show whether the service is healthy and fast. They complement model-quality metrics. I/O means input/output.

    What NVIDIA says (1)

    “Examples include: Latency IO/memory/disk use System reliability (uptime) Auditability”

    — A Guide to Monitoring Machine Learning Models in Production

  3. Reliable pipelines need retries, run history and clear failure reports. An orchestrator provides them. MLOps means machine learning operations.

    What NVIDIA says (1)

    “Orchestration: Prefect flows chain the stages together, track runs, and handle failures”

    — RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

Key terms: Prefect MLOps Model monitoring

Practice 6.1 (3 questions) Objective page

6.2 Tracking experiments

Official objective: “Managing and tracking experiments with MLflow, Weights & Biases, and custom tools”

MLflow, Weights & Biases and why every run must be traceable.

Key points

  1. Experiment tracking records what was run, with which settings, and how it scored. That lets you compare runs and reproduce the best one. ML means machine learning. XGBoost means eXtreme Gradient Boosting.

    What NVIDIA says (1)

    “Experiment tracking: MLflow logs hyperparameters, model artifacts, and evaluation metrics for every run”

    — RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

  2. Weights & Biases (W&B) is an experiment tracking platform. Like MLflow, it records runs so you can debug and reproduce models. ML means machine learning.

    What NVIDIA says (1)

    “Weights & Biases : W&B’s tools help many ML users build better models faster, debugging and reproducing their models with just a few lines of code”

    — What is MLOps?

  3. MLOps means machine learning operations: practices for running AI reliably in production. Tracking ties each model to its data, code and settings.

    What NVIDIA says (1)

    “AI models require careful tracking through cycles of experiments, tuning and retraining.”

    — What is MLOps?

Key terms: MLflow Weights & Biases MLOps Experiment tracking Model artifact

Practice 6.2 (3 questions) Objective page

6.3 Saving, loading and predicting

Official objective: “Model saving, loading, and prediction generation”

pickle and joblib, the pickle security risk, and ONNX export.

Key points

  1. Serialization turns a model into bytes that can be stored and loaded later. All single-GPU cuML estimators support the standard Python methods.

    What NVIDIA says (2)

    “Single GPU Model Serialization # All single-GPU cuML estimators support serialization using standard Python libraries.”

    — cuML: Pickling cuML models for persistence

    “This notebook demonstrates how to save and load cuML models using various serialization methods, including pickle, joblib, and cross-platform deployment strategies.”

    — cuML: Pickling cuML models for persistence

  2. Loading a pickle can execute code embedded in the file. Treat model files like programs.

    What NVIDIA says (2)

    “Security Warning # Only unpickle or deserialize models from trusted sources.”

    — cuML: Pickling cuML models for persistence

    “Malicious pickle data can execute arbitrary code during deserialization, potentially compromising your entire system.”

    — cuML: Pickling cuML models for persistence

  3. ONNX is a portable model format. Exporting to it removes the cuML dependency at inference time.

    What NVIDIA says (2)

    “Use as_sklearn() to convert the cuML model to a scikit-learn estimator, then pass it to skl2onnx.convert_sklearn() .”

    — cuML: Pickling cuML models for persistence

    “The resulting .onnx file can be loaded with ONNX Runtime for inference on both CPU and GPU, with no cuML dependency at inference time.”

    — cuML: Pickling cuML models for persistence

Key terms: Model serialization ONNX

Practice 6.3 (3 questions) Objective page

6.4 Detecting drift

Official objective: “Monitoring production models for drift and performance degradation”

Data drift, prediction drift and monitoring without ground truth.

Key points

  1. Drift means the world the model sees has changed. Statistical tests on feature distributions flag it early.

    What NVIDIA says (2)

    “Data drift : Changes in distribution between the training data and production data can be monitored to check for drift: this is done by detecting changes in the statistical properties of feature values over time.”

    — A Guide to Monitoring Machine Learning Models in Production

    “Statistical tests should be used to detect drift, and predictive performance should be monitored to evaluate the model’s performance over time.”

    — A Guide to Monitoring Machine Learning Models in Production

  2. Ground truth means the correct label. When it is delayed, the distribution of the model's own outputs is a useful proxy.

    What NVIDIA says (2)

    “Prediction drift: When it is not possible to acquire ground truth labels, predictions must be monitored.”

    — A Guide to Monitoring Machine Learning Models in Production

    “If there is a drastic change in the distribution of predictions, something has potentially gone wrong.”

    — A Guide to Monitoring Machine Learning Models in Production

Key terms: Model monitoring Data drift Prediction drift

Practice 6.4 (2 questions) Objective page

6.5 Artifacts and configurations for reproducibility

Official objective: “Managing model artifacts and configurations for reproducibility”

Model registries, champion aliases and configs in version control.

Key points

  1. A model registry stores model versions with names and aliases. An alias like 'champion' is a stable pointer that moves when a better model wins. MLOps means machine learning operations.

    What NVIDIA says (1)

    “The champion alias always points to the current production model: In a typical production cycle, start with a full pipeline run to establish a baseline champion.”

    — RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

  2. Version control records every change to files over time. Configs in version control can rebuild an environment exactly. MLOps means machine learning operations.

    What NVIDIA says (1)

    “Storing infrastructure configurations in version control so production environments can be quickly replicated, whether for fault-tolerance or to create a replica environment for development.”

    — What Is MLOps? (NVIDIA Glossary)

Key terms: MLflow Model registry Model artifact

Practice 6.5 (2 questions) Objective page

6.6 Benchmarking and choosing hardware

Official objective: “Benchmarking workflows and selecting optimal hardware”

Timing CPU and GPU runs fairly, and Multi-Instance GPU for small jobs.

Key points

  1. Benchmarking means timing the same workload on different hardware. Wall-clock time is the real elapsed time. HPO means hyperparameter optimization.

    What NVIDIA says (1)

    “In the notebook demo below, we compare benchmarking results to show how GPU can accelerate HPO tuning jobs relative to CPU.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU vs CPU benchmark

  2. MIG partitions one GPU into isolated instances, each with its own resources. It suits many small workloads; trade-offs include no NVLink between instances. NVLink means NVIDIA's high-speed GPU interconnect.

    What NVIDIA says (1)

    “Multi-Instance GPU is a technology that allows partitioning a single GPU into multiple instances, making each one seem as a completely independent GPU.”

    — RAPIDS Deployment: Multi-Instance GPU (MIG)

Key terms: Multi-Instance GPU Benchmarking

Practice 6.6 (2 questions) Objective page