6.4 Detecting drift

NCA-ADS · Introductory MLOps Practices (10% of the exam) · Official objective: “Monitoring production models for drift and performance degradation”

Data drift, prediction drift and monitoring without ground truth.

Key points

  1. Drift means the world the model sees has changed. Statistical tests on feature distributions flag it early.

    What NVIDIA says (2)

    “Data drift : Changes in distribution between the training data and production data can be monitored to check for drift: this is done by detecting changes in the statistical properties of feature values over time.”

    — A Guide to Monitoring Machine Learning Models in Production

    “Statistical tests should be used to detect drift, and predictive performance should be monitored to evaluate the model’s performance over time.”

    — A Guide to Monitoring Machine Learning Models in Production

  2. Ground truth means the correct label. When it is delayed, the distribution of the model's own outputs is a useful proxy.

    What NVIDIA says (2)

    “Prediction drift: When it is not possible to acquire ground truth labels, predictions must be monitored.”

    — A Guide to Monitoring Machine Learning Models in Production

    “If there is a drastic change in the distribution of predictions, something has potentially gone wrong.”

    — A Guide to Monitoring Machine Learning Models in Production

Key terms

Sample question

What is data drift, and how do you detect it?

Show the answer

Answer: A change in the distribution of features between training and production, detected by tracking their statistical properties over time

Drift means the world the model sees has changed. Statistical tests on feature distributions flag it early.

What NVIDIA says (2)

“Data drift : Changes in distribution between the training data and production data can be monitored to check for drift: this is done by detecting changes in the statistical properties of feature values over time.”

— A Guide to Monitoring Machine Learning Models in Production

“Statistical tests should be used to detect drift, and predictive performance should be monitored to evaluate the model’s performance over time.”

— A Guide to Monitoring Machine Learning Models in Production

Practice 6.4 (2 questions) Full Introductory MLOps Practices guide

← 6.3 Saving, loading and predicting · 6.5 Artifacts and configurations for reproducibility →