6.3 Saving, loading and predicting

NCA-ADS · Introductory MLOps Practices (10% of the exam) · Official objective: “Model saving, loading, and prediction generation”

pickle and joblib, the pickle security risk, and ONNX export.

Key points

  1. Serialization turns a model into bytes that can be stored and loaded later. All single-GPU cuML estimators support the standard Python methods.

    What NVIDIA says (2)

    “Single GPU Model Serialization # All single-GPU cuML estimators support serialization using standard Python libraries.”

    — cuML: Pickling cuML models for persistence

    “This notebook demonstrates how to save and load cuML models using various serialization methods, including pickle, joblib, and cross-platform deployment strategies.”

    — cuML: Pickling cuML models for persistence

  2. Loading a pickle can execute code embedded in the file. Treat model files like programs.

    What NVIDIA says (2)

    “Security Warning # Only unpickle or deserialize models from trusted sources.”

    — cuML: Pickling cuML models for persistence

    “Malicious pickle data can execute arbitrary code during deserialization, potentially compromising your entire system.”

    — cuML: Pickling cuML models for persistence

  3. ONNX is a portable model format. Exporting to it removes the cuML dependency at inference time.

    What NVIDIA says (2)

    “Use as_sklearn() to convert the cuML model to a scikit-learn estimator, then pass it to skl2onnx.convert_sklearn() .”

    — cuML: Pickling cuML models for persistence

    “The resulting .onnx file can be loaded with ONNX Runtime for inference on both CPU and GPU, with no cuML dependency at inference time.”

    — cuML: Pickling cuML models for persistence

Key terms

Sample question

How can you save a trained single-GPU cuML model and load it later to make predictions?

Show the answer

Answer: Serialize it with pickle or joblib, then load it and call predict()

Serialization turns a model into bytes that can be stored and loaded later. All single-GPU cuML estimators support the standard Python methods.

What NVIDIA says (2)

“Single GPU Model Serialization # All single-GPU cuML estimators support serialization using standard Python libraries.”

— cuML: Pickling cuML models for persistence

“This notebook demonstrates how to save and load cuML models using various serialization methods, including pickle, joblib, and cross-platform deployment strategies.”

— cuML: Pickling cuML models for persistence

Practice 6.3 (3 questions) Full Introductory MLOps Practices guide

← 6.2 Tracking experiments · 6.4 Detecting drift →