1.7 The AI development and deployment life cycle

NCA-AIIO · Essential AI Knowledge (38% of the exam) · Official objective: “Describe the software components related to the life cycle of AI development and deployment.”

The software that supports each step from data to a model in production.

Key points

  1. The AI life cycle runs from data preparation to model development, deployment and monitoring. TensorRT sits between training and deployment. NVIDIA says it optimizes models trained on all major frameworks, calibrates them for lower precision with high accuracy, and deploys them from data centers to edge devices.

    What NVIDIA says (4)

    “It takes trained models from frameworks such as PyTorch and ONNX and compiles them into engines, which are optimized executable artifacts for a specific deployment configuration.”

    — NVIDIA TensorRT Documentation

    “TensorRT supports mixed precision”

    — NVIDIA TensorRT Documentation

    “TensorRT includes libraries that optimize neural network models trained on all major frameworks, calibrate them for lower precision with high accuracy, and deploy them to hyperscale data centers, workstations, laptops, and edge devices.”

    — NVIDIA TensorRT (developer page)

    “throughout the entire ML lifecycle—from exploratory analysis, data preparation, and model development to deployment, monitoring, and ongoing optimization.”

    — What Is MLOps? (NVIDIA Glossary)

  2. TensorRT is NVIDIA's SDK for optimizing deep learning inference on NVIDIA GPUs. Quantization stores numbers with fewer bits. Fusion merges several operations into one. Kernel tuning picks the fastest GPU code for each step. NVIDIA says TensorRT uses all three to optimize inference.

    What NVIDIA says (2)

    “is an SDK for optimizing deep learning inference on NVIDIA GPUs.”

    — NVIDIA TensorRT Documentation

    “optimizes inference using quantization, layer and tensor fusion, and kernel tuning techniques.”

    — NVIDIA TensorRT (developer page)

  3. MLOps (machine learning operations) applies DevOps ideas to machine learning. NVIDIA defines it as practices and principles that streamline the development, deployment and maintenance of ML models in production. It covers the whole life cycle, from exploratory analysis and data preparation to deployment, monitoring and ongoing optimization.

    What NVIDIA says (2)

    “short for machine learning operations, is a set of practices and principles that aims to streamline the development, deployment, and maintenance of machine learning (ML) models in production environments.”

    — What Is MLOps? (NVIDIA Glossary)

    “MLOps is an extension of the existing discipline of DevOps, the modern practice of efficiently writing, deploying, and running enterprise applications.”

    — What Is MLOps? (NVIDIA Glossary)

  4. Training from scratch needs a lot of representative data and compute. NVIDIA says a pretrained model is trained on large datasets and can be used as is or fine-tuned to fit an application's needs. That reuses earlier work instead of repeating it.

    What NVIDIA says (3)

    “You can use the pretrained models for inference or fine-tune them with transfer learning.”

    — NGC Catalog User Guide

    “It can be used as is or further fine-tuned to fit an application’s specific needs.”

    — What Is a Pretrained AI Model?

    “That model needs a lot of representative data to learn from.”

    — What Is a Pretrained AI Model?

Key terms

Try it

Sample question

In the AI life cycle, which step does TensorRT belong to?

Show the answer

Answer: Optimizing a trained model, including calibrating it for lower precision, before deployment.

The AI life cycle runs from data preparation to model development, deployment and monitoring. TensorRT sits between training and deployment. NVIDIA says it optimizes models trained on all major frameworks, calibrates them for lower precision with high accuracy, and deploys them from data centers to edge devices.

What NVIDIA says (4)

“It takes trained models from frameworks such as PyTorch and ONNX and compiles them into engines, which are optimized executable artifacts for a specific deployment configuration.”

— NVIDIA TensorRT Documentation

“TensorRT supports mixed precision”

— NVIDIA TensorRT Documentation

“TensorRT includes libraries that optimize neural network models trained on all major frameworks, calibrate them for lower precision with high accuracy, and deploy them to hyperscale data centers, workstations, laptops, and edge devices.”

— NVIDIA TensorRT (developer page)

“throughout the entire ML lifecycle—from exploratory analysis, data preparation, and model development to deployment, monitoring, and ongoing optimization.”

— What Is MLOps? (NVIDIA Glossary)

Practice 1.7 (4 questions) Full Essential AI Knowledge guide

← 1.6 NVIDIA solutions and what they are for · 1.8 GPU vs CPU architecture →