1.7 The AI development and deployment life cycle
The software that supports each step from data to a model in production.
Key points
The AI life cycle runs from data preparation to model development, deployment and monitoring. TensorRT sits between training and deployment. NVIDIA says it optimizes models trained on all major frameworks, calibrates them for lower precision with high accuracy, and deploys them from data centers to edge devices.
What NVIDIA says (4)
“It takes trained models from frameworks such as PyTorch and ONNX and compiles them into engines, which are optimized executable artifacts for a specific deployment configuration.”
“TensorRT supports mixed precision”
“TensorRT includes libraries that optimize neural network models trained on all major frameworks, calibrate them for lower precision with high accuracy, and deploy them to hyperscale data centers, workstations, laptops, and edge devices.”
“throughout the entire ML lifecycle—from exploratory analysis, data preparation, and model development to deployment, monitoring, and ongoing optimization.”
TensorRT is NVIDIA's SDK for optimizing deep learning inference on NVIDIA GPUs. Quantization stores numbers with fewer bits. Fusion merges several operations into one. Kernel tuning picks the fastest GPU code for each step. NVIDIA says TensorRT uses all three to optimize inference.
What NVIDIA says (2)
“is an SDK for optimizing deep learning inference on NVIDIA GPUs.”
“optimizes inference using quantization, layer and tensor fusion, and kernel tuning techniques.”
MLOps (machine learning operations) applies DevOps ideas to machine learning. NVIDIA defines it as practices and principles that streamline the development, deployment and maintenance of ML models in production. It covers the whole life cycle, from exploratory analysis and data preparation to deployment, monitoring and ongoing optimization.
What NVIDIA says (2)
“short for machine learning operations, is a set of practices and principles that aims to streamline the development, deployment, and maintenance of machine learning (ML) models in production environments.”
“MLOps is an extension of the existing discipline of DevOps, the modern practice of efficiently writing, deploying, and running enterprise applications.”
Training from scratch needs a lot of representative data and compute. NVIDIA says a pretrained model is trained on large datasets and can be used as is or fine-tuned to fit an application's needs. That reuses earlier work instead of repeating it.
What NVIDIA says (3)
“You can use the pretrained models for inference or fine-tune them with transfer learning.”
“It can be used as is or further fine-tuned to fit an application’s specific needs.”
“That model needs a lot of representative data to learn from.”
Key terms
- Pretrained model: A model already trained on a large dataset that you can use as is or fine-tune for your task.
- MLOps: Practices for building, deploying and maintaining ML models in production, extending DevOps.
Try it
Sample question
In the AI life cycle, which step does TensorRT belong to?
Show the answer
Answer: Optimizing a trained model, including calibrating it for lower precision, before deployment.
The AI life cycle runs from data preparation to model development, deployment and monitoring. TensorRT sits between training and deployment. NVIDIA says it optimizes models trained on all major frameworks, calibrates them for lower precision with high accuracy, and deploys them from data centers to edge devices.
What NVIDIA says (4)
“It takes trained models from frameworks such as PyTorch and ONNX and compiles them into engines, which are optimized executable artifacts for a specific deployment configuration.”
“TensorRT supports mixed precision”
“TensorRT includes libraries that optimize neural network models trained on all major frameworks, calibrate them for lower precision with high accuracy, and deploy them to hyperscale data centers, workstations, laptops, and edge devices.”
“throughout the entire ML lifecycle—from exploratory analysis, data preparation, and model development to deployment, monitoring, and ongoing optimization.”
Practice 1.7 (4 questions) Full Essential AI Knowledge guide
← 1.6 NVIDIA solutions and what they are for · 1.8 GPU vs CPU architecture →