2.8 Energy-efficient, trustworthy models
Quantizing the costly part, curating less data, and being transparent.
Key points
Quantization stores and computes with fewer bits, such as INT8, to save time and energy. NVIDIA uses ModelOpt to calibrate and quantize the SDXL UNet. VAE means variational autoencoder. INT8 means 8-bit integer.
What NVIDIA says (1)
“The UNet part typically consumes >95% of the e2e Stable Diffusion latency.”
Less data means fewer training steps and less energy. Removing near-duplicates keeps most of the useful signal.
What NVIDIA says (1)
“it can remove up to 50% of the data with minimal performance loss.”
Trustworthy AI puts safety and transparency first. Being transparent includes sharing benchmarks and dataset descriptions.
What NVIDIA says (1)
“They’re also transparent — providing information such as accuracy benchmarks or a description of the training dataset”
Key terms
- Semantic deduplication: Removing data pairs that mean nearly the same thing, found by clustering embeddings and comparing cosine similarity.
- Quantization: Running a model with fewer bits, such as INT8, to cut latency and energy.
- Trustworthy AI: An approach to AI development that puts safety and transparency first.
Sample question
To make Stable Diffusion XL cheaper to run, which part should you quantize first, and why?
Show the answer
Answer: The UNet, because it takes most of the end-to-end latency
Quantization stores and computes with fewer bits, such as INT8, to save time and energy. NVIDIA uses ModelOpt to calibrate and quantize the SDXL UNet. VAE means variational autoencoder. INT8 means 8-bit integer.
What NVIDIA says (1)
“The UNet part typically consumes >95% of the e2e Stable Diffusion latency.”
Practice 2.8 (3 questions) Full Core Machine Learning and AI Knowledge guide
← 2.7 Emerging multimodal trends · 2.9 Prompt engineering principles →