2.8 Energy-efficient, trustworthy models

NCA-GENM · Core Machine Learning and AI Knowledge (20% of the exam) · Official objective: “Contribute to the design, development, and deployment of energy-efficient, trustworthy multimodal AI models.”

Quantizing the costly part, curating less data, and being transparent.

Key points

  1. Quantization stores and computes with fewer bits, such as INT8, to save time and energy. NVIDIA uses ModelOpt to calibrate and quantize the SDXL UNet. VAE means variational autoencoder. INT8 means 8-bit integer.

    What NVIDIA says (1)

    “The UNet part typically consumes >95% of the e2e Stable Diffusion latency.”

    — NeMo Framework 24.09: Stable Diffusion XL Int8 Quantization

  2. Less data means fewer training steps and less energy. Removing near-duplicates keeps most of the useful signal.

    What NVIDIA says (1)

    “it can remove up to 50% of the data with minimal performance loss.”

    — NeMo Curator (NeMo Framework 24.09): Semantic Deduplication

  3. Trustworthy AI puts safety and transparency first. Being transparent includes sharing benchmarks and dataset descriptions.

    What NVIDIA says (1)

    “They’re also transparent — providing information such as accuracy benchmarks or a description of the training dataset”

    — What Is Trustworthy AI?

Key terms

Sample question

To make Stable Diffusion XL cheaper to run, which part should you quantize first, and why?

Show the answer

Answer: The UNet, because it takes most of the end-to-end latency

Quantization stores and computes with fewer bits, such as INT8, to save time and energy. NVIDIA uses ModelOpt to calibrate and quantize the SDXL UNet. VAE means variational autoencoder. INT8 means 8-bit integer.

What NVIDIA says (1)

“The UNet part typically consumes >95% of the e2e Stable Diffusion latency.”

— NeMo Framework 24.09: Stable Diffusion XL Int8 Quantization

Practice 2.8 (3 questions) Full Core Machine Learning and AI Knowledge guide

← 2.7 Emerging multimodal trends · 2.9 Prompt engineering principles →