2.2 Multimodal loss functions

NCA-GENM · Core Machine Learning and AI Knowledge (20% of the exam) · Official objective: “Develop content introducing multimodal loss functions.”

Contrastive loss, the diffusion denoising loss and DreamBooth prior preservation.

Key points

  1. A loss function scores how wrong a model is; training lowers it. CLIP's loss compares every image with every caption in a batch. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”

    — NeMo Framework 24.09: CLIP

  2. A diffusion model learns to remove noise step by step. Its denoiser, often a U-Net, is trained with mean squared error (MSE). CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “Training a denoiser network (typically a U-Net) with basic loss—the mean square error between its output and the clean target—achieves precisely this result.”

    — NVIDIA Technical Blog: Demystifying Diffusion-Based Models

  3. Language drift is when the model forgets what a general word means after narrow fine-tuning. Prior preservation loss guides the model with its own generated samples of the general class. Its weight is set with model.prior_loss_weight. VAE means variational autoencoder. KL means Kullback-Leibler (a distance between distributions). L1 means absolute-difference.

    What NVIDIA says (2)

    “problems like language drift and decreased output variety often arise.”

    — NeMo Framework 24.09: DreamBooth

    “it guides the model using its self-generated samples and incorporates the discrepancy between the model-predicted noise on these samples.”

    — NeMo Framework 24.09: DreamBooth

  4. CLIP compares every image with every caption, which builds a large matrix across GPUs. local_loss avoids building the full global matrix. This helps memory when training on many devices. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “local_loss : If set to True , the loss is calculated with local features at a global level, avoiding the need to realize the full global matrix.”

    — NeMo Framework 24.09: CLIP

Key terms

Try it

Sample question

What does CLIP's contrastive loss reward?

Show the answer

Answer: High similarity for correct image-text pairs and low similarity for incorrect pairs

A loss function scores how wrong a model is; training lowers it. CLIP's loss compares every image with every caption in a batch. CLIP means Contrastive Language-Image Pre-training.

What NVIDIA says (1)

“maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”

— NeMo Framework 24.09: CLIP

Practice 2.2 (4 questions) Full Core Machine Learning and AI Knowledge guide

← 2.1 Training stability · 2.3 Machine learning fundamentals →