2.2 Multimodal loss functions
Contrastive loss, the diffusion denoising loss and DreamBooth prior preservation.
Key points
A loss function scores how wrong a model is; training lowers it. CLIP's loss compares every image with every caption in a batch. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”
A diffusion model learns to remove noise step by step. Its denoiser, often a U-Net, is trained with mean squared error (MSE). CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“Training a denoiser network (typically a U-Net) with basic loss—the mean square error between its output and the clean target—achieves precisely this result.”
Language drift is when the model forgets what a general word means after narrow fine-tuning. Prior preservation loss guides the model with its own generated samples of the general class. Its weight is set with model.prior_loss_weight. VAE means variational autoencoder. KL means Kullback-Leibler (a distance between distributions). L1 means absolute-difference.
What NVIDIA says (2)
“problems like language drift and decreased output variety often arise.”
“it guides the model using its self-generated samples and incorporates the discrepancy between the model-predicted noise on these samples.”
CLIP compares every image with every caption, which builds a large matrix across GPUs. local_loss avoids building the full global matrix. This helps memory when training on many devices. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“local_loss : If set to True , the loss is calculated with local features at a global level, avoiding the need to realize the full global matrix.”
Key terms
- Contrastive loss: A loss that rewards high similarity for matching pairs and low similarity for mismatched pairs.
- Prior preservation loss: A DreamBooth loss term that uses the model's own samples to stop language drift and loss of variety.
Try it
Sample question
What does CLIP's contrastive loss reward?
Show the answer
Answer: High similarity for correct image-text pairs and low similarity for incorrect pairs
A loss function scores how wrong a model is; training lowers it. CLIP's loss compares every image with every caption in a batch. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“maximizing the similarity between the correct (image, text) pairs while minimizing the similarity between incorrect pairs.”
Practice 2.2 (4 questions) Full Core Machine Learning and AI Knowledge guide
← 2.1 Training stability · 2.3 Machine learning fundamentals →