4.3 Prompting generative models
Text-to-image prompts, DreamBooth identifiers, top-k, system prompts and Perfusion gating.
Key points
A text-to-image prompt is a description of the image you want. Its embeddings condition each denoising step. Clear, specific prompts give better guidance. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“The text-encoder, typically a simple transformer like CLIP, converts input prompts into embeddings, which guides the U-Net’s denoising process.”
DreamBooth teaches a model a new subject from a few images. It binds a unique identifier token to the subject. Using that token in a prompt recalls it in new scenes.
What NVIDIA says (1)
“so that it learns to bind a unique identifier with a special subject.”
Top-k limits the next-token choice to the k most likely tokens. Top-p instead keeps tokens whose probabilities add up to p. LLM means large language model.
What NVIDIA says (1)
“Top-k tells the model that it has to keep the top k highest probability tokens, from which the next token is selected at random. Lower values reduce randomness”
A system prompt is hidden setup text that frames every conversation. It sets role, tone and rules. LLM means large language model.
What NVIDIA says (1)
“This approach involves adding a system-level prompt in addition to the user prompt to provide specific and detailed instructions to the LLMs to behave as intended.”
Personalization teaches a model a specific concept, such as your own teddy bear. Perfusion adds a gate that controls how strongly the concept shows, without retraining. VAE means variational autoencoder.
What NVIDIA says (1)
“We also added a gating mechanism to regulate how strongly the learned concept is considered”
Key terms
- Text encoder: A network that turns a prompt into embeddings that guide a generative model.
- DreamBooth: A fine-tuning method that teaches a text-to-image model a specific subject from a few images, bound to a unique identifier.
- Prompt engineering: Writing model inputs to get the output you want.
- Top-k sampling: Sampling the next token only from the k most likely tokens.
- System prompt: Developer-written instructions sent with every user prompt to set the model's behavior.
Try it
Sample question
In Stable Diffusion, how does your text prompt affect the image?
Show the answer
Answer: A text encoder such as CLIP turns the prompt into embeddings that guide the U-Net's denoising
A text-to-image prompt is a description of the image you want. Its embeddings condition each denoising step. Clear, specific prompts give better guidance. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“The text-encoder, typically a simple transformer like CLIP, converts input prompts into embeddings, which guides the U-Net’s denoising process.”
Practice 4.3 (5 questions) Full Software Development guide
← 4.2 Best practices and software quality · 4.4 U-Nets: from noise, and as an autoencoder →