4.3 Prompting generative models

NCA-GENM · Software Development (15% of the exam) · Official objective: “Use prompt engineering to better influence the output of generative AI models.”

Text-to-image prompts, DreamBooth identifiers, top-k, system prompts and Perfusion gating.

Key points

  1. A text-to-image prompt is a description of the image you want. Its embeddings condition each denoising step. Clear, specific prompts give better guidance. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “The text-encoder, typically a simple transformer like CLIP, converts input prompts into embeddings, which guides the U-Net’s denoising process.”

    — NeMo Framework 24.09: Stable Diffusion

  2. DreamBooth teaches a model a new subject from a few images. It binds a unique identifier token to the subject. Using that token in a prompt recalls it in new scenes.

    What NVIDIA says (1)

    “so that it learns to bind a unique identifier with a special subject.”

    — NeMo Framework 24.09: DreamBooth

  3. Top-k limits the next-token choice to the k most likely tokens. Top-p instead keeps tokens whose probabilities add up to p. LLM means large language model.

    What NVIDIA says (1)

    “Top-k tells the model that it has to keep the top k highest probability tokens, from which the next token is selected at random. Lower values reduce randomness”

    — How to Get Better Outputs from Your Large Language Model

  4. A system prompt is hidden setup text that frames every conversation. It sets role, tone and rules. LLM means large language model.

    What NVIDIA says (1)

    “This approach involves adding a system-level prompt in addition to the user prompt to provide specific and detailed instructions to the LLMs to behave as intended.”

    — Mastering LLM Techniques: Customization

  5. Personalization teaches a model a specific concept, such as your own teddy bear. Perfusion adds a gate that controls how strongly the concept shows, without retraining. VAE means variational autoencoder.

    What NVIDIA says (1)

    “We also added a gating mechanism to regulate how strongly the learned concept is considered”

    — NVIDIA Technical Blog: Personalizing Text-to-Image Models

Key terms

Try it

Sample question

In Stable Diffusion, how does your text prompt affect the image?

Show the answer

Answer: A text encoder such as CLIP turns the prompt into embeddings that guide the U-Net's denoising

A text-to-image prompt is a description of the image you want. Its embeddings condition each denoising step. Clear, specific prompts give better guidance. CLIP means Contrastive Language-Image Pre-training.

What NVIDIA says (1)

“The text-encoder, typically a simple transformer like CLIP, converts input prompts into embeddings, which guides the U-Net’s denoising process.”

— NeMo Framework 24.09: Stable Diffusion

Practice 4.3 (5 questions) Full Software Development guide

← 4.2 Best practices and software quality · 4.4 U-Nets: from noise, and as an autoencoder →