4.4 U-Nets: from noise, and as an autoencoder

NCA-GENM · Software Development (15% of the exam) · Official objective: “Build a U-Net to generate images from pure noise and as a type of autoencoder.”

How denoising diffusion works, what the U-Net predicts and the VAE autoencoder.

Key points

  1. A denoiser removes noise. Applied many times to pure noise, it reveals a new image. The training data decides what kinds of images appear.

    What NVIDIA says (1)

    “first draw a random image of pure white noise, and then chip away at the noise level”

    — NVIDIA Technical Blog: Demystifying Diffusion-Based Models

  2. The timestep tells the U-Net how noisy the input is. Subtracting the predicted noise moves the latent toward a clean image. VAE means variational autoencoder.

    What NVIDIA says (1)

    “The Unet processes the noisy latents (x) to predict the noise, utilizing a conditional model which also incorporates the timestep (t) and text embedding for guidance.”

    — NeMo Framework 24.09: Stable Diffusion

  3. At high noise, many clean images are possible, so the best guess is their average. Repeated small steps sharpen this into one image.

    What NVIDIA says (1)

    “the denoiser must output the blurry average of all possible clean images that could have been hiding under the noise.”

    — NVIDIA Technical Blog: Demystifying Diffusion-Based Models

  4. An autoencoder compresses data with an encoder and rebuilds it with a decoder. In Stable Diffusion the VAE plays this role around the U-Net. VAE means variational autoencoder.

    What NVIDIA says (1)

    “Subsequently, during inference, the decoder reverses this process by transforming denoised latent representations back into their original, tangible image forms.”

    — NeMo Framework 24.09: Stable Diffusion

  5. This is a cascade: one base model and super-resolution models in sequence. Each stage is a diffusion U-Net. GAN means generative adversarial network.

    What NVIDIA says (1)

    “Imagen first generates an image at a 64x64 resolution and then upsamples the generated image to 256x256 and 1024x1024 resolutions, all using diffusion models.”

    — NeMo Framework 24.09: Imagen

Key terms

Try it

Sample question

What is the core idea of denoising diffusion?

Show the answer

Answer: Start from pure random noise and repeatedly denoise it until a clean image appears

A denoiser removes noise. Applied many times to pure noise, it reveals a new image. The training data decides what kinds of images appear.

What NVIDIA says (1)

“first draw a random image of pure white noise, and then chip away at the noise level”

— NVIDIA Technical Blog: Demystifying Diffusion-Based Models

Practice 4.4 (5 questions) Full Software Development guide

← 4.3 Prompting generative models · 4.5 CLIP for text-to-image →