4.4 U-Nets: from noise, and as an autoencoder
How denoising diffusion works, what the U-Net predicts and the VAE autoencoder.
Key points
A denoiser removes noise. Applied many times to pure noise, it reveals a new image. The training data decides what kinds of images appear.
What NVIDIA says (1)
“first draw a random image of pure white noise, and then chip away at the noise level”
The timestep tells the U-Net how noisy the input is. Subtracting the predicted noise moves the latent toward a clean image. VAE means variational autoencoder.
What NVIDIA says (1)
“The Unet processes the noisy latents (x) to predict the noise, utilizing a conditional model which also incorporates the timestep (t) and text embedding for guidance.”
At high noise, many clean images are possible, so the best guess is their average. Repeated small steps sharpen this into one image.
What NVIDIA says (1)
“the denoiser must output the blurry average of all possible clean images that could have been hiding under the noise.”
An autoencoder compresses data with an encoder and rebuilds it with a decoder. In Stable Diffusion the VAE plays this role around the U-Net. VAE means variational autoencoder.
What NVIDIA says (1)
“Subsequently, during inference, the decoder reverses this process by transforming denoised latent representations back into their original, tangible image forms.”
This is a cascade: one base model and super-resolution models in sequence. Each stage is a diffusion U-Net. GAN means generative adversarial network.
What NVIDIA says (1)
“Imagen first generates an image at a 64x64 resolution and then upsamples the generated image to 256x256 and 1024x1024 resolutions, all using diffusion models.”
Key terms
- Stable Diffusion: A latent text-to-image diffusion model made of a U-Net, a VAE image encoder and a CLIP text encoder.
- Denoising diffusion: A way to generate data by starting from pure noise and removing noise step by step with a trained denoiser.
- U-Net: A convolutional network with a down-sampling path and an up-sampling path joined by skip connections, used as the denoiser in diffusion models.
- Variational autoencoder: An encoder-decoder network that compresses images into a smaller latent space and rebuilds them.
- Latent space: A compressed representation of data in which a model can work more cheaply than on raw pixels.
- Imagen: A cascaded text-to-image diffusion model that generates at 64x64 and then upsamples to 256x256 and 1024x1024.
Try it
Sample question
What is the core idea of denoising diffusion?
Show the answer
Answer: Start from pure random noise and repeatedly denoise it until a clean image appears
A denoiser removes noise. Applied many times to pure noise, it reveals a new image. The training data decides what kinds of images appear.
What NVIDIA says (1)
“first draw a random image of pure white noise, and then chip away at the noise level”
Practice 4.4 (5 questions) Full Software Development guide
← 4.3 Prompting generative models · 4.5 CLIP for text-to-image →