2.4 Nonsequential networks and residual connections
Skip connections in ResNet and U-Nets, and ControlNet side branches.
Key points
A residual connection adds a block's input to its output, skipping the layers in between. A nonsequential network has paths like this that do not run strictly one layer after another. The skip path lets gradients reach early layers.
What NVIDIA says (1)
“ResNet allows deep neural networks to be trained thanks to the residual, or skip, connections, which let the gradient to flow through many network layers without vanishing.”
A U-Net has down-sampling and up-sampling paths joined by skip connections. Efficient UNet moves parameters to low-resolution blocks and scales the skips. NVIDIA says it converges faster and uses memory better.
What NVIDIA says (2)
“adding more residual blocks for the lower resolutions”
“Scaling skip connection by 1/sqrt(2)”
ControlNet is a side branch beside the original network. The locked copy preserves the pretrained model. The trainable copy learns the new condition.
What NVIDIA says (1)
“It copies the weights of neural network blocks into a “locked” copy and a “trainable” copy.”
Key terms
- U-Net: A convolutional network with a down-sampling path and an up-sampling path joined by skip connections, used as the denoiser in diffusion models.
- ControlNet: A side network that adds conditions such as edge maps to a diffusion model through a locked copy and a trainable copy of its blocks.
- Residual (skip) connection: A path that adds a block's input to its output, so gradients can flow through deep networks.
Sample question
Why do residual (skip) connections let very deep networks train?
Show the answer
Answer: They let the gradient flow through many layers without vanishing
A residual connection adds a block's input to its output, skipping the layers in between. A nonsequential network has paths like this that do not run strictly one layer after another. The skip path lets gradients reach early layers.
What NVIDIA says (1)
“ResNet allows deep neural networks to be trained thanks to the residual, or skip, connections, which let the gradient to flow through many network layers without vanishing.”
Practice 2.4 (3 questions) Full Core Machine Learning and AI Knowledge guide
← 2.3 Machine learning fundamentals · 2.5 Statistics for evaluating pipelines →