2.4 Nonsequential networks and residual connections

NCA-GENM · Core Machine Learning and AI Knowledge (20% of the exam) · Official objective: “Explain nonsequential neural networks and residual connections.”

Skip connections in ResNet and U-Nets, and ControlNet side branches.

Key points

  1. A residual connection adds a block's input to its output, skipping the layers in between. A nonsequential network has paths like this that do not run strictly one layer after another. The skip path lets gradients reach early layers.

    What NVIDIA says (1)

    “ResNet allows deep neural networks to be trained thanks to the residual, or skip, connections, which let the gradient to flow through many network layers without vanishing.”

    — NVIDIA Technical Blog: Accelerating AI Training with MLPerf Containers and Models from NGC

  2. A U-Net has down-sampling and up-sampling paths joined by skip connections. Efficient UNet moves parameters to low-resolution blocks and scales the skips. NVIDIA says it converges faster and uses memory better.

    What NVIDIA says (2)

    “adding more residual blocks for the lower resolutions”

    — NeMo Framework 24.09: Imagen

    “Scaling skip connection by 1/sqrt(2)”

    — NeMo Framework 24.09: Imagen

  3. ControlNet is a side branch beside the original network. The locked copy preserves the pretrained model. The trainable copy learns the new condition.

    What NVIDIA says (1)

    “It copies the weights of neural network blocks into a “locked” copy and a “trainable” copy.”

    — NeMo Framework 24.09: ControlNet

Key terms

Sample question

Why do residual (skip) connections let very deep networks train?

Show the answer

Answer: They let the gradient flow through many layers without vanishing

A residual connection adds a block's input to its output, skipping the layers in between. A nonsequential network has paths like this that do not run strictly one layer after another. The skip path lets gradients reach early layers.

What NVIDIA says (1)

“ResNet allows deep neural networks to be trained thanks to the residual, or skip, connections, which let the gradient to flow through many network layers without vanishing.”

— NVIDIA Technical Blog: Accelerating AI Training with MLPerf Containers and Models from NGC

Practice 2.4 (3 questions) Full Core Machine Learning and AI Knowledge guide

← 2.3 Machine learning fundamentals · 2.5 Statistics for evaluating pipelines →