Skip to content
CS-702 (B) · Deep & Reinforcement Learning/Quick Revision Short Notes

Deep & Reinforcement Learning (CS-702 (B)) - Unit 3 Short Notes

How unit 3 is examined

This unit covers pre-training, activations, initialization, CNNs and their famous architectures, and visualization methods; the marks sit in greedy layerwise pre-training, guided backpropagation and Deep Dream (14 marks each), then weight initialization and CNN architecture (7 marks each).

Greedy Layerwise Pre-training

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Greedy layerwise pre-training trains a deep network one layer at a time, each layer learned without labels on the output of the layer below, and the stacked weights are then used to initialize the network for supervised fine-tuning.</mark>

Key points.

  1. Layer 1 is trained as an autoencoder or RBM on the raw input, then frozen, and its hidden output becomes the training data for layer 2, and so on.
  2. It is called greedy because each layer is optimized alone for its own local objective, not for the final task loss.
  3. The stage is unsupervised, so it uses plentiful unlabelled data and gives an unsupervised initialization that starts the weights in a good region.
  4. After stacking, a classifier layer is added and the whole network is fine-tuned with backpropagation on labelled data.
  5. It gave deep belief networks (stacked RBMs, Hinton 2006) and stacked autoencoders their success, because plain random initialization with sigmoid units suffered vanishing gradients and poor local minima.
  6. It also acts as a regularizer, since the initial weights already capture structure of the input distribution.
  7. It has largely been replaced by ReLU, better initialization, batch normalization and large labelled data, but is still used when labels are scarce.

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-01" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="Layer-by-layer training; X input, H1-H3 hidden layers, Y softmax output added for fine-tuning"><style>#dsfig-u3-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-01 .t{fill:#16181D;font-weight:500}#dsfig-u3-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-01 .dot{fill:#16181D}#dsfig-u3-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-01 .ah{fill:#454C5A}#dsfig-u3-01 .ah.hi{fill:#2340B8}#dsfig-u3-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-01 .e{stroke:#B1B7C3}html.dark #dsfig-u3-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-01 .t{fill:#E6E8ED}html.dark #dsfig-u3-01 .t.inv{fill:#0F1115}html.dark #dsfig-u3-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-01 .dot{fill:#E6E8ED}html.dark #dsfig-u3-01 .ann{fill:#8FA3FF}html.dark #dsfig-u3-01 .lbl{fill:#858D9C}html.dark #dsfig-u3-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-01 .ah{fill:#B1B7C3}html.dark #dsfig-u3-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah6)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah6)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah6)"/><path class="e" d="M446,40 L535,40" marker-end="url(#ah6)"/><g class="wl"><rect x="77.4" y="31" width="54.3" height="18" rx="9"/><text class="t" x="104.5" y="40" dy=".35em" text-anchor="middle">train1</text></g><g class="wl"><rect x="206.4" y="31" width="54.3" height="18" rx="9"/><text class="t" x="233.5" y="40" dy=".35em" text-anchor="middle">train2</text></g><g class="wl"><rect x="335.4" y="31" width="54.3" height="18" rx="9"/><text class="t" x="362.5" y="40" dy=".35em" text-anchor="middle">train3</text></g><g class="wl"><rect x="457.2" y="31" width="68.7" height="18" rx="9"/><text class="t" x="491.5" y="40" dy=".35em" text-anchor="middle">finetune</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">X</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">H1</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">H2</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">H3</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">Y</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Layer-by-layer training; X input, H1-H3 hidden layers, Y softmax output added for fine-tuning</figcaption></figure>

Better activation functions (recent years).

  1. ReLU, $\max(0,x)$, does not saturate for $x>0$, so gradients do not vanish and training is fast, but neurons with $x<0$ can die.
  2. Leaky ReLU, $\max(0.01x,x)$, and PReLU keep a small slope for negatives, so no neuron dies.
  3. ELU, $x$ for $x>0$ and $\alpha(e^{x}-1)$ otherwise, gives smooth negative outputs and mean activation near zero.
  4. GELU, $x\,\Phi(x)$, weights the input by the Gaussian CDF and is used in transformers.
  5. Swish, $x\,\sigma(x)$, and Mish, $x\tanh(\mathrm{softplus}(x))$, are smooth, non-monotonic and often beat ReLU in deep nets.

Answer frame. Open with the definition; draw the layer-by-layer figure; develop points 1-5 (procedure, greedy, unsupervised, fine-tune, DBN role); then list the five activations with formula and advantage (non-saturation, faster training, vanishing-gradient relief); close with "pre-training is now replaced by ReLU-family activations and good initialization".

Asked: [14 marks] (Jun 2025) Discuss the concept of greedy layerwise pre-training. What are some better activation functions introduced in recent years?

Better activation functions

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. An activation function adds non-linearity to a neuron, and a better one avoids saturation and vanishing gradients.

Key points.

  1. Sigmoid and tanh saturate at large $|x|$, so their gradients shrink layer by layer.
  2. ReLU, Leaky ReLU, ELU, GELU, Swish and Mish (formulas above) give non-saturating gradients for positive inputs and faster convergence.
  3. Softmax is used only at the output for class probabilities: $e^{z_i}/\sum_j e^{z_j}$.
  4. Choose ReLU or Leaky ReLU as the default hidden activation, and GELU in transformers.

Better weight initialization methods

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Weight initialization sets the starting weights so that the variance of activations and gradients stays roughly constant across layers, which prevents vanishing and exploding gradients.</mark>

Key points.

  1. Zero initialization makes every neuron in a layer compute and learn the same thing (symmetry), so it must never be used for weights.
  2. Small random weights break symmetry but the signal shrinks layer by layer in deep nets (vanishing), while large random weights make it blow up (exploding) or saturate sigmoid and tanh.
  3. Xavier (Glorot) initialization suits tanh and sigmoid and keeps variance equal in forward and backward passes: $\mathrm{Var}(w)=\dfrac{2}{n_{in}+n_{out}}$.
  4. He initialization suits ReLU, which zeroes half the inputs, so the variance is doubled: $\mathrm{Var}(w)=\dfrac{2}{n_{in}}$.
  5. LeCun initialization uses $\mathrm{Var}(w)=\dfrac{1}{n_{in}}$ and suits SELU and older tanh nets.
  6. Orthogonal initialization uses an orthogonal matrix so the norm of the signal is preserved, and it is common in RNNs.
  7. Biases are usually set to zero, and pre-trained weights (transfer learning or layerwise pre-training) are another good starting point.

Answer frame. Open with why initialization matters for vanishing and exploding gradients; develop points 1-2 (zero, random), then Xavier, He, LeCun with their variances, then orthogonal and pretrained; close with "match the initializer to the activation: Xavier for tanh, He for ReLU".

Asked: [7 marks] (Dec 2020, Nov 2023) Explain Better Weight Initialization Methods.

Learning Vectorial Representations Of Words

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Word embeddings map each word to a dense real vector, learned so that words in similar contexts get similar vectors.

Key points.

  1. One-hot vectors are huge and sparse and carry no similarity, so dense vectors of 100-300 dimensions are learned instead.
  2. Word2Vec has two models: CBOW predicts a word from its context, and skip-gram predicts the context from the word.
  3. GloVe factorizes the word co-occurrence matrix, and both give analogies such as $king-man+woman\approx queen$.
  4. Similarity is measured with cosine similarity, and the vectors feed RNNs and other models.

Convolutional Neural Networks

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>A convolutional neural network is a feed-forward network for grid-like data such as images that uses shared local filters (convolution), pooling and fully connected layers to learn a hierarchy of features.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-02" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="CNN: Input image, Convolution, ReLU, Pooling (Conv-ReLU-Pool repeated), Flatten, Fully connected with softmax output"><style>#dsfig-u3-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-02 .t{fill:#16181D;font-weight:500}#dsfig-u3-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-02 .dot{fill:#16181D}#dsfig-u3-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-02 .ah{fill:#454C5A}#dsfig-u3-02 .ah.hi{fill:#2340B8}#dsfig-u3-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-02 .e{stroke:#B1B7C3}html.dark #dsfig-u3-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-02 .t{fill:#E6E8ED}html.dark #dsfig-u3-02 .t.inv{fill:#0F1115}html.dark #dsfig-u3-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-02 .dot{fill:#E6E8ED}html.dark #dsfig-u3-02 .ann{fill:#8FA3FF}html.dark #dsfig-u3-02 .lbl{fill:#858D9C}html.dark #dsfig-u3-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-02 .ah{fill:#B1B7C3}html.dark #dsfig-u3-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L122.2,40" marker-end="url(#ah7)"/><path class="e" d="M162.2,40 L225.4,40" marker-end="url(#ah7)"/><path class="e" d="M265.4,40 L328.6,40" marker-end="url(#ah7)"/><path class="e" d="M368.6,40 L431.8,40" marker-end="url(#ah7)"/><path class="e" d="M471.8,40 L535,40" marker-end="url(#ah7)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">In</text><circle class="n" cx="143.2" cy="40" r="18"/><text class="t" x="143.2" y="40" dy=".35em" text-anchor="middle">Cv</text><circle class="n" cx="246.4" cy="40" r="18"/><text class="t" x="246.4" y="40" dy=".35em" text-anchor="middle">Rl</text><circle class="n" cx="349.6" cy="40" r="18"/><text class="t" x="349.6" y="40" dy=".35em" text-anchor="middle">Pl</text><circle class="n" cx="452.8" cy="40" r="18"/><text class="t" x="452.8" y="40" dy=".35em" text-anchor="middle">Fl</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">FC</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">CNN: Input image, Convolution, ReLU, Pooling (Conv-ReLU-Pool repeated), Flatten, Fully connected with softmax output</figcaption></figure>

Key points.

  1. The convolution layer slides small filters (for example 3x3) over the input, and each filter produces one feature map that shows where its pattern occurs.
  2. Filters share weights across positions and see only a local patch, so parameters are far fewer than in a dense network.
  3. Stride is the step of the filter and padding adds border zeros, and the output size is $\dfrac{W-F+2P}{S}+1$; for example $W=32,F=5,P=0,S=1$ gives 28.
  4. ReLU is applied after convolution to add non-linearity and avoid vanishing gradients.
  5. Pooling (max or average, usually 2x2 with stride 2) halves the size, cuts computation and gives small translation invariance.
  6. Early layers learn edges, middle layers learn textures and parts, and deep layers learn whole objects (hierarchical features).
  7. Flattened features go to fully connected layers and a softmax that outputs class probabilities, and the whole network is trained by backpropagation; LeNet-5 and AlexNet follow exactly this flow.

Answer frame. Open with the definition; draw the block figure with labelled layers and data flow; develop points 1-5 in layer order; then hierarchy and the classifier; close with applications (image classification, detection, medical imaging).

Asked: [7 marks] (Dec 2020, Nov 2023) Draw and explain the architecture of Convolutional Network.

LeNet

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. LeNet-5 (LeCun, 1998) is the first successful CNN, built for handwritten digit recognition on 32x32 grayscale images.

Key points.

  1. The layers are C1 (6 maps, 5x5, 28x28), S2 (average pool, 14x14), C3 (16 maps, 10x10), S4 (pool, 5x5), C5 (120), F6 (84) and an output of 10 classes.
  2. It has about 60,000 parameters and used tanh or sigmoid with average pooling.
  3. It showed that convolution and pooling learned by backpropagation beat hand-made features, and it was used to read bank cheques.

AlexNet

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. AlexNet (Krizhevsky, 2012) is an eight-layer CNN that won ImageNet 2012 by a large margin and started the deep learning boom.

Key points.

  1. It has 5 convolution layers and 3 fully connected layers, with about 60 million parameters and a 224x224x3 input.
  2. It used ReLU instead of tanh, which trained several times faster.
  3. Dropout in the fully connected layers and data augmentation reduced overfitting.
  4. It used overlapping max pooling, local response normalization, and two GPUs for training.

ZF-Net

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ZF-Net (Zeiler and Fergus, 2013) is an improved AlexNet that won ImageNet 2013 and introduced deconvnet visualization.

Key points.

  1. It reduced the first filter from 11x11 to 7x7 and the stride from 4 to 2, which kept more detail.
  2. It used a deconvolution network to project activations back to pixels and see what each layer detects.
  3. These visualizations guided the design changes, showing that visualization can improve architecture.

VGGNet

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. VGGNet (Simonyan and Zisserman, 2014) is a very deep CNN built only from stacked 3x3 convolutions with 2x2 max pooling.

Key points.

  1. VGG16 has 13 convolution and 3 fully connected layers (16 weight layers), and VGG19 has 19.
  2. Two stacked 3x3 layers see a 5x5 region and three see 7x7, with fewer parameters and more non-linearity than one large filter.
  3. The design is simple and uniform, but it has about 138 million parameters and is heavy in memory.

GoogLeNet

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. GoogLeNet (Inception v1, 2014) is a 22-layer CNN built from Inception modules that apply 1x1, 3x3, 5x5 convolutions and pooling in parallel and concatenate the results.

Key points.

  1. Parallel filters capture features at several scales in one layer.
  2. 1x1 convolutions reduce the channel depth before the costly 3x3 and 5x5 filters, cutting computation.
  3. It replaces the large fully connected layers with global average pooling, so it has only about 5-7 million parameters.
  4. Auxiliary classifiers in the middle help gradients flow, and it won ImageNet 2014.

ResNet

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. ==ResNet (He et al., 2015) is a very deep CNN of residual blocks in which a skip connection adds the input to the block output, $y=F(x)+x$.==

Key points.

  1. Plain networks get worse when made deeper (the degradation problem), because the layers struggle to learn even an identity mapping.
  2. A residual block only learns the change $F(x)=H(x)-x$, and if the identity is best it pushes $F(x)$ to zero.
  3. The skip path gives the gradient a direct route, since $\partial y/\partial x=\partial F/\partial x+1$, which relieves vanishing gradients.
  4. ResNet-152 won ImageNet 2015 with 3.57% top-5 error, using batch normalization and bottleneck blocks.
  5. Compared with LeNet-5 (5 layers, digits, 60 thousand parameters), ResNet is over 100 layers deep, classifies ImageNet, and is used as a backbone in detection.

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-03" viewBox="0 0 596 166" width="596" height="166" role="img" aria-label="Residual block; X input, W1 and W2 weight layers, Add adds F(x) and x, Y output"><style>#dsfig-u3-03 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-03 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-03 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-03 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-03 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-03 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-03 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-03 .t{fill:#16181D;font-weight:500}#dsfig-u3-03 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-03 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-03 .dot{fill:#16181D}#dsfig-u3-03 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-03 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-03 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-03 .ah{fill:#454C5A}#dsfig-u3-03 .ah.hi{fill:#2340B8}#dsfig-u3-03 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-03 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-03 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-03 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-03 .e{stroke:#B1B7C3}html.dark #dsfig-u3-03 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-03 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-03 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-03 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-03 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-03 .t{fill:#E6E8ED}html.dark #dsfig-u3-03 .t.inv{fill:#0F1115}html.dark #dsfig-u3-03 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-03 .dot{fill:#E6E8ED}html.dark #dsfig-u3-03 .ann{fill:#8FA3FF}html.dark #dsfig-u3-03 .lbl{fill:#858D9C}html.dark #dsfig-u3-03 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-03 .ah{fill:#B1B7C3}html.dark #dsfig-u3-03 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-03 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-03 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-03 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.8,115.5 L151.5,51.6" marker-end="url(#ah8)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah8)"/><path class="e" d="M313.8,50.5 L409.5,114.4" marker-end="url(#ah8)"/><path class="e" d="M59,126 L406,126" marker-end="url(#ah8)"/><path class="e" d="M446,126 L535,126" marker-end="url(#ah8)"/><g class="wl"><rect x="342.1" y="74" width="40.8" height="18" rx="9"/><text class="t" x="362.5" y="83" dy=".35em" text-anchor="middle">F(x)</text></g><g class="wl"><rect x="213.1" y="117" width="40.8" height="18" rx="9"/><text class="t" x="233.5" y="126" dy=".35em" text-anchor="middle">skip</text></g><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">X</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">W1</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">W2</text><circle class="n" cx="427" cy="126" r="18"/><text class="t" x="427" y="126" dy=".35em" text-anchor="middle">Add</text><circle class="n" cx="556" cy="126" r="18"/><text class="t" x="556" y="126" dy=".35em" text-anchor="middle">Y</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Residual block; X input, W1 and W2 weight layers, Add adds F(x) and x, Y output</figcaption></figure>

Answer frame. Give LeNet-5 first (see LeNet: layers and digit task), then the residual block figure with $y=F(x)+x$; explain degradation and gradient relief; close with a depth and use-case comparison.

Asked: [7 marks] (Dec 2020) Explain ResNet and LeNet in detail.

Visualizing Convolutional Neural Networks

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. CNN visualization shows what the filters and layers respond to, so that the network is not a black box.

Key points.

  1. Feature maps and filters can be plotted directly, showing edges in early layers and object parts in later ones.
  2. Saliency maps use the gradient of the class score with respect to the input pixels to show which pixels matter.
  3. Deconvnet, guided backpropagation, activation maximization and Deep Dream are the main techniques.

Guided Backpropagation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Guided backpropagation is a visualization method that backpropagates from a neuron to the input image but lets only positive gradients pass through each ReLU, and only where the forward activation was also positive.</mark>

Key points.

  1. Ordinary backpropagation passes the gradient through a ReLU wherever the forward input was positive.
  2. Deconvnet passes it wherever the incoming gradient is positive, and guided backpropagation applies both conditions: $R_l=\mathbb{1}[f_l>0]\cdot\mathbb{1}[R_{l+1}>0]\cdot R_{l+1}$.
  3. Masking the negative gradients removes evidence that lowers the neuron's activation, so only pixels that excite it remain.
  4. The result is a sharp, clean image of edges and shapes that the neuron or class responds to.
  5. It is used for saliency, explaining predictions and debugging, and it needs no change to the network.

Dataset augmentation. It enlarges the training set by applying label-preserving transforms such as flip, rotation, crop, scale, shift, brightness change and noise, which reduces overfitting and improves generalization; it is standard in image tasks.

LSTM. Long Short-Term Memory is an RNN cell with a cell state $c_t$ that carries information over long spans and three gates: the forget gate $f_t$ decides what to erase, the input gate $i_t$ decides what to write, and the output gate $o_t$ decides what to expose as $h_t$. Because the cell state is updated additively, $c_t=f_t\odot c_{t-1}+i_t\odot\tilde c_t$, gradients do not vanish easily, so it handles long dependencies in language, speech and time series.

Answer frame. Give three short parts of about equal length: definition, mechanism and application for each of the three terms, with the LSTM gate equation and the guided-backprop masking rule written out.

Asked: [14 marks] (Dec 2020) Explain following term: i) Guided Back propagation ii) Dataset augmentation iii) LSTM

Deep Dream

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Deep Dream is a technique that uses gradient ascent on the input image to maximize the activations of a chosen layer, so the network amplifies the patterns it detects and shows what it has learned.</mark>

Key points.

  1. Networks are hard to interpret, so visualization is needed to trust, debug and improve them.
  2. Deep Dream feeds an image through a trained CNN and picks a layer, and its loss is the sum of squared activations of that layer.
  3. It then updates the image, not the weights: $x\leftarrow x+\eta\,\partial L/\partial x$, and repeats.
  4. Low layers make edges and textures, while high layers make eyes, faces and animal parts, which shows the feature hierarchy.
  5. It is run at several image scales (octaves) and blended back into the original so the result stays natural.
  6. It is mainly used for interpretation and art, and it shows what a layer has learned to detect.

How visualization helps understanding.

  1. Deep Dream shows what a layer or neuron responds to, so we learn what features form at each depth.
  2. Guided backpropagation shows which input pixels drive a particular prediction, giving saliency and attribution.
  3. Together they reveal the hierarchy from edges to objects and expose wrong cues (for example the network looking at the background), so they serve as a debugging aid.

Answer frame. Open with the need for interpretability; define Deep Dream with its update rule; define guided backpropagation with the masking rule; end with a three-line "what we learn" list and a closing sentence that the two are complementary, one visualizing features and one attributing decisions.

Asked: [14 marks] (Jun 2025) How do visualization techniques like Deep Dream and Guided Backpropagation help in understanding neural networks?

Deep Art

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Deep Art (neural style transfer) combines the content of one image with the style of another using a pre-trained CNN.

Key points.

  1. Content is taken from the feature maps of a deep layer, and style from the Gram matrix $G_{ij}=\sum_k F_{ik}F_{jk}$ of feature maps.
  2. The loss is $\alpha L_{content}+\beta L_{style}$.
  3. The generated image is optimized by gradient descent, with the network weights fixed.

Recent Trends in Deep Learning Architectures

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Recent architectures replace or extend CNNs and RNNs with attention-based and efficient designs.

Key points.

  1. Transformers use self-attention, $\mathrm{softmax}(QK^{T}/\sqrt{d_k})V$, and process sequences in parallel.
  2. Vision Transformers split an image into patches and treat them as tokens.
  3. EfficientNet scales depth, width and resolution together, and MobileNet uses depthwise separable convolutions.
  4. Other trends are GANs, diffusion models, graph networks and large pre-trained foundation models.

Last-minute revision

  • Greedy layerwise pre-training: train layer by layer without labels, then fine-tune the whole network.
  • Pre-training was used in deep belief networks (RBMs) and stacked autoencoders.
  • ReLU is $\max(0,x)$, ELU is $\alpha(e^x-1)$ for $x<0$, GELU is $x\Phi(x)$, Swish is $x\sigma(x)$, Mish is $x\tanh(\mathrm{softplus}(x))$.
  • Xavier: $2/(n_{in}+n_{out})$ for tanh; He: $2/n_{in}$ for ReLU; LeCun: $1/n_{in}$.
  • CNN output size is $(W-F+2P)/S+1$.
  • LeNet-5: 32x32 input, C1 6@28x28, S2, C3 16@10x10, S4, C5 120, F6 84, output 10; about 60k parameters.
  • AlexNet: 8 layers, ReLU, dropout, 60M parameters; VGG16: 3x3 filters, 138M; GoogLeNet: Inception and 1x1, 22 layers.
  • ResNet: $y=F(x)+x$ solves degradation and vanishing gradients.
  • Guided backpropagation passes a gradient only if both the forward input and the incoming gradient are positive.
  • Deep Dream is gradient ascent on the image to maximize layer activations.

Memory hooks

  • Greedy = one layer at a time, like building a tower floor by floor.
  • Xavier for the X-shaped (tanh, sigmoid) curves; He for the Hard-cut ReLU, so twice the variance.
  • Inception = "we need to go deeper", with 1x1 to slim the channels.
  • ResNet: add the input back, so "shortcut keeps gradient alive".
  • Guided = Gradient positive AND activation positive (two masks).

Coverage checklist

  • Greedy Layerwise Pre-training: Jun 2025 pre-training and activation question.
  • Better activation functions: covered with the Jun 2025 question.
  • Better weight initialization methods: Dec 2020, Nov 2023 explain question.
  • Learning Vectorial Representations Of Words: no past question.
  • Convolutional Neural Networks: Dec 2020, Nov 2023 draw architecture.
  • LeNet: covered with the Dec 2020 ResNet and LeNet question.
  • AlexNet: no past question.
  • ZF-Net: no past question.
  • VGGNet: no past question.
  • GoogLeNet: no past question.
  • ResNet: Dec 2020 ResNet and LeNet question.
  • Visualizing Convolutional Neural Networks: no past question.
  • Guided Backpropagation: Dec 2020 three-term question; Jun 2025 visualization question.
  • Deep Dream: Jun 2025 visualization question.
  • Deep Art: no past question.
  • Recent Trends in Deep Learning Architectures: no past question.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in