How unit 3 is examined
This unit covers pre-training, activations, initialization, CNNs and their famous architectures, and visualization methods; the marks sit in greedy layerwise pre-training, guided backpropagation and Deep Dream (14 marks each), then weight initialization and CNN architecture (7 marks each).
Greedy Layerwise Pre-training
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Greedy layerwise pre-training trains a deep network one layer at a time, each layer learned without labels on the output of the layer below, and the stacked weights are then used to initialize the network for supervised fine-tuning.</mark>
Key points.
- Layer 1 is trained as an autoencoder or RBM on the raw input, then frozen, and its hidden output becomes the training data for layer 2, and so on.
- It is called greedy because each layer is optimized alone for its own local objective, not for the final task loss.
- The stage is unsupervised, so it uses plentiful unlabelled data and gives an unsupervised initialization that starts the weights in a good region.
- After stacking, a classifier layer is added and the whole network is fine-tuned with backpropagation on labelled data.
- It gave deep belief networks (stacked RBMs, Hinton 2006) and stacked autoencoders their success, because plain random initialization with sigmoid units suffered vanishing gradients and poor local minima.
- It also acts as a regularizer, since the initial weights already capture structure of the input distribution.
- It has largely been replaced by ReLU, better initialization, batch normalization and large labelled data, but is still used when labels are scarce.
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-01" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="Layer-by-layer training; X input, H1-H3 hidden layers, Y softmax output added for fine-tuning"><style>#dsfig-u3-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-01 .t{fill:#16181D;font-weight:500}#dsfig-u3-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-01 .dot{fill:#16181D}#dsfig-u3-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-01 .ah{fill:#454C5A}#dsfig-u3-01 .ah.hi{fill:#2340B8}#dsfig-u3-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-01 .e{stroke:#B1B7C3}html.dark #dsfig-u3-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-01 .t{fill:#E6E8ED}html.dark #dsfig-u3-01 .t.inv{fill:#0F1115}html.dark #dsfig-u3-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-01 .dot{fill:#E6E8ED}html.dark #dsfig-u3-01 .ann{fill:#8FA3FF}html.dark #dsfig-u3-01 .lbl{fill:#858D9C}html.dark #dsfig-u3-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-01 .ah{fill:#B1B7C3}html.dark #dsfig-u3-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah6)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah6)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah6)"/><path class="e" d="M446,40 L535,40" marker-end="url(#ah6)"/><g class="wl"><rect x="77.4" y="31" width="54.3" height="18" rx="9"/><text class="t" x="104.5" y="40" dy=".35em" text-anchor="middle">train1</text></g><g class="wl"><rect x="206.4" y="31" width="54.3" height="18" rx="9"/><text class="t" x="233.5" y="40" dy=".35em" text-anchor="middle">train2</text></g><g class="wl"><rect x="335.4" y="31" width="54.3" height="18" rx="9"/><text class="t" x="362.5" y="40" dy=".35em" text-anchor="middle">train3</text></g><g class="wl"><rect x="457.2" y="31" width="68.7" height="18" rx="9"/><text class="t" x="491.5" y="40" dy=".35em" text-anchor="middle">finetune</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">X</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">H1</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">H2</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">H3</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">Y</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Layer-by-layer training; X input, H1-H3 hidden layers, Y softmax output added for fine-tuning</figcaption></figure>
Better activation functions (recent years).
- ReLU, $\max(0,x)$, does not saturate for $x>0$, so gradients do not vanish and training is fast, but neurons with $x<0$ can die.
- Leaky ReLU, $\max(0.01x,x)$, and PReLU keep a small slope for negatives, so no neuron dies.
- ELU, $x$ for $x>0$ and $\alpha(e^{x}-1)$ otherwise, gives smooth negative outputs and mean activation near zero.
- GELU, $x\,\Phi(x)$, weights the input by the Gaussian CDF and is used in transformers.
- Swish, $x\,\sigma(x)$, and Mish, $x\tanh(\mathrm{softplus}(x))$, are smooth, non-monotonic and often beat ReLU in deep nets.
Answer frame. Open with the definition; draw the layer-by-layer figure; develop points 1-5 (procedure, greedy, unsupervised, fine-tune, DBN role); then list the five activations with formula and advantage (non-saturation, faster training, vanishing-gradient relief); close with "pre-training is now replaced by ReLU-family activations and good initialization".
Asked: [14 marks] (Jun 2025) Discuss the concept of greedy layerwise pre-training. What are some better activation functions introduced in recent years?
Better activation functions
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. An activation function adds non-linearity to a neuron, and a better one avoids saturation and vanishing gradients.
Key points.
- Sigmoid and tanh saturate at large $|x|$, so their gradients shrink layer by layer.
- ReLU, Leaky ReLU, ELU, GELU, Swish and Mish (formulas above) give non-saturating gradients for positive inputs and faster convergence.
- Softmax is used only at the output for class probabilities: $e^{z_i}/\sum_j e^{z_j}$.
- Choose ReLU or Leaky ReLU as the default hidden activation, and GELU in transformers.
Better weight initialization methods
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Weight initialization sets the starting weights so that the variance of activations and gradients stays roughly constant across layers, which prevents vanishing and exploding gradients.</mark>
Key points.
- Zero initialization makes every neuron in a layer compute and learn the same thing (symmetry), so it must never be used for weights.
- Small random weights break symmetry but the signal shrinks layer by layer in deep nets (vanishing), while large random weights make it blow up (exploding) or saturate sigmoid and tanh.
- Xavier (Glorot) initialization suits tanh and sigmoid and keeps variance equal in forward and backward passes: $\mathrm{Var}(w)=\dfrac{2}{n_{in}+n_{out}}$.
- He initialization suits ReLU, which zeroes half the inputs, so the variance is doubled: $\mathrm{Var}(w)=\dfrac{2}{n_{in}}$.
- LeCun initialization uses $\mathrm{Var}(w)=\dfrac{1}{n_{in}}$ and suits SELU and older tanh nets.
- Orthogonal initialization uses an orthogonal matrix so the norm of the signal is preserved, and it is common in RNNs.
- Biases are usually set to zero, and pre-trained weights (transfer learning or layerwise pre-training) are another good starting point.
Answer frame. Open with why initialization matters for vanishing and exploding gradients; develop points 1-2 (zero, random), then Xavier, He, LeCun with their variances, then orthogonal and pretrained; close with "match the initializer to the activation: Xavier for tanh, He for ReLU".
Asked: [7 marks] (Dec 2020, Nov 2023) Explain Better Weight Initialization Methods.
Learning Vectorial Representations Of Words
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Word embeddings map each word to a dense real vector, learned so that words in similar contexts get similar vectors.
Key points.
- One-hot vectors are huge and sparse and carry no similarity, so dense vectors of 100-300 dimensions are learned instead.
- Word2Vec has two models: CBOW predicts a word from its context, and skip-gram predicts the context from the word.
- GloVe factorizes the word co-occurrence matrix, and both give analogies such as $king-man+woman\approx queen$.
- Similarity is measured with cosine similarity, and the vectors feed RNNs and other models.
Convolutional Neural Networks
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>A convolutional neural network is a feed-forward network for grid-like data such as images that uses shared local filters (convolution), pooling and fully connected layers to learn a hierarchy of features.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-02" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="CNN: Input image, Convolution, ReLU, Pooling (Conv-ReLU-Pool repeated), Flatten, Fully connected with softmax output"><style>#dsfig-u3-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-02 .t{fill:#16181D;font-weight:500}#dsfig-u3-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-02 .dot{fill:#16181D}#dsfig-u3-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-02 .ah{fill:#454C5A}#dsfig-u3-02 .ah.hi{fill:#2340B8}#dsfig-u3-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-02 .e{stroke:#B1B7C3}html.dark #dsfig-u3-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-02 .t{fill:#E6E8ED}html.dark #dsfig-u3-02 .t.inv{fill:#0F1115}html.dark #dsfig-u3-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-02 .dot{fill:#E6E8ED}html.dark #dsfig-u3-02 .ann{fill:#8FA3FF}html.dark #dsfig-u3-02 .lbl{fill:#858D9C}html.dark #dsfig-u3-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-02 .ah{fill:#B1B7C3}html.dark #dsfig-u3-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L122.2,40" marker-end="url(#ah7)"/><path class="e" d="M162.2,40 L225.4,40" marker-end="url(#ah7)"/><path class="e" d="M265.4,40 L328.6,40" marker-end="url(#ah7)"/><path class="e" d="M368.6,40 L431.8,40" marker-end="url(#ah7)"/><path class="e" d="M471.8,40 L535,40" marker-end="url(#ah7)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">In</text><circle class="n" cx="143.2" cy="40" r="18"/><text class="t" x="143.2" y="40" dy=".35em" text-anchor="middle">Cv</text><circle class="n" cx="246.4" cy="40" r="18"/><text class="t" x="246.4" y="40" dy=".35em" text-anchor="middle">Rl</text><circle class="n" cx="349.6" cy="40" r="18"/><text class="t" x="349.6" y="40" dy=".35em" text-anchor="middle">Pl</text><circle class="n" cx="452.8" cy="40" r="18"/><text class="t" x="452.8" y="40" dy=".35em" text-anchor="middle">Fl</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">FC</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">CNN: Input image, Convolution, ReLU, Pooling (Conv-ReLU-Pool repeated), Flatten, Fully connected with softmax output</figcaption></figure>
Key points.
- The convolution layer slides small filters (for example 3x3) over the input, and each filter produces one feature map that shows where its pattern occurs.
- Filters share weights across positions and see only a local patch, so parameters are far fewer than in a dense network.
- Stride is the step of the filter and padding adds border zeros, and the output size is $\dfrac{W-F+2P}{S}+1$; for example $W=32,F=5,P=0,S=1$ gives 28.
- ReLU is applied after convolution to add non-linearity and avoid vanishing gradients.
- Pooling (max or average, usually 2x2 with stride 2) halves the size, cuts computation and gives small translation invariance.
- Early layers learn edges, middle layers learn textures and parts, and deep layers learn whole objects (hierarchical features).
- Flattened features go to fully connected layers and a softmax that outputs class probabilities, and the whole network is trained by backpropagation; LeNet-5 and AlexNet follow exactly this flow.
Answer frame. Open with the definition; draw the block figure with labelled layers and data flow; develop points 1-5 in layer order; then hierarchy and the classifier; close with applications (image classification, detection, medical imaging).
Asked: [7 marks] (Dec 2020, Nov 2023) Draw and explain the architecture of Convolutional Network.
LeNet
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. LeNet-5 (LeCun, 1998) is the first successful CNN, built for handwritten digit recognition on 32x32 grayscale images.
Key points.
- The layers are C1 (6 maps, 5x5, 28x28), S2 (average pool, 14x14), C3 (16 maps, 10x10), S4 (pool, 5x5), C5 (120), F6 (84) and an output of 10 classes.
- It has about 60,000 parameters and used tanh or sigmoid with average pooling.
- It showed that convolution and pooling learned by backpropagation beat hand-made features, and it was used to read bank cheques.
AlexNet
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. AlexNet (Krizhevsky, 2012) is an eight-layer CNN that won ImageNet 2012 by a large margin and started the deep learning boom.
Key points.
- It has 5 convolution layers and 3 fully connected layers, with about 60 million parameters and a 224x224x3 input.
- It used ReLU instead of tanh, which trained several times faster.
- Dropout in the fully connected layers and data augmentation reduced overfitting.
- It used overlapping max pooling, local response normalization, and two GPUs for training.
ZF-Net
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ZF-Net (Zeiler and Fergus, 2013) is an improved AlexNet that won ImageNet 2013 and introduced deconvnet visualization.
Key points.
- It reduced the first filter from 11x11 to 7x7 and the stride from 4 to 2, which kept more detail.
- It used a deconvolution network to project activations back to pixels and see what each layer detects.
- These visualizations guided the design changes, showing that visualization can improve architecture.
VGGNet
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. VGGNet (Simonyan and Zisserman, 2014) is a very deep CNN built only from stacked 3x3 convolutions with 2x2 max pooling.
Key points.
- VGG16 has 13 convolution and 3 fully connected layers (16 weight layers), and VGG19 has 19.
- Two stacked 3x3 layers see a 5x5 region and three see 7x7, with fewer parameters and more non-linearity than one large filter.
- The design is simple and uniform, but it has about 138 million parameters and is heavy in memory.
GoogLeNet
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. GoogLeNet (Inception v1, 2014) is a 22-layer CNN built from Inception modules that apply 1x1, 3x3, 5x5 convolutions and pooling in parallel and concatenate the results.
Key points.
- Parallel filters capture features at several scales in one layer.
- 1x1 convolutions reduce the channel depth before the costly 3x3 and 5x5 filters, cutting computation.
- It replaces the large fully connected layers with global average pooling, so it has only about 5-7 million parameters.
- Auxiliary classifiers in the middle help gradients flow, and it won ImageNet 2014.
ResNet
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. ==ResNet (He et al., 2015) is a very deep CNN of residual blocks in which a skip connection adds the input to the block output, $y=F(x)+x$.==
Key points.
- Plain networks get worse when made deeper (the degradation problem), because the layers struggle to learn even an identity mapping.
- A residual block only learns the change $F(x)=H(x)-x$, and if the identity is best it pushes $F(x)$ to zero.
- The skip path gives the gradient a direct route, since $\partial y/\partial x=\partial F/\partial x+1$, which relieves vanishing gradients.
- ResNet-152 won ImageNet 2015 with 3.57% top-5 error, using batch normalization and bottleneck blocks.
- Compared with LeNet-5 (5 layers, digits, 60 thousand parameters), ResNet is over 100 layers deep, classifies ImageNet, and is used as a backbone in detection.
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-03" viewBox="0 0 596 166" width="596" height="166" role="img" aria-label="Residual block; X input, W1 and W2 weight layers, Add adds F(x) and x, Y output"><style>#dsfig-u3-03 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-03 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-03 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-03 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-03 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-03 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-03 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-03 .t{fill:#16181D;font-weight:500}#dsfig-u3-03 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-03 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-03 .dot{fill:#16181D}#dsfig-u3-03 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-03 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-03 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-03 .ah{fill:#454C5A}#dsfig-u3-03 .ah.hi{fill:#2340B8}#dsfig-u3-03 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-03 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-03 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-03 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-03 .e{stroke:#B1B7C3}html.dark #dsfig-u3-03 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-03 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-03 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-03 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-03 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-03 .t{fill:#E6E8ED}html.dark #dsfig-u3-03 .t.inv{fill:#0F1115}html.dark #dsfig-u3-03 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-03 .dot{fill:#E6E8ED}html.dark #dsfig-u3-03 .ann{fill:#8FA3FF}html.dark #dsfig-u3-03 .lbl{fill:#858D9C}html.dark #dsfig-u3-03 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-03 .ah{fill:#B1B7C3}html.dark #dsfig-u3-03 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-03 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-03 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-03 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.8,115.5 L151.5,51.6" marker-end="url(#ah8)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah8)"/><path class="e" d="M313.8,50.5 L409.5,114.4" marker-end="url(#ah8)"/><path class="e" d="M59,126 L406,126" marker-end="url(#ah8)"/><path class="e" d="M446,126 L535,126" marker-end="url(#ah8)"/><g class="wl"><rect x="342.1" y="74" width="40.8" height="18" rx="9"/><text class="t" x="362.5" y="83" dy=".35em" text-anchor="middle">F(x)</text></g><g class="wl"><rect x="213.1" y="117" width="40.8" height="18" rx="9"/><text class="t" x="233.5" y="126" dy=".35em" text-anchor="middle">skip</text></g><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">X</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">W1</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">W2</text><circle class="n" cx="427" cy="126" r="18"/><text class="t" x="427" y="126" dy=".35em" text-anchor="middle">Add</text><circle class="n" cx="556" cy="126" r="18"/><text class="t" x="556" y="126" dy=".35em" text-anchor="middle">Y</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Residual block; X input, W1 and W2 weight layers, Add adds F(x) and x, Y output</figcaption></figure>
Answer frame. Give LeNet-5 first (see LeNet: layers and digit task), then the residual block figure with $y=F(x)+x$; explain degradation and gradient relief; close with a depth and use-case comparison.
Asked: [7 marks] (Dec 2020) Explain ResNet and LeNet in detail.
Visualizing Convolutional Neural Networks
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. CNN visualization shows what the filters and layers respond to, so that the network is not a black box.
Key points.
- Feature maps and filters can be plotted directly, showing edges in early layers and object parts in later ones.
- Saliency maps use the gradient of the class score with respect to the input pixels to show which pixels matter.
- Deconvnet, guided backpropagation, activation maximization and Deep Dream are the main techniques.
Guided Backpropagation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Guided backpropagation is a visualization method that backpropagates from a neuron to the input image but lets only positive gradients pass through each ReLU, and only where the forward activation was also positive.</mark>
Key points.
- Ordinary backpropagation passes the gradient through a ReLU wherever the forward input was positive.
- Deconvnet passes it wherever the incoming gradient is positive, and guided backpropagation applies both conditions: $R_l=\mathbb{1}[f_l>0]\cdot\mathbb{1}[R_{l+1}>0]\cdot R_{l+1}$.
- Masking the negative gradients removes evidence that lowers the neuron's activation, so only pixels that excite it remain.
- The result is a sharp, clean image of edges and shapes that the neuron or class responds to.
- It is used for saliency, explaining predictions and debugging, and it needs no change to the network.
Dataset augmentation. It enlarges the training set by applying label-preserving transforms such as flip, rotation, crop, scale, shift, brightness change and noise, which reduces overfitting and improves generalization; it is standard in image tasks.
LSTM. Long Short-Term Memory is an RNN cell with a cell state $c_t$ that carries information over long spans and three gates: the forget gate $f_t$ decides what to erase, the input gate $i_t$ decides what to write, and the output gate $o_t$ decides what to expose as $h_t$. Because the cell state is updated additively, $c_t=f_t\odot c_{t-1}+i_t\odot\tilde c_t$, gradients do not vanish easily, so it handles long dependencies in language, speech and time series.
Answer frame. Give three short parts of about equal length: definition, mechanism and application for each of the three terms, with the LSTM gate equation and the guided-backprop masking rule written out.
Asked: [14 marks] (Dec 2020) Explain following term: i) Guided Back propagation ii) Dataset augmentation iii) LSTM
Deep Dream
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Deep Dream is a technique that uses gradient ascent on the input image to maximize the activations of a chosen layer, so the network amplifies the patterns it detects and shows what it has learned.</mark>
Key points.
- Networks are hard to interpret, so visualization is needed to trust, debug and improve them.
- Deep Dream feeds an image through a trained CNN and picks a layer, and its loss is the sum of squared activations of that layer.
- It then updates the image, not the weights: $x\leftarrow x+\eta\,\partial L/\partial x$, and repeats.
- Low layers make edges and textures, while high layers make eyes, faces and animal parts, which shows the feature hierarchy.
- It is run at several image scales (octaves) and blended back into the original so the result stays natural.
- It is mainly used for interpretation and art, and it shows what a layer has learned to detect.
How visualization helps understanding.
- Deep Dream shows what a layer or neuron responds to, so we learn what features form at each depth.
- Guided backpropagation shows which input pixels drive a particular prediction, giving saliency and attribution.
- Together they reveal the hierarchy from edges to objects and expose wrong cues (for example the network looking at the background), so they serve as a debugging aid.
Answer frame. Open with the need for interpretability; define Deep Dream with its update rule; define guided backpropagation with the masking rule; end with a three-line "what we learn" list and a closing sentence that the two are complementary, one visualizing features and one attributing decisions.
Asked: [14 marks] (Jun 2025) How do visualization techniques like Deep Dream and Guided Backpropagation help in understanding neural networks?
Deep Art
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Deep Art (neural style transfer) combines the content of one image with the style of another using a pre-trained CNN.
Key points.
- Content is taken from the feature maps of a deep layer, and style from the Gram matrix $G_{ij}=\sum_k F_{ik}F_{jk}$ of feature maps.
- The loss is $\alpha L_{content}+\beta L_{style}$.
- The generated image is optimized by gradient descent, with the network weights fixed.
Recent Trends in Deep Learning Architectures
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Recent architectures replace or extend CNNs and RNNs with attention-based and efficient designs.
Key points.
- Transformers use self-attention, $\mathrm{softmax}(QK^{T}/\sqrt{d_k})V$, and process sequences in parallel.
- Vision Transformers split an image into patches and treat them as tokens.
- EfficientNet scales depth, width and resolution together, and MobileNet uses depthwise separable convolutions.
- Other trends are GANs, diffusion models, graph networks and large pre-trained foundation models.
Last-minute revision
- Greedy layerwise pre-training: train layer by layer without labels, then fine-tune the whole network.
- Pre-training was used in deep belief networks (RBMs) and stacked autoencoders.
- ReLU is $\max(0,x)$, ELU is $\alpha(e^x-1)$ for $x<0$, GELU is $x\Phi(x)$, Swish is $x\sigma(x)$, Mish is $x\tanh(\mathrm{softplus}(x))$.
- Xavier: $2/(n_{in}+n_{out})$ for tanh; He: $2/n_{in}$ for ReLU; LeCun: $1/n_{in}$.
- CNN output size is $(W-F+2P)/S+1$.
- LeNet-5: 32x32 input, C1 6@28x28, S2, C3 16@10x10, S4, C5 120, F6 84, output 10; about 60k parameters.
- AlexNet: 8 layers, ReLU, dropout, 60M parameters; VGG16: 3x3 filters, 138M; GoogLeNet: Inception and 1x1, 22 layers.
- ResNet: $y=F(x)+x$ solves degradation and vanishing gradients.
- Guided backpropagation passes a gradient only if both the forward input and the incoming gradient are positive.
- Deep Dream is gradient ascent on the image to maximize layer activations.
Memory hooks
- Greedy = one layer at a time, like building a tower floor by floor.
- Xavier for the X-shaped (tanh, sigmoid) curves; He for the Hard-cut ReLU, so twice the variance.
- Inception = "we need to go deeper", with 1x1 to slim the channels.
- ResNet: add the input back, so "shortcut keeps gradient alive".
- Guided = Gradient positive AND activation positive (two masks).
Coverage checklist
- Greedy Layerwise Pre-training: Jun 2025 pre-training and activation question.
- Better activation functions: covered with the Jun 2025 question.
- Better weight initialization methods: Dec 2020, Nov 2023 explain question.
- Learning Vectorial Representations Of Words: no past question.
- Convolutional Neural Networks: Dec 2020, Nov 2023 draw architecture.
- LeNet: covered with the Dec 2020 ResNet and LeNet question.
- AlexNet: no past question.
- ZF-Net: no past question.
- VGGNet: no past question.
- GoogLeNet: no past question.
- ResNet: Dec 2020 ResNet and LeNet question.
- Visualizing Convolutional Neural Networks: no past question.
- Guided Backpropagation: Dec 2020 three-term question; Jun 2025 visualization question.
- Deep Dream: Jun 2025 visualization question.
- Deep Art: no past question.
- Recent Trends in Deep Learning Architectures: no past question.