Skip to content
AD-601 · Deep Learning/Quick Revision Short Notes

Deep Learning (AD-601) - Unit 3 Short Notes

How unit 3 is examined

This unit covers convolution and its variants, structured outputs, fast convolution, random features, LeNet and AlexNet; the convolution operation, CNN basics with pooling, structured outputs and LeNet carry the marks.

The Convolution Operation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Convolution is a linear operation that slides a small kernel over the input and sums the element-wise products at each position to produce a feature map.</mark>

Formula. Continuous: $s(t)=\int x(a)\,w(t-a)\,da$. Discrete 1-D: $s(t)=\sum_a x(a)\,w(t-a)$. For an image $I$ and kernel $K$: $S(i,j)=\sum_m\sum_n I(i+m,j+n)\,K(m,n)$ (CNN libraries actually compute this cross-correlation without flipping the kernel). Output size: $\frac{W-F+2P}{S}+1$.

Key points.

  1. A kernel (filter) is a small learnable weight matrix, for example 3x3, that acts as a feature detector such as an edge or a texture.
  2. Sparse interactions: each output depends only on a small local patch, so far fewer connections are needed than in a dense layer.
  3. Parameter sharing: the same kernel is reused at every position, so parameters are few and independent of image size.
  4. Equivariance: shifting the input shifts the feature map by the same amount, which gives translation invariance for objects.
  5. Stride is the step of the kernel; padding adds border zeros so the size is preserved and edge pixels are used.
  6. Depth is the number of kernels; each gives one feature map, and stacked layers build a hierarchy from edges to parts to objects.
  7. Pooling downsamples each map: max pooling keeps the largest value in the window, average pooling keeps the mean. It cuts size and parameters, adds small-shift invariance and reduces overfitting.
  8. A CNN is convolution, activation and pooling layers followed by fully connected layers.

Example. Vertical-edge kernel $K=\begin{bmatrix}1&0&-1\\1&0&-1\\1&0&-1\end{bmatrix}$ on a 4x4 input, stride 1, no padding, so the output is 2x2.

Input rows 1 2 0 1 / 3 1 2 0 / 0 2 1 3 / 1 0 2 1
$S(0,0)$ (1+3+0)-(0+2+1) = 1
$S(0,1)$ (2+1+2)-(1+0+3) = 1
$S(1,0)$ (3+0+1)-(2+1+2) = -1
$S(1,1)$ (1+2+0)-(0+3+1) = -1

Result: [[1, 1], [-1, -1]]. Pooling a 4x4 map [[1,3,2,4],[5,6,1,2],[7,2,9,0],[3,4,1,8]] with 2x2, stride 2: max gives [[6,4],[7,9]], average gives [[3.75,2.25],[4,4.5]].

CNN vs RNN.

Basis CNN RNN
Data Grid data such as images Sequences such as text, speech
Connections Feedforward, local Recurrent loops
Memory None Hidden state
Key layers Convolution, pooling Recurrent cells (LSTM, GRU)
Weight sharing Across space Across time
Use Classification, detection Translation, speech

Answer frame. Open with the definition and formula; draw kernel sliding on the input giving a feature map; develop points 1-6, then the edge example; close with "convolution gives parameter sharing and translation-equivariant features". For CNN or pooling questions, list layers (point 8) then point 7. For a kernel-role answer, use points 1, 3, 5, 6 with kernel size, stride and number.

Asked: [8 marks] (May 2023, May 2024, Jun 2025) What is Convolution Operation? Explain in detail. Why do we use convolution in deep learning? Significance and how it captures spatial features. Asked: [7 marks] (Jun 2025) Difference between Convolutional Neural Network and Recurrent Neural Network. Asked: [7 marks] (Jun 2025) What is the role of filters or kernels in CNNs? Asked: [7 marks] (May 2024) Briefly discuss Convolutional Neural Network with Deep Learning. Asked: [7 marks] (Jun 2025) Role of pooling layers in CNNs; explain max pooling and average pooling.

Variants of the Basic Convolution Function

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. Variants change how the kernel is applied: padding, stride, dilation, channels or weight sharing.

Key points.

  1. Valid convolution uses no padding, so the output shrinks to $W-F+1$; same convolution pads so output size equals input size; full convolution pads $F-1$ on each side so every overlap counts.
  2. Strided convolution moves the kernel by more than one pixel, which downsamples and reduces computation.
  3. Dilated convolution spaces kernel taps apart, enlarging the receptive field without more parameters.
  4. Separable (depthwise) convolution splits a kernel into per-channel and 1x1 steps, cutting cost.
  5. Locally connected layers use a different kernel at each position, with no sharing; tiled convolution cycles through a few kernels as a middle way.

Asked: [7 marks] (May 2023) Explain the variants of the basic convolution function in detail.

Structured Outputs

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Structured output means the network outputs a high-dimensional structured object, such as a per-pixel label map, instead of one class label.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-01" viewBox="0 0 474 80" width="474" height="80" role="img" aria-label="Structured output. In = input image, Conv = convolution layers (fully convolutional, no dense layer), Up = upsampling, Map = per-pixel label map"><style>#dsfig-u3-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-01 .t{fill:#16181D;font-weight:500}#dsfig-u3-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-01 .dot{fill:#16181D}#dsfig-u3-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-01 .ah{fill:#454C5A}#dsfig-u3-01 .ah.hi{fill:#2340B8}#dsfig-u3-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-01 .e{stroke:#B1B7C3}html.dark #dsfig-u3-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-01 .t{fill:#E6E8ED}html.dark #dsfig-u3-01 .t.inv{fill:#0F1115}html.dark #dsfig-u3-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-01 .dot{fill:#E6E8ED}html.dark #dsfig-u3-01 .ann{fill:#8FA3FF}html.dark #dsfig-u3-01 .lbl{fill:#858D9C}html.dark #dsfig-u3-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-01 .ah{fill:#B1B7C3}html.dark #dsfig-u3-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L141,40" marker-end="url(#ah6)"/><path class="e" d="M195,40 L277,40" marker-end="url(#ah6)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah6)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">In</text><rect class="n" x="144" y="25" width="50" height="30" rx="15"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">Conv</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Up</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">Map</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Structured output. In = input image, Conv = convolution layers (fully convolutional, no dense layer), Up = upsampling, Map = per-pixel label map</figcaption></figure>

Key points.

  1. The output tensor $Y_{i,j,c}$ gives the probability that pixel $(i,j)$ belongs to class $c$, as in image segmentation.
  2. The network is fully convolutional: convolutions keep the spatial layout, with no dense layer that would destroy it.
  3. Pooling shrinks the map, so upsampling (transposed convolution) or a stride-1, no-pooling design restores the resolution.
  4. Recurrent refinement can feed the earlier label estimate back as input so neighbouring pixel labels become consistent.
  5. Labels can be pooled into regions or grids to give dense predictions without a full-size output.
  6. Applications: semantic segmentation, depth estimation, image-to-image translation and bounding-box maps.
  7. Advantage: one forward pass labels the whole image, and shared weights keep parameters low.

Answer frame. Open with the definition; draw the diagram; develop points 1-5, then applications; close with the advantage. For "any two", pair this with AlexNet (own section below).

Asked: [14 marks] (May 2024, Jun 2025) Explain any two: Structured Output in Convolutional Networks, AlexNet, Deep Boltzmann Machines, Self Organizing Maps; or Deep Reinforcement, Autoencoder Architecture.

Efficient Convolution Algorithms

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Efficient algorithms compute the same convolution with fewer operations.

Key points.

  1. FFT convolution transforms input and kernel to the frequency domain, multiplies them and transforms back, since $x*w=\mathcal{F}^{-1}(\mathcal{F}(x)\cdot\mathcal{F}(w))$.
  2. Winograd reduces multiplications for small kernels such as 3x3.
  3. im2col rewrites convolution as one large matrix multiplication, which suits GPUs.
  4. Separable kernels turn a $k\times k$ cost into $2k$ per pixel.

Random or Unsupervised Features

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Convolution kernels can be set without supervised training, saving the cost of learning them end to end.

Key points.

  1. Random kernels often work surprisingly well, because pooling and the architecture supply much of the invariance.
  2. Kernels can be learned unsupervised, for example with k-means clustering on image patches, sparse coding or autoencoders.
  3. Only the final classifier is then trained, so training is cheap and layers can be pretrained one by one.
  4. Backpropagation on all layers is now standard, but this approach is useful with little labelled data.

LeNet

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>LeNet-5 (Yann LeCun, 1998) is the pioneering CNN for handwritten digit recognition, trained end to end with backpropagation.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-02" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="LeNet-5. In = 32x32 image, C1 = 6 maps 28x28, S2 = 6 maps 14x14, C3 = 16 maps 10x10, S4 = 16 maps 5x5, F = FC 120 and 84, Out = 10 classes"><style>#dsfig-u3-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-02 .t{fill:#16181D;font-weight:500}#dsfig-u3-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-02 .dot{fill:#16181D}#dsfig-u3-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-02 .ah{fill:#454C5A}#dsfig-u3-02 .ah.hi{fill:#2340B8}#dsfig-u3-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-02 .e{stroke:#B1B7C3}html.dark #dsfig-u3-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-02 .t{fill:#E6E8ED}html.dark #dsfig-u3-02 .t.inv{fill:#0F1115}html.dark #dsfig-u3-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-02 .dot{fill:#E6E8ED}html.dark #dsfig-u3-02 .ann{fill:#8FA3FF}html.dark #dsfig-u3-02 .lbl{fill:#858D9C}html.dark #dsfig-u3-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-02 .ah{fill:#B1B7C3}html.dark #dsfig-u3-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L105,40" marker-end="url(#ah7)"/><path class="e" d="M145,40 L191,40" marker-end="url(#ah7)"/><path class="e" d="M231,40 L277,40" marker-end="url(#ah7)"/><path class="e" d="M317,40 L363,40" marker-end="url(#ah7)"/><path class="e" d="M403,40 L449,40" marker-end="url(#ah7)"/><path class="e" d="M489,40 L535,40" marker-end="url(#ah7)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">In</text><circle class="n" cx="126" cy="40" r="18"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">C1</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">S2</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">C3</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">S4</text><circle class="n" cx="470" cy="40" r="18"/><text class="t" x="470" y="40" dy=".35em" text-anchor="middle">F</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">Out</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">LeNet-5. In = 32x32 image, C1 = 6 maps 28x28, S2 = 6 maps 14x14, C3 = 16 maps 10x10, S4 = 16 maps 5x5, F = FC 120 and 84, Out = 10 classes</figcaption></figure>

Key points.

  1. Input is a 32x32 grayscale image; C1 uses six 5x5 kernels, giving 6@28x28.
  2. S2 is 2x2 subsampling (average pooling), giving 6@14x14; C3 has sixteen 5x5 kernels, giving 16@10x10; S4 subsamples to 16@5x5.
  3. Fully connected layers of 120 and 84 units follow, then 10 output units for digits 0-9; tanh or sigmoid activations were used.
  4. It has about 60,000 parameters; for example C1 has $6\times(25+1)=156$.
  5. Role: it showed that convolution, pooling and dense layers trained end to end beat hand-made features, and it was used to read bank cheques and postal codes.
  6. Limitation: it was small, grayscale and limited by the data and hardware of its time.

Answer frame. Open with LeCun 1998 and digit recognition; draw the diagram; list layers in order (points 1-3); close with its role as the ancestor of modern CNNs.

Asked: [7 marks] (May 2023, May 2024) Briefly discuss the role of LeNet in Convolutional Networks. What is LeNet? Explain its usage in object recognition.

AlexNet

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. AlexNet (Krizhevsky, Sutskever and Hinton, 2012) is a deep CNN that won ImageNet 2012 with a large margin.

Key points.

  1. It has five convolution layers and three fully connected layers, with about 60 million parameters.
  2. It used ReLU activations, which train faster than tanh.
  3. Dropout in the dense layers and data augmentation reduced overfitting.
  4. Overlapping max pooling and local response normalisation were used, and two GPUs trained it.
  5. It started the deep learning boom in vision.

Last-minute revision

  • Convolution: $S(i,j)=\sum_m\sum_n I(i+m,j+n)K(m,n)$; output size $(W-F+2P)/S+1$.
  • Three ideas: sparse interactions, parameter sharing, equivariance.
  • Valid shrinks, same keeps size, full enlarges; dilated widens receptive field.
  • Max pooling keeps the largest value; average pooling keeps the mean.
  • CNN handles grid data with no memory; RNN handles sequences with a hidden state.
  • Structured output is a per-pixel label map from a fully convolutional network.
  • FFT, Winograd and im2col speed up convolution.
  • LeNet-5: 32x32 input, C1 6@28x28, S2, C3 16@10x10, S4, FC 120, 84, output 10; about 60k parameters.
  • AlexNet: 2012, 5 conv and 3 FC layers, ReLU, dropout, 60M parameters.

Memory hooks

  • SPE: Sparse, Parameter sharing, Equivariance.
  • Valid shrinks, Same stays, Full grows.
  • LeNet digits 1998 (small); AlexNet ImageNet 2012 (big).
  • LeNet layers: C-S-C-S-F-F-O.

Coverage checklist

  • The Convolution Operation: 8-mark convolution question, CNN vs RNN, role of filters, CNN overview, pooling.
  • Variants of the Basic Convolution Function: variants question.
  • Structured Outputs: any-two 14-mark question.
  • Efficient Convolution Algorithms: no past question.
  • Random or Unsupervised Features: no past question.
  • LeNet: both LeNet questions.
  • AlexNet: covered here, and as an option in the any-two question.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in