How unit 3 is examined
This unit covers convolution and its variants, structured outputs, fast convolution, random features, LeNet and AlexNet; the convolution operation, CNN basics with pooling, structured outputs and LeNet carry the marks.
The Convolution Operation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Convolution is a linear operation that slides a small kernel over the input and sums the element-wise products at each position to produce a feature map.</mark>
Formula. Continuous: $s(t)=\int x(a)\,w(t-a)\,da$. Discrete 1-D: $s(t)=\sum_a x(a)\,w(t-a)$. For an image $I$ and kernel $K$: $S(i,j)=\sum_m\sum_n I(i+m,j+n)\,K(m,n)$ (CNN libraries actually compute this cross-correlation without flipping the kernel). Output size: $\frac{W-F+2P}{S}+1$.
Key points.
- A kernel (filter) is a small learnable weight matrix, for example 3x3, that acts as a feature detector such as an edge or a texture.
- Sparse interactions: each output depends only on a small local patch, so far fewer connections are needed than in a dense layer.
- Parameter sharing: the same kernel is reused at every position, so parameters are few and independent of image size.
- Equivariance: shifting the input shifts the feature map by the same amount, which gives translation invariance for objects.
- Stride is the step of the kernel; padding adds border zeros so the size is preserved and edge pixels are used.
- Depth is the number of kernels; each gives one feature map, and stacked layers build a hierarchy from edges to parts to objects.
- Pooling downsamples each map: max pooling keeps the largest value in the window, average pooling keeps the mean. It cuts size and parameters, adds small-shift invariance and reduces overfitting.
- A CNN is convolution, activation and pooling layers followed by fully connected layers.
Example. Vertical-edge kernel $K=\begin{bmatrix}1&0&-1\\1&0&-1\\1&0&-1\end{bmatrix}$ on a 4x4 input, stride 1, no padding, so the output is 2x2.
| Input rows | 1 2 0 1 / 3 1 2 0 / 0 2 1 3 / 1 0 2 1 |
|---|---|
| $S(0,0)$ | (1+3+0)-(0+2+1) = 1 |
| $S(0,1)$ | (2+1+2)-(1+0+3) = 1 |
| $S(1,0)$ | (3+0+1)-(2+1+2) = -1 |
| $S(1,1)$ | (1+2+0)-(0+3+1) = -1 |
Result: [[1, 1], [-1, -1]]. Pooling a 4x4 map [[1,3,2,4],[5,6,1,2],[7,2,9,0],[3,4,1,8]] with 2x2, stride 2: max gives [[6,4],[7,9]], average gives [[3.75,2.25],[4,4.5]].
CNN vs RNN.
| Basis | CNN | RNN |
|---|---|---|
| Data | Grid data such as images | Sequences such as text, speech |
| Connections | Feedforward, local | Recurrent loops |
| Memory | None | Hidden state |
| Key layers | Convolution, pooling | Recurrent cells (LSTM, GRU) |
| Weight sharing | Across space | Across time |
| Use | Classification, detection | Translation, speech |
Answer frame. Open with the definition and formula; draw kernel sliding on the input giving a feature map; develop points 1-6, then the edge example; close with "convolution gives parameter sharing and translation-equivariant features". For CNN or pooling questions, list layers (point 8) then point 7. For a kernel-role answer, use points 1, 3, 5, 6 with kernel size, stride and number.
Asked: [8 marks] (May 2023, May 2024, Jun 2025) What is Convolution Operation? Explain in detail. Why do we use convolution in deep learning? Significance and how it captures spatial features. Asked: [7 marks] (Jun 2025) Difference between Convolutional Neural Network and Recurrent Neural Network. Asked: [7 marks] (Jun 2025) What is the role of filters or kernels in CNNs? Asked: [7 marks] (May 2024) Briefly discuss Convolutional Neural Network with Deep Learning. Asked: [7 marks] (Jun 2025) Role of pooling layers in CNNs; explain max pooling and average pooling.
Variants of the Basic Convolution Function
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. Variants change how the kernel is applied: padding, stride, dilation, channels or weight sharing.
Key points.
- Valid convolution uses no padding, so the output shrinks to $W-F+1$; same convolution pads so output size equals input size; full convolution pads $F-1$ on each side so every overlap counts.
- Strided convolution moves the kernel by more than one pixel, which downsamples and reduces computation.
- Dilated convolution spaces kernel taps apart, enlarging the receptive field without more parameters.
- Separable (depthwise) convolution splits a kernel into per-channel and 1x1 steps, cutting cost.
- Locally connected layers use a different kernel at each position, with no sharing; tiled convolution cycles through a few kernels as a middle way.
Asked: [7 marks] (May 2023) Explain the variants of the basic convolution function in detail.
Structured Outputs
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Structured output means the network outputs a high-dimensional structured object, such as a per-pixel label map, instead of one class label.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-01" viewBox="0 0 474 80" width="474" height="80" role="img" aria-label="Structured output. In = input image, Conv = convolution layers (fully convolutional, no dense layer), Up = upsampling, Map = per-pixel label map"><style>#dsfig-u3-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-01 .t{fill:#16181D;font-weight:500}#dsfig-u3-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-01 .dot{fill:#16181D}#dsfig-u3-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-01 .ah{fill:#454C5A}#dsfig-u3-01 .ah.hi{fill:#2340B8}#dsfig-u3-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-01 .e{stroke:#B1B7C3}html.dark #dsfig-u3-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-01 .t{fill:#E6E8ED}html.dark #dsfig-u3-01 .t.inv{fill:#0F1115}html.dark #dsfig-u3-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-01 .dot{fill:#E6E8ED}html.dark #dsfig-u3-01 .ann{fill:#8FA3FF}html.dark #dsfig-u3-01 .lbl{fill:#858D9C}html.dark #dsfig-u3-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-01 .ah{fill:#B1B7C3}html.dark #dsfig-u3-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L141,40" marker-end="url(#ah6)"/><path class="e" d="M195,40 L277,40" marker-end="url(#ah6)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah6)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">In</text><rect class="n" x="144" y="25" width="50" height="30" rx="15"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">Conv</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Up</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">Map</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Structured output. In = input image, Conv = convolution layers (fully convolutional, no dense layer), Up = upsampling, Map = per-pixel label map</figcaption></figure>
Key points.
- The output tensor $Y_{i,j,c}$ gives the probability that pixel $(i,j)$ belongs to class $c$, as in image segmentation.
- The network is fully convolutional: convolutions keep the spatial layout, with no dense layer that would destroy it.
- Pooling shrinks the map, so upsampling (transposed convolution) or a stride-1, no-pooling design restores the resolution.
- Recurrent refinement can feed the earlier label estimate back as input so neighbouring pixel labels become consistent.
- Labels can be pooled into regions or grids to give dense predictions without a full-size output.
- Applications: semantic segmentation, depth estimation, image-to-image translation and bounding-box maps.
- Advantage: one forward pass labels the whole image, and shared weights keep parameters low.
Answer frame. Open with the definition; draw the diagram; develop points 1-5, then applications; close with the advantage. For "any two", pair this with AlexNet (own section below).
Asked: [14 marks] (May 2024, Jun 2025) Explain any two: Structured Output in Convolutional Networks, AlexNet, Deep Boltzmann Machines, Self Organizing Maps; or Deep Reinforcement, Autoencoder Architecture.
Efficient Convolution Algorithms
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Efficient algorithms compute the same convolution with fewer operations.
Key points.
- FFT convolution transforms input and kernel to the frequency domain, multiplies them and transforms back, since $x*w=\mathcal{F}^{-1}(\mathcal{F}(x)\cdot\mathcal{F}(w))$.
- Winograd reduces multiplications for small kernels such as 3x3.
- im2col rewrites convolution as one large matrix multiplication, which suits GPUs.
- Separable kernels turn a $k\times k$ cost into $2k$ per pixel.
Random or Unsupervised Features
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Convolution kernels can be set without supervised training, saving the cost of learning them end to end.
Key points.
- Random kernels often work surprisingly well, because pooling and the architecture supply much of the invariance.
- Kernels can be learned unsupervised, for example with k-means clustering on image patches, sparse coding or autoencoders.
- Only the final classifier is then trained, so training is cheap and layers can be pretrained one by one.
- Backpropagation on all layers is now standard, but this approach is useful with little labelled data.
LeNet
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>LeNet-5 (Yann LeCun, 1998) is the pioneering CNN for handwritten digit recognition, trained end to end with backpropagation.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-02" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="LeNet-5. In = 32x32 image, C1 = 6 maps 28x28, S2 = 6 maps 14x14, C3 = 16 maps 10x10, S4 = 16 maps 5x5, F = FC 120 and 84, Out = 10 classes"><style>#dsfig-u3-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-02 .t{fill:#16181D;font-weight:500}#dsfig-u3-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-02 .dot{fill:#16181D}#dsfig-u3-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-02 .ah{fill:#454C5A}#dsfig-u3-02 .ah.hi{fill:#2340B8}#dsfig-u3-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-02 .e{stroke:#B1B7C3}html.dark #dsfig-u3-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-02 .t{fill:#E6E8ED}html.dark #dsfig-u3-02 .t.inv{fill:#0F1115}html.dark #dsfig-u3-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-02 .dot{fill:#E6E8ED}html.dark #dsfig-u3-02 .ann{fill:#8FA3FF}html.dark #dsfig-u3-02 .lbl{fill:#858D9C}html.dark #dsfig-u3-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-02 .ah{fill:#B1B7C3}html.dark #dsfig-u3-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L105,40" marker-end="url(#ah7)"/><path class="e" d="M145,40 L191,40" marker-end="url(#ah7)"/><path class="e" d="M231,40 L277,40" marker-end="url(#ah7)"/><path class="e" d="M317,40 L363,40" marker-end="url(#ah7)"/><path class="e" d="M403,40 L449,40" marker-end="url(#ah7)"/><path class="e" d="M489,40 L535,40" marker-end="url(#ah7)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">In</text><circle class="n" cx="126" cy="40" r="18"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">C1</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">S2</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">C3</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">S4</text><circle class="n" cx="470" cy="40" r="18"/><text class="t" x="470" y="40" dy=".35em" text-anchor="middle">F</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">Out</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">LeNet-5. In = 32x32 image, C1 = 6 maps 28x28, S2 = 6 maps 14x14, C3 = 16 maps 10x10, S4 = 16 maps 5x5, F = FC 120 and 84, Out = 10 classes</figcaption></figure>
Key points.
- Input is a 32x32 grayscale image; C1 uses six 5x5 kernels, giving 6@28x28.
- S2 is 2x2 subsampling (average pooling), giving 6@14x14; C3 has sixteen 5x5 kernels, giving 16@10x10; S4 subsamples to 16@5x5.
- Fully connected layers of 120 and 84 units follow, then 10 output units for digits 0-9; tanh or sigmoid activations were used.
- It has about 60,000 parameters; for example C1 has $6\times(25+1)=156$.
- Role: it showed that convolution, pooling and dense layers trained end to end beat hand-made features, and it was used to read bank cheques and postal codes.
- Limitation: it was small, grayscale and limited by the data and hardware of its time.
Answer frame. Open with LeCun 1998 and digit recognition; draw the diagram; list layers in order (points 1-3); close with its role as the ancestor of modern CNNs.
Asked: [7 marks] (May 2023, May 2024) Briefly discuss the role of LeNet in Convolutional Networks. What is LeNet? Explain its usage in object recognition.
AlexNet
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. AlexNet (Krizhevsky, Sutskever and Hinton, 2012) is a deep CNN that won ImageNet 2012 with a large margin.
Key points.
- It has five convolution layers and three fully connected layers, with about 60 million parameters.
- It used ReLU activations, which train faster than tanh.
- Dropout in the dense layers and data augmentation reduced overfitting.
- Overlapping max pooling and local response normalisation were used, and two GPUs trained it.
- It started the deep learning boom in vision.
Last-minute revision
- Convolution: $S(i,j)=\sum_m\sum_n I(i+m,j+n)K(m,n)$; output size $(W-F+2P)/S+1$.
- Three ideas: sparse interactions, parameter sharing, equivariance.
- Valid shrinks, same keeps size, full enlarges; dilated widens receptive field.
- Max pooling keeps the largest value; average pooling keeps the mean.
- CNN handles grid data with no memory; RNN handles sequences with a hidden state.
- Structured output is a per-pixel label map from a fully convolutional network.
- FFT, Winograd and im2col speed up convolution.
- LeNet-5: 32x32 input, C1 6@28x28, S2, C3 16@10x10, S4, FC 120, 84, output 10; about 60k parameters.
- AlexNet: 2012, 5 conv and 3 FC layers, ReLU, dropout, 60M parameters.
Memory hooks
- SPE: Sparse, Parameter sharing, Equivariance.
- Valid shrinks, Same stays, Full grows.
- LeNet digits 1998 (small); AlexNet ImageNet 2012 (big).
- LeNet layers: C-S-C-S-F-F-O.
Coverage checklist
- The Convolution Operation: 8-mark convolution question, CNN vs RNN, role of filters, CNN overview, pooling.
- Variants of the Basic Convolution Function: variants question.
- Structured Outputs: any-two 14-mark question.
- Efficient Convolution Algorithms: no past question.
- Random or Unsupervised Features: no past question.
- LeNet: both LeNet questions.
- AlexNet: covered here, and as an option in the any-two question.