How unit 3 is examined
This unit covers the CNN pipeline and its building blocks (padding, stride, pooling, dense and loss layers), Inception, transfer and one-shot learning, dimension reduction with PCA, and TensorFlow/Keras implementation; CNN, dimension reduction, padding, Inception and TensorFlow carry the marks.
Convolutional neural network
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>A Convolutional Neural Network (CNN) is a feedforward network that uses shared, small learnable filters to convolve over grid-like data such as images, so it learns a hierarchy of features from edges to objects before a fully connected classifier decides.</mark>
Types of neural networks. Feedforward (MLP) for tabular data, CNN for images, RNN/LSTM for sequences, and autoencoders for compression and unsupervised learning.
Diagram.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-01" viewBox="0 0 517 80" width="517" height="80" role="img" aria-label="CNN pipeline. In = input image, Conv = convolution layer, ReLU = activation, Pool = pooling layer, FC = dense layer, Loss = loss layer. Conv-ReLU-Pool is repeated several times."><style>#dsfig-u3-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-01 .t{fill:#16181D;font-weight:500}#dsfig-u3-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-01 .dot{fill:#16181D}#dsfig-u3-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-01 .ah{fill:#454C5A}#dsfig-u3-01 .ah.hi{fill:#2340B8}#dsfig-u3-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-01 .e{stroke:#B1B7C3}html.dark #dsfig-u3-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-01 .t{fill:#E6E8ED}html.dark #dsfig-u3-01 .t.inv{fill:#0F1115}html.dark #dsfig-u3-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-01 .dot{fill:#E6E8ED}html.dark #dsfig-u3-01 .ann{fill:#8FA3FF}html.dark #dsfig-u3-01 .lbl{fill:#858D9C}html.dark #dsfig-u3-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-01 .ah{fill:#B1B7C3}html.dark #dsfig-u3-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L98,40" marker-end="url(#ah4)"/><path class="e" d="M152,40 L184,40" marker-end="url(#ah4)"/><path class="e" d="M238,40 L270,40" marker-end="url(#ah4)"/><path class="e" d="M324,40 L363,40" marker-end="url(#ah4)"/><path class="e" d="M403,40 L442,40" marker-end="url(#ah4)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">In</text><rect class="n" x="101" y="25" width="50" height="30" rx="15"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">Conv</text><rect class="n" x="187" y="25" width="50" height="30" rx="15"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">ReLU</text><rect class="n" x="273" y="25" width="50" height="30" rx="15"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Pool</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">FC</text><rect class="n" x="445" y="25" width="50" height="30" rx="15"/><text class="t" x="470" y="40" dy=".35em" text-anchor="middle">Loss</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">CNN pipeline. In = input image, Conv = convolution layer, ReLU = activation, Pool = pooling layer, FC = dense layer, Loss = loss layer. Conv-ReLU-Pool is repeated several times.</figcaption></figure>
Key points.
- The convolution layer slides small filters over the image and produces feature maps, so it detects local patterns such as edges and textures.
- An activation (ReLU) is applied to each feature map to add non-linearity; without it stacked convolutions collapse into one linear map.
- Pooling downsamples each feature map, which cuts computation and gives small translation invariance.
- Flattening and dense (fully connected) layers combine the extracted features and produce class scores, and the loss layer (softmax with cross-entropy) measures the error for backpropagation.
- Local connectivity means each neuron sees only a small patch, and weight sharing means one filter is reused at every position, so parameters are far fewer than in an MLP.
- Early layers learn edges, middle layers learn shapes and parts, and deep layers learn whole objects: this is hierarchical feature learning.
- Why the image is downscaled while filters increase: pooling and stride shrink height and width to cut computation and enlarge the receptive field, while more filters are needed because deep features are more abstract and varied (many parts, many combinations), and the extra depth compensates for the spatial information lost.
- Applications include image classification, object detection, face recognition and medical imaging.
Overfitting and underfitting (May 2023).
| Aspect | Underfitting | Overfitting |
|---|---|---|
| Training loss | High | Very low |
| Validation loss | High | High, rising while training falls |
| Cause | Model too simple, too little training | Model too complex, little data |
Detection: split the data into train, validation and test sets, plot training and validation loss (learning curves) per epoch, and compare accuracy. A large train-validation gap means overfitting; both poor means underfitting. Solutions: for overfitting use dropout, L1/L2 regularization, data augmentation, early stopping, batch normalization, a smaller network and more data; for underfitting use a bigger model, more epochs, better features and a lower regularization.
Answer frame. Open with the definition; draw the pipeline diagram with full layer names; develop points 1-6 in order, then the downscaling reason (point 7); close with applications. For the overfitting question write the table, the detection steps, then the solutions list.
Pitfall: Saying pooling adds parameters; pooling has no learnable weights.
Asked: [7 marks] (Dec 2020) What are the different types of Neural networks? Explain the convolution neural network model in detail. Asked: [7 marks] (May 2023, Dec 2024, Jun 2026) Break down how CNN operates; why is the image downscaled and the number of filters increased towards the output? / Explain CNN architecture and how it extracts hierarchical features through layers. / Explain convolution, pooling, dense and loss layers in CNN. Asked: [7 marks] (May 2023) Describe the procedure to identify overfitting and underfitting in a CNN model, with potential solutions.
Flattening
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Flattening converts the final 3-D stack of feature maps into a single 1-D vector so it can enter the dense layer.
Key points.
- A $7\times7\times64$ output becomes a vector of $7\times7\times64=3136$ values.
- Flattening has no parameters; it only reshapes.
- It sits between the last pooling layer and the first dense layer.
Subsampling
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Subsampling (pooling) reduces the height and width of a feature map by summarising each small window with one value, keeping the strongest features while discarding detail.</mark>
Key points.
- Steps: choose a window (e.g. $2\times2$) and stride (2), slide it over each feature map, and take the maximum (max pooling) or the mean (average pooling) of each window.
- A $28\times28$ map becomes $14\times14$, so computation and memory fall to a quarter.
- Depth is unchanged because each channel is pooled independently.
- Subsampling makes features robust to small shifts and reduces overfitting.
- Keras features: a modular, high-level API on TensorFlow (earlier also Theano/CNTK backends), with ready layers, simple
compile,fitandevaluate, and easy CPU/GPU running.
Asked: [7 marks] (Dec 2020) Explain the process of sub-sampling of input data in a neural network model. Some of the features of Keras framework for implementing neural network models.
Padding
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Padding adds extra border pixels (usually zeros) around the input so that the filter can cover edge pixels and the output size can be controlled.</mark>
Formula. For input $n$, kernel $k$, padding $p$, stride $s$:
$$\text{output}=\left\lfloor\frac{n+2p-k}{s}\right\rfloor+1$$
Key points.
- Without padding, every convolution shrinks the map and edge pixels are used less often, so edge information is lost.
- Valid padding adds no zeros ($p=0$), so the output is $n-k+1$ (for $s=1$) and the map shrinks.
- Same padding adds enough zeros ($p=(k-1)/2$ for odd $k$) so the output size equals the input size (for $s=1$).
- Full padding adds $k-1$ zeros on each side, so every pixel meets every filter position and the output is $n+k-1$; it is rarely used.
- Zero padding is the common choice because zeros add no false signal.
- Padding therefore preserves spatial resolution, lets us build deep networks, and controls output dimensions.
Example. $n=32$, $k=5$: valid gives $28$; same ($p=2$) gives $32$; full ($p=4$) gives $36$.
Other terms in the Jun 2026 question.
| Term | Meaning |
|---|---|
| Stride | Step of the filter; larger stride shrinks output |
| Flattening | Reshape feature maps to a 1-D vector |
| Subsampling | Pooling to reduce height and width |
| Input channels | Depth of input (3 for RGB); each filter has $k\times k\times C$ weights |
| $1\times1$ convolution | Mixes channels per pixel; changes depth cheaply |
Answer frame. Open with the definition and draw a $5\times5$ grid with a zero border; explain valid, same and full with the formula and the numbers above; close with why padding preserves edge information. For the Jun 2026 question add one line each from the table.
Asked: [7 marks] (May 2024, Jun 2025) List and explain the types of padding used in CNNs. / Define padding. How does padding work in CNN? Asked: [7 marks] (Jun 2026) Discuss padding, stride, flattening, subsampling, input channels, and $1\times1$ convolution.
Stride
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Stride is the number of pixels the filter moves at each step.
Key points.
- Stride 1 moves one pixel and keeps detail; stride 2 halves the output size.
- Output $=\lfloor (n+2p-k)/s\rfloor+1$; for $n=7$, $k=3$, $p=0$, $s=2$ the output is $3$.
- A larger stride is a cheap way to downsample without a pooling layer.
Convolution layer
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>The convolution layer applies learnable filters that slide over the input and compute dot products, producing feature maps that detect local features.</mark>
Key points.
- Each filter of size $k\times k\times C$ produces one feature map, so $F$ filters give $F$ maps.
- Parameters per layer $=(k\cdot k\cdot C+1)\cdot F$; for $5\times5$, $C=3$, $F=64$ this is $4864$.
- Weights are shared across positions and connectivity is local.
- Overview of layers (Dec 2020): convolution extracts features, pooling downsamples, dense classifies and loss measures the error.
Asked: [7 marks] (Dec 2020) Explain the different layers in a neural network. What do you mean by convolution layer, pooling layer, loss layer, dense layer? Describe each in brief.
Pooling layer
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>A pooling layer downsamples each feature map by replacing every window with its maximum or average, reducing spatial size without learnable weights.</mark>
Key points.
- Max pooling keeps the largest value in the window and preserves the strongest activation; average pooling keeps the mean and gives smoother output.
- With a $2\times2$ window and stride 2, a $28\times28$ map becomes $14\times14$: output $=\lfloor (n-f)/s\rfloor+1$.
- Because each window is summarised by one value, the next layer has fewer inputs and parameters, so computation and overfitting fall.
- A small shift in the input leaves the window maximum unchanged, which gives translation invariance.
- The number of channels is unchanged.
Asked: [7 marks] (Dec 2024) Explain how pooling layers reduce the spatial dimensions of feature maps.
Loss layer
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. The loss layer is the last layer; it compares predictions with true labels and outputs the error that backpropagation minimises.
Key points.
- Classification uses softmax with cross-entropy $L=-\sum y_i\log\hat y_i$.
- Regression uses mean squared error.
- Its gradient starts the backward pass that updates the filters.
Dance layer
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. The "dance layer" is the dense (fully connected) layer, in which every neuron connects to every neuron of the previous layer.
Key points.
- It takes the flattened feature vector and combines features for classification.
- Output $=f(Wx+b)$; the last dense layer has one unit per class with softmax.
- It holds most of the CNN parameters, e.g. $3136\times128$ weights.
1x1 convolution
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A $1\times1$ convolution applies a filter of size $1\times1\times C$ at each pixel, mixing channels without looking at neighbours.
Key points.
- It reduces or increases depth, e.g. $256\to64$ channels, cheaply.
- It costs $1\cdot1\cdot256\cdot64=16384$ weights against $25\cdot256\cdot64=409600$ for a $5\times5$ filter.
- It adds non-linearity through the following ReLU and is used as a bottleneck in Inception.
Inception network
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>The Inception network (GoogLeNet) is a CNN built from Inception modules that apply $1\times1$, $3\times3$ and $5\times5$ convolutions and pooling in parallel and concatenate the results.</mark>
Diagram.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-02" viewBox="0 0 338 338" width="338" height="338" role="img" aria-label="Inception module. A = 1x1 conv, B = 1x1 then 3x3, C = 1x1 then 5x5, D = 3x3 max pool then 1x1. Cat = channel concatenation."><style>#dsfig-u3-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-02 .t{fill:#16181D;font-weight:500}#dsfig-u3-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-02 .dot{fill:#16181D}#dsfig-u3-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-02 .ah{fill:#454C5A}#dsfig-u3-02 .ah.hi{fill:#2340B8}#dsfig-u3-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-02 .e{stroke:#B1B7C3}html.dark #dsfig-u3-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-02 .t{fill:#E6E8ED}html.dark #dsfig-u3-02 .t.inv{fill:#0F1115}html.dark #dsfig-u3-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-02 .dot{fill:#E6E8ED}html.dark #dsfig-u3-02 .ann{fill:#8FA3FF}html.dark #dsfig-u3-02 .lbl{fill:#858D9C}html.dark #dsfig-u3-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-02 .ah{fill:#B1B7C3}html.dark #dsfig-u3-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M53.4,155.6 L154.2,54.8" marker-end="url(#ah5)"/><path class="e" d="M58,163 L149.1,132.6" marker-end="url(#ah5)"/><path class="e" d="M58,175 L149.1,205.4" marker-end="url(#ah5)"/><path class="e" d="M53.4,182.4 L154.2,283.2" marker-end="url(#ah5)"/><path class="e" d="M182.4,53.4 L283.2,154.2" marker-end="url(#ah5)"/><path class="e" d="M187,132 L278.1,162.4" marker-end="url(#ah5)"/><path class="e" d="M187,206 L278.1,175.6" marker-end="url(#ah5)"/><path class="e" d="M182.4,284.6 L283.2,183.8" marker-end="url(#ah5)"/><circle class="n" cx="40" cy="169" r="18"/><text class="t" x="40" y="169" dy=".35em" text-anchor="middle">In</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">A</text><circle class="n" cx="169" cy="126" r="18"/><text class="t" x="169" y="126" dy=".35em" text-anchor="middle">B</text><circle class="n" cx="169" cy="212" r="18"/><text class="t" x="169" y="212" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="169" cy="298" r="18"/><text class="t" x="169" y="298" dy=".35em" text-anchor="middle">D</text><circle class="n" cx="298" cy="169" r="18"/><text class="t" x="298" y="169" dy=".35em" text-anchor="middle">Cat</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Inception module. A = 1x1 conv, B = 1x1 then 3x3, C = 1x1 then 5x5, D = 3x3 max pool then 1x1. Cat = channel concatenation.</figcaption></figure>
Key points.
- Each module lets the network look at several scales at once, so it does not have to choose one filter size.
- $1\times1$ convolutions before the $3\times3$ and $5\times5$ branches reduce channels, which cuts computation greatly.
- Outputs of all branches are concatenated along the depth axis.
- GoogLeNet has 22 layers with nine stacked modules, and uses global average pooling instead of large dense layers, so it has about 5 million parameters.
- Auxiliary classifiers in middle layers give extra gradient and reduce vanishing gradients.
- It won ILSVRC 2014.
Transfer learning link. Transfer learning reuses a model trained on a large dataset (e.g. Inception on ImageNet) for a new task; the early layers hold general features (edges, textures) that transfer. Benefits: less data needed, faster training and better generalisation. Applications: medical image diagnosis, plant disease detection, product recognition. Limitation: negative transfer if the domains differ greatly.
Answer frame. Open with the definition of Inception; draw the module; develop points 1-5; close with the transfer learning definition, benefits, applications and limitation.
Asked: [7 marks] (May 2023, Jun 2026) Describe the benefits of transfer learning features that can be transferred. Explain Inception net architecture in detail. / Explain the Inception Network and transfer learning with suitable applications.
Input channels
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Input channels are the depth of the input: 1 for greyscale, 3 for RGB; later layers receive the previous layer's feature maps as channels.
Key points.
- A filter has the same depth as the input, $k\times k\times C$, and sums over all channels into one output map.
- Number of filters equals the number of output channels.
- A $32\times32\times3$ image with 10 filters of $5\times5\times3$ gives $28\times28\times10$ (valid).
Transfer learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Transfer learning reuses a model pre-trained on a large source task as the starting point for a new target task.</mark>
Key points.
- Feature extraction freezes the pre-trained weights, uses the network as a fixed feature extractor, and trains only a new final classifier; it is fast and needs little data, so use it when the domains are similar and data is small.
- Fine-tuning initialises with pre-trained weights and unfreezes some or all layers, retraining with a small learning rate; it is more effective but needs more data.
- Domain adaptation transfers knowledge between a source and a target domain with different data distributions.
- Choose feature extraction for similar domains with limited data, and fine-tuning when the target data is sufficient.
Asked: [7 marks] (May 2024) Describe the types of transfer learning, including feature extraction and fine-tuning.
One shot learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>One-shot learning classifies new categories after seeing only one labelled example of each.</mark>
Key points.
- Instead of learning classes, the network learns a similarity function, so it compares a new input with the single stored example.
- Siamese networks pass two inputs through twin CNNs with shared weights and compare the embeddings by distance; training uses contrastive or triplet loss.
- Use cases: face recognition and verification, signature verification and rare-object recognition.
- It works because the embedding learned from many classes transfers to unseen classes.
Asked: [7 marks] (Jun 2026) Discuss one-shot learning, dimension reduction and CNN implementation using TensorFlow and Keras.
Dimension reductions
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Dimensionality reduction is the process of reducing the number of input features while keeping as much useful information (variance) as possible.</mark>
Why high-dimensional data is hard.
- Curse of dimensionality: volume grows exponentially with dimensions, so data becomes sparse and points are isolated.
- Overfitting: with more features than samples the model memorises noise instead of learning patterns.
- Computational cost: training becomes slower and needs more memory.
- Distances become nearly equal in high dimensions, so distance-based methods (KNN, clustering) lose meaning.
- Solutions: PCA, LDA, t-SNE, feature selection and regularisation.
Steps of PCA.
Step 1: Standardise each feature (subtract the mean, divide by standard deviation).
Step 2: Compute the covariance matrix C.
Step 3: Find eigenvalues and eigenvectors of C.
Step 4: Sort eigenvectors by decreasing eigenvalue; keep the top k as principal components.
Step 5: Project the data: Z = X_centred . W_k.
Example. Points $(2,1),(3,5),(4,3),(5,6),(6,7),(7,8)$: means $(4.5,5)$; $\text{var}_x=3.5$, $\text{var}_y=6.8$, $\text{cov}=4.4$. Eigenvalues are $9.85$ and $0.45$; the first eigenvector is $(0.57,0.82)$. PC1 keeps $9.85/10.3=95.6\%$ of the variance, so 2-D reduces to 1-D.
Key points on PCA.
- PCA gives orthogonal, uncorrelated components ordered by variance.
- Select $k$ so the cumulative explained variance $\sum_{i\le k}\lambda_i/\sum\lambda_i$ reaches about 90-95%.
- Benefits: visualisation in 2-D/3-D, faster training, less noise and storage. Data can be approximately reconstructed by $Z W_k^T$.
- Other techniques: LDA (supervised, maximises class separation) and t-SNE (non-linear, for visualisation). Applications: image compression, face recognition (eigenfaces), genomics.
Answer frame. For PCA, open with the definition, write the five steps, add the small example and close with benefits. For the "why hard" question, list points 1-5 in order. For the definition question, give definition, need (curse), techniques (PCA, LDA, t-SNE) and applications.
Pitfall: Forgetting to standardise or centre the data before computing the covariance matrix.
Asked: [7 marks] (Dec 2020, May 2022) Describe how principal component analysis is carried out to reduce the dimensionality of data sets. / Explain in detail principal component analysis for dimension reduction. Asked: [7 marks] (May 2024) Explain why high-dimensional data can be challenging for machine learning algorithms. Asked: [7 marks] (Jun 2025) What do you mean by dimension reduction? Discuss in detail.
Implementation of CNN like TensorFlow
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>TensorFlow is an open-source library that represents computation as a graph of tensor operations and provides ready CNN layers, automatic differentiation, GPU training and deployment tools.</mark>
Key points.
- Computational graph: operations are nodes and tensors flow along the edges, so TensorFlow can optimise and differentiate the graph automatically.
- Layers such as
Conv2D,MaxPooling2D,FlattenandDenseare built in through the Keras API. - GPU/TPU support makes training fast, and
tf.databuilds input pipelines (loading, batching, augmentation). - Workflow: load data, build the model, compile, train, evaluate, and save or deploy (SavedModel, TF Lite).
import tensorflow as tf
from tensorflow.keras import layers, models
model = models.Sequential([
layers.Conv2D(32, 3, activation="relu", input_shape=(28, 28, 1)),
layers.MaxPooling2D(2),
layers.Flatten(),
layers.Dense(10, activation="softmax")])
model.compile(optimizer="adam", loss="sparse_categorical_crossentropy", metrics=["accuracy"])
model.fit(x_train, y_train, epochs=5, validation_split=0.1)
model.evaluate(x_test, y_test) # reports test loss and accuracy
Answer frame. Open with the role of TensorFlow; list points 1-3; write the code above with a comment per line; close with the workflow and deployment.
Asked: [7 marks] (Dec 2024, Jun 2025) Describe the role of TensorFlow in facilitating the implementation and training of CNNs. / Explain the process of implementing CNN in TensorFlow.
Keras
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Keras is a high-level, user-friendly deep learning API that runs on top of TensorFlow.
Key points.
- It is modular: layers, losses and optimisers are plug-in blocks.
Sequentialand the functional API build models in a few lines.- It has simple
compile,fitandevaluatecalls and supports CPU and GPU.
Last-minute revision
- CNN = convolution, ReLU, pooling, then flatten, dense and loss layers; filters are shared and connections local.
- Output size $=\lfloor (n+2p-k)/s\rfloor+1$; same padding $p=(k-1)/2$; valid gives $n-k+1$; full gives $n+k-1$.
- Downscale to cut computation and grow the receptive field; add filters to hold more abstract features.
- Pooling has no weights; max pooling keeps the strongest value; a $2\times2$ stride-2 pool halves height and width.
- Parameters of a conv layer $=(k^2C+1)F$.
- $1\times1$ convolution changes depth cheaply (bottleneck); Inception runs $1\times1$, $3\times3$, $5\times5$ and pooling in parallel and concatenates.
- Feature extraction freezes weights; fine-tuning retrains with a small learning rate.
- One-shot learning learns similarity (Siamese network) from one example per class.
- PCA: standardise, covariance, eigenvectors, sort, project; keep about 90-95% variance.
- Curse of dimensionality: sparse data, overfitting, slow computation, useless distances.
- Overfitting = low train loss, high validation loss; fix with dropout, augmentation, early stopping, regularisation.
- TensorFlow flow: load, build, compile, fit, evaluate, deploy.
Memory hooks
- CRPFD: Convolve, ReLU, Pool, Flatten, Dense (then loss).
- "Valid shrinks, Same stays, Full grows."
- Inception = "many eyes, one output": parallel filters, then concatenate.
- Feature extraction = freeze, fine-tuning = thaw slowly.
- PCA = Standardise, Covariance, Eigen, Keep k, Project (SCEKP).
Coverage checklist
- Convolutional neural network: Q1, Q2, Q3 (Dec 2020, May 2023, Dec 2024, Jun 2026).
- flattening: covered in the Jun 2026 padding question (Q12).
- subsampling: Q14.
- padding: Q11, Q12.
- stride: Q12.
- convolution layer: Q4.
- pooling layer: Q13, Q4.
- loss layer: Q2, Q4.
- dance layer: Q2, Q4.
- 1x1 convolution: Q12, Q9.
- inception network: Q9.
- input channels: Q12.
- transfer learning: Q15, Q9.
- one shot learning: Q10.
- dimension reductions: Q5, Q6, Q7, Q10.
- implementation of CNN like tensor flow: Q8, Q10.
- keras: Q14, Q10.