How unit 2 is examined
This unit covers the building blocks of neural-network training; the marks sit in loss functions and optimizers (14 marks), activation functions, gradient descent, backpropagation, the multilayer perceptron, autoencoders, L1/L2 regularization and hyperparameter tuning.
Linearity vs non linearity
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. A linear model computes $y = w^T x + b$, so its output changes in proportion to its input and its decision boundary is a straight line or plane. A non-linear model passes the weighted sum through a non-linear function so it can bend the boundary.
Key points.
- A model is made non-linear by applying a non-linear activation function (sigmoid, tanh, ReLU) to the weighted sum of every hidden neuron.
- Without activations, stacking layers gives $W_2(W_1 x) = (W_2 W_1)x$, so any number of linear layers collapses into one linear layer and depth adds nothing.
- A purely linear model cannot learn XOR or any pattern that is not linearly separable; non-linear hidden layers can approximate any continuous function.
- For gradient descent, a linear model with squared error has a convex loss, so there is one global minimum and convergence is easy but the model is weak.
- A non-linear network has a non-convex loss with many local minima and saddle points, so gradient descent needs a good learning rate and only reaches a good, not guaranteed best, solution.
<mark>A network with only linear activations is equivalent to a single linear model, so non-linear activation functions are what give a neural network its power.</mark>
Answer frame. Open with the definition of linear versus non-linear; show the collapse $W_2W_1$ in one line; then activations as the fix; close with convex versus non-convex effect on gradient descent.
Asked: [7 marks] (May 2023) How can we make the model non-linear? If we only use linearity, how will that affect Gradient Descent?
activation functions like sigmoid
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. An activation function is a non-linear function applied to a neuron's weighted sum $z = w^Tx + b$ to produce its output. A perceptron is the single neuron that computes $z$ and applies a threshold or activation.
Formula. $\sigma(z) = \dfrac{1}{1+e^{-z}}$, and its derivative is $\sigma'(z) = \sigma(z)(1-\sigma(z))$.
Key points.
- Sigmoid is S-shaped and squashes any real input into the range $(0,1)$, with $\sigma(0)=0.5$.
- Its output can be read as a probability, so it is used in the output neuron of binary classification.
- It is smooth and differentiable, and the derivative reuses the output.
- The maximum derivative is 0.25 at $z=0$; for large $|z|$ the curve is flat and the gradient is near zero (saturation), which causes the vanishing gradient problem in deep networks.
- Its output is not zero-centred, which slows convergence.
| Property | Sigmoid | ReLU |
|---|---|---|
| Formula | $1/(1+e^{-z})$ | $\max(0,z)$ |
| Output range | $(0,1)$ | $[0,\infty)$ |
| Vanishing gradient | Yes, gradient below 0.25 and near 0 at extremes | No for $z>0$, gradient is 1 |
| Drawback | Saturation, not zero-centred | Dying ReLU, gradient 0 for $z<0$ |
Classification vs regression. Classification predicts a discrete class (spam or not, logistic regression, SVM, sigmoid or softmax output, cross-entropy loss); regression predicts a continuous value (house price, linear regression, linear output, MSE loss).
<mark>Sigmoid squashes any input into (0,1) so it acts as a probability, but its gradient never exceeds 0.25 and vanishes at both ends.</mark>
Answer frame. Open with the formula and range; draw the S-curve with 0, 0.5, 1 marked; develop points 1-5, then the sigmoid versus ReLU table; close with vanishing gradient as the reason ReLU is preferred in hidden layers.
Asked: [7 marks] (Dec 2020) Explain the concept of perceptron, back propagation and sigmoid activation function in brief. Differentiate between classification and regression. Asked: [7 marks] (May 2024) Discuss the advantages and limitations of sigmoid and ReLU activation functions in terms of vanishing gradient problem and output range. Asked: [5 marks] (Jun 2025) Discuss the sigmoid activation function in detail.
ReLU
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. The rectified linear unit is $f(z)=\max(0,z)$: it passes positive inputs unchanged and outputs zero for negative inputs.
Key points.
- Its gradient is 1 for $z>0$ and 0 for $z<0$, so it does not saturate for positive inputs and avoids vanishing gradients.
- It is very cheap to compute and gives sparse activations and faster training.
- A neuron stuck with negative inputs always outputs 0 and stops learning (dying ReLU); Leaky ReLU $\max(0.01z, z)$ fixes it.
weights and bias
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. Weights $w$ are learnable numbers that scale each input and decide how strongly it influences the neuron; bias $b$ is a learnable offset added to the weighted sum, $z = \sum w_ix_i + b$.
Key points.
- Weights are the parameters the network learns; training adjusts them to reduce the loss.
- Bias shifts the activation threshold, like the intercept of a line.
- In the forward pass inputs are multiplied by weights and passed through activations to the output; in the backward pass the loss gradient updates $w$ and $b$ by gradient descent.
- Together: the forward pass gives the prediction, the loss measures the error, and gradient descent uses the backward pass to update $w$ and $b$.
Asked: [7 marks] (Jun 2026) Explain weights, bias, loss function and gradient descent in neural networks.
loss function
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. A loss function measures how far the network's prediction $\hat y$ is from the true value $y$ for one example; the cost function is its average over the data, and training minimises it.
Key points.
- Mean squared error $L=\frac1n\sum(y-\hat y)^2$ is used for regression; it punishes large errors heavily and is smooth.
- Mean absolute error $\frac1n\sum|y-\hat y|$ suits regression with outliers.
- Binary cross-entropy $-[y\log\hat y+(1-y)\log(1-\hat y)]$ is used for two-class problems with a sigmoid output.
- Categorical cross-entropy $-\sum y_k\log\hat y_k$ is used for multi-class problems with a softmax output.
- Hinge loss $\max(0,1-y\hat y)$ with $y=\pm1$ is used for SVM-style margin classification.
| Loss | Task | Use case |
|---|---|---|
| MSE | Regression | Price prediction |
| MAE | Regression | Data with outliers |
| Cross-entropy | Classification | Probabilistic outputs |
| Hinge | Classification | Maximum-margin, SVM |
| Optimizer | Data per update | Computation |
| --- | --- | --- |
| Batch GD | All samples | Slow per step |
| Stochastic GD | One sample | Fast per step |
| Mini-batch GD | Small subset (32-256) | Balanced |
| RMSprop | Mini-batch, divides step by running average of squared gradients | Slightly more |
| Adam | Momentum plus RMSprop | Slightly more |
<mark>A loss function converts the prediction error into a single number that gradient descent then minimises.</mark>
Answer frame. Part (i): define loss, then the loss table with formula and use case; part (ii): one-line update rule, then the optimizer table; close by naming Adam as the common default for its adaptive rate and speed.
Pitfall. Using MSE for classification or cross-entropy for regression; match the loss to the task.
Asked: [14 marks] (May 2023) Differentiate: i) Various loss functions ii) Types of Gradient Descent Optimizers
gradient descent
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. Gradient descent is an iterative optimization algorithm that minimises a loss function by repeatedly moving the parameters in the direction opposite to the gradient.
Formula. $\theta \leftarrow \theta - \alpha\,\nabla J(\theta)$, where $\alpha$ is the learning rate.
Key points.
- The gradient points towards steepest increase of the loss, so subtracting it moves the parameters downhill.
- A small learning rate $\alpha$ converges slowly; a large one overshoots the minimum and may diverge.
- Batch GD uses all samples per update, giving a smooth, stable path but slow steps and high memory.
- Stochastic GD uses one sample per update, so it is fast and can escape shallow minima but its path is noisy.
- Mini-batch GD uses a small batch, balancing speed, stability and hardware efficiency, and is the standard choice.
- It stops when the change in loss is below a tolerance or after a fixed number of epochs.
Answer frame. Open with the definition and update rule; explain the role of $\alpha$; then the three types in the table above with pros and cons; close with mini-batch as the practical default.
Asked: [7 marks] (May 2022, May 2024) What do you mean by Gradient Descent? Define gradient descent in the context of optimization algorithms used in machine learning. Asked: [7 marks] (Jun 2025) Explain the types of Gradient descent.
multilayer network
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. A multilayer perceptron (MLP) is a feed-forward neural network with an input layer, one or more hidden layers and an output layer, where every neuron is connected to all neurons of the next layer with weights.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-01" viewBox="0 0 424 338" width="424" height="338" role="img" aria-label="MLP: X1,X2 input layer, H1-H3 hidden layer (weights and activation), O1,O2 output layer"><style>#dsfig-u2-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-01 .t{fill:#16181D;font-weight:500}#dsfig-u2-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-01 .dot{fill:#16181D}#dsfig-u2-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-01 .ah{fill:#454C5A}#dsfig-u2-01 .ah.hi{fill:#2340B8}#dsfig-u2-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-01 .e{stroke:#B1B7C3}html.dark #dsfig-u2-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-01 .t{fill:#E6E8ED}html.dark #dsfig-u2-01 .t.inv{fill:#0F1115}html.dark #dsfig-u2-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-01 .dot{fill:#E6E8ED}html.dark #dsfig-u2-01 .ann{fill:#8FA3FF}html.dark #dsfig-u2-01 .lbl{fill:#858D9C}html.dark #dsfig-u2-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-01 .ah{fill:#B1B7C3}html.dark #dsfig-u2-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M58.4,78.4 L191.6,45.1" marker-end="url(#ah2)"/><path class="e" d="M57,91.5 L193.2,159.6" marker-end="url(#ah2)"/><path class="e" d="M51.9,97.8 L198.9,281.6" marker-end="url(#ah2)"/><path class="e" d="M51.9,240.2 L198.9,56.4" marker-end="url(#ah2)"/><path class="e" d="M57,246.5 L193.2,178.4" marker-end="url(#ah2)"/><path class="e" d="M58.4,259.6 L191.6,292.9" marker-end="url(#ah2)"/><path class="e" d="M229,48.5 L365.2,116.6" marker-end="url(#ah2)"/><path class="e" d="M223.9,54.8 L370.9,238.6" marker-end="url(#ah2)"/><path class="e" d="M230.4,164.4 L363.6,131.1" marker-end="url(#ah2)"/><path class="e" d="M229,177.5 L365.2,245.6" marker-end="url(#ah2)"/><path class="e" d="M225.4,284.6 L369.2,140.8" marker-end="url(#ah2)"/><path class="e" d="M230.4,293.4 L363.6,260.1" marker-end="url(#ah2)"/><circle class="n" cx="40" cy="83" r="18"/><text class="t" x="40" y="83" dy=".35em" text-anchor="middle">X1</text><circle class="n" cx="40" cy="255" r="18"/><text class="t" x="40" y="255" dy=".35em" text-anchor="middle">X2</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">H1</text><circle class="n" cx="212" cy="169" r="18"/><text class="t" x="212" y="169" dy=".35em" text-anchor="middle">H2</text><circle class="n" cx="212" cy="298" r="18"/><text class="t" x="212" y="298" dy=".35em" text-anchor="middle">H3</text><circle class="n" cx="384" cy="126" r="18"/><text class="t" x="384" y="126" dy=".35em" text-anchor="middle">O1</text><circle class="n" cx="384" cy="255" r="18"/><text class="t" x="384" y="255" dy=".35em" text-anchor="middle">O2</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">MLP: X1,X2 input layer, H1-H3 hidden layer (weights and activation), O1,O2 output layer</figcaption></figure>
Key points.
- The input layer only passes the features on; it does no computation.
- Each hidden neuron computes $h=f(Wx+b)$ with a non-linear activation, which lets the network learn non-linear patterns such as XOR.
- The output layer gives the prediction, using sigmoid or softmax for classification or a linear unit for regression.
- In the forward pass data flows layer by layer from input to output.
- Training uses backpropagation to compute the error gradients and gradient descent to update the weights.
- More hidden layers give more power but more risk of overfitting.
Answer frame. Open with the definition; draw the labelled diagram with input, hidden and output layers, weights on arrows; develop points 1-6; close with training by backpropagation, then a line each on weight initialization, training and testing if the question lists them.
Asked: [7 marks] (Dec 2020, Jun 2026) Explain the multilayer perceptron model in detail with neat diagram. Discuss multilayer neural networks, backpropagation, weight initialization, training and testing.
backpropagation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. Backpropagation is the algorithm that computes the gradient of the loss with respect to every weight by applying the chain rule backwards from the output layer to the input layer, and then updates the weights by gradient descent.
Formula. Output delta $\delta_o=(t-o)\,o(1-o)$; hidden delta $\delta_h=h(1-h)\sum w\,\delta_o$; update $\Delta w=\eta\,\delta\,x_{in}$.
Step 1: Initialise all weights and biases with small random values.
Step 2: Forward pass: compute each layer's output with the activation function up to the output o.
Step 3: Compute the error t - o and the output delta.
Step 4: Propagate deltas backwards to each hidden layer using the chain rule.
Step 5: Update every weight: w = w + eta * delta * input.
Step 6: Repeat for all patterns until the error is small or the epoch limit is reached.
Key points.
- The chain rule breaks the gradient of a nested function into a product of local derivatives, layer by layer.
- Reusing the deltas of later layers makes one backward pass cost about the same as a forward pass, instead of computing each weight's gradient separately.
Example (May 2023). Given $x=[0,1]$, $t=1$, $\eta=0.3$. $z_1$ net $=0.5(0)+0.6(1)+0.2=0.8$, $h_1=0.6900$; $z_2$ net $=-0.1(0)+0.7(1)+0.4=1.1$, $h_2=0.7503$; $y$ net $=0.3(0.6900)+0.2(0.7503)+0.4=0.7571$, $o=0.6807$.
| Quantity | Working | Value |
|---|---|---|
| $\delta_o$ | $(1-0.6807)(0.6807)(0.3193)$ | 0.0694 |
| $\delta_1$, $\delta_2$ | $h(1-h)\,\delta_o\,w$ | 0.00445, 0.00260 |
| $w_{z1y}$ | $0.3+0.3(0.0694)(0.6900)$ | 0.3144 |
| $w_{z2y}$ | $0.2+0.3(0.0694)(0.7503)$ | 0.2156 |
| $w_{x2z1}$, $w_{x2z2}$ | $w+0.3\,\delta_h(1)$ | 0.6013, 0.7008 |
New weights: $w_{x1z1}=0.5$, $w_{x1z2}=-0.1$ (unchanged as $x_1=0$), $w_{x2z1}=0.6013$, $w_{x2z2}=0.7008$, $w_{z1y}=0.3144$, $w_{z2y}=0.2156$.
<mark>Backpropagation applies the chain rule backwards through the network so all weight gradients are found in one pass.</mark>
Answer frame. Algorithm question: open with the definition, write the six steps, close with the stopping criterion. Chain-rule question: state the rule, show one layer's product, stress efficiency. Numerical: forward pass, deltas, updates, boxed weights.
Asked: [7 marks] (May 2022) Write the algorithm for back propagation. Asked: [7 marks] (May 2023) Replace the old weights (not the bias) in the network depicted in the figure using back-propagation. Input [0,1], desired output 1, sigmoid activation, learning rate 0.3. Asked: [7 marks] (Dec 2024) Discuss the importance of the chain rule in the back propagation algorithm for computing gradients efficiently.
weight initialization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Weight initialization sets the starting values of the weights before training so that signals and gradients neither vanish nor explode.
Key points.
- All-zero weights make every neuron in a layer identical, so they learn the same thing; random values break this symmetry.
- Xavier (Glorot) initialization draws weights with variance $2/(n_{in}+n_{out})$ and suits sigmoid and tanh.
- He initialization uses variance $2/n_{in}$ and suits ReLU.
- Good initialization speeds convergence and reduces unstable gradients.
training
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. Training is the process of feeding the labelled training data through the network, computing the loss and updating the weights repeatedly until the model fits the data.
| Basis | Training data | Testing data |
|---|---|---|
| Purpose | Learn the weights | Evaluate the final model |
| Use | Repeatedly, in every epoch | Once, after training |
| Labels | Used to compute the loss | Used only to compare with predictions |
| Role | Learning | Measuring generalization |
| Typical share | 70-80% | 20-30% |
Key points.
- A common split is 80:20 or 70:30, with a separate validation set to tune hyperparameters.
- High training accuracy with low test accuracy signals overfitting, meaning poor generalization.
Asked: [7 marks] (May 2022) Differentiate between Training data and Testing data.
testing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. Testing evaluates the trained model on unseen data; the confusion matrix tabulates its predictions against the true classes.
| Predicted positive | Predicted negative | |
|---|---|---|
| Actual positive | TP | FN |
| Actual negative | FP | TN |
Formula. Accuracy $=\dfrac{TP+TN}{TP+TN+FP+FN}$; Precision $=\dfrac{TP}{TP+FP}$; Recall $=\dfrac{TP}{TP+FN}$; F1 $=\dfrac{2PR}{P+R}$.
Key points.
- On imbalanced data (for example 99% negatives) accuracy misleads, so precision, recall and F1 are used.
- F1 is the harmonic mean of precision and recall.
Asked: [7 marks] (May 2023) Clarify the meaning of Confusion metrics in the context of machine learning. What other metrics might be derived from the metric of confusion?
unstable gradient problem
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. In deep networks the gradients shrink (vanishing) or grow (exploding) exponentially as they are multiplied through many layers during backpropagation.
Key points.
- Vanishing gradients occur when derivatives are below 1, as with sigmoid (maximum 0.25), so early layers barely learn.
- Exploding gradients occur when weights or derivatives are above 1, giving huge updates and unstable loss.
- Remedies: ReLU, careful initialization (Xavier, He), batch normalization, gradient clipping and residual connections.
auto encoders
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. An autoencoder is an unsupervised neural network trained to reproduce its input at its output, by compressing the input into a small code and rebuilding it.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-02" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="Autoencoder: Input, Encoder, Z bottleneck (latent code), Decoder, Output (reconstruction)"><style>#dsfig-u2-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-02 .t{fill:#16181D;font-weight:500}#dsfig-u2-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-02 .dot{fill:#16181D}#dsfig-u2-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-02 .ah{fill:#454C5A}#dsfig-u2-02 .ah.hi{fill:#2340B8}#dsfig-u2-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-02 .e{stroke:#B1B7C3}html.dark #dsfig-u2-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-02 .t{fill:#E6E8ED}html.dark #dsfig-u2-02 .t.inv{fill:#0F1115}html.dark #dsfig-u2-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-02 .dot{fill:#E6E8ED}html.dark #dsfig-u2-02 .ann{fill:#8FA3FF}html.dark #dsfig-u2-02 .lbl{fill:#858D9C}html.dark #dsfig-u2-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-02 .ah{fill:#B1B7C3}html.dark #dsfig-u2-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah3" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh3" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah3)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah3)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah3)"/><path class="e" d="M446,40 L535,40" marker-end="url(#ah3)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">In</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">Enc</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Z</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">Dec</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">Out</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Autoencoder: Input, Encoder, Z bottleneck (latent code), Decoder, Output (reconstruction)</figcaption></figure>
Key points.
- The encoder maps the input to a lower-dimensional code $z=f(x)$, and the decoder reconstructs $\hat x=g(z)$.
- The bottleneck layer has fewer neurons than the input, so the network cannot copy the input and must keep only the most important features.
- Training minimises the reconstruction loss $\|x-\hat x\|^2$ without labels.
- Sparse autoencoders add a sparsity penalty; denoising autoencoders rebuild clean input from noisy input; variational autoencoders learn a probability distribution for generating new data; undercomplete or vanilla is the basic form.
- Uses: dimensionality reduction, denoising, anomaly detection.
Answer frame. Open with the definition; draw the encoder-bottleneck-decoder block; develop points 1-3; list the types with one line each; close with applications.
Asked: [7 marks] (Dec 2024) Describe the role of the bottleneck layer in auto encoders and how it captures the essential features of the input data. Asked: [7 marks] (Jun 2025) What are Auto Encoders? Explain the types of Auto Encoders.
batch normalization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. Batch normalization normalises the inputs of a layer over each mini-batch to zero mean and unit variance, then rescales them with learnable parameters.
Formula. $\hat x=\dfrac{x-\mu_B}{\sqrt{\sigma_B^2+\epsilon}}$, then $y=\gamma\hat x+\beta$.
Key points.
- It is placed after the linear (or convolution) layer and before the activation.
- It allows higher learning rates and faster training and reduces sensitivity to weight initialization.
- Batch noise gives a mild regularizing effect.
Asked: [5 marks] (Jun 2025) Explain the following concepts in detail: Batch normalization.
dropout
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Dropout randomly switches off a fraction $p$ of neurons in each training step so the network cannot rely on any single neuron.
Key points.
- It reduces overfitting by acting like averaging many thinned networks.
- It is used only during training; at test time all neurons are active with weights scaled by $1-p$.
L1 regularization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. Regularization adds a penalty on the size of the weights to the loss to reduce overfitting; L1 (Lasso) penalises the sum of absolute weights.
Formula. L1: $J=L+\lambda\sum|\theta_i|$; L2: $J=L+\lambda\sum\theta_i^2$.
| Basis | L1 (Lasso) | L2 (Ridge) |
|---|---|---|
| Penalty | $\lambda\sum|\theta|$ | $\lambda\sum\theta^2$ |
| Effect on weights | Drives some exactly to zero | Shrinks all, none exactly zero |
| Model | Sparse | Dense, all features kept |
| Feature selection | Yes, automatic | No |
| Differentiability | Not at 0 | Smooth everywhere |
| Best when | Many irrelevant features | All features matter, stable |
Key points.
- $\lambda$ controls the strength: too large underfits, too small does not prevent overfitting.
- L1 gives sparse models and built-in feature selection, useful in high-dimensional data.
- L2 shrinks weights smoothly and is more stable when features are correlated.
<mark>L1 adds the absolute size of the weights as a penalty and produces sparse weights; L2 adds the squared size and only shrinks them.</mark>
Answer frame. Open with why regularization is needed; write both penalty formulas; give the comparison table; close with sparsity versus shrinkage. For the Jun 2026 combined question, add one short paragraph each on momentum, hyperparameter tuning and autoencoders from their sections.
Asked: [7 marks] (May 2024) Explain how L1 regularization (Lasso) and L2 regularization (Ridge) work to penalize the magnitude of model parameters differently. Asked: [7 marks] (Jun 2026) Compare L1 and L2 regularization. Discuss momentum, hyperparameter tuning and autoencoders.
L2 regularization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. L2 (Ridge, weight decay) adds $\lambda\sum\theta^2$ to the loss, penalising large weights.
Key points.
- The gradient adds $2\lambda\theta$, so each update shrinks every weight a little, which is why it is called weight decay.
- It keeps all features with smaller weights; larger $\lambda$ means a simpler model.
momentum
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Momentum is an optimizer that adds a fraction of the previous update to the current one, so the parameters keep moving in a consistent direction.
Formula. $v \leftarrow \beta v - \alpha\nabla J$, $\theta \leftarrow \theta + v$, with $\beta\approx0.9$.
Key points.
- It smooths noisy gradients and damps oscillation across steep directions.
- It speeds progress along shallow, consistent directions and helps cross flat regions and small local minima.
tuning hyper parameters
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. Hyperparameters are settings chosen before training, such as learning rate, batch size, number of layers and neurons, epochs, dropout rate and $\lambda$; parameters (weights and biases) are learned from data. Tuning is the search for the hyperparameter values that give the best validation performance.
Key points.
- Grid search tries every combination of a fixed set of values; it is thorough but cost grows exponentially with the number of hyperparameters.
- Random search samples combinations randomly and often finds good values faster than grid search.
- Bayesian optimization builds a model of past results to choose the next promising trial, so it needs fewer trials.
- Performance is sensitive to them: too high a learning rate diverges, too low a rate is slow, too many layers or neurons overfit, and too few underfit.
- Tuning must use a validation set (or cross-validation), never the test set.
- Challenges are heavy computation cost, a large search space and overfitting to the validation set.
<mark>Hyperparameters are set before training and control how well the model learns, so careful tuning moves it between underfitting and overfitting.</mark>
Answer frame. Open with hyperparameters versus parameters and examples; develop the three search methods; give an overfitting/underfitting example (deeper network overfits, shallow underfits); close with the challenges.
Asked: [7 marks] (May 2023, Dec 2024) "Model performance can be greatly improved through careful hyperparameter tuning". Justify the statement. Discuss the techniques and challenges of hyper parameter tuning.
Last-minute revision
- Sigmoid $1/(1+e^{-z})$, range (0,1), derivative $\sigma(1-\sigma)\le0.25$; ReLU $\max(0,z)$.
- Stacked linear layers collapse to one linear layer; non-linear activation avoids it.
- Gradient descent: $\theta\leftarrow\theta-\alpha\nabla J$; batch, stochastic, mini-batch.
- MSE for regression, cross-entropy for classification, hinge for SVM.
- Backprop: $\delta_o=(t-o)o(1-o)$, $\Delta w=\eta\delta x$; May 2023 gives $o=0.6807$, $\delta_o=0.0694$.
- Precision $TP/(TP+FP)$, recall $TP/(TP+FN)$, F1 $=2PR/(P+R)$.
- Autoencoder = encoder, bottleneck, decoder; L1 sparse, L2 shrinks.
- Xavier for sigmoid/tanh, He for ReLU; hyperparameter search: grid, random, Bayesian.
Memory hooks
- "S-shaped Sigmoid saturates": flat ends kill the gradient.
- "Lasso lassoes weights to zero, Ridge only rounds them down."
- "Batch, Stochastic, Mini": all, one, some samples per step.
- "Forward to predict, Backward to blame": backpropagation.
- "Xavier for S, He for ReLU."
Coverage checklist
- Linearity vs non linearity: May 2023.
- activation functions like sigmoid: Dec 2020, May 2024, Jun 2025.
- ReLU: covered (May 2024 comparison).
- weights and bias: Jun 2026.
- loss function: May 2023.
- gradient descent: May 2022, May 2024, Jun 2025.
- multilayer network: Dec 2020, Jun 2026.
- backpropagation: May 2022, May 2023, Dec 2024.
- weight initialization: covered.
- training: May 2022.
- testing: May 2023.
- unstable gradient problem: covered.
- auto encoders: Dec 2024, Jun 2025.
- batch normalization: Jun 2025.
- dropout: covered.
- L1 regularization: May 2024, Jun 2026.
- L2 regularization: covered.
- momentum: covered.
- tuning hyper parameters: May 2023, Dec 2024.