How unit 2 is examined
This unit covers autoencoders and their regularized variants, then general regularization and normalization; the marks sit in sparse and contractive autoencoders (7 marks) and in regularization with dropout and batch normalization (14 marks).
Autoencoders and relation to PCA
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>An autoencoder is a neural network trained to reproduce its own input, passing it through a narrow hidden layer so that the hidden code is a compressed representation.</mark>
Diagram.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-01" viewBox="0 0 424 80" width="424" height="80" role="img" aria-label="Autoencoder. X input, H code (bottleneck), R reconstruction of X"><style>#dsfig-u2-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-01 .t{fill:#16181D;font-weight:500}#dsfig-u2-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-01 .dot{fill:#16181D}#dsfig-u2-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-01 .ah{fill:#454C5A}#dsfig-u2-01 .ah.hi{fill:#2340B8}#dsfig-u2-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-01 .e{stroke:#B1B7C3}html.dark #dsfig-u2-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-01 .t{fill:#E6E8ED}html.dark #dsfig-u2-01 .t.inv{fill:#0F1115}html.dark #dsfig-u2-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-01 .dot{fill:#E6E8ED}html.dark #dsfig-u2-01 .ann{fill:#8FA3FF}html.dark #dsfig-u2-01 .lbl{fill:#858D9C}html.dark #dsfig-u2-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-01 .ah{fill:#B1B7C3}html.dark #dsfig-u2-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L191,40" marker-end="url(#ah4)"/><path class="e" d="M231,40 L363,40" marker-end="url(#ah4)"/><g class="wl"><rect x="95.3" y="31" width="61.5" height="18" rx="9"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">encoder</text></g><g class="wl"><rect x="267.3" y="31" width="61.5" height="18" rx="9"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">decoder</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">X</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">H</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">R</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Autoencoder. X input, H code (bottleneck), R reconstruction of X</figcaption></figure>
Key points.
- The encoder maps input $x$ to code $h=f(Wx+b)$ and the decoder maps it back to $\hat{x}=g(W'h+b')$.
- Training minimises the reconstruction loss $L=\lVert x-\hat{x}\rVert^2$ with no labels, so it is unsupervised.
- A linear autoencoder with one hidden layer and squared error learns the same subspace as PCA, spanning the top-$k$ principal components.
- A nonlinear autoencoder is more powerful than PCA because it can capture curved structure.
Regularization in autoencoders
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Regularization in an autoencoder adds a penalty to the reconstruction loss, $L(x,\hat{x})+\Omega(h)$, so the network learns useful features instead of copying the input.</mark>
Key points.
- An overcomplete autoencoder (code larger than input) can learn the identity function and learn nothing useful.
- The penalty $\Omega$ restricts the code or the weights, forcing it to keep only the important structure of the data.
- Sparse, denoising and contractive autoencoders differ only in the choice of this penalty or corruption.
- Weight decay $\lambda\lVert W\rVert^2$ is the simplest such penalty.
Denoising autoencoders
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A denoising autoencoder is fed a corrupted input $\tilde{x}$ and trained to reconstruct the clean input $x$.</mark>
Key points.
- The loss is $L=\lVert x-g(f(\tilde{x}))\rVert^2$, where $\tilde{x}$ is $x$ with Gaussian noise added or some inputs masked to zero.
- Because the target is the clean input, the network cannot just copy and must learn the structure of the data.
- It learns features that are robust to noise and works even when the code is overcomplete.
- It can be used to clean noisy images.
Sparse autoencoders
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>A sparse autoencoder adds a sparsity penalty on the hidden activations, so only a few hidden units are active for any input.</mark>
Diagram.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-02" viewBox="0 0 424 80" width="424" height="80" role="img" aria-label="Autoencoder with penalty on H. Sparse AE penalises activations of H; contractive AE penalises the Jacobian of H with respect to X"><style>#dsfig-u2-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-02 .t{fill:#16181D;font-weight:500}#dsfig-u2-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-02 .dot{fill:#16181D}#dsfig-u2-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-02 .ah{fill:#454C5A}#dsfig-u2-02 .ah.hi{fill:#2340B8}#dsfig-u2-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-02 .e{stroke:#B1B7C3}html.dark #dsfig-u2-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-02 .t{fill:#E6E8ED}html.dark #dsfig-u2-02 .t.inv{fill:#0F1115}html.dark #dsfig-u2-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-02 .dot{fill:#E6E8ED}html.dark #dsfig-u2-02 .ann{fill:#8FA3FF}html.dark #dsfig-u2-02 .lbl{fill:#858D9C}html.dark #dsfig-u2-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-02 .ah{fill:#B1B7C3}html.dark #dsfig-u2-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L191,40" marker-end="url(#ah5)"/><path class="e" d="M231,40 L363,40" marker-end="url(#ah5)"/><g class="wl"><rect x="98.9" y="31" width="54.3" height="18" rx="9"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">encode</text></g><g class="wl"><rect x="270.9" y="31" width="54.3" height="18" rx="9"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">decode</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">X</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">H</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">R</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Autoencoder with penalty on H. Sparse AE penalises activations of H; contractive AE penalises the Jacobian of H with respect to X</figcaption></figure>
Key points.
- The reconstruction objective is $L(x,\hat{x})=\lVert x-\hat{x}\rVert^2$, and both types add a regularizer to it.
- The sparse loss is $L+\beta\sum_j \mathrm{KL}(\rho\,\Vert\,\hat{\rho}_j)$, where $\rho$ is a small target average activation such as 0.05 and $\hat{\rho}_j$ is the actual mean activation of hidden unit $j$.
- The KL penalty is $\rho\log\frac{\rho}{\hat{\rho}_j}+(1-\rho)\log\frac{1-\rho}{1-\hat{\rho}_j}$; it grows as $\hat{\rho}_j$ moves away from $\rho$, so most units stay near zero. An L1 penalty $\lambda\sum|h_j|$ also works.
- Each hidden unit becomes a specialised feature detector, so sparse codes can be larger than the input and still learn features.
- A contractive autoencoder adds $\lambda\lVert J_f(x)\rVert_F^2$, the squared Frobenius norm of the Jacobian of the code with respect to the input, $J=\partial h/\partial x$.
- This makes the code change very little for small changes in input, so features are locally invariant and robust.
| Point | Sparse AE | Contractive AE |
|---|---|---|
| Penalty | KL divergence or L1 on activations $h$ | Frobenius norm of Jacobian $\partial h/\partial x$ |
| Effect | Few hidden units fire | Code insensitive to small input changes |
| Goal | Feature learning, specialised units | Robustness, local invariance |
| Code size | Can be overcomplete | Can be overcomplete |
| Hyperparameter | $\beta$, target $\rho$ | $\lambda$ |
Answer frame. Open by defining an autoencoder and its reconstruction loss and why regularization is needed; draw the encoder-code-decoder figure; develop points 2-4 for sparse, then 5-6 for contractive; end with the comparison table and the closing line that sparse learns selective features while contractive learns robust ones.
Asked: [7 marks] (Dec 2020, Nov 2023) Explain sparse and contractive auto encoders.
Contractive autoencoders
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A contractive autoencoder penalises the Frobenius norm of the Jacobian of the encoder, $\lambda\sum_{ij}(\partial h_j/\partial x_i)^2$, to make the code robust to small input changes.</mark>
Key points.
- The total loss is reconstruction error plus $\lambda\lVert J_f(x)\rVert_F^2$.
- Reconstruction keeps information while the penalty contracts the space around training points.
- It learns features that vary only along the directions of the data manifold.
- It is covered with the sparse autoencoder above.
Regularization: Bias Variance Tradeoff
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Regularization is any technique that reduces a model's generalization error, the error on unseen data, without a matching increase in training error.</mark>
Key points.
- Bias is the error from wrong assumptions, so high bias means underfitting and high error on both training and test data.
- Variance is the sensitivity to the particular training set, so high variance means overfitting: low training error but high test error.
- Expected error is $\text{Bias}^2+\text{Variance}+\text{Noise}$; making a model more complex lowers bias but raises variance, and the best model sits at the balance point.
- Deep networks have huge capacity and overfit small datasets, so regularization is needed to cut variance at a small cost in bias.
- Dropout randomly sets each hidden unit to zero with probability $p$ during training, so no unit can depend on specific others (co-adaptation) and the network behaves like an ensemble of many thinned sub-networks; at test time all units are used with weights scaled by $1-p$.
- Batch normalization normalises each layer's inputs over the mini-batch as $\hat{x}=\frac{x-\mu_B}{\sqrt{\sigma_B^2+\epsilon}}$, then applies learnable scale and shift $y=\gamma\hat{x}+\beta$, which reduces internal covariate shift, the change in the distribution of layer inputs during training.
- Batch normalization allows higher learning rates, gives faster convergence, reduces sensitivity to initialization, and its batch noise gives a mild regularizing effect.
- Together they improve generalization (dropout), training stability and speed (batch norm), and final accuracy; other regularizers are L2, early stopping, augmentation and noise.
| Model | Bias | Variance | Symptom |
|---|---|---|---|
| Too simple | High | Low | Underfitting |
| Balanced | Low | Low | Good generalization |
| Too complex | Low | High | Overfitting |
Answer frame. Open by defining overfitting and why regularization is needed; sketch training and validation error against epochs or complexity; develop points 1-4 first, then dropout (5) and batch normalization (6-7) with their formulas; close with points 8, that regularization lowers variance and improves generalization, convergence and performance.
Pitfall: Do not say dropout is used at test time; at test time all units stay on with scaled weights.
Asked: [14 marks] (Jun 2025) Explain the role of regularization techniques in deep learning. How do dropout and batch normalization contribute to improved model performance?
L2 regularization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==L2 regularization adds the squared weight norm to the loss: $\tilde{L}=L+\frac{\lambda}{2}\lVert w\rVert_2^2$.==
Key points.
- The gradient step becomes $w\leftarrow(1-\eta\lambda)w-\eta\nabla L$, so weights shrink each step; hence the name weight decay.
- It keeps weights small, giving smoother functions and lower variance.
- It is equivalent to ridge regression and a Gaussian prior on the weights.
- A larger $\lambda$ means stronger regularization and more bias.
Early stopping
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Early stopping halts training when the validation error starts to rise, and returns the parameters from the best validation epoch.</mark>
Key points.
- Training error keeps falling, but validation error falls and then rises as the model starts to overfit.
- Keep a copy of the best weights and stop after a patience of several epochs with no improvement.
- It is cheap and needs no change to the loss, and acts much like L2 by limiting how far weights grow.
- It needs a held-out validation set.
Dataset augmentation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Dataset augmentation creates extra training examples by applying label-preserving transformations to existing ones.</mark>
Key points.
- For images use flips, rotations, crops, scaling and colour changes.
- More varied data reduces overfitting and makes the model invariant to these changes.
- The transformation must not change the label, for example a flip changes 6 into 9 only if rotated by 180 degrees.
- Noise added to input or hidden units is also a form of augmentation.
Parameter sharing and tying
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Parameter sharing forces different parts of a model to use the same weights, and parameter tying forces two models' parameters to be close or equal.</mark>
Key points.
- A convolutional layer applies the same filter at every position, so it has far fewer parameters.
- Fewer free parameters means lower capacity and less overfitting.
- Tying is done by a penalty such as $\lVert w^{(A)}-w^{(B)}\rVert^2$ between two models that solve similar tasks.
- In an autoencoder the decoder weights are often tied as $W'=W^{T}$.
Injecting noise at input
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Noise injection adds random noise to the inputs, and sometimes to weights or hidden units, during training as a regularizer.</mark>
Key points.
- Training with $x+\epsilon$, where $\epsilon\sim N(0,\sigma^2)$, makes the network robust to small input changes.
- For small noise it is equivalent to a penalty on the weights, similar to L2.
- It is the basis of the denoising autoencoder.
- Noise is applied only during training, not at test time.
Ensemble methods
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>An ensemble combines the predictions of several models, by averaging or voting, so that their individual errors cancel.</mark>
Key points.
- Bagging trains each model on a bootstrap sample of the data and averages them, which cuts variance.
- Boosting trains models in sequence, each focusing on previous mistakes, which cuts bias.
- If $k$ models have independent errors of variance $v$, the averaged error variance is $v/k$.
- Ensembles of deep networks are costly to train, which motivates dropout as a cheap approximation.
Dropout
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Dropout randomly sets each hidden unit to zero with probability $p$ at every training step, and uses the full network with scaled weights at test time.</mark>
Key points.
- It prevents co-adaptation because a unit cannot rely on any particular other unit being present.
- It trains an exponential number of thinned sub-networks that share weights, giving an ensemble effect.
- At test time all units are on and weights are multiplied by the keep probability $1-p$ (or inverted dropout scales by $1/(1-p)$ in training).
- Typical rates are 0.5 for hidden layers and 0.2 for inputs.
Batch Normalization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Batch normalization normalises each unit's activations to zero mean and unit variance over the mini-batch, then rescales them with learnable $\gamma$ and $\beta$.</mark>
Formula.
$$\mu_B=\frac1m\sum_i x_i,\quad \sigma_B^2=\frac1m\sum_i(x_i-\mu_B)^2,\quad \hat{x}_i=\frac{x_i-\mu_B}{\sqrt{\sigma_B^2+\epsilon}},\quad y_i=\gamma\hat{x}_i+\beta$$
Key points.
- It reduces internal covariate shift, so each layer sees a more stable input distribution.
- It allows larger learning rates and faster training.
- At test time it uses running averages of $\mu$ and $\sigma^2$ collected during training.
- It works poorly with very small batches, since the batch statistics become noisy.
Instance Normalization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Instance normalization normalises each channel of each single sample over its own spatial dimensions, independent of the batch.</mark>
Key points.
- Mean and variance are computed per sample and per channel over height and width.
- It does not depend on batch size and behaves the same at training and test time.
- It removes instance-specific contrast, so it is used in style transfer.
- It is a special case of group normalization with one channel per group.
Group Normalization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Group normalization divides the channels into groups and normalises within each group for each sample, independent of the batch.</mark>
Key points.
- Statistics are computed per sample over the channels in a group and over height and width.
- It works well with small batches, where batch normalization fails.
- With one group it becomes layer normalization, and with one channel per group it becomes instance normalization.
- A typical choice is 32 groups.
Last-minute revision
- Autoencoder: encoder $h=f(x)$, decoder $\hat{x}=g(h)$, loss $\lVert x-\hat{x}\rVert^2$; linear AE equals PCA subspace.
- Denoising AE: corrupted input, clean target.
- Sparse AE: penalty $\beta\sum \mathrm{KL}(\rho\Vert\hat{\rho}_j)$ or L1 on activations.
- Contractive AE: penalty $\lambda\lVert\partial h/\partial x\rVert_F^2$.
- Expected error = Bias$^2$ + Variance + Noise; underfit is high bias, overfit is high variance.
- L2: $L+\frac{\lambda}{2}\lVert w\rVert^2$, weight decay $w\leftarrow(1-\eta\lambda)w$.
- Early stopping: stop at minimum validation error, keep best weights.
- Dropout: zero units with probability $p$ in training, scale at test time; ensemble effect.
- Batch norm: $\hat{x}=(x-\mu_B)/\sqrt{\sigma_B^2+\epsilon}$, then $\gamma\hat{x}+\beta$.
- Batch norm over batch; instance norm per sample per channel; group norm per sample per channel group.
Memory hooks
- Sparse = Selective units; Contractive = Calm code (Jacobian small).
- Bias is Blind (too simple); Variance is Volatile (too complex).
- Dropout = many thin nets in one; test time use them all.
- BN, IN, GN: normalise over Batch, Instance, Group.
- Denoising: dirty in, clean out.
Coverage checklist
- Autoencoders and relation to PCA: definition, PCA link.
- Regularization in autoencoders: penalty on the reconstruction loss.
- Denoising autoencoders: corrupted input, clean target.
- Sparse autoencoders: covers Q2 (7 marks, Dec 2020, Nov 2023) with contractive comparison.
- Contractive autoencoders: Jacobian penalty, Q2.
- Regularization: Bias Variance Tradeoff: covers Q1 (14 marks, Jun 2025).
- L2 regularization: formula and weight decay.
- Early stopping: validation rule.
- Dataset augmentation: label-preserving transforms.
- Parameter sharing and tying: shared weights, tied penalty.
- Injecting noise at input: noise as regularizer.
- Ensemble methods: bagging and boosting.
- Dropout: mechanism, ensemble effect, Q1.
- Batch Normalization: steps and covariate shift, Q1.
- Instance Normalization: per sample per channel.
- Group Normalization: per channel group.