Skip to content
CS-702 (B) · Deep & Reinforcement Learning/Quick Revision Short Notes

Deep & Reinforcement Learning (CS-702 (B)) - Unit 2 Short Notes

How unit 2 is examined

This unit covers autoencoders and their regularized variants, then general regularization and normalization; the marks sit in sparse and contractive autoencoders (7 marks) and in regularization with dropout and batch normalization (14 marks).

Autoencoders and relation to PCA

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>An autoencoder is a neural network trained to reproduce its own input, passing it through a narrow hidden layer so that the hidden code is a compressed representation.</mark>

Diagram.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-01" viewBox="0 0 424 80" width="424" height="80" role="img" aria-label="Autoencoder. X input, H code (bottleneck), R reconstruction of X"><style>#dsfig-u2-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-01 .t{fill:#16181D;font-weight:500}#dsfig-u2-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-01 .dot{fill:#16181D}#dsfig-u2-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-01 .ah{fill:#454C5A}#dsfig-u2-01 .ah.hi{fill:#2340B8}#dsfig-u2-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-01 .e{stroke:#B1B7C3}html.dark #dsfig-u2-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-01 .t{fill:#E6E8ED}html.dark #dsfig-u2-01 .t.inv{fill:#0F1115}html.dark #dsfig-u2-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-01 .dot{fill:#E6E8ED}html.dark #dsfig-u2-01 .ann{fill:#8FA3FF}html.dark #dsfig-u2-01 .lbl{fill:#858D9C}html.dark #dsfig-u2-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-01 .ah{fill:#B1B7C3}html.dark #dsfig-u2-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L191,40" marker-end="url(#ah4)"/><path class="e" d="M231,40 L363,40" marker-end="url(#ah4)"/><g class="wl"><rect x="95.3" y="31" width="61.5" height="18" rx="9"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">encoder</text></g><g class="wl"><rect x="267.3" y="31" width="61.5" height="18" rx="9"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">decoder</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">X</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">H</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">R</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Autoencoder. X input, H code (bottleneck), R reconstruction of X</figcaption></figure>

Key points.

  1. The encoder maps input $x$ to code $h=f(Wx+b)$ and the decoder maps it back to $\hat{x}=g(W'h+b')$.
  2. Training minimises the reconstruction loss $L=\lVert x-\hat{x}\rVert^2$ with no labels, so it is unsupervised.
  3. A linear autoencoder with one hidden layer and squared error learns the same subspace as PCA, spanning the top-$k$ principal components.
  4. A nonlinear autoencoder is more powerful than PCA because it can capture curved structure.

Regularization in autoencoders

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Regularization in an autoencoder adds a penalty to the reconstruction loss, $L(x,\hat{x})+\Omega(h)$, so the network learns useful features instead of copying the input.</mark>

Key points.

  1. An overcomplete autoencoder (code larger than input) can learn the identity function and learn nothing useful.
  2. The penalty $\Omega$ restricts the code or the weights, forcing it to keep only the important structure of the data.
  3. Sparse, denoising and contractive autoencoders differ only in the choice of this penalty or corruption.
  4. Weight decay $\lambda\lVert W\rVert^2$ is the simplest such penalty.

Denoising autoencoders

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A denoising autoencoder is fed a corrupted input $\tilde{x}$ and trained to reconstruct the clean input $x$.</mark>

Key points.

  1. The loss is $L=\lVert x-g(f(\tilde{x}))\rVert^2$, where $\tilde{x}$ is $x$ with Gaussian noise added or some inputs masked to zero.
  2. Because the target is the clean input, the network cannot just copy and must learn the structure of the data.
  3. It learns features that are robust to noise and works even when the code is overcomplete.
  4. It can be used to clean noisy images.

Sparse autoencoders

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>A sparse autoencoder adds a sparsity penalty on the hidden activations, so only a few hidden units are active for any input.</mark>

Diagram.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-02" viewBox="0 0 424 80" width="424" height="80" role="img" aria-label="Autoencoder with penalty on H. Sparse AE penalises activations of H; contractive AE penalises the Jacobian of H with respect to X"><style>#dsfig-u2-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-02 .t{fill:#16181D;font-weight:500}#dsfig-u2-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-02 .dot{fill:#16181D}#dsfig-u2-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-02 .ah{fill:#454C5A}#dsfig-u2-02 .ah.hi{fill:#2340B8}#dsfig-u2-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-02 .e{stroke:#B1B7C3}html.dark #dsfig-u2-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-02 .t{fill:#E6E8ED}html.dark #dsfig-u2-02 .t.inv{fill:#0F1115}html.dark #dsfig-u2-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-02 .dot{fill:#E6E8ED}html.dark #dsfig-u2-02 .ann{fill:#8FA3FF}html.dark #dsfig-u2-02 .lbl{fill:#858D9C}html.dark #dsfig-u2-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-02 .ah{fill:#B1B7C3}html.dark #dsfig-u2-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L191,40" marker-end="url(#ah5)"/><path class="e" d="M231,40 L363,40" marker-end="url(#ah5)"/><g class="wl"><rect x="98.9" y="31" width="54.3" height="18" rx="9"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">encode</text></g><g class="wl"><rect x="270.9" y="31" width="54.3" height="18" rx="9"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">decode</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">X</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">H</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">R</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Autoencoder with penalty on H. Sparse AE penalises activations of H; contractive AE penalises the Jacobian of H with respect to X</figcaption></figure>

Key points.

  1. The reconstruction objective is $L(x,\hat{x})=\lVert x-\hat{x}\rVert^2$, and both types add a regularizer to it.
  2. The sparse loss is $L+\beta\sum_j \mathrm{KL}(\rho\,\Vert\,\hat{\rho}_j)$, where $\rho$ is a small target average activation such as 0.05 and $\hat{\rho}_j$ is the actual mean activation of hidden unit $j$.
  3. The KL penalty is $\rho\log\frac{\rho}{\hat{\rho}_j}+(1-\rho)\log\frac{1-\rho}{1-\hat{\rho}_j}$; it grows as $\hat{\rho}_j$ moves away from $\rho$, so most units stay near zero. An L1 penalty $\lambda\sum|h_j|$ also works.
  4. Each hidden unit becomes a specialised feature detector, so sparse codes can be larger than the input and still learn features.
  5. A contractive autoencoder adds $\lambda\lVert J_f(x)\rVert_F^2$, the squared Frobenius norm of the Jacobian of the code with respect to the input, $J=\partial h/\partial x$.
  6. This makes the code change very little for small changes in input, so features are locally invariant and robust.
Point Sparse AE Contractive AE
Penalty KL divergence or L1 on activations $h$ Frobenius norm of Jacobian $\partial h/\partial x$
Effect Few hidden units fire Code insensitive to small input changes
Goal Feature learning, specialised units Robustness, local invariance
Code size Can be overcomplete Can be overcomplete
Hyperparameter $\beta$, target $\rho$ $\lambda$

Answer frame. Open by defining an autoencoder and its reconstruction loss and why regularization is needed; draw the encoder-code-decoder figure; develop points 2-4 for sparse, then 5-6 for contractive; end with the comparison table and the closing line that sparse learns selective features while contractive learns robust ones.

Asked: [7 marks] (Dec 2020, Nov 2023) Explain sparse and contractive auto encoders.

Contractive autoencoders

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A contractive autoencoder penalises the Frobenius norm of the Jacobian of the encoder, $\lambda\sum_{ij}(\partial h_j/\partial x_i)^2$, to make the code robust to small input changes.</mark>

Key points.

  1. The total loss is reconstruction error plus $\lambda\lVert J_f(x)\rVert_F^2$.
  2. Reconstruction keeps information while the penalty contracts the space around training points.
  3. It learns features that vary only along the directions of the data manifold.
  4. It is covered with the sparse autoencoder above.

Regularization: Bias Variance Tradeoff

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Regularization is any technique that reduces a model's generalization error, the error on unseen data, without a matching increase in training error.</mark>

Key points.

  1. Bias is the error from wrong assumptions, so high bias means underfitting and high error on both training and test data.
  2. Variance is the sensitivity to the particular training set, so high variance means overfitting: low training error but high test error.
  3. Expected error is $\text{Bias}^2+\text{Variance}+\text{Noise}$; making a model more complex lowers bias but raises variance, and the best model sits at the balance point.
  4. Deep networks have huge capacity and overfit small datasets, so regularization is needed to cut variance at a small cost in bias.
  5. Dropout randomly sets each hidden unit to zero with probability $p$ during training, so no unit can depend on specific others (co-adaptation) and the network behaves like an ensemble of many thinned sub-networks; at test time all units are used with weights scaled by $1-p$.
  6. Batch normalization normalises each layer's inputs over the mini-batch as $\hat{x}=\frac{x-\mu_B}{\sqrt{\sigma_B^2+\epsilon}}$, then applies learnable scale and shift $y=\gamma\hat{x}+\beta$, which reduces internal covariate shift, the change in the distribution of layer inputs during training.
  7. Batch normalization allows higher learning rates, gives faster convergence, reduces sensitivity to initialization, and its batch noise gives a mild regularizing effect.
  8. Together they improve generalization (dropout), training stability and speed (batch norm), and final accuracy; other regularizers are L2, early stopping, augmentation and noise.
Model Bias Variance Symptom
Too simple High Low Underfitting
Balanced Low Low Good generalization
Too complex Low High Overfitting

Answer frame. Open by defining overfitting and why regularization is needed; sketch training and validation error against epochs or complexity; develop points 1-4 first, then dropout (5) and batch normalization (6-7) with their formulas; close with points 8, that regularization lowers variance and improves generalization, convergence and performance.

Pitfall: Do not say dropout is used at test time; at test time all units stay on with scaled weights.

Asked: [14 marks] (Jun 2025) Explain the role of regularization techniques in deep learning. How do dropout and batch normalization contribute to improved model performance?

L2 regularization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==L2 regularization adds the squared weight norm to the loss: $\tilde{L}=L+\frac{\lambda}{2}\lVert w\rVert_2^2$.==

Key points.

  1. The gradient step becomes $w\leftarrow(1-\eta\lambda)w-\eta\nabla L$, so weights shrink each step; hence the name weight decay.
  2. It keeps weights small, giving smoother functions and lower variance.
  3. It is equivalent to ridge regression and a Gaussian prior on the weights.
  4. A larger $\lambda$ means stronger regularization and more bias.

Early stopping

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Early stopping halts training when the validation error starts to rise, and returns the parameters from the best validation epoch.</mark>

Key points.

  1. Training error keeps falling, but validation error falls and then rises as the model starts to overfit.
  2. Keep a copy of the best weights and stop after a patience of several epochs with no improvement.
  3. It is cheap and needs no change to the loss, and acts much like L2 by limiting how far weights grow.
  4. It needs a held-out validation set.

Dataset augmentation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Dataset augmentation creates extra training examples by applying label-preserving transformations to existing ones.</mark>

Key points.

  1. For images use flips, rotations, crops, scaling and colour changes.
  2. More varied data reduces overfitting and makes the model invariant to these changes.
  3. The transformation must not change the label, for example a flip changes 6 into 9 only if rotated by 180 degrees.
  4. Noise added to input or hidden units is also a form of augmentation.

Parameter sharing and tying

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Parameter sharing forces different parts of a model to use the same weights, and parameter tying forces two models' parameters to be close or equal.</mark>

Key points.

  1. A convolutional layer applies the same filter at every position, so it has far fewer parameters.
  2. Fewer free parameters means lower capacity and less overfitting.
  3. Tying is done by a penalty such as $\lVert w^{(A)}-w^{(B)}\rVert^2$ between two models that solve similar tasks.
  4. In an autoencoder the decoder weights are often tied as $W'=W^{T}$.

Injecting noise at input

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Noise injection adds random noise to the inputs, and sometimes to weights or hidden units, during training as a regularizer.</mark>

Key points.

  1. Training with $x+\epsilon$, where $\epsilon\sim N(0,\sigma^2)$, makes the network robust to small input changes.
  2. For small noise it is equivalent to a penalty on the weights, similar to L2.
  3. It is the basis of the denoising autoencoder.
  4. Noise is applied only during training, not at test time.

Ensemble methods

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>An ensemble combines the predictions of several models, by averaging or voting, so that their individual errors cancel.</mark>

Key points.

  1. Bagging trains each model on a bootstrap sample of the data and averages them, which cuts variance.
  2. Boosting trains models in sequence, each focusing on previous mistakes, which cuts bias.
  3. If $k$ models have independent errors of variance $v$, the averaged error variance is $v/k$.
  4. Ensembles of deep networks are costly to train, which motivates dropout as a cheap approximation.

Dropout

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Dropout randomly sets each hidden unit to zero with probability $p$ at every training step, and uses the full network with scaled weights at test time.</mark>

Key points.

  1. It prevents co-adaptation because a unit cannot rely on any particular other unit being present.
  2. It trains an exponential number of thinned sub-networks that share weights, giving an ensemble effect.
  3. At test time all units are on and weights are multiplied by the keep probability $1-p$ (or inverted dropout scales by $1/(1-p)$ in training).
  4. Typical rates are 0.5 for hidden layers and 0.2 for inputs.

Batch Normalization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Batch normalization normalises each unit's activations to zero mean and unit variance over the mini-batch, then rescales them with learnable $\gamma$ and $\beta$.</mark>

Formula.

$$\mu_B=\frac1m\sum_i x_i,\quad \sigma_B^2=\frac1m\sum_i(x_i-\mu_B)^2,\quad \hat{x}_i=\frac{x_i-\mu_B}{\sqrt{\sigma_B^2+\epsilon}},\quad y_i=\gamma\hat{x}_i+\beta$$

Key points.

  1. It reduces internal covariate shift, so each layer sees a more stable input distribution.
  2. It allows larger learning rates and faster training.
  3. At test time it uses running averages of $\mu$ and $\sigma^2$ collected during training.
  4. It works poorly with very small batches, since the batch statistics become noisy.

Instance Normalization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Instance normalization normalises each channel of each single sample over its own spatial dimensions, independent of the batch.</mark>

Key points.

  1. Mean and variance are computed per sample and per channel over height and width.
  2. It does not depend on batch size and behaves the same at training and test time.
  3. It removes instance-specific contrast, so it is used in style transfer.
  4. It is a special case of group normalization with one channel per group.

Group Normalization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Group normalization divides the channels into groups and normalises within each group for each sample, independent of the batch.</mark>

Key points.

  1. Statistics are computed per sample over the channels in a group and over height and width.
  2. It works well with small batches, where batch normalization fails.
  3. With one group it becomes layer normalization, and with one channel per group it becomes instance normalization.
  4. A typical choice is 32 groups.

Last-minute revision

  • Autoencoder: encoder $h=f(x)$, decoder $\hat{x}=g(h)$, loss $\lVert x-\hat{x}\rVert^2$; linear AE equals PCA subspace.
  • Denoising AE: corrupted input, clean target.
  • Sparse AE: penalty $\beta\sum \mathrm{KL}(\rho\Vert\hat{\rho}_j)$ or L1 on activations.
  • Contractive AE: penalty $\lambda\lVert\partial h/\partial x\rVert_F^2$.
  • Expected error = Bias$^2$ + Variance + Noise; underfit is high bias, overfit is high variance.
  • L2: $L+\frac{\lambda}{2}\lVert w\rVert^2$, weight decay $w\leftarrow(1-\eta\lambda)w$.
  • Early stopping: stop at minimum validation error, keep best weights.
  • Dropout: zero units with probability $p$ in training, scale at test time; ensemble effect.
  • Batch norm: $\hat{x}=(x-\mu_B)/\sqrt{\sigma_B^2+\epsilon}$, then $\gamma\hat{x}+\beta$.
  • Batch norm over batch; instance norm per sample per channel; group norm per sample per channel group.

Memory hooks

  • Sparse = Selective units; Contractive = Calm code (Jacobian small).
  • Bias is Blind (too simple); Variance is Volatile (too complex).
  • Dropout = many thin nets in one; test time use them all.
  • BN, IN, GN: normalise over Batch, Instance, Group.
  • Denoising: dirty in, clean out.

Coverage checklist

  • Autoencoders and relation to PCA: definition, PCA link.
  • Regularization in autoencoders: penalty on the reconstruction loss.
  • Denoising autoencoders: corrupted input, clean target.
  • Sparse autoencoders: covers Q2 (7 marks, Dec 2020, Nov 2023) with contractive comparison.
  • Contractive autoencoders: Jacobian penalty, Q2.
  • Regularization: Bias Variance Tradeoff: covers Q1 (14 marks, Jun 2025).
  • L2 regularization: formula and weight decay.
  • Early stopping: validation rule.
  • Dataset augmentation: label-preserving transforms.
  • Parameter sharing and tying: shared weights, tied penalty.
  • Injecting noise at input: noise as regularizer.
  • Ensemble methods: bagging and boosting.
  • Dropout: mechanism, ensemble effect, Q1.
  • Batch Normalization: steps and covariate shift, Q1.
  • Instance Normalization: per sample per channel.
  • Group Normalization: per channel group.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in