Skip to content
AD-601 · Deep Learning/Quick Revision Short Notes

Deep Learning (AD-601) - Unit 2 Short Notes

How unit 2 is examined

This unit covers the multilayer perceptron, gradient descent and backpropagation (with weight initialization), empirical risk minimization, regularization and autoencoders; the marks sit in MLP and gradient descent.

Multilayer Perceptron

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>A multilayer perceptron (MLP) is a feedforward neural network with an input layer, one or more hidden layers and an output layer, where every neuron is fully connected to the next layer and hidden neurons use a non-linear activation function.</mark>

Diagram.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-01" viewBox="0 0 424 338" width="424" height="338" role="img" aria-label="MLP: input x1, x2; one hidden layer h1-h3; output y (signals flow left to right)"><style>#dsfig-u2-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-01 .t{fill:#16181D;font-weight:500}#dsfig-u2-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-01 .dot{fill:#16181D}#dsfig-u2-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-01 .ah{fill:#454C5A}#dsfig-u2-01 .ah.hi{fill:#2340B8}#dsfig-u2-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-01 .e{stroke:#B1B7C3}html.dark #dsfig-u2-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-01 .t{fill:#E6E8ED}html.dark #dsfig-u2-01 .t.inv{fill:#0F1115}html.dark #dsfig-u2-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-01 .dot{fill:#E6E8ED}html.dark #dsfig-u2-01 .ann{fill:#8FA3FF}html.dark #dsfig-u2-01 .lbl{fill:#858D9C}html.dark #dsfig-u2-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-01 .ah{fill:#B1B7C3}html.dark #dsfig-u2-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah3" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh3" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M58.4,78.4 L191.6,45.1" marker-end="url(#ah3)"/><path class="e" d="M57,91.5 L193.2,159.6" marker-end="url(#ah3)"/><path class="e" d="M51.9,97.8 L198.9,281.6" marker-end="url(#ah3)"/><path class="e" d="M51.9,240.2 L198.9,56.4" marker-end="url(#ah3)"/><path class="e" d="M57,246.5 L193.2,178.4" marker-end="url(#ah3)"/><path class="e" d="M58.4,259.6 L191.6,292.9" marker-end="url(#ah3)"/><path class="e" d="M227.2,51.4 L367.2,156.4" marker-end="url(#ah3)"/><path class="e" d="M231,169 L363,169" marker-end="url(#ah3)"/><path class="e" d="M227.2,286.6 L367.2,181.6" marker-end="url(#ah3)"/><circle class="n" cx="40" cy="83" r="18"/><text class="t" x="40" y="83" dy=".35em" text-anchor="middle">x1</text><circle class="n" cx="40" cy="255" r="18"/><text class="t" x="40" y="255" dy=".35em" text-anchor="middle">x2</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">h1</text><circle class="n" cx="212" cy="169" r="18"/><text class="t" x="212" y="169" dy=".35em" text-anchor="middle">h2</text><circle class="n" cx="212" cy="298" r="18"/><text class="t" x="212" y="298" dy=".35em" text-anchor="middle">h3</text><circle class="n" cx="384" cy="169" r="18"/><text class="t" x="384" y="169" dy=".35em" text-anchor="middle">y</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">MLP: input x1, x2; one hidden layer h1-h3; output y (signals flow left to right)</figcaption></figure>

Key points.

  1. A linear (single-layer) perceptron computes $y=f(\mathbf{w}\cdot\mathbf{x}+b)$ with a step function, so its decision boundary is a single straight line (hyperplane).
  2. It can therefore only classify linearly separable data; it fails on XOR because no single line separates the two classes.
  3. An MLP adds hidden layers, so each hidden neuron draws one line and the output neuron combines these lines into a non-linear decision region.
  4. Hidden neurons use non-linear activations such as sigmoid, tanh or ReLU; without them the whole network collapses into one linear map.
  5. Training uses forward propagation to get the output and backpropagation with gradient descent to update the weights.
  6. Depth means the number of hidden layers: early layers learn simple features (edges), later layers combine them into complex ones (shapes, objects), which is hierarchical feature learning.
  7. More depth increases representational capacity and usually accuracy, but it brings vanishing or exploding gradients, overfitting and higher computation, which are countered by ReLU, good initialization, regularization and more data.
  8. Applications include classification, regression, pattern recognition and speech and image tasks.

Example. XOR with step units: $h_1=\text{step}(x_1+x_2-0.5)$ (OR), $h_2=\text{step}(x_1+x_2-1.5)$ (AND), $y=\text{step}(h_1-h_2-0.5)$. This gives 0,1,1,0 for inputs 00,01,10,11.

Point Linear perceptron MLP
Layers Input and output only Input, hidden, output
Boundary Straight line Non-linear region
Solves XOR No Yes
Training Perceptron rule Backpropagation

Answer frame. Open with the definition; draw the MLP figure and, beside it, the single perceptron; develop points 1-5 with the XOR example and the table; close with applications. For the depth question, define depth, then develop points 6-7 and end with the trade-offs.

Asked: [7 marks] (May 2024, Jun 2025) Write about Linear and Multilayer Perceptron in detail. What is an MLP? How does it overcome the limitations of single-layer perceptrons? Asked: [7 marks] (Jun 2025) Explain the significance of depth in deep neural networks. How does increasing the number of layers affect model performance?

Gradient Descent

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Gradient descent is an iterative optimization algorithm that minimizes the cost function $J(\mathbf{w})$ by repeatedly moving the weights in the direction opposite to the gradient.</mark>

Formula.

$$\mathbf{w}\leftarrow\mathbf{w}-\eta\,\nabla_{\mathbf{w}}J(\mathbf{w})$$

Here $\eta$ is the learning rate. Example: $J=w^2$, $w=3$, $\eta=0.1$ gives gradient $6$, so $w=3-0.6=\mathbf{2.4}$.

Key points.

  1. The gradient points towards the steepest increase of the cost, so stepping against it lowers the cost.
  2. The learning rate sets the step size: too small is slow, too large overshoots or diverges.
  3. It is important because it is how neural networks are trained: backpropagation supplies the gradient and gradient descent uses it to update every weight.
  4. Batch gradient descent uses the whole dataset per update, so it is stable but slow and memory-hungry on large data.
  5. Stochastic gradient descent (SGD) updates after each single sample, so it is fast and noisy and can escape shallow local minima, but it fluctuates around the minimum.
  6. Mini-batch gradient descent uses small batches (32-256), balancing speed and stability, and is the practice in deep learning.
  7. Adam combines momentum (running mean of gradients) with adaptive per-weight learning rates (running mean of squared gradients).
  8. Adam update: $m=\beta_1 m+(1-\beta_1)g$, $v=\beta_2 v+(1-\beta_2)g^2$, $w\leftarrow w-\eta\,\hat m/(\sqrt{\hat v}+\epsilon)$, with defaults $\beta_1=0.9$, $\beta_2=0.999$.
Basis Batch GD Stochastic GD
Data per update Whole dataset One sample
Speed per update Slow Fast
Path to minimum Smooth, stable Noisy, zig-zag
Memory High Low
Convergence Exact minimum (convex case) Hovers near the minimum
Best for Small datasets Large or online data

Answer frame. Open with the definition and the update rule; draw a bowl-shaped cost curve with steps (or skip it); develop points 1-3 for importance, then 4-6 for variants; for the optimizer question add Adam (points 7-8); for the comparison question give the table and end with mini-batch as the trade-off.

Asked: [7 marks] (May 2023) Explain the importance of Gradient Descent algorithm in Deep Learning. Asked: [7 marks] (Jun 2025) Discuss different optimization techniques used in training feedforward networks, such as gradient descent, stochastic gradient descent, and Adam optimizer. Asked: [7 marks] (Jun 2025) Discuss how batch gradient descent and stochastic gradient descent are different.

Backpropagation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Backpropagation is the algorithm that trains an MLP by computing the loss in a forward pass and then propagating the error backwards with the chain rule to get the gradient for every weight.</mark>

Diagram.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-02" viewBox="0 0 424 80" width="424" height="80" role="img" aria-label="Forward pass computes output and loss; backward pass sends error gradients to each layer"><style>#dsfig-u2-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-02 .t{fill:#16181D;font-weight:500}#dsfig-u2-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-02 .dot{fill:#16181D}#dsfig-u2-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-02 .ah{fill:#454C5A}#dsfig-u2-02 .ah.hi{fill:#2340B8}#dsfig-u2-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-02 .e{stroke:#B1B7C3}html.dark #dsfig-u2-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-02 .t{fill:#E6E8ED}html.dark #dsfig-u2-02 .t.inv{fill:#0F1115}html.dark #dsfig-u2-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-02 .dot{fill:#E6E8ED}html.dark #dsfig-u2-02 .ann{fill:#8FA3FF}html.dark #dsfig-u2-02 .lbl{fill:#858D9C}html.dark #dsfig-u2-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-02 .ah{fill:#B1B7C3}html.dark #dsfig-u2-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M57.8,46.6 Q126,72 192.3,47.3" marker-end="url(#ah4)"/><path class="e" d="M229.8,46.6 Q298,72 364.3,47.3" marker-end="url(#ah4)"/><path class="e" d="M366.2,33.4 Q298,8 231.7,32.7" marker-end="url(#ah4)"/><path class="e" d="M194.2,33.4 Q126,8 59.7,32.7" marker-end="url(#ah4)"/><g class="wl"><rect x="108.7" y="50.5" width="33.6" height="18" rx="9"/><text class="t" x="125.5" y="59.5" dy=".35em" text-anchor="middle">fwd</text></g><g class="wl"><rect x="280.7" y="50.5" width="33.6" height="18" rx="9"/><text class="t" x="297.5" y="59.5" dy=".35em" text-anchor="middle">fwd</text></g><g class="wl"><rect x="281.7" y="11.5" width="33.6" height="18" rx="9"/><text class="t" x="298.5" y="20.5" dy=".35em" text-anchor="middle">err</text></g><g class="wl"><rect x="109.7" y="11.5" width="33.6" height="18" rx="9"/><text class="t" x="126.5" y="20.5" dy=".35em" text-anchor="middle">err</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">In</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">Hid</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">Out</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Forward pass computes output and loss; backward pass sends error gradients to each layer</figcaption></figure>

Steps.

Step 1: Initialize weights with small random values.
Step 2: Forward pass: a = f(Wx + b) layer by layer to get output y.
Step 3: Compute loss L(y, t) against target t.
Step 4: Output error delta_out = dL/dz; hidden delta = (W^T delta_next) * f'(z).
Step 5: Gradient dL/dW = delta * (input to layer)^T.
Step 6: Update W = W - eta * dL/dW; repeat for all samples and epochs.

Key points.

  1. It uses the chain rule, $\frac{\partial L}{\partial w}=\frac{\partial L}{\partial a}\frac{\partial a}{\partial z}\frac{\partial z}{\partial w}$, to reuse computed terms and avoid recomputing each gradient.
  2. It is efficient because one backward pass gives all gradients at a cost similar to a forward pass.
  3. Issues: vanishing or exploding gradients, local minima, slow convergence and overfitting.

Weight initialization. Weights must start small and random, never equal, and the scale should fit the layer size.

  1. Zero (or equal) initialization makes all neurons in a layer compute and learn the same thing, so symmetry is never broken.
  2. Too large weights saturate sigmoid or tanh and explode activations and gradients; too small weights make signals and gradients vanish.
  3. Xavier (Glorot), for sigmoid or tanh, uses variance $\frac{2}{n_{in}+n_{out}}$; He, for ReLU, uses variance $\frac{2}{n_{in}}$; biases are set to zero.
  4. Proper initialization keeps activation variance stable across layers, so training converges faster and gradients neither vanish nor explode.

Answer frame. Open by defining MLP, draw the layered figure, then the forward pass, backward pass with the chain rule and weight update, and close with advantages and training issues. For initialization, list methods then explain why it matters.

Asked: [7 marks] (May 2023) What is Multilayer Perceptron? Explain Backpropagation algorithm in detail. Asked: [7 marks] (Jun 2025) Explain the process of initializing weights in a neural network. Why is proper initialization important?

Empirical Risk Minimization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Empirical risk minimization (ERM) chooses the model that minimizes the average loss on the training data, as a substitute for the true risk, which is the expected loss over the unknown data distribution.</mark>

Formula.

$$R(f)=\mathbb{E}_{(x,y)\sim P}[L(f(x),y)],\qquad \hat R(f)=\frac1n\sum_{i=1}^{n}L(f(x_i),y_i),\qquad \hat f=\arg\min_{f\in\mathcal F}\hat R(f)$$

Key points.

  1. The objective is a model that generalizes, meaning low true risk, but since $P$ is unknown we minimize the empirical risk on the sample.
  2. Training a neural network is ERM: the loss (cross-entropy or squared error) averaged over the training set is minimized by gradient descent.
  3. By the law of large numbers, empirical risk approaches true risk as $n$ grows.
  4. Pure ERM with a very flexible model overfits: training risk is low but true risk is high.
  5. This is controlled by regularization, giving structural risk minimization: minimize $\hat R(f)+\lambda\,\Omega(f)$.

Asked: [6 marks] (May 2024) What is the objective of the empirical risk minimization? Explain the principle of Empirical Risk Minimization.

Regularization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Regularization is any technique that reduces overfitting by constraining the model so it generalizes better to unseen data.

Key points.

  1. L2 (weight decay) adds $\frac{\lambda}{2}\lVert\mathbf w\rVert^2$ to the loss and shrinks weights smoothly; L1 adds $\lambda\sum|w_i|$ and drives some weights to zero, giving sparsity.
  2. Dropout randomly switches off hidden neurons (for example with probability 0.5) during training, so neurons cannot co-adapt; at test time all are used.
  3. Early stopping halts training when validation error starts rising.
  4. Data augmentation and more data also reduce overfitting.

Autoencoders

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. ==An autoencoder is an unsupervised feedforward network trained to reconstruct its own input, by an encoder $z=f(x)$ that compresses it to a code and a decoder $\hat x=g(z)$ that rebuilds it, minimizing $\lVert x-\hat x\rVert^2$.==

Diagram.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-03" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="X input, Enc encoder, Z code (bottleneck), Dec decoder, Xh reconstruction"><style>#dsfig-u2-03 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-03 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-03 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-03 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-03 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-03 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-03 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-03 .t{fill:#16181D;font-weight:500}#dsfig-u2-03 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-03 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-03 .dot{fill:#16181D}#dsfig-u2-03 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-03 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-03 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-03 .ah{fill:#454C5A}#dsfig-u2-03 .ah.hi{fill:#2340B8}#dsfig-u2-03 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-03 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-03 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-03 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-03 .e{stroke:#B1B7C3}html.dark #dsfig-u2-03 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-03 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-03 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-03 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-03 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-03 .t{fill:#E6E8ED}html.dark #dsfig-u2-03 .t.inv{fill:#0F1115}html.dark #dsfig-u2-03 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-03 .dot{fill:#E6E8ED}html.dark #dsfig-u2-03 .ann{fill:#8FA3FF}html.dark #dsfig-u2-03 .lbl{fill:#858D9C}html.dark #dsfig-u2-03 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-03 .ah{fill:#B1B7C3}html.dark #dsfig-u2-03 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-03 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-03 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-03 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah5)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah5)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah5)"/><path class="e" d="M446,40 L535,40" marker-end="url(#ah5)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">X</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">Enc</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Z</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">Dec</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">Xh</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">X input, Enc encoder, Z code (bottleneck), Dec decoder, Xh reconstruction</figcaption></figure>

Types.

  1. Vanilla (undercomplete): the bottleneck is smaller than the input, forcing a compressed representation.
  2. Sparse: a sparsity penalty keeps most hidden units inactive, so each learns a specific feature.
  3. Denoising: the input is corrupted with noise and the network must reconstruct the clean original, learning robust features.
  4. Contractive: a penalty on the Jacobian of the encoder makes the code insensitive to small input changes.
  5. Variational (VAE): the encoder outputs a mean and variance of a distribution; sampling from it lets the decoder generate new data, and the loss is reconstruction plus KL divergence.

Applications. Dimensionality reduction, denoising, anomaly detection and data generation.

Asked: [7 marks] (May 2024) Explain different types of Autoencoders.

Last-minute revision

  • MLP = input + hidden + output layers, non-linear activation, fully connected, trained by backpropagation.
  • Single perceptron gives a linear boundary and cannot solve XOR; an MLP with one hidden layer can.
  • Gradient descent: $w\leftarrow w-\eta\nabla J$; example $w=3,\eta=0.1,J=w^2$ gives $2.4$.
  • Batch GD uses all data, SGD one sample, mini-batch a small batch.
  • Adam = momentum + adaptive learning rate; $\beta_1=0.9$, $\beta_2=0.999$.
  • Backpropagation = forward pass, loss, backward pass with the chain rule, weight update.
  • Xavier variance $2/(n_{in}+n_{out})$; He variance $2/n_{in}$; zero init breaks nothing (symmetry).
  • ERM: minimize $\frac1n\sum L$ instead of the true expected loss.
  • L2 shrinks weights, L1 gives sparsity, dropout drops neurons.
  • Autoencoder types: vanilla, sparse, denoising, contractive, variational.

Memory hooks

  • XOR needs hidden layers: "one line is not enough, so add more lines".
  • Batch = Big and Better path; Stochastic = Speedy but Shaky.
  • Backprop: "Forward for the loss, backward for the blame".
  • Autoencoder types: "V-S-D-C-V" for vanilla, sparse, denoising, contractive, variational.
  • Xavier for tanh, He for ReLU ("He needs ReLU").

Coverage checklist

  • Multilayer Perceptron: linear vs multilayer perceptron (May 2024, Jun 2025); significance of depth (Jun 2025).
  • Gradient Descent: importance of gradient descent (May 2023); optimization techniques GD, SGD, Adam (Jun 2025); batch vs stochastic GD (Jun 2025).
  • Backpropagation: MLP and backpropagation (May 2023); weight initialization (Jun 2025).
  • Empirical Risk Minimization: objective and principle (May 2024).
  • regularization: no past questions.
  • auto encoders: types of autoencoders (May 2024).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in