Skip to content
AD-601 · Deep Learning/Quick Revision Short Notes

Deep Learning (AD-601) - Unit 5 Short Notes

How unit 5 is examined

This unit covers energy-based generative networks (Boltzmann machines, RBMs, DBNs, DBMs), the sampling methods that train them, and applications; MCMC and Gibbs sampling and Deep Boltzmann Machines carry the marks.

Boltzmann Machines

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. A Boltzmann machine is a network of symmetrically connected stochastic binary units that learns a probability distribution over its inputs by minimising an energy function.

Key points.

  1. Every unit is visible or hidden, and each unit switches on with a probability, not deterministically.
  2. Weights are symmetric ($w_{ij}=w_{ji}$) and every pair of units may be connected, so learning is slow.
  3. Energy of a state: $E(x)=-\sum_{i<j} w_{ij}x_ix_j-\sum_i b_ix_i$, and $P(x)=e^{-E(x)}/Z$, where $Z$ sums $e^{-E}$ over all states.
  4. Low-energy states are the most probable, so training lowers the energy of the training data.

Restricted Boltzmann Machines

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. An RBM is a Boltzmann machine restricted to two layers, visible and hidden, with connections only between layers (a bipartite graph) and none within a layer.

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-01" viewBox="0 0 338 338" width="338" height="338" role="img" aria-label="RBM: V = visible units, H = hidden units, no links inside a layer"><style>#dsfig-u5-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-01 .t{fill:#16181D;font-weight:500}#dsfig-u5-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-01 .dot{fill:#16181D}#dsfig-u5-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-01 .ah{fill:#454C5A}#dsfig-u5-01 .ah.hi{fill:#2340B8}#dsfig-u5-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-01 .e{stroke:#B1B7C3}html.dark #dsfig-u5-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-01 .t{fill:#E6E8ED}html.dark #dsfig-u5-01 .t.inv{fill:#0F1115}html.dark #dsfig-u5-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-01 .dot{fill:#E6E8ED}html.dark #dsfig-u5-01 .ann{fill:#8FA3FF}html.dark #dsfig-u5-01 .lbl{fill:#858D9C}html.dark #dsfig-u5-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-01 .ah{fill:#B1B7C3}html.dark #dsfig-u5-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M58.5,44.3 L279.5,95.9"/><path class="e" d="M55.1,51.6 L282.9,226.2"/><path class="e" d="M58.4,164.1 L279.6,105.1"/><path class="e" d="M58.4,173.9 L279.6,232.9"/><path class="e" d="M55.1,286.4 L282.9,111.8"/><path class="e" d="M58.5,293.7 L279.5,242.1"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">V1</text><circle class="n" cx="40" cy="169" r="18"/><text class="t" x="40" y="169" dy=".35em" text-anchor="middle">V2</text><circle class="n" cx="40" cy="298" r="18"/><text class="t" x="40" y="298" dy=".35em" text-anchor="middle">V3</text><circle class="n" cx="298" cy="100.2" r="18"/><text class="t" x="298" y="100.2" dy=".35em" text-anchor="middle">H1</text><circle class="n" cx="298" cy="237.8" r="18"/><text class="t" x="298" y="237.8" dy=".35em" text-anchor="middle">H2</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">RBM: V = visible units, H = hidden units, no links inside a layer</figcaption></figure>

Key points.

  1. Energy is $E(v,h)=-b^Tv-c^Th-v^TWh$, and $P(v,h)=e^{-E(v,h)}/Z$.
  2. Because there are no within-layer links, hidden units are independent given $v$: $P(h_j=1\mid v)=\sigma(c_j+\sum_i v_iW_{ij})$, and likewise $P(v_i=1\mid h)$, which makes sampling fast.
  3. It is trained without labels by contrastive divergence, so it is an unsupervised generative model.
Basis Autoencoder RBM
Principle Encoder compresses input, decoder reconstructs it Energy-based model of the joint distribution of $v,h$
Architecture Directed feed-forward, encoder then decoder Undirected bipartite, one weight matrix used both ways
Units Deterministic Stochastic binary
Training Backpropagation on reconstruction error Contrastive divergence lowering energy
Use Compression, denoising, feature learning Generative modelling, pre-training DBNs

Asked: [7 marks] (May 2024) What is the difference between Autoencoders and RBM? Discuss.

Introduction to MCMC and Gibbs Sampling

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. Markov Chain Monte Carlo (MCMC) draws samples from a distribution that is hard to sample directly by building a Markov chain whose long-run (stationary) distribution is the target distribution. <mark>Gibbs sampling is an MCMC method that updates one variable at a time by sampling it from its conditional distribution given all the other variables.</mark>

Key points.

  1. Sampling is needed in deep learning because models such as RBMs have an intractable normaliser $Z$, so expectations must be approximated by averages over samples.
  2. A Markov chain moves to the next state depending only on the current state, and after a burn-in period its states behave like samples from the target.
  3. Monte Carlo means estimating an expectation by averaging over samples: $E[f(x)]\approx \frac1N\sum_t f(x^{(t)})$.
  4. Metropolis-Hastings proposes a new state and accepts it with probability $\min\!\left(1,\frac{p(x')q(x\mid x')}{p(x)q(x'\mid x)}\right)$, otherwise it keeps the old state.
  5. Gibbs sampling is a special case where the proposal is always accepted, because each variable is drawn from its exact conditional.
  6. Consecutive samples are correlated and early samples depend on the start, so the first samples are discarded (burn-in) and every k-th is kept (thinning).
  7. Applications: training RBMs and DBMs, Bayesian inference, image denoising and estimating intractable integrals.

Steps.

Step 1: Start from an arbitrary state x = (x1, ..., xn).
Step 2: For i = 1 to n, draw xi from p(xi | all other variables) using current values.
Step 3: One full sweep gives a new sample; repeat Step 2 many times.
Step 4: Discard the burn-in samples and use the rest as samples of p(x).

Example. In an RBM, Gibbs sampling alternates $h\sim P(h\mid v)$ and $v\sim P(v\mid h)$; after many alternations $(v,h)$ is a sample from the model.

Answer frame. Open with the definition of MCMC; for Gibbs, write the four steps and the RBM example; develop points 1-7 in order; close with applications. For "any two" (Q2) pick Gibbs sampling and Deep generative models (defined under Deep Boltzmann Machines); each needs definition, working, one feature and applications.

Pitfall: Do not say Gibbs sampling gives independent samples; successive samples are correlated.

Asked: [14 marks] (May 2023) Explain any two: a) Deep Generic Models b) Gibbs Sampling c) Feed forward Network d) Artificial Intelligence versus Deep Learning Asked: [7 marks] (May 2024) Give an overview of MCMC Sampling.

Gradient computations in RBMs

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. The log-likelihood gradient of an RBM is the difference between a data-driven and a model-driven expectation, and contrastive divergence (CD) approximates it cheaply.

Key points.

  1. Exact gradient: $\partial \log p(v)/\partial W_{ij}=\langle v_ih_j\rangle_{data}-\langle v_ih_j\rangle_{model}$; the model term needs the intractable $Z$.
  2. CD-1 replaces the model term with one Gibbs step: $v\to h\to v'\to h'$.
  3. Update: $\Delta W_{ij}=\epsilon\,(\langle v_ih_j\rangle_{data}-\langle v_ih_j\rangle_{recon})$, with $\epsilon$ the learning rate.

Deep Belief Networks

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. A Deep Belief Network is a generative model built by stacking RBMs, in which the top two layers form an undirected RBM and the lower layers are directed top-down.

Key points.

  1. It is trained greedily: train the first RBM on data, freeze it, and use its hidden activations as input to train the next RBM.
  2. This layer-wise generative pre-training gives good initial weights, which are then fine-tuned with labels using backpropagation.
  3. Directed lower layers allow one fast top-down pass to generate data.

Deep Boltzmann Machines

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. A Deep Boltzmann Machine (DBM) is a generative model with several hidden layers stacked from RBMs in which all connections are undirected and only adjacent layers are linked. <mark>A DBM is a stack of RBMs with fully undirected connections, so each hidden layer is influenced by both the layer below and the layer above.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-02" viewBox="0 0 424 80" width="424" height="80" role="img" aria-label="DBM: V = visible layer, H1 and H2 = hidden layers, all links undirected"><style>#dsfig-u5-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-02 .t{fill:#16181D;font-weight:500}#dsfig-u5-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-02 .dot{fill:#16181D}#dsfig-u5-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-02 .ah{fill:#454C5A}#dsfig-u5-02 .ah.hi{fill:#2340B8}#dsfig-u5-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-02 .e{stroke:#B1B7C3}html.dark #dsfig-u5-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-02 .t{fill:#E6E8ED}html.dark #dsfig-u5-02 .t.inv{fill:#0F1115}html.dark #dsfig-u5-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-02 .dot{fill:#E6E8ED}html.dark #dsfig-u5-02 .ann{fill:#8FA3FF}html.dark #dsfig-u5-02 .lbl{fill:#858D9C}html.dark #dsfig-u5-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-02 .ah{fill:#B1B7C3}html.dark #dsfig-u5-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah11" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh11" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L193,40"/><path class="e" d="M231,40 L365,40"/><g class="wl"><rect x="112.8" y="31" width="26.4" height="18" rx="9"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">W1</text></g><g class="wl"><rect x="284.8" y="31" width="26.4" height="18" rx="9"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">W2</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">V</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">H1</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">H2</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">DBM: V = visible layer, H1 and H2 = hidden layers, all links undirected</figcaption></figure>

Key points.

  1. Energy for two hidden layers: $E(v,h^1,h^2)=-v^TW^1h^1-h^{1T}W^2h^2$ (bias terms omitted), and $P=e^{-E}/Z$.
  2. Because every link is undirected, inference for a hidden layer uses both neighbours, unlike a DBN, so it is more accurate but harder.
  3. Training uses greedy layer-wise RBM pre-training, then joint fine-tuning with mean-field inference for the data term and Gibbs (MCMC) sampling for the model term.
  4. Compared with a DBN, whose lower links are directed, a DBM is a single Boltzmann machine with no directed part.
  5. DBMs learn increasingly abstract features layer by layer, useful for recognition and generation.

Deep Belief Network (part b). A DBN has an undirected RBM on top and directed layers below; it is trained by greedy layer-wise pre-training.

Supervised learning (part c). Learning from labelled pairs $(x,y)$ to learn a mapping $x\to y$, by minimising a loss on the predictions (classification, regression). Generative pre-training uses unlabelled data to model $p(x)$, while discriminative learning directly models $p(y\mid x)$ with labels.

Deep generative models. They are networks that learn the data distribution $p(x)$ and can generate new samples. Types: GANs (generator versus discriminator), VAEs (encoder-decoder with a latent distribution), autoregressive models (predict each element from earlier ones), and Boltzmann machines. Applications: image synthesis and super-resolution, text and speech generation, data augmentation, anomaly detection and drug discovery (generating molecules).

Answer frame. For Q1 give each of the three parts a definition, one diagram (DBM) and two features, and end with the generative versus discriminative comparison. For Q5 open with the definition, list the three types, state that all learn $p(x)$, then applications.

Asked: [14 marks] (May 2023) Give an overview of: a) Deep Boltzmann Machines b) Deep Belief Network c) Supervised Learning Asked: [7 marks] (Jun 2025) How are deep generative models used in machine learning, and what are some of their key applications?

Image Processing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Image processing with deep learning uses networks, mostly convolutional, to classify, detect or generate images from raw pixels.

Key points.

  1. Networks learn features (edges, textures, objects) automatically instead of using hand-made features.
  2. Uses: image classification, object detection, segmentation and denoising.
  3. Generative models produce or complete images.

Speech Recognition

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Speech recognition (ASR) converts a spoken audio signal into text.

Key points.

  1. Audio is converted to spectrogram or MFCC features.
  2. Deep networks replaced Gaussian mixture acoustic models and map features to phonemes.
  3. Recurrent networks (LSTM) model the time sequence.

Natural Language Processing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. NLP uses deep networks to understand and generate human language text.

Key points.

  1. Words are converted to vectors (embeddings) before entering the network.
  2. Recurrent networks and LSTMs handle word sequences.
  3. Tasks: translation, sentiment analysis, question answering and text generation.

Last-minute revision

  • Boltzmann machine: stochastic symmetric binary units, $P(x)=e^{-E(x)}/Z$.
  • RBM: bipartite, visible and hidden layers, no links inside a layer.
  • RBM conditionals: $P(h_j=1\mid v)=\sigma(c_j+\sum_i v_iW_{ij})$.
  • MCMC builds a Markov chain whose stationary distribution is the target.
  • Gibbs sampling updates one variable at a time from its conditional.
  • Metropolis-Hastings accepts with $\min(1,\text{ratio})$.
  • CD-1: $\Delta W=\epsilon(\langle vh\rangle_{data}-\langle vh\rangle_{recon})$.
  • DBN: stacked RBMs, top undirected, lower directed; DBM: all undirected.
  • Autoencoder is deterministic and uses backpropagation; RBM is stochastic and energy-based.
  • Deep generative models: GAN, VAE, autoregressive; learn $p(x)$.

Memory hooks

  • "DBN = Directed Below, Non-directed on top" while DBM is all undirected.
  • "Gibbs = one at a time"; "Metropolis = propose, then accept".
  • "Restricted = no links within a layer".
  • CD: "data minus dream".

Coverage checklist

  • Boltzmann Machines: definition, energy, probability.
  • Restricted Boltzmann Machines: Q4 (Autoencoders versus RBM).
  • Introduction to MCMC and Gibbs Sampling: Q2 (May 2023), Q3 (May 2024).
  • Gradient computations in RBMs: gradient and CD-1.
  • Deep Belief Networks: definition and training.
  • Deep Boltzmann Machines: Q1 (May 2023), Q5 (Jun 2025).
  • Image Processing: definition and uses.
  • Speech Recognition: definition and features.
  • Natural Language Processing: definition and tasks.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in