How unit 5 is examined
This unit covers energy-based generative networks (Boltzmann machines, RBMs, DBNs, DBMs), the sampling methods that train them, and applications; MCMC and Gibbs sampling and Deep Boltzmann Machines carry the marks.
Boltzmann Machines
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A Boltzmann machine is a network of symmetrically connected stochastic binary units that learns a probability distribution over its inputs by minimising an energy function.
Key points.
- Every unit is visible or hidden, and each unit switches on with a probability, not deterministically.
- Weights are symmetric ($w_{ij}=w_{ji}$) and every pair of units may be connected, so learning is slow.
- Energy of a state: $E(x)=-\sum_{i<j} w_{ij}x_ix_j-\sum_i b_ix_i$, and $P(x)=e^{-E(x)}/Z$, where $Z$ sums $e^{-E}$ over all states.
- Low-energy states are the most probable, so training lowers the energy of the training data.
Restricted Boltzmann Machines
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. An RBM is a Boltzmann machine restricted to two layers, visible and hidden, with connections only between layers (a bipartite graph) and none within a layer.
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-01" viewBox="0 0 338 338" width="338" height="338" role="img" aria-label="RBM: V = visible units, H = hidden units, no links inside a layer"><style>#dsfig-u5-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-01 .t{fill:#16181D;font-weight:500}#dsfig-u5-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-01 .dot{fill:#16181D}#dsfig-u5-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-01 .ah{fill:#454C5A}#dsfig-u5-01 .ah.hi{fill:#2340B8}#dsfig-u5-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-01 .e{stroke:#B1B7C3}html.dark #dsfig-u5-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-01 .t{fill:#E6E8ED}html.dark #dsfig-u5-01 .t.inv{fill:#0F1115}html.dark #dsfig-u5-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-01 .dot{fill:#E6E8ED}html.dark #dsfig-u5-01 .ann{fill:#8FA3FF}html.dark #dsfig-u5-01 .lbl{fill:#858D9C}html.dark #dsfig-u5-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-01 .ah{fill:#B1B7C3}html.dark #dsfig-u5-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M58.5,44.3 L279.5,95.9"/><path class="e" d="M55.1,51.6 L282.9,226.2"/><path class="e" d="M58.4,164.1 L279.6,105.1"/><path class="e" d="M58.4,173.9 L279.6,232.9"/><path class="e" d="M55.1,286.4 L282.9,111.8"/><path class="e" d="M58.5,293.7 L279.5,242.1"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">V1</text><circle class="n" cx="40" cy="169" r="18"/><text class="t" x="40" y="169" dy=".35em" text-anchor="middle">V2</text><circle class="n" cx="40" cy="298" r="18"/><text class="t" x="40" y="298" dy=".35em" text-anchor="middle">V3</text><circle class="n" cx="298" cy="100.2" r="18"/><text class="t" x="298" y="100.2" dy=".35em" text-anchor="middle">H1</text><circle class="n" cx="298" cy="237.8" r="18"/><text class="t" x="298" y="237.8" dy=".35em" text-anchor="middle">H2</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">RBM: V = visible units, H = hidden units, no links inside a layer</figcaption></figure>
Key points.
- Energy is $E(v,h)=-b^Tv-c^Th-v^TWh$, and $P(v,h)=e^{-E(v,h)}/Z$.
- Because there are no within-layer links, hidden units are independent given $v$: $P(h_j=1\mid v)=\sigma(c_j+\sum_i v_iW_{ij})$, and likewise $P(v_i=1\mid h)$, which makes sampling fast.
- It is trained without labels by contrastive divergence, so it is an unsupervised generative model.
| Basis | Autoencoder | RBM |
|---|---|---|
| Principle | Encoder compresses input, decoder reconstructs it | Energy-based model of the joint distribution of $v,h$ |
| Architecture | Directed feed-forward, encoder then decoder | Undirected bipartite, one weight matrix used both ways |
| Units | Deterministic | Stochastic binary |
| Training | Backpropagation on reconstruction error | Contrastive divergence lowering energy |
| Use | Compression, denoising, feature learning | Generative modelling, pre-training DBNs |
Asked: [7 marks] (May 2024) What is the difference between Autoencoders and RBM? Discuss.
Introduction to MCMC and Gibbs Sampling
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. Markov Chain Monte Carlo (MCMC) draws samples from a distribution that is hard to sample directly by building a Markov chain whose long-run (stationary) distribution is the target distribution. <mark>Gibbs sampling is an MCMC method that updates one variable at a time by sampling it from its conditional distribution given all the other variables.</mark>
Key points.
- Sampling is needed in deep learning because models such as RBMs have an intractable normaliser $Z$, so expectations must be approximated by averages over samples.
- A Markov chain moves to the next state depending only on the current state, and after a burn-in period its states behave like samples from the target.
- Monte Carlo means estimating an expectation by averaging over samples: $E[f(x)]\approx \frac1N\sum_t f(x^{(t)})$.
- Metropolis-Hastings proposes a new state and accepts it with probability $\min\!\left(1,\frac{p(x')q(x\mid x')}{p(x)q(x'\mid x)}\right)$, otherwise it keeps the old state.
- Gibbs sampling is a special case where the proposal is always accepted, because each variable is drawn from its exact conditional.
- Consecutive samples are correlated and early samples depend on the start, so the first samples are discarded (burn-in) and every k-th is kept (thinning).
- Applications: training RBMs and DBMs, Bayesian inference, image denoising and estimating intractable integrals.
Steps.
Step 1: Start from an arbitrary state x = (x1, ..., xn).
Step 2: For i = 1 to n, draw xi from p(xi | all other variables) using current values.
Step 3: One full sweep gives a new sample; repeat Step 2 many times.
Step 4: Discard the burn-in samples and use the rest as samples of p(x).
Example. In an RBM, Gibbs sampling alternates $h\sim P(h\mid v)$ and $v\sim P(v\mid h)$; after many alternations $(v,h)$ is a sample from the model.
Answer frame. Open with the definition of MCMC; for Gibbs, write the four steps and the RBM example; develop points 1-7 in order; close with applications. For "any two" (Q2) pick Gibbs sampling and Deep generative models (defined under Deep Boltzmann Machines); each needs definition, working, one feature and applications.
Pitfall: Do not say Gibbs sampling gives independent samples; successive samples are correlated.
Asked: [14 marks] (May 2023) Explain any two: a) Deep Generic Models b) Gibbs Sampling c) Feed forward Network d) Artificial Intelligence versus Deep Learning Asked: [7 marks] (May 2024) Give an overview of MCMC Sampling.
Gradient computations in RBMs
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. The log-likelihood gradient of an RBM is the difference between a data-driven and a model-driven expectation, and contrastive divergence (CD) approximates it cheaply.
Key points.
- Exact gradient: $\partial \log p(v)/\partial W_{ij}=\langle v_ih_j\rangle_{data}-\langle v_ih_j\rangle_{model}$; the model term needs the intractable $Z$.
- CD-1 replaces the model term with one Gibbs step: $v\to h\to v'\to h'$.
- Update: $\Delta W_{ij}=\epsilon\,(\langle v_ih_j\rangle_{data}-\langle v_ih_j\rangle_{recon})$, with $\epsilon$ the learning rate.
Deep Belief Networks
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A Deep Belief Network is a generative model built by stacking RBMs, in which the top two layers form an undirected RBM and the lower layers are directed top-down.
Key points.
- It is trained greedily: train the first RBM on data, freeze it, and use its hidden activations as input to train the next RBM.
- This layer-wise generative pre-training gives good initial weights, which are then fine-tuned with labels using backpropagation.
- Directed lower layers allow one fast top-down pass to generate data.
Deep Boltzmann Machines
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. A Deep Boltzmann Machine (DBM) is a generative model with several hidden layers stacked from RBMs in which all connections are undirected and only adjacent layers are linked. <mark>A DBM is a stack of RBMs with fully undirected connections, so each hidden layer is influenced by both the layer below and the layer above.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-02" viewBox="0 0 424 80" width="424" height="80" role="img" aria-label="DBM: V = visible layer, H1 and H2 = hidden layers, all links undirected"><style>#dsfig-u5-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-02 .t{fill:#16181D;font-weight:500}#dsfig-u5-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-02 .dot{fill:#16181D}#dsfig-u5-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-02 .ah{fill:#454C5A}#dsfig-u5-02 .ah.hi{fill:#2340B8}#dsfig-u5-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-02 .e{stroke:#B1B7C3}html.dark #dsfig-u5-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-02 .t{fill:#E6E8ED}html.dark #dsfig-u5-02 .t.inv{fill:#0F1115}html.dark #dsfig-u5-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-02 .dot{fill:#E6E8ED}html.dark #dsfig-u5-02 .ann{fill:#8FA3FF}html.dark #dsfig-u5-02 .lbl{fill:#858D9C}html.dark #dsfig-u5-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-02 .ah{fill:#B1B7C3}html.dark #dsfig-u5-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah11" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh11" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L193,40"/><path class="e" d="M231,40 L365,40"/><g class="wl"><rect x="112.8" y="31" width="26.4" height="18" rx="9"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">W1</text></g><g class="wl"><rect x="284.8" y="31" width="26.4" height="18" rx="9"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">W2</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">V</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">H1</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">H2</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">DBM: V = visible layer, H1 and H2 = hidden layers, all links undirected</figcaption></figure>
Key points.
- Energy for two hidden layers: $E(v,h^1,h^2)=-v^TW^1h^1-h^{1T}W^2h^2$ (bias terms omitted), and $P=e^{-E}/Z$.
- Because every link is undirected, inference for a hidden layer uses both neighbours, unlike a DBN, so it is more accurate but harder.
- Training uses greedy layer-wise RBM pre-training, then joint fine-tuning with mean-field inference for the data term and Gibbs (MCMC) sampling for the model term.
- Compared with a DBN, whose lower links are directed, a DBM is a single Boltzmann machine with no directed part.
- DBMs learn increasingly abstract features layer by layer, useful for recognition and generation.
Deep Belief Network (part b). A DBN has an undirected RBM on top and directed layers below; it is trained by greedy layer-wise pre-training.
Supervised learning (part c). Learning from labelled pairs $(x,y)$ to learn a mapping $x\to y$, by minimising a loss on the predictions (classification, regression). Generative pre-training uses unlabelled data to model $p(x)$, while discriminative learning directly models $p(y\mid x)$ with labels.
Deep generative models. They are networks that learn the data distribution $p(x)$ and can generate new samples. Types: GANs (generator versus discriminator), VAEs (encoder-decoder with a latent distribution), autoregressive models (predict each element from earlier ones), and Boltzmann machines. Applications: image synthesis and super-resolution, text and speech generation, data augmentation, anomaly detection and drug discovery (generating molecules).
Answer frame. For Q1 give each of the three parts a definition, one diagram (DBM) and two features, and end with the generative versus discriminative comparison. For Q5 open with the definition, list the three types, state that all learn $p(x)$, then applications.
Asked: [14 marks] (May 2023) Give an overview of: a) Deep Boltzmann Machines b) Deep Belief Network c) Supervised Learning Asked: [7 marks] (Jun 2025) How are deep generative models used in machine learning, and what are some of their key applications?
Image Processing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Image processing with deep learning uses networks, mostly convolutional, to classify, detect or generate images from raw pixels.
Key points.
- Networks learn features (edges, textures, objects) automatically instead of using hand-made features.
- Uses: image classification, object detection, segmentation and denoising.
- Generative models produce or complete images.
Speech Recognition
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Speech recognition (ASR) converts a spoken audio signal into text.
Key points.
- Audio is converted to spectrogram or MFCC features.
- Deep networks replaced Gaussian mixture acoustic models and map features to phonemes.
- Recurrent networks (LSTM) model the time sequence.
Natural Language Processing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. NLP uses deep networks to understand and generate human language text.
Key points.
- Words are converted to vectors (embeddings) before entering the network.
- Recurrent networks and LSTMs handle word sequences.
- Tasks: translation, sentiment analysis, question answering and text generation.
Last-minute revision
- Boltzmann machine: stochastic symmetric binary units, $P(x)=e^{-E(x)}/Z$.
- RBM: bipartite, visible and hidden layers, no links inside a layer.
- RBM conditionals: $P(h_j=1\mid v)=\sigma(c_j+\sum_i v_iW_{ij})$.
- MCMC builds a Markov chain whose stationary distribution is the target.
- Gibbs sampling updates one variable at a time from its conditional.
- Metropolis-Hastings accepts with $\min(1,\text{ratio})$.
- CD-1: $\Delta W=\epsilon(\langle vh\rangle_{data}-\langle vh\rangle_{recon})$.
- DBN: stacked RBMs, top undirected, lower directed; DBM: all undirected.
- Autoencoder is deterministic and uses backpropagation; RBM is stochastic and energy-based.
- Deep generative models: GAN, VAE, autoregressive; learn $p(x)$.
Memory hooks
- "DBN = Directed Below, Non-directed on top" while DBM is all undirected.
- "Gibbs = one at a time"; "Metropolis = propose, then accept".
- "Restricted = no links within a layer".
- CD: "data minus dream".
Coverage checklist
- Boltzmann Machines: definition, energy, probability.
- Restricted Boltzmann Machines: Q4 (Autoencoders versus RBM).
- Introduction to MCMC and Gibbs Sampling: Q2 (May 2023), Q3 (May 2024).
- Gradient computations in RBMs: gradient and CD-1.
- Deep Belief Networks: definition and training.
- Deep Boltzmann Machines: Q1 (May 2023), Q5 (Jun 2025).
- Image Processing: definition and uses.
- Speech Recognition: definition and features.
- Natural Language Processing: definition and tasks.