Skip to content
AD-601 · Deep Learning/Quick Revision Short Notes

Deep Learning (AD-601) - Unit 4 Short Notes

How unit 4 is examined

This unit covers bidirectional RNNs, deep (stacked) RNNs, recursive networks, the LSTM and other gated RNNs such as the GRU; Bidirectional RNNs and the LSTM carry the marks.

Bidirectional RNNs

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. ==A bidirectional RNN (BRNN) processes the sequence in both directions with two separate hidden layers, a forward layer reading from $t=1$ to $T$ and a backward layer reading from $T$ to $1$, and combines both states to produce each output.==

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-01" viewBox="0 0 424 424" width="424" height="424" role="img" aria-label="BRNN. X = input, F = forward hidden state, B = backward hidden state, Y = output at each time step"><style>#dsfig-u4-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-01 .t{fill:#16181D;font-weight:500}#dsfig-u4-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-01 .dot{fill:#16181D}#dsfig-u4-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-01 .ah{fill:#454C5A}#dsfig-u4-01 .ah.hi{fill:#2340B8}#dsfig-u4-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-01 .e{stroke:#B1B7C3}html.dark #dsfig-u4-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-01 .t{fill:#E6E8ED}html.dark #dsfig-u4-01 .t.inv{fill:#0F1115}html.dark #dsfig-u4-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-01 .dot{fill:#E6E8ED}html.dark #dsfig-u4-01 .ann{fill:#8FA3FF}html.dark #dsfig-u4-01 .lbl{fill:#858D9C}html.dark #dsfig-u4-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-01 .ah{fill:#B1B7C3}html.dark #dsfig-u4-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M40,365 L40,319" marker-end="url(#ah8)"/><path class="e" d="M212,365 L212,319" marker-end="url(#ah8)"/><path class="e" d="M384,365 L384,319" marker-end="url(#ah8)"/><path class="e" d="M59,298 L191,298" marker-end="url(#ah8)"/><path class="e" d="M231,298 L363,298" marker-end="url(#ah8)"/><path class="e" d="M40,365 L40,190" marker-end="url(#ah8)"/><path class="e" d="M212,365 L212,190" marker-end="url(#ah8)"/><path class="e" d="M384,365 L384,190" marker-end="url(#ah8)"/><path class="e" d="M365,169 L233,169" marker-end="url(#ah8)"/><path class="e" d="M193,169 L61,169" marker-end="url(#ah8)"/><path class="e" d="M40,279 L40,61" marker-end="url(#ah8)"/><path class="e" d="M212,279 L212,61" marker-end="url(#ah8)"/><path class="e" d="M384,279 L384,61" marker-end="url(#ah8)"/><path class="e" d="M40,150 L40,61" marker-end="url(#ah8)"/><path class="e" d="M212,150 L212,61" marker-end="url(#ah8)"/><path class="e" d="M384,150 L384,61" marker-end="url(#ah8)"/><circle class="n" cx="40" cy="384" r="18"/><text class="t" x="40" y="384" dy=".35em" text-anchor="middle">X1</text><circle class="n" cx="212" cy="384" r="18"/><text class="t" x="212" y="384" dy=".35em" text-anchor="middle">X2</text><circle class="n" cx="384" cy="384" r="18"/><text class="t" x="384" y="384" dy=".35em" text-anchor="middle">X3</text><circle class="n" cx="40" cy="298" r="18"/><text class="t" x="40" y="298" dy=".35em" text-anchor="middle">F1</text><circle class="n" cx="212" cy="298" r="18"/><text class="t" x="212" y="298" dy=".35em" text-anchor="middle">F2</text><circle class="n" cx="384" cy="298" r="18"/><text class="t" x="384" y="298" dy=".35em" text-anchor="middle">F3</text><circle class="n" cx="40" cy="169" r="18"/><text class="t" x="40" y="169" dy=".35em" text-anchor="middle">B1</text><circle class="n" cx="212" cy="169" r="18"/><text class="t" x="212" y="169" dy=".35em" text-anchor="middle">B2</text><circle class="n" cx="384" cy="169" r="18"/><text class="t" x="384" y="169" dy=".35em" text-anchor="middle">B3</text><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Y1</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">Y2</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">Y3</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">BRNN. X = input, F = forward hidden state, B = backward hidden state, Y = output at each time step</figcaption></figure>

Key points.

  1. A plain RNN output at time $t$ depends only on past inputs, but many tasks need future context too, for example a word's meaning depends on the words after it.
  2. The forward layer computes $\overrightarrow{h}_t = f(W_x x_t + W_h \overrightarrow{h}_{t-1} + b)$ using the past context.
  3. The backward layer computes $\overleftarrow{h}_t = f(W_x' x_t + W_h' \overleftarrow{h}_{t+1} + b')$ using the future context, so the two layers have separate weights and are not connected to each other.
  4. The output combines both states: $y_t = g(W_y[\overrightarrow{h}_t ; \overleftarrow{h}_t] + b_y)$, by concatenation, sum or average.
  5. Training uses backpropagation through time (BPTT) over both layers, and the whole sequence must be available before any output is computed.
  6. This is why a BRNN suits offline tasks, but not real-time prediction where the future is unknown.
  7. Applications include speech recognition, named-entity recognition, POS tagging, machine translation encoders and handwriting recognition.
Basis RNN Bidirectional RNN
Hidden layers One, forward only Two, forward and backward
Context used Past only Past and future
Parameters $W_x, W_h, W_y$ Two sets of hidden weights, roughly double
Training and cost BPTT one way, cheaper BPTT through both layers, costlier
Online prediction Possible Not possible, needs full sequence
Accuracy on tagging Lower Higher

Answer frame. Open with the definition; draw the two-layer diagram with arrows in opposite directions; then develop points 1-7 in order (need, forward equation, backward equation, output combination, BPTT, limitation, applications); close with "a BRNN gives each output the whole sequence as context". For the comparison, draw both networks side by side and give the six-row table.

Pitfall: Do not say the forward and backward layers share weights or feed each other; they are independent and only meet at the output.

Asked: [7 marks] (May 2023) What are Bidirectional RNNs? Explain in detail. Asked: [7 marks] (May 2024) Explain the difference between RNN and Bidirectional RNN.

Deep Recurrent Networks

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. A recurrent network has cyclic connections so that the hidden state $h_t = f(W_h h_{t-1} + W_x x_t + b)$ acts as memory of earlier inputs. <mark>A deep recurrent network stacks several recurrent layers, so the hidden sequence of one layer is the input sequence of the next.</mark>

Key points.

  1. Unfolding in time turns the loop into a chain of copies of the same cell with shared weights, which is what BPTT differentiates.
  2. In layer $l$ the state is $h_t^{(l)} = f(W^{(l)} h_t^{(l-1)} + U^{(l)} h_{t-1}^{(l)} + b^{(l)})$, with $h_t^{(0)} = x_t$.
  3. Lower layers learn short-term, low-level features and upper layers learn longer-term, abstract ones, giving a hierarchy of temporal features.
  4. Depth can also be added in the input-to-hidden, hidden-to-hidden or hidden-to-output transitions using small MLPs.
  5. The cost is harder training, so skip connections, LSTM or GRU cells and gradient clipping are used.

Asked: [7 marks] (May 2023) Define Recurrent Network? Discuss about Deep Recurrent Networks.

Recursive Neural Networks

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A recursive neural network applies the same weights repeatedly over a tree structure, combining child representations into a parent representation from the leaves up to the root.</mark>

Key points.

  1. The parent vector is $p = f(W[c_1 ; c_2] + b)$ with the same $W$ at every node, so a recursive net generalises a chain-shaped RNN to a tree.
  2. It suits data with hierarchical structure, such as parse trees of sentences, and is used for sentiment analysis and sentence meaning.
  3. Depth is $O(\log n)$ for a balanced tree against $n$ for a chain, which shortens the path gradients travel.
  4. It needs the tree structure to be given or parsed beforehand.

The Long Short-Term Memory

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>An LSTM is a gated recurrent cell that keeps a separate cell state $c_t$ and uses forget, input and output gates to control what is erased, written and read, so that it learns long-term dependencies.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-02" viewBox="0 0 553 338" width="553" height="338" role="img" aria-label="LSTM cell. Xh = x_t with h_(t-1), F = forget gate, I = input gate, G = candidate memory, O = output gate, C = cell state c_t, H = hidden state h_t"><style>#dsfig-u4-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-02 .t{fill:#16181D;font-weight:500}#dsfig-u4-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-02 .dot{fill:#16181D}#dsfig-u4-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-02 .ah{fill:#454C5A}#dsfig-u4-02 .ah.hi{fill:#2340B8}#dsfig-u4-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-02 .e{stroke:#B1B7C3}html.dark #dsfig-u4-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-02 .t{fill:#E6E8ED}html.dark #dsfig-u4-02 .t.inv{fill:#0F1115}html.dark #dsfig-u4-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-02 .dot{fill:#E6E8ED}html.dark #dsfig-u4-02 .ann{fill:#8FA3FF}html.dark #dsfig-u4-02 .lbl{fill:#858D9C}html.dark #dsfig-u4-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-02 .ah{fill:#B1B7C3}html.dark #dsfig-u4-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M53.4,155.6 L154.2,54.8" marker-end="url(#ah9)"/><path class="e" d="M58,163 L149.1,132.6" marker-end="url(#ah9)"/><path class="e" d="M58,175 L149.1,205.4" marker-end="url(#ah9)"/><path class="e" d="M53.4,182.4 L154.2,283.2" marker-end="url(#ah9)"/><path class="e" d="M186,48.5 L322.2,116.6" marker-end="url(#ah9)"/><path class="e" d="M188,126 L320,126" marker-end="url(#ah9)"/><path class="e" d="M186,203.5 L322.2,135.4" marker-end="url(#ah9)"/><path class="e" d="M358,134.5 L494.2,202.6" marker-end="url(#ah9)"/><path class="e" d="M187.4,293.4 L492.6,217.1" marker-end="url(#ah9)"/><circle class="n" cx="40" cy="169" r="18"/><text class="t" x="40" y="169" dy=".35em" text-anchor="middle">Xh</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">F</text><circle class="n" cx="169" cy="126" r="18"/><text class="t" x="169" y="126" dy=".35em" text-anchor="middle">I</text><circle class="n" cx="169" cy="212" r="18"/><text class="t" x="169" y="212" dy=".35em" text-anchor="middle">G</text><circle class="n" cx="169" cy="298" r="18"/><text class="t" x="169" y="298" dy=".35em" text-anchor="middle">O</text><circle class="n" cx="341" cy="126" r="18"/><text class="t" x="341" y="126" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="513" cy="212" r="18"/><text class="t" x="513" y="212" dy=".35em" text-anchor="middle">H</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">LSTM cell. Xh = x_t with h_(t-1), F = forget gate, I = input gate, G = candidate memory, O = output gate, C = cell state c_t, H = hidden state h_t</figcaption></figure>

Key points.

  1. A simple RNN suffers from vanishing gradients, because repeated multiplication by $W_h$ and $f'$ shrinks the gradient, so it forgets distant inputs; the LSTM cell state gives gradients a nearly additive path.
  2. Forget gate: $f_t = \sigma(W_f[h_{t-1}, x_t] + b_f)$ decides how much of $c_{t-1}$ to keep.
  3. Input gate: $i_t = \sigma(W_i[h_{t-1}, x_t] + b_i)$ decides how much new information to write, with candidate $\tilde{c}_t = \tanh(W_c[h_{t-1}, x_t] + b_c)$.
  4. Cell update: $c_t = f_t \odot c_{t-1} + i_t \odot \tilde{c}_t$, an elementwise sum that carries information across many steps.
  5. Output gate: $o_t = \sigma(W_o[h_{t-1}, x_t] + b_o)$ and hidden state $h_t = o_t \odot \tanh(c_t)$.
  6. Sigmoid gates output values in $(0,1)$ and act as soft switches, with 0 blocking and 1 passing fully.
  7. Applications include language modelling, machine translation, speech recognition, handwriting recognition and time-series forecasting.

Answer frame. Open with the definition and the vanishing-gradient need; draw the cell with the cell-state line on top and three gates; then develop points 2-5 with one equation each; close with the applications and "the additive cell state lets gradients flow over long spans".

Asked: [7 marks] (May 2024, Jun 2025) What is Long Short Term Memory in Neural Network? Explain.

Other Gated RNNs

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Gated RNNs use learned gates to control information flow through time, which eases vanishing gradients and lets the network keep long-term dependencies; the LSTM and the GRU are the main types.</mark>

Key points.

  1. The LSTM uses input, forget and output gates with a separate cell state, as described above.
  2. The Gated Recurrent Unit (GRU) merges cell and hidden state and uses only two gates, so it has fewer parameters and trains faster.
  3. Update gate: $z_t = \sigma(W_z[h_{t-1}, x_t])$ and reset gate: $r_t = \sigma(W_r[h_{t-1}, x_t])$.
  4. Candidate $\tilde{h}_t = \tanh(W[r_t \odot h_{t-1}, x_t])$ and new state $h_t = (1-z_t)\odot h_{t-1} + z_t \odot \tilde{h}_t$.
  5. Uses include translation, speech recognition and text and music modelling.

Asked: [6 marks] (May 2023) Write about Gated RNNs.

Last-minute revision

  • A BRNN has a forward and a backward hidden layer and uses past and future context.
  • A BRNN needs the whole sequence, so it cannot do online prediction.
  • Deep RNN: stacked recurrent layers, $h_t^{(l)}$ takes $h_t^{(l-1)}$ as input.
  • Recursive NN: shared weights over a tree, $p = f(W[c_1;c_2]+b)$.
  • LSTM has three gates: forget, input, output, plus cell state $c_t$.
  • $c_t = f_t \odot c_{t-1} + i_t \odot \tilde{c}_t$ and $h_t = o_t \odot \tanh(c_t)$.
  • Sigmoid gates give values in (0,1); the candidate uses tanh.
  • GRU has two gates: update and reset, and no separate cell state.
  • LSTM fixes vanishing gradient through the additive cell-state path.
  • Unfolding in time turns the RNN loop into a chain with shared weights.

Memory hooks

  • BRNN: "look back and look ahead", two layers meeting at the output.
  • LSTM gates FIO: Forget, Input, Output.
  • GRU = LSTM minus one gate, two gates: update and reset.
  • Recursive is a tree, recurrent is a chain.
  • Deep RNN: stacking floors, low floors short-term, high floors long-term.

Coverage checklist

  • Bidirectional RNNs: May 2023 (7), May 2024 (7).
  • Deep Recurrent Networks: May 2023 (7).
  • Recursive Neural Networks: not asked recently, taught by definition and key points.
  • The Long Short-Term Memory: May 2024, Jun 2025 (7).
  • Other Gated RNNs: May 2023 (6).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in