How unit 4 is examined
This unit covers bidirectional RNNs, deep (stacked) RNNs, recursive networks, the LSTM and other gated RNNs such as the GRU; Bidirectional RNNs and the LSTM carry the marks.
Bidirectional RNNs
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. ==A bidirectional RNN (BRNN) processes the sequence in both directions with two separate hidden layers, a forward layer reading from $t=1$ to $T$ and a backward layer reading from $T$ to $1$, and combines both states to produce each output.==
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-01" viewBox="0 0 424 424" width="424" height="424" role="img" aria-label="BRNN. X = input, F = forward hidden state, B = backward hidden state, Y = output at each time step"><style>#dsfig-u4-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-01 .t{fill:#16181D;font-weight:500}#dsfig-u4-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-01 .dot{fill:#16181D}#dsfig-u4-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-01 .ah{fill:#454C5A}#dsfig-u4-01 .ah.hi{fill:#2340B8}#dsfig-u4-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-01 .e{stroke:#B1B7C3}html.dark #dsfig-u4-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-01 .t{fill:#E6E8ED}html.dark #dsfig-u4-01 .t.inv{fill:#0F1115}html.dark #dsfig-u4-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-01 .dot{fill:#E6E8ED}html.dark #dsfig-u4-01 .ann{fill:#8FA3FF}html.dark #dsfig-u4-01 .lbl{fill:#858D9C}html.dark #dsfig-u4-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-01 .ah{fill:#B1B7C3}html.dark #dsfig-u4-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M40,365 L40,319" marker-end="url(#ah8)"/><path class="e" d="M212,365 L212,319" marker-end="url(#ah8)"/><path class="e" d="M384,365 L384,319" marker-end="url(#ah8)"/><path class="e" d="M59,298 L191,298" marker-end="url(#ah8)"/><path class="e" d="M231,298 L363,298" marker-end="url(#ah8)"/><path class="e" d="M40,365 L40,190" marker-end="url(#ah8)"/><path class="e" d="M212,365 L212,190" marker-end="url(#ah8)"/><path class="e" d="M384,365 L384,190" marker-end="url(#ah8)"/><path class="e" d="M365,169 L233,169" marker-end="url(#ah8)"/><path class="e" d="M193,169 L61,169" marker-end="url(#ah8)"/><path class="e" d="M40,279 L40,61" marker-end="url(#ah8)"/><path class="e" d="M212,279 L212,61" marker-end="url(#ah8)"/><path class="e" d="M384,279 L384,61" marker-end="url(#ah8)"/><path class="e" d="M40,150 L40,61" marker-end="url(#ah8)"/><path class="e" d="M212,150 L212,61" marker-end="url(#ah8)"/><path class="e" d="M384,150 L384,61" marker-end="url(#ah8)"/><circle class="n" cx="40" cy="384" r="18"/><text class="t" x="40" y="384" dy=".35em" text-anchor="middle">X1</text><circle class="n" cx="212" cy="384" r="18"/><text class="t" x="212" y="384" dy=".35em" text-anchor="middle">X2</text><circle class="n" cx="384" cy="384" r="18"/><text class="t" x="384" y="384" dy=".35em" text-anchor="middle">X3</text><circle class="n" cx="40" cy="298" r="18"/><text class="t" x="40" y="298" dy=".35em" text-anchor="middle">F1</text><circle class="n" cx="212" cy="298" r="18"/><text class="t" x="212" y="298" dy=".35em" text-anchor="middle">F2</text><circle class="n" cx="384" cy="298" r="18"/><text class="t" x="384" y="298" dy=".35em" text-anchor="middle">F3</text><circle class="n" cx="40" cy="169" r="18"/><text class="t" x="40" y="169" dy=".35em" text-anchor="middle">B1</text><circle class="n" cx="212" cy="169" r="18"/><text class="t" x="212" y="169" dy=".35em" text-anchor="middle">B2</text><circle class="n" cx="384" cy="169" r="18"/><text class="t" x="384" y="169" dy=".35em" text-anchor="middle">B3</text><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Y1</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">Y2</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">Y3</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">BRNN. X = input, F = forward hidden state, B = backward hidden state, Y = output at each time step</figcaption></figure>
Key points.
- A plain RNN output at time $t$ depends only on past inputs, but many tasks need future context too, for example a word's meaning depends on the words after it.
- The forward layer computes $\overrightarrow{h}_t = f(W_x x_t + W_h \overrightarrow{h}_{t-1} + b)$ using the past context.
- The backward layer computes $\overleftarrow{h}_t = f(W_x' x_t + W_h' \overleftarrow{h}_{t+1} + b')$ using the future context, so the two layers have separate weights and are not connected to each other.
- The output combines both states: $y_t = g(W_y[\overrightarrow{h}_t ; \overleftarrow{h}_t] + b_y)$, by concatenation, sum or average.
- Training uses backpropagation through time (BPTT) over both layers, and the whole sequence must be available before any output is computed.
- This is why a BRNN suits offline tasks, but not real-time prediction where the future is unknown.
- Applications include speech recognition, named-entity recognition, POS tagging, machine translation encoders and handwriting recognition.
| Basis | RNN | Bidirectional RNN |
|---|---|---|
| Hidden layers | One, forward only | Two, forward and backward |
| Context used | Past only | Past and future |
| Parameters | $W_x, W_h, W_y$ | Two sets of hidden weights, roughly double |
| Training and cost | BPTT one way, cheaper | BPTT through both layers, costlier |
| Online prediction | Possible | Not possible, needs full sequence |
| Accuracy on tagging | Lower | Higher |
Answer frame. Open with the definition; draw the two-layer diagram with arrows in opposite directions; then develop points 1-7 in order (need, forward equation, backward equation, output combination, BPTT, limitation, applications); close with "a BRNN gives each output the whole sequence as context". For the comparison, draw both networks side by side and give the six-row table.
Pitfall: Do not say the forward and backward layers share weights or feed each other; they are independent and only meet at the output.
Asked: [7 marks] (May 2023) What are Bidirectional RNNs? Explain in detail. Asked: [7 marks] (May 2024) Explain the difference between RNN and Bidirectional RNN.
Deep Recurrent Networks
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. A recurrent network has cyclic connections so that the hidden state $h_t = f(W_h h_{t-1} + W_x x_t + b)$ acts as memory of earlier inputs. <mark>A deep recurrent network stacks several recurrent layers, so the hidden sequence of one layer is the input sequence of the next.</mark>
Key points.
- Unfolding in time turns the loop into a chain of copies of the same cell with shared weights, which is what BPTT differentiates.
- In layer $l$ the state is $h_t^{(l)} = f(W^{(l)} h_t^{(l-1)} + U^{(l)} h_{t-1}^{(l)} + b^{(l)})$, with $h_t^{(0)} = x_t$.
- Lower layers learn short-term, low-level features and upper layers learn longer-term, abstract ones, giving a hierarchy of temporal features.
- Depth can also be added in the input-to-hidden, hidden-to-hidden or hidden-to-output transitions using small MLPs.
- The cost is harder training, so skip connections, LSTM or GRU cells and gradient clipping are used.
Asked: [7 marks] (May 2023) Define Recurrent Network? Discuss about Deep Recurrent Networks.
Recursive Neural Networks
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A recursive neural network applies the same weights repeatedly over a tree structure, combining child representations into a parent representation from the leaves up to the root.</mark>
Key points.
- The parent vector is $p = f(W[c_1 ; c_2] + b)$ with the same $W$ at every node, so a recursive net generalises a chain-shaped RNN to a tree.
- It suits data with hierarchical structure, such as parse trees of sentences, and is used for sentiment analysis and sentence meaning.
- Depth is $O(\log n)$ for a balanced tree against $n$ for a chain, which shortens the path gradients travel.
- It needs the tree structure to be given or parsed beforehand.
The Long Short-Term Memory
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>An LSTM is a gated recurrent cell that keeps a separate cell state $c_t$ and uses forget, input and output gates to control what is erased, written and read, so that it learns long-term dependencies.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-02" viewBox="0 0 553 338" width="553" height="338" role="img" aria-label="LSTM cell. Xh = x_t with h_(t-1), F = forget gate, I = input gate, G = candidate memory, O = output gate, C = cell state c_t, H = hidden state h_t"><style>#dsfig-u4-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-02 .t{fill:#16181D;font-weight:500}#dsfig-u4-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-02 .dot{fill:#16181D}#dsfig-u4-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-02 .ah{fill:#454C5A}#dsfig-u4-02 .ah.hi{fill:#2340B8}#dsfig-u4-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-02 .e{stroke:#B1B7C3}html.dark #dsfig-u4-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-02 .t{fill:#E6E8ED}html.dark #dsfig-u4-02 .t.inv{fill:#0F1115}html.dark #dsfig-u4-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-02 .dot{fill:#E6E8ED}html.dark #dsfig-u4-02 .ann{fill:#8FA3FF}html.dark #dsfig-u4-02 .lbl{fill:#858D9C}html.dark #dsfig-u4-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-02 .ah{fill:#B1B7C3}html.dark #dsfig-u4-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M53.4,155.6 L154.2,54.8" marker-end="url(#ah9)"/><path class="e" d="M58,163 L149.1,132.6" marker-end="url(#ah9)"/><path class="e" d="M58,175 L149.1,205.4" marker-end="url(#ah9)"/><path class="e" d="M53.4,182.4 L154.2,283.2" marker-end="url(#ah9)"/><path class="e" d="M186,48.5 L322.2,116.6" marker-end="url(#ah9)"/><path class="e" d="M188,126 L320,126" marker-end="url(#ah9)"/><path class="e" d="M186,203.5 L322.2,135.4" marker-end="url(#ah9)"/><path class="e" d="M358,134.5 L494.2,202.6" marker-end="url(#ah9)"/><path class="e" d="M187.4,293.4 L492.6,217.1" marker-end="url(#ah9)"/><circle class="n" cx="40" cy="169" r="18"/><text class="t" x="40" y="169" dy=".35em" text-anchor="middle">Xh</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">F</text><circle class="n" cx="169" cy="126" r="18"/><text class="t" x="169" y="126" dy=".35em" text-anchor="middle">I</text><circle class="n" cx="169" cy="212" r="18"/><text class="t" x="169" y="212" dy=".35em" text-anchor="middle">G</text><circle class="n" cx="169" cy="298" r="18"/><text class="t" x="169" y="298" dy=".35em" text-anchor="middle">O</text><circle class="n" cx="341" cy="126" r="18"/><text class="t" x="341" y="126" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="513" cy="212" r="18"/><text class="t" x="513" y="212" dy=".35em" text-anchor="middle">H</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">LSTM cell. Xh = x_t with h_(t-1), F = forget gate, I = input gate, G = candidate memory, O = output gate, C = cell state c_t, H = hidden state h_t</figcaption></figure>
Key points.
- A simple RNN suffers from vanishing gradients, because repeated multiplication by $W_h$ and $f'$ shrinks the gradient, so it forgets distant inputs; the LSTM cell state gives gradients a nearly additive path.
- Forget gate: $f_t = \sigma(W_f[h_{t-1}, x_t] + b_f)$ decides how much of $c_{t-1}$ to keep.
- Input gate: $i_t = \sigma(W_i[h_{t-1}, x_t] + b_i)$ decides how much new information to write, with candidate $\tilde{c}_t = \tanh(W_c[h_{t-1}, x_t] + b_c)$.
- Cell update: $c_t = f_t \odot c_{t-1} + i_t \odot \tilde{c}_t$, an elementwise sum that carries information across many steps.
- Output gate: $o_t = \sigma(W_o[h_{t-1}, x_t] + b_o)$ and hidden state $h_t = o_t \odot \tanh(c_t)$.
- Sigmoid gates output values in $(0,1)$ and act as soft switches, with 0 blocking and 1 passing fully.
- Applications include language modelling, machine translation, speech recognition, handwriting recognition and time-series forecasting.
Answer frame. Open with the definition and the vanishing-gradient need; draw the cell with the cell-state line on top and three gates; then develop points 2-5 with one equation each; close with the applications and "the additive cell state lets gradients flow over long spans".
Asked: [7 marks] (May 2024, Jun 2025) What is Long Short Term Memory in Neural Network? Explain.
Other Gated RNNs
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Gated RNNs use learned gates to control information flow through time, which eases vanishing gradients and lets the network keep long-term dependencies; the LSTM and the GRU are the main types.</mark>
Key points.
- The LSTM uses input, forget and output gates with a separate cell state, as described above.
- The Gated Recurrent Unit (GRU) merges cell and hidden state and uses only two gates, so it has fewer parameters and trains faster.
- Update gate: $z_t = \sigma(W_z[h_{t-1}, x_t])$ and reset gate: $r_t = \sigma(W_r[h_{t-1}, x_t])$.
- Candidate $\tilde{h}_t = \tanh(W[r_t \odot h_{t-1}, x_t])$ and new state $h_t = (1-z_t)\odot h_{t-1} + z_t \odot \tilde{h}_t$.
- Uses include translation, speech recognition and text and music modelling.
Asked: [6 marks] (May 2023) Write about Gated RNNs.
Last-minute revision
- A BRNN has a forward and a backward hidden layer and uses past and future context.
- A BRNN needs the whole sequence, so it cannot do online prediction.
- Deep RNN: stacked recurrent layers, $h_t^{(l)}$ takes $h_t^{(l-1)}$ as input.
- Recursive NN: shared weights over a tree, $p = f(W[c_1;c_2]+b)$.
- LSTM has three gates: forget, input, output, plus cell state $c_t$.
- $c_t = f_t \odot c_{t-1} + i_t \odot \tilde{c}_t$ and $h_t = o_t \odot \tanh(c_t)$.
- Sigmoid gates give values in (0,1); the candidate uses tanh.
- GRU has two gates: update and reset, and no separate cell state.
- LSTM fixes vanishing gradient through the additive cell-state path.
- Unfolding in time turns the RNN loop into a chain with shared weights.
Memory hooks
- BRNN: "look back and look ahead", two layers meeting at the output.
- LSTM gates FIO: Forget, Input, Output.
- GRU = LSTM minus one gate, two gates: update and reset.
- Recursive is a tree, recurrent is a chain.
- Deep RNN: stacking floors, low floors short-term, high floors long-term.
Coverage checklist
- Bidirectional RNNs: May 2023 (7), May 2024 (7).
- Deep Recurrent Networks: May 2023 (7).
- Recursive Neural Networks: not asked recently, taught by definition and key points.
- The Long Short-Term Memory: May 2024, Jun 2025 (7).
- Other Gated RNNs: May 2023 (6).