How unit 5 is examined
This unit covers batch and deep value-based RL, policy gradients, actor-critic, imitation and inverse RL; the marks sit in Actor-Critic, Recent Trends (GAIL) and Policy Gradient for full RL.
Fitted Q
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Fitted Q iteration (FQI) is batch RL that repeatedly regresses a Q-function onto Bellman targets computed from a fixed dataset of transitions.</mark>
Key points.
- The dataset is a fixed set of tuples $(s,a,r,s')$ collected beforehand, so no new interaction is needed.
- Each round builds targets $y_i = r_i + \gamma \max_{a'} Q_k(s'_i,a')$.
- A supervised regressor (trees, neural net) is fitted to give $Q_{k+1}(s_i,a_i)\approx y_i$, and the loop repeats until $Q$ stabilises.
- It is sample-efficient and off-policy, but errors from the regressor can compound.
Deep Q-Learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Deep Q-learning (DQN) approximates $Q(s,a;\theta)$ with a deep network trained by Q-learning on minibatches drawn from a replay buffer, using a frozen target network.</mark>
Key points.
- Loss: $L(\theta)=\mathbb{E}\big[(r+\gamma\max_{a'}Q(s',a';\theta^-)-Q(s,a;\theta))^2\big]$.
- Experience replay stores past transitions and samples them randomly, which breaks the correlation between consecutive samples.
- The target network $\theta^-$ is copied from $\theta$ every few thousand steps, which stops the target from chasing itself.
- Atari DQN used raw pixels and $\epsilon$-greedy exploration.
Advanced Q-learning algorithms
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Advanced Q-learning algorithms are fixes to DQN's overestimation, sample inefficiency and instability, combined in Rainbow.</mark>
Key points.
- Double DQN chooses the action with the online network and evaluates it with the target network, which reduces overestimation.
- Dueling DQN splits the network into a state value $V(s)$ and an advantage $A(s,a)$, with $Q=V+A-\text{mean}(A)$.
- Prioritised replay samples transitions with large TD error more often.
- Rainbow combines these with multi-step returns, noisy nets and distributional RL.
Learning policies by imitating optimal controllers
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Imitation learning trains a policy to copy an expert or optimal controller from demonstrations, without a reward function.</mark>
Key points.
- Behavioural cloning treats demonstrations $(s,a^*)$ as supervised data and minimises $\|\pi_\theta(s)-a^*\|^2$.
- Its weakness is distribution shift: small errors take the agent to states the expert never showed, and errors compound.
- DAgger fixes this by letting the learner drive while the expert relabels the visited states.
- Guided policy search uses optimal controllers (e.g. LQR) to generate the data the network imitates.
DQN & Policy Gradient
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>DQN is value-based and off-policy, while policy gradient is policy-based and on-policy; hybrids such as actor-critic combine the two.</mark>
Key points.
- DQN learns $Q$ and acts greedily, so it suits discrete actions; policy gradient outputs $\pi_\theta(a|s)$ and handles continuous actions.
- DQN is sample-efficient because of replay, but policy gradient is more stable and directly optimises return.
- Hybrids include actor-critic, DDPG and SAC.
Policy Gradient Algorithms for Full RL
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Policy gradient optimises a parametrised policy $\pi_\theta(a|s)$ directly by gradient ascent on the expected return $J(\theta)$.</mark>
Key points.
- The full RL setting is an MDP in which the agent sees states and delayed rewards and must learn from sampled trajectories $\tau$ without knowing the model.
- The objective is $J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}[R(\tau)]=\sum_\tau P(\tau;\theta)R(\tau)$.
- Trajectories are sampled by Monte-Carlo rollouts of the current policy, so the expectation is replaced by an average.
- A baseline $b(s)$ (usually $V(s)$) is subtracted from the return; it lowers variance and leaves the gradient unbiased.
Derivation. Use the log-derivative trick $\nabla_\theta P=P\,\nabla_\theta\log P$; the unknown dynamics terms do not depend on $\theta$ and drop out:
$$\nabla_\theta J=\sum_\tau P(\tau;\theta)\nabla_\theta\log P(\tau;\theta)R(\tau)=\mathbb{E}_\tau\Big[\sum_{t}\nabla_\theta\log\pi_\theta(a_t|s_t)\,R(\tau)\Big]$$
Steps (REINFORCE).
Step 1: Initialise policy parameters theta.
Step 2: Sample N trajectories by running pi_theta.
Step 3: Compute G_t = sum of discounted rewards from step t, minus baseline b(s_t).
Step 4: Estimate grad J = (1/N) sum_i sum_t grad log pi(a_t|s_t) * (G_t - b).
Step 5: Update theta <- theta + alpha * grad J; repeat.
Answer frame. Open with the definition of a parametrised policy and $J(\theta)$; write the log-derivative derivation; give the REINFORCE steps; close with the baseline and the note that variance is high, which motivates actor-critic.
Asked: [7 marks] (Dec 2020, Nov 2023) Explain Policy Gradient Algorithm for Full RL.
Hierarchical RL
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Hierarchical RL decomposes a task into sub-policies (options or skills) chosen by a higher-level policy, so long tasks become short decisions.</mark>
Key points.
- An option is a triple $(I,\pi,\beta)$: initiation set, intra-option policy and termination condition.
- The high-level policy picks an option and the low-level policy runs until it terminates (a semi-MDP).
- It speeds up learning on sparse-reward tasks and allows skills to be reused.
POMDPs
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A POMDP is an MDP in which the agent sees only partial observations $o$ rather than the true state, so it must act on a belief.</mark>
Key points.
- It is the tuple $(S,A,T,R,\Omega,O,\gamma)$, with observation space $\Omega$ and observation function $O$.
- The belief $b(s)$ is a probability distribution over states, updated by Bayes rule after each action and observation.
- Deep RL handles it with recurrent policies (LSTM) that carry a memory of the history.
Actor-Critic Method
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Actor-critic combines a policy network (actor) that chooses actions with a value network (critic) that scores them, so the critic's TD error guides the actor's policy-gradient update.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-01" viewBox="0 0 252 252" width="252" height="252" role="img" aria-label="Actor-critic loop. Act = actor (policy), Cri = critic (value), Env = environment"><style>#dsfig-u5-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-01 .t{fill:#16181D;font-weight:500}#dsfig-u5-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-01 .dot{fill:#16181D}#dsfig-u5-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-01 .ah{fill:#454C5A}#dsfig-u5-01 .ah.hi{fill:#2340B8}#dsfig-u5-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-01 .e{stroke:#B1B7C3}html.dark #dsfig-u5-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-01 .t{fill:#E6E8ED}html.dark #dsfig-u5-01 .t.inv{fill:#0F1115}html.dark #dsfig-u5-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-01 .dot{fill:#E6E8ED}html.dark #dsfig-u5-01 .ann{fill:#8FA3FF}html.dark #dsfig-u5-01 .lbl{fill:#858D9C}html.dark #dsfig-u5-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-01 .ah{fill:#B1B7C3}html.dark #dsfig-u5-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M57,117.5 L193.2,49.4" marker-end="url(#ah10)"/><path class="e" d="M212,59 L212,191" marker-end="url(#ah10)"/><path class="e" d="M195,203.5 L58.8,135.4" marker-end="url(#ah10)"/><g class="wl"><rect x="98.9" y="74" width="54.3" height="18" rx="9"/><text class="t" x="126" y="83" dy=".35em" text-anchor="middle">action</text></g><g class="wl"><rect x="163.7" y="117" width="96.6" height="18" rx="9"/><text class="t" x="212" y="126" dy=".35em" text-anchor="middle">state,reward</text></g><g class="wl"><rect x="91.7" y="160" width="68.7" height="18" rx="9"/><text class="t" x="126" y="169" dy=".35em" text-anchor="middle">TD_error</text></g><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">Act</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">Env</text><circle class="n" cx="212" cy="212" r="18"/><text class="t" x="212" y="212" dy=".35em" text-anchor="middle">Cri</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Actor-critic loop. Act = actor (policy), Cri = critic (value), Env = environment</figcaption></figure>
Formula. TD error $\delta_t=r_{t+1}+\gamma V_w(s_{t+1})-V_w(s_t)$ estimates the advantage. Critic: $w\leftarrow w+\alpha_w\,\delta_t\nabla_w V_w(s_t)$. Actor: $\theta\leftarrow\theta+\alpha_\theta\,\delta_t\nabla_\theta\log\pi_\theta(a_t|s_t)$.
Key points.
- The actor $\pi_\theta(a|s)$ is the policy; it selects actions and is trained by policy gradient.
- The critic $V_w(s)$ (or $Q_w$) evaluates the actor's states and is trained by temporal-difference learning.
- The advantage $A(s,a)=Q(s,a)-V(s)$ tells how much better an action is than average, and $\delta_t$ is its one-step estimate.
- Replacing the Monte-Carlo return with the critic's bootstrapped estimate lowers variance, at the cost of some bias.
- Updates happen every step, so it learns online without waiting for episodes to end.
- It handles continuous actions naturally, which Q-learning and DQN cannot.
- Variants are A2C/A3C (parallel workers), DDPG, PPO and SAC.
| Aspect | REINFORCE | Q-learning / DQN | Actor-critic |
|---|---|---|---|
| Learns | Policy only | Value only | Policy and value |
| Update | End of episode | Every step | Every step |
| Variance | High | Low | Low to medium |
| Bias | None | From bootstrapping | Some, from critic |
| Continuous actions | Yes | Hard (needs max) | Yes |
| Stability | Noisy | Can diverge | More stable |
Other parts of Q1. Face recognition: a CNN maps a face to an embedding, trained with triplet or softmax loss so the same person's faces lie close; used in phone unlock, attendance and security. Group Normalization: channels are split into groups and normalised within each group per sample, so it is independent of batch size and works with small batches; used in detection and segmentation.
Answer frame. Open with the definition; draw the actor-critic loop; state the two roles, the TD error and advantage; give the comparison table; close with the advantages (lower variance, continuous control, stability, online updates). For the short note, write points 1-6 in about ten lines.
Pitfall: Do not say actor-critic removes bias; it trades variance for critic bias.
Asked: [14 marks] (Nov 2023) Write short notes: a) Face Recognition Application b) Actor-Critic Method c) Group Normalization Asked: [14 marks] (Jun 2025) Compare the Actor-Critic method with traditional reinforcement learning techniques. What advantages does it provide?
Inverse reinforcement learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Inverse RL infers the reward function that explains an expert's demonstrated behaviour.</mark>
Key points.
- It assumes the expert is (near) optimal for some unknown reward $R(s)$, often linear: $R=w^\top\phi(s)$.
- It matches the expert's feature expectations $\mu(\pi)=\mathbb{E}[\sum_t\gamma^t\phi(s_t)]$.
- The problem is ill-posed, since many rewards explain the same behaviour, so extra criteria are needed.
- The learned reward generalises better than copied actions.
Maximum Entropy Deep Inverse Reinforcement Learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>MaxEnt IRL picks, among all reward-consistent trajectory distributions, the one with maximum entropy, with a deep network representing the reward.</mark>
Key points.
- Trajectory probability is $P(\tau)\propto\exp(R_\theta(\tau))$, so better trajectories are exponentially more likely.
- This resolves reward ambiguity without adding bias beyond matching feature expectations.
- The gradient is expert feature counts minus the learner's expected feature counts.
- A deep network for $R_\theta$ replaces the linear reward and captures nonlinear structure.
Generative Adversarial Imitation Learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>GAIL trains a policy (generator) so that a discriminator cannot tell its trajectories from the expert's.</mark>
Key points.
- The discriminator $D(s,a)$ is trained to separate expert pairs from policy pairs.
- The policy is trained by RL (e.g. TRPO) with reward $-\log D(s,a)$, so it is rewarded for looking like the expert.
- It learns a policy directly, with no explicit reward recovery, and needs fewer demonstrations than cloning.
Recent Trends in RL Architectures
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Recent trends in RL architectures scale RL with deep networks and add distributional, model-based, offline and adversarial-imitation ideas to improve sample efficiency and stability.</mark>
Key points.
- Deep RL (DQN, A3C, PPO) uses neural networks for value and policy from raw inputs.
- Distributional RL (C51, QR-DQN) learns the whole return distribution rather than only its mean, which gives richer learning signals.
- Model-based RL (Dyna, MBPO, Dreamer) learns a dynamics model and plans or trains in imagination, which greatly cuts real samples.
- Offline RL learns from a fixed dataset; hierarchical RL and transformer-based RL (Decision Transformer) treat RL as sequence modelling.
- GAIL frames imitation as a GAN: the generator is the policy, the discriminator separates policy from expert trajectories, and the discriminator output is the reward; it needs no reward function or expert queries.
- Applications are robotics, games (AlphaGo, Atari), autonomous driving and recommendation.
- Challenges are sample inefficiency, training instability, reward design, safety and sim-to-real transfer.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-02" viewBox="0 0 424 252" width="424" height="252" role="img" aria-label="GAIL. Exp = expert demos, Pol = policy (generator), Dis = discriminator, Rew = reward"><style>#dsfig-u5-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-02 .t{fill:#16181D;font-weight:500}#dsfig-u5-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-02 .dot{fill:#16181D}#dsfig-u5-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-02 .ah{fill:#454C5A}#dsfig-u5-02 .ah.hi{fill:#2340B8}#dsfig-u5-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-02 .e{stroke:#B1B7C3}html.dark #dsfig-u5-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-02 .t{fill:#E6E8ED}html.dark #dsfig-u5-02 .t.inv{fill:#0F1115}html.dark #dsfig-u5-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-02 .dot{fill:#E6E8ED}html.dark #dsfig-u5-02 .ann{fill:#8FA3FF}html.dark #dsfig-u5-02 .lbl{fill:#858D9C}html.dark #dsfig-u5-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-02 .ah{fill:#B1B7C3}html.dark #dsfig-u5-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah11" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh11" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M57,48.5 L193.2,116.6" marker-end="url(#ah11)"/><path class="e" d="M57,203.5 L193.2,135.4" marker-end="url(#ah11)"/><path class="e" d="M231,126 L363,126" marker-end="url(#ah11)"/><path class="e" d="M365.6,130.6 L60.4,206.9" marker-end="url(#ah11)"/><g class="wl"><rect x="98.9" y="74" width="54.3" height="18" rx="9"/><text class="t" x="126" y="83" dy=".35em" text-anchor="middle">expert</text></g><g class="wl"><rect x="95.3" y="160" width="61.5" height="18" rx="9"/><text class="t" x="126" y="169" dy=".35em" text-anchor="middle">samples</text></g><g class="wl"><rect x="270.9" y="117" width="54.3" height="18" rx="9"/><text class="t" x="298" y="126" dy=".35em" text-anchor="middle">D(s,a)</text></g><g class="wl"><rect x="184.9" y="160" width="54.3" height="18" rx="9"/><text class="t" x="212" y="169" dy=".35em" text-anchor="middle">update</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Exp</text><circle class="n" cx="40" cy="212" r="18"/><text class="t" x="40" y="212" dy=".35em" text-anchor="middle">Pol</text><circle class="n" cx="212" cy="126" r="18"/><text class="t" x="212" y="126" dy=".35em" text-anchor="middle">Dis</text><circle class="n" cx="384" cy="126" r="18"/><text class="t" x="384" y="126" dy=".35em" text-anchor="middle">Rew</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">GAIL. Exp = expert demos, Pol = policy (generator), Dis = discriminator, Rew = reward</figcaption></figure>
Answer frame. Open with one line on why architectures evolved; list the trends 1-4 in one line each; draw the GAIL diagram and explain it; close with applications and challenges.
Asked: [14 marks] (Jun 2025) Discuss recent trends in reinforcement learning architectures, including Generative Adversarial Imitation Learning.
Last-minute revision
- FQI: regress $Q_{k+1}$ onto $r+\gamma\max Q_k(s',a')$ on a fixed batch.
- DQN needs replay buffer and target network; loss is squared TD error.
- Double DQN cuts overestimation; Dueling splits $V$ and $A$; Rainbow combines all.
- Behavioural cloning suffers distribution shift; DAgger fixes it.
- $\nabla_\theta J=\mathbb{E}[\sum_t\nabla_\theta\log\pi_\theta(a_t|s_t)R(\tau)]$.
- A baseline lowers variance without adding bias.
- Actor = policy, critic = value; TD error $\delta=r+\gamma V(s')-V(s)$ is the advantage estimate.
- Option = $(I,\pi,\beta)$; belief = distribution over states in a POMDP.
- MaxEnt IRL: $P(\tau)\propto\exp R(\tau)$.
- GAIL reward is $-\log D(s,a)$.
Memory hooks
- "Actor acts, critic criticises" gives the two roles.
- Replay and target network are the two DQN stabilisers: "shuffle and freeze".
- Double DQN: one network picks, the other judges.
- GAIL is a GAN where the policy plays the forger and the expert data is the real thing.
- REINFORCE = high variance, so add a baseline, which becomes a critic.
Coverage checklist
- Fitted Q: no past questions.
- Deep Q-Learning: no past questions.
- Advanced Q-learning algorithms: no past questions.
- Learning policies by imitating optimal controllers: no past questions.
- DQN & Policy Gradient: no past questions.
- Policy Gradient Algorithms for Full RL: Dec 2020, Nov 2023 (7 marks).
- Hierarchical RL: no past questions.
- POMDPs: no past questions.
- Actor-Critic Method: Nov 2023 short notes, Jun 2025 comparison (14 marks each).
- Inverse reinforcement learning: no past questions.
- Maximum Entropy Deep Inverse Reinforcement Learning: no past questions.
- Generative Adversarial Imitation Learning: covered under Recent Trends (Jun 2025).
- Recent Trends in RL Architectures: Jun 2025 (14 marks).