Skip to content
CS-702 (B) · Deep & Reinforcement Learning/Quick Revision Short Notes

Deep & Reinforcement Learning (CS-702 (B)) - Unit 5 Short Notes

How unit 5 is examined

This unit covers batch and deep value-based RL, policy gradients, actor-critic, imitation and inverse RL; the marks sit in Actor-Critic, Recent Trends (GAIL) and Policy Gradient for full RL.

Fitted Q

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Fitted Q iteration (FQI) is batch RL that repeatedly regresses a Q-function onto Bellman targets computed from a fixed dataset of transitions.</mark>

Key points.

  1. The dataset is a fixed set of tuples $(s,a,r,s')$ collected beforehand, so no new interaction is needed.
  2. Each round builds targets $y_i = r_i + \gamma \max_{a'} Q_k(s'_i,a')$.
  3. A supervised regressor (trees, neural net) is fitted to give $Q_{k+1}(s_i,a_i)\approx y_i$, and the loop repeats until $Q$ stabilises.
  4. It is sample-efficient and off-policy, but errors from the regressor can compound.

Deep Q-Learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Deep Q-learning (DQN) approximates $Q(s,a;\theta)$ with a deep network trained by Q-learning on minibatches drawn from a replay buffer, using a frozen target network.</mark>

Key points.

  1. Loss: $L(\theta)=\mathbb{E}\big[(r+\gamma\max_{a'}Q(s',a';\theta^-)-Q(s,a;\theta))^2\big]$.
  2. Experience replay stores past transitions and samples them randomly, which breaks the correlation between consecutive samples.
  3. The target network $\theta^-$ is copied from $\theta$ every few thousand steps, which stops the target from chasing itself.
  4. Atari DQN used raw pixels and $\epsilon$-greedy exploration.

Advanced Q-learning algorithms

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Advanced Q-learning algorithms are fixes to DQN's overestimation, sample inefficiency and instability, combined in Rainbow.</mark>

Key points.

  1. Double DQN chooses the action with the online network and evaluates it with the target network, which reduces overestimation.
  2. Dueling DQN splits the network into a state value $V(s)$ and an advantage $A(s,a)$, with $Q=V+A-\text{mean}(A)$.
  3. Prioritised replay samples transitions with large TD error more often.
  4. Rainbow combines these with multi-step returns, noisy nets and distributional RL.

Learning policies by imitating optimal controllers

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Imitation learning trains a policy to copy an expert or optimal controller from demonstrations, without a reward function.</mark>

Key points.

  1. Behavioural cloning treats demonstrations $(s,a^*)$ as supervised data and minimises $\|\pi_\theta(s)-a^*\|^2$.
  2. Its weakness is distribution shift: small errors take the agent to states the expert never showed, and errors compound.
  3. DAgger fixes this by letting the learner drive while the expert relabels the visited states.
  4. Guided policy search uses optimal controllers (e.g. LQR) to generate the data the network imitates.

DQN & Policy Gradient

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>DQN is value-based and off-policy, while policy gradient is policy-based and on-policy; hybrids such as actor-critic combine the two.</mark>

Key points.

  1. DQN learns $Q$ and acts greedily, so it suits discrete actions; policy gradient outputs $\pi_\theta(a|s)$ and handles continuous actions.
  2. DQN is sample-efficient because of replay, but policy gradient is more stable and directly optimises return.
  3. Hybrids include actor-critic, DDPG and SAC.

Policy Gradient Algorithms for Full RL

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Policy gradient optimises a parametrised policy $\pi_\theta(a|s)$ directly by gradient ascent on the expected return $J(\theta)$.</mark>

Key points.

  1. The full RL setting is an MDP in which the agent sees states and delayed rewards and must learn from sampled trajectories $\tau$ without knowing the model.
  2. The objective is $J(\theta)=\mathbb{E}_{\tau\sim\pi_\theta}[R(\tau)]=\sum_\tau P(\tau;\theta)R(\tau)$.
  3. Trajectories are sampled by Monte-Carlo rollouts of the current policy, so the expectation is replaced by an average.
  4. A baseline $b(s)$ (usually $V(s)$) is subtracted from the return; it lowers variance and leaves the gradient unbiased.

Derivation. Use the log-derivative trick $\nabla_\theta P=P\,\nabla_\theta\log P$; the unknown dynamics terms do not depend on $\theta$ and drop out:

$$\nabla_\theta J=\sum_\tau P(\tau;\theta)\nabla_\theta\log P(\tau;\theta)R(\tau)=\mathbb{E}_\tau\Big[\sum_{t}\nabla_\theta\log\pi_\theta(a_t|s_t)\,R(\tau)\Big]$$

Steps (REINFORCE).

Step 1: Initialise policy parameters theta.
Step 2: Sample N trajectories by running pi_theta.
Step 3: Compute G_t = sum of discounted rewards from step t, minus baseline b(s_t).
Step 4: Estimate grad J = (1/N) sum_i sum_t grad log pi(a_t|s_t) * (G_t - b).
Step 5: Update theta <- theta + alpha * grad J; repeat.

Answer frame. Open with the definition of a parametrised policy and $J(\theta)$; write the log-derivative derivation; give the REINFORCE steps; close with the baseline and the note that variance is high, which motivates actor-critic.

Asked: [7 marks] (Dec 2020, Nov 2023) Explain Policy Gradient Algorithm for Full RL.

Hierarchical RL

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Hierarchical RL decomposes a task into sub-policies (options or skills) chosen by a higher-level policy, so long tasks become short decisions.</mark>

Key points.

  1. An option is a triple $(I,\pi,\beta)$: initiation set, intra-option policy and termination condition.
  2. The high-level policy picks an option and the low-level policy runs until it terminates (a semi-MDP).
  3. It speeds up learning on sparse-reward tasks and allows skills to be reused.

POMDPs

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A POMDP is an MDP in which the agent sees only partial observations $o$ rather than the true state, so it must act on a belief.</mark>

Key points.

  1. It is the tuple $(S,A,T,R,\Omega,O,\gamma)$, with observation space $\Omega$ and observation function $O$.
  2. The belief $b(s)$ is a probability distribution over states, updated by Bayes rule after each action and observation.
  3. Deep RL handles it with recurrent policies (LSTM) that carry a memory of the history.

Actor-Critic Method

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Actor-critic combines a policy network (actor) that chooses actions with a value network (critic) that scores them, so the critic's TD error guides the actor's policy-gradient update.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-01" viewBox="0 0 252 252" width="252" height="252" role="img" aria-label="Actor-critic loop. Act = actor (policy), Cri = critic (value), Env = environment"><style>#dsfig-u5-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-01 .t{fill:#16181D;font-weight:500}#dsfig-u5-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-01 .dot{fill:#16181D}#dsfig-u5-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-01 .ah{fill:#454C5A}#dsfig-u5-01 .ah.hi{fill:#2340B8}#dsfig-u5-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-01 .e{stroke:#B1B7C3}html.dark #dsfig-u5-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-01 .t{fill:#E6E8ED}html.dark #dsfig-u5-01 .t.inv{fill:#0F1115}html.dark #dsfig-u5-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-01 .dot{fill:#E6E8ED}html.dark #dsfig-u5-01 .ann{fill:#8FA3FF}html.dark #dsfig-u5-01 .lbl{fill:#858D9C}html.dark #dsfig-u5-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-01 .ah{fill:#B1B7C3}html.dark #dsfig-u5-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M57,117.5 L193.2,49.4" marker-end="url(#ah10)"/><path class="e" d="M212,59 L212,191" marker-end="url(#ah10)"/><path class="e" d="M195,203.5 L58.8,135.4" marker-end="url(#ah10)"/><g class="wl"><rect x="98.9" y="74" width="54.3" height="18" rx="9"/><text class="t" x="126" y="83" dy=".35em" text-anchor="middle">action</text></g><g class="wl"><rect x="163.7" y="117" width="96.6" height="18" rx="9"/><text class="t" x="212" y="126" dy=".35em" text-anchor="middle">state,reward</text></g><g class="wl"><rect x="91.7" y="160" width="68.7" height="18" rx="9"/><text class="t" x="126" y="169" dy=".35em" text-anchor="middle">TD_error</text></g><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">Act</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">Env</text><circle class="n" cx="212" cy="212" r="18"/><text class="t" x="212" y="212" dy=".35em" text-anchor="middle">Cri</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Actor-critic loop. Act = actor (policy), Cri = critic (value), Env = environment</figcaption></figure>

Formula. TD error $\delta_t=r_{t+1}+\gamma V_w(s_{t+1})-V_w(s_t)$ estimates the advantage. Critic: $w\leftarrow w+\alpha_w\,\delta_t\nabla_w V_w(s_t)$. Actor: $\theta\leftarrow\theta+\alpha_\theta\,\delta_t\nabla_\theta\log\pi_\theta(a_t|s_t)$.

Key points.

  1. The actor $\pi_\theta(a|s)$ is the policy; it selects actions and is trained by policy gradient.
  2. The critic $V_w(s)$ (or $Q_w$) evaluates the actor's states and is trained by temporal-difference learning.
  3. The advantage $A(s,a)=Q(s,a)-V(s)$ tells how much better an action is than average, and $\delta_t$ is its one-step estimate.
  4. Replacing the Monte-Carlo return with the critic's bootstrapped estimate lowers variance, at the cost of some bias.
  5. Updates happen every step, so it learns online without waiting for episodes to end.
  6. It handles continuous actions naturally, which Q-learning and DQN cannot.
  7. Variants are A2C/A3C (parallel workers), DDPG, PPO and SAC.
Aspect REINFORCE Q-learning / DQN Actor-critic
Learns Policy only Value only Policy and value
Update End of episode Every step Every step
Variance High Low Low to medium
Bias None From bootstrapping Some, from critic
Continuous actions Yes Hard (needs max) Yes
Stability Noisy Can diverge More stable

Other parts of Q1. Face recognition: a CNN maps a face to an embedding, trained with triplet or softmax loss so the same person's faces lie close; used in phone unlock, attendance and security. Group Normalization: channels are split into groups and normalised within each group per sample, so it is independent of batch size and works with small batches; used in detection and segmentation.

Answer frame. Open with the definition; draw the actor-critic loop; state the two roles, the TD error and advantage; give the comparison table; close with the advantages (lower variance, continuous control, stability, online updates). For the short note, write points 1-6 in about ten lines.

Pitfall: Do not say actor-critic removes bias; it trades variance for critic bias.

Asked: [14 marks] (Nov 2023) Write short notes: a) Face Recognition Application b) Actor-Critic Method c) Group Normalization Asked: [14 marks] (Jun 2025) Compare the Actor-Critic method with traditional reinforcement learning techniques. What advantages does it provide?

Inverse reinforcement learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Inverse RL infers the reward function that explains an expert's demonstrated behaviour.</mark>

Key points.

  1. It assumes the expert is (near) optimal for some unknown reward $R(s)$, often linear: $R=w^\top\phi(s)$.
  2. It matches the expert's feature expectations $\mu(\pi)=\mathbb{E}[\sum_t\gamma^t\phi(s_t)]$.
  3. The problem is ill-posed, since many rewards explain the same behaviour, so extra criteria are needed.
  4. The learned reward generalises better than copied actions.

Maximum Entropy Deep Inverse Reinforcement Learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>MaxEnt IRL picks, among all reward-consistent trajectory distributions, the one with maximum entropy, with a deep network representing the reward.</mark>

Key points.

  1. Trajectory probability is $P(\tau)\propto\exp(R_\theta(\tau))$, so better trajectories are exponentially more likely.
  2. This resolves reward ambiguity without adding bias beyond matching feature expectations.
  3. The gradient is expert feature counts minus the learner's expected feature counts.
  4. A deep network for $R_\theta$ replaces the linear reward and captures nonlinear structure.

Generative Adversarial Imitation Learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>GAIL trains a policy (generator) so that a discriminator cannot tell its trajectories from the expert's.</mark>

Key points.

  1. The discriminator $D(s,a)$ is trained to separate expert pairs from policy pairs.
  2. The policy is trained by RL (e.g. TRPO) with reward $-\log D(s,a)$, so it is rewarded for looking like the expert.
  3. It learns a policy directly, with no explicit reward recovery, and needs fewer demonstrations than cloning.

Recent Trends in RL Architectures

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Recent trends in RL architectures scale RL with deep networks and add distributional, model-based, offline and adversarial-imitation ideas to improve sample efficiency and stability.</mark>

Key points.

  1. Deep RL (DQN, A3C, PPO) uses neural networks for value and policy from raw inputs.
  2. Distributional RL (C51, QR-DQN) learns the whole return distribution rather than only its mean, which gives richer learning signals.
  3. Model-based RL (Dyna, MBPO, Dreamer) learns a dynamics model and plans or trains in imagination, which greatly cuts real samples.
  4. Offline RL learns from a fixed dataset; hierarchical RL and transformer-based RL (Decision Transformer) treat RL as sequence modelling.
  5. GAIL frames imitation as a GAN: the generator is the policy, the discriminator separates policy from expert trajectories, and the discriminator output is the reward; it needs no reward function or expert queries.
  6. Applications are robotics, games (AlphaGo, Atari), autonomous driving and recommendation.
  7. Challenges are sample inefficiency, training instability, reward design, safety and sim-to-real transfer.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-02" viewBox="0 0 424 252" width="424" height="252" role="img" aria-label="GAIL. Exp = expert demos, Pol = policy (generator), Dis = discriminator, Rew = reward"><style>#dsfig-u5-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-02 .t{fill:#16181D;font-weight:500}#dsfig-u5-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-02 .dot{fill:#16181D}#dsfig-u5-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-02 .ah{fill:#454C5A}#dsfig-u5-02 .ah.hi{fill:#2340B8}#dsfig-u5-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-02 .e{stroke:#B1B7C3}html.dark #dsfig-u5-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-02 .t{fill:#E6E8ED}html.dark #dsfig-u5-02 .t.inv{fill:#0F1115}html.dark #dsfig-u5-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-02 .dot{fill:#E6E8ED}html.dark #dsfig-u5-02 .ann{fill:#8FA3FF}html.dark #dsfig-u5-02 .lbl{fill:#858D9C}html.dark #dsfig-u5-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-02 .ah{fill:#B1B7C3}html.dark #dsfig-u5-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah11" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh11" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M57,48.5 L193.2,116.6" marker-end="url(#ah11)"/><path class="e" d="M57,203.5 L193.2,135.4" marker-end="url(#ah11)"/><path class="e" d="M231,126 L363,126" marker-end="url(#ah11)"/><path class="e" d="M365.6,130.6 L60.4,206.9" marker-end="url(#ah11)"/><g class="wl"><rect x="98.9" y="74" width="54.3" height="18" rx="9"/><text class="t" x="126" y="83" dy=".35em" text-anchor="middle">expert</text></g><g class="wl"><rect x="95.3" y="160" width="61.5" height="18" rx="9"/><text class="t" x="126" y="169" dy=".35em" text-anchor="middle">samples</text></g><g class="wl"><rect x="270.9" y="117" width="54.3" height="18" rx="9"/><text class="t" x="298" y="126" dy=".35em" text-anchor="middle">D(s,a)</text></g><g class="wl"><rect x="184.9" y="160" width="54.3" height="18" rx="9"/><text class="t" x="212" y="169" dy=".35em" text-anchor="middle">update</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Exp</text><circle class="n" cx="40" cy="212" r="18"/><text class="t" x="40" y="212" dy=".35em" text-anchor="middle">Pol</text><circle class="n" cx="212" cy="126" r="18"/><text class="t" x="212" y="126" dy=".35em" text-anchor="middle">Dis</text><circle class="n" cx="384" cy="126" r="18"/><text class="t" x="384" y="126" dy=".35em" text-anchor="middle">Rew</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">GAIL. Exp = expert demos, Pol = policy (generator), Dis = discriminator, Rew = reward</figcaption></figure>

Answer frame. Open with one line on why architectures evolved; list the trends 1-4 in one line each; draw the GAIL diagram and explain it; close with applications and challenges.

Asked: [14 marks] (Jun 2025) Discuss recent trends in reinforcement learning architectures, including Generative Adversarial Imitation Learning.

Last-minute revision

  • FQI: regress $Q_{k+1}$ onto $r+\gamma\max Q_k(s',a')$ on a fixed batch.
  • DQN needs replay buffer and target network; loss is squared TD error.
  • Double DQN cuts overestimation; Dueling splits $V$ and $A$; Rainbow combines all.
  • Behavioural cloning suffers distribution shift; DAgger fixes it.
  • $\nabla_\theta J=\mathbb{E}[\sum_t\nabla_\theta\log\pi_\theta(a_t|s_t)R(\tau)]$.
  • A baseline lowers variance without adding bias.
  • Actor = policy, critic = value; TD error $\delta=r+\gamma V(s')-V(s)$ is the advantage estimate.
  • Option = $(I,\pi,\beta)$; belief = distribution over states in a POMDP.
  • MaxEnt IRL: $P(\tau)\propto\exp R(\tau)$.
  • GAIL reward is $-\log D(s,a)$.

Memory hooks

  • "Actor acts, critic criticises" gives the two roles.
  • Replay and target network are the two DQN stabilisers: "shuffle and freeze".
  • Double DQN: one network picks, the other judges.
  • GAIL is a GAN where the policy plays the forger and the expert data is the real thing.
  • REINFORCE = high variance, so add a baseline, which becomes a critic.

Coverage checklist

  • Fitted Q: no past questions.
  • Deep Q-Learning: no past questions.
  • Advanced Q-learning algorithms: no past questions.
  • Learning policies by imitating optimal controllers: no past questions.
  • DQN & Policy Gradient: no past questions.
  • Policy Gradient Algorithms for Full RL: Dec 2020, Nov 2023 (7 marks).
  • Hierarchical RL: no past questions.
  • POMDPs: no past questions.
  • Actor-Critic Method: Nov 2023 short notes, Jun 2025 comparison (14 marks each).
  • Inverse reinforcement learning: no past questions.
  • Maximum Entropy Deep Inverse Reinforcement Learning: no past questions.
  • Generative Adversarial Imitation Learning: covered under Recent Trends (Jun 2025).
  • Recent Trends in RL Architectures: Jun 2025 (14 marks).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in