How unit 5 is examined
This unit covers actor-critic learning, inverse RL, MaxEnt deep IRL, GAIL and recent RL trends; only Inverse reinforcement learning has been asked (7 marks, Jun 2025), so learn it fully and know the other four as short definitions with formulas.
Actor-Critic Method
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Actor-critic is a policy-gradient method with two components: the actor, a parameterised policy that chooses actions, and the critic, a learned value function that evaluates those actions, so the critic's TD error tells the actor how to change its policy.</mark>
Key points.
- The actor is the policy $\pi_\theta(a|s)$ with parameters $\theta$; it decides which action is taken in each state.
- The critic is a value estimate $V_w(s)$ (or $Q_w(s,a)$) with parameters $w$; it does not act, it only scores how good the visited states are.
- The critic computes the TD error $\delta = r + \gamma V_w(s') - V_w(s)$, which is positive when the outcome was better than expected and negative when it was worse.
- The actor is updated in the direction that makes actions with positive $\delta$ more likely, and the critic is updated by TD learning to reduce $\delta^2$.
- Pure policy gradient (REINFORCE) uses the full Monte-Carlo return, which has high variance; replacing it by the critic's TD error lowers the variance and allows learning at every step, not only at episode end.
- The price is bias, because the critic's estimate is itself imperfect and learned at the same time as the actor.
- A2C (advantage actor-critic) uses the advantage $A(s,a)=Q(s,a)-V(s)$ and runs synchronously; A3C runs several workers asynchronously, each with its own copy of the environment, and updates shared parameters.
Formula.
$$\delta_t = r_{t+1} + \gamma V_w(s_{t+1}) - V_w(s_t)$$
$$\theta \leftarrow \theta + \alpha_\theta\,\delta_t\,\nabla_\theta \log \pi_\theta(a_t|s_t), \qquad w \leftarrow w + \alpha_w\,\delta_t\,\nabla_w V_w(s_t)$$
Diagram.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-01" viewBox="0 0 338 252" width="338" height="252" role="img" aria-label="Actor-critic loop. Act = actor (policy), Env = environment, Cri = critic (value function). The actor acts, the critic turns the reward into a TD error and feeds it back to the actor."><style>#dsfig-u5-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-01 .t{fill:#16181D;font-weight:500}#dsfig-u5-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-01 .dot{fill:#16181D}#dsfig-u5-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-01 .ah{fill:#454C5A}#dsfig-u5-01 .ah.hi{fill:#2340B8}#dsfig-u5-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-01 .e{stroke:#B1B7C3}html.dark #dsfig-u5-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-01 .t{fill:#E6E8ED}html.dark #dsfig-u5-01 .t.inv{fill:#0F1115}html.dark #dsfig-u5-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-01 .dot{fill:#E6E8ED}html.dark #dsfig-u5-01 .ann{fill:#8FA3FF}html.dark #dsfig-u5-01 .lbl{fill:#858D9C}html.dark #dsfig-u5-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-01 .ah{fill:#B1B7C3}html.dark #dsfig-u5-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L277,40" marker-end="url(#ah4)"/><path class="e" d="M286.6,55.2 L181.6,195.2" marker-end="url(#ah4)"/><path class="e" d="M157.6,196.8 L52.6,56.8" marker-end="url(#ah4)"/><g class="wl"><rect x="141.9" y="31" width="54.3" height="18" rx="9"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">action</text></g><g class="wl"><rect x="185.2" y="117" width="96.6" height="18" rx="9"/><text class="t" x="233.5" y="126" dy=".35em" text-anchor="middle">state,reward</text></g><g class="wl"><rect x="70.2" y="117" width="68.7" height="18" rx="9"/><text class="t" x="104.5" y="126" dy=".35em" text-anchor="middle">TD-error</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Act</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Env</text><circle class="n" cx="169" cy="212" r="18"/><text class="t" x="169" y="212" dy=".35em" text-anchor="middle">Cri</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Actor-critic loop. Act = actor (policy), Env = environment, Cri = critic (value function). The actor acts, the critic turns the reward into a TD error and feeds it back to the actor.</figcaption></figure>
Answer frame. Open with the definition of actor and critic; draw the loop of actor, environment and critic; then develop points 1-5 in order; close with the update formulas and the sentence that the critic reduces variance while the actor keeps the policy directly.
Inverse reinforcement learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Inverse reinforcement learning (IRL) is the problem of inferring the reward function that an expert is optimising, given the expert's observed behaviour (demonstrations) in an environment, instead of being given the reward as in ordinary RL.</mark>
Key points.
- In forward RL the reward $R$ is given and the agent learns an optimal policy $\pi^*$; in IRL the expert policy or trajectories are given and the aim is to recover the reward $R$ that makes them optimal.
- The input is a set of expert trajectories $\{(s_0,a_0,s_1,a_1,\dots)\}$ together with the MDP dynamics; the output is a reward function, usually $R(s)=\theta^\top\phi(s)$ with state features $\phi$.
- The reward is the most compact and transferable description of a task, so it is often easier to demonstrate a task than to write a reward by hand, and a learned reward generalises to new situations better than copied actions.
- The problem is ill-posed because many reward functions (including $R=0$) make the same behaviour optimal, so a method must add a preference to pick one reward.
- Feature matching: the learned policy should have the same expected feature counts $\mu(\pi)=\mathbb{E}\big[\sum_t\gamma^t\phi(s_t)\big]$ as the expert, $\mu(\pi)=\mu_E$.
- Methods: apprenticeship learning (Abbeel and Ng) with feature matching, max-margin IRL, Bayesian IRL and maximum-entropy IRL, which handles imperfect experts.
- A typical algorithm loops: guess a reward, solve the forward RL problem for the optimal policy, compare its behaviour with the expert's, and update the reward until they match.
- Applications: autonomous driving styles, robot manipulation and navigation, game playing, and modelling human preferences.
Comparison.
| Aspect | Forward RL | Inverse RL |
|---|---|---|
| Given | Reward function | Expert demonstrations |
| Learned | Policy | Reward function |
| Direction | Reward to behaviour | Behaviour to reward |
| Uniqueness | Optimal policy is well defined | Many rewards fit, so ill-posed |
| Cost | One RL solve | Repeated RL solves inside a loop |
| Use | Task with known goal | Task where the goal is hard to write |
Diagram.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-02" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="IRL pipeline. Exp = expert demonstrations, IRL = inverse RL step, Rew = learned reward, RL = forward RL that trains the agent policy."><style>#dsfig-u5-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-02 .t{fill:#16181D;font-weight:500}#dsfig-u5-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-02 .dot{fill:#16181D}#dsfig-u5-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-02 .ah{fill:#454C5A}#dsfig-u5-02 .ah.hi{fill:#2340B8}#dsfig-u5-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-02 .e{stroke:#B1B7C3}html.dark #dsfig-u5-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-02 .t{fill:#E6E8ED}html.dark #dsfig-u5-02 .t.inv{fill:#0F1115}html.dark #dsfig-u5-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-02 .dot{fill:#E6E8ED}html.dark #dsfig-u5-02 .ann{fill:#8FA3FF}html.dark #dsfig-u5-02 .lbl{fill:#858D9C}html.dark #dsfig-u5-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-02 .ah{fill:#B1B7C3}html.dark #dsfig-u5-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L191,40" marker-end="url(#ah5)"/><path class="e" d="M231,40 L363,40" marker-end="url(#ah5)"/><path class="e" d="M403,40 L535,40" marker-end="url(#ah5)"/><g class="wl"><rect x="102.5" y="31" width="47.1" height="18" rx="9"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">demos</text></g><g class="wl"><rect x="274.5" y="31" width="47.1" height="18" rx="9"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">infer</text></g><g class="wl"><rect x="442.9" y="31" width="54.3" height="18" rx="9"/><text class="t" x="470" y="40" dy=".35em" text-anchor="middle">reward</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Exp</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">IRL</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">Rew</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">RL</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">IRL pipeline. Exp = expert demonstrations, IRL = inverse RL step, Rew = learned reward, RL = forward RL that trains the agent policy.</figcaption></figure>
Answer frame. Open with the definition (reward inferred from behaviour); draw the pipeline and the comparison table with forward RL; then develop points 1-5 and the method list in order; close with applications and the sentence that IRL is ill-posed, which MaxEnt resolves.
Pitfall: Do not describe IRL as learning a policy directly; that is behavioural cloning, and IRL recovers the reward first.
Asked: [7 marks] (Jun 2025) Explain inverse reinforcement learning.
Maximum Entropy Deep Inverse Reinforcement Learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Maximum entropy IRL chooses, among all reward functions whose policy matches the expert's feature expectations, the one whose trajectory distribution has the highest entropy, and the deep version represents the reward with a neural network.</mark>
Key points.
- Many rewards explain the demonstrations, so the principle of maximum entropy removes the ambiguity by assuming nothing beyond the data: the expert is only exponentially more likely to choose better trajectories.
- The trajectory probability is $P(\tau|\theta)=\dfrac{1}{Z}\exp\big(R_\theta(\tau)\big)$, where $Z$ is the normalising partition function over all trajectories.
- Training maximises the log-likelihood of the expert trajectories $\sum_i \log P(\tau_i|\theta)$ by gradient ascent.
- The gradient is the expert's state-visitation frequency minus the learner's expected state-visitation frequency: $\nabla_\theta L = \mu_E - \mathbb{E}_{P}[\mu]$, and the learner's term needs a forward RL solve each step.
- In the deep version the reward is a neural network $R_\theta(s)$, so non-linear features are learned from raw inputs such as images instead of being hand-designed.
- It copes with suboptimal or noisy experts, because the probability of a worse trajectory is smaller but not zero.
Answer frame. Open with the entropy principle; write $P(\tau)\propto\exp(R)$; give the gradient; close with the advantage of a neural reward.
Generative Adversarial Imitation Learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>GAIL (Ho and Ermon) is an imitation-learning method in which a policy acts as a generator and a discriminator tries to tell the policy's state-action pairs from the expert's, so the policy imitates the expert without first recovering a reward function.</mark>
Key points.
- It applies the GAN idea to imitation: the generator is the policy $\pi_\theta$ and the discriminator $D_w(s,a)$ outputs the probability that a pair came from the expert.
- The discriminator is trained to maximise its ability to separate expert data from policy data, and the policy is trained to make its data indistinguishable.
- The policy receives a surrogate reward such as $-\log D_w(s,a)$ and is optimised by a policy-gradient method such as TRPO or PPO.
- The objective is $\min_\pi\max_D\ \mathbb{E}_\pi[\log D(s,a)]+\mathbb{E}_{\pi_E}[\log(1-D(s,a))]$.
- It avoids the expensive inner RL loop of IRL and needs only expert samples and interaction with the environment, not a reward model.
- Its drawback is that it needs many environment interactions and, like any GAN, can be unstable to train.
Answer frame. Open with the generator and discriminator roles; write the min-max objective; close with the advantage over IRL.
Recent Trends in RL Architectures
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Recent RL architectures combine deep networks with off-policy actor-critic learning, learned world models, sequence models and human feedback to make RL more stable, sample efficient and general.</mark>
Key points.
- Stable actor-critic algorithms such as PPO (clipped policy updates) and SAC (soft actor-critic with an entropy bonus) are now the default for continuous control.
- Model-based RL learns a model of the environment and plans or generates imagined experience in it, which cuts the number of real interactions.
- Offline (batch) RL learns from a fixed dataset without further interaction, which is useful where exploration is costly or unsafe.
- Transformer-based RL, such as the Decision Transformer, treats trajectories as sequences and predicts actions conditioned on the desired return.
- RLHF fine-tunes large language models with a reward model learned from human preferences, which is IRL-style reward learning at scale.
- Other directions are hierarchical, multi-agent and meta-RL, which aim at reuse and transfer across tasks.
Answer frame. Open with the aim of stability and sample efficiency; list points 1-5 with one example each; close with RLHF.
Last-minute revision
- Actor = policy $\pi_\theta$; critic = value function $V_w$; the critic's TD error trains the actor.
- TD error: $\delta = r + \gamma V(s') - V(s)$; actor update: $\theta\leftarrow\theta+\alpha\,\delta\,\nabla\log\pi_\theta$.
- Critic reduces the variance of the policy gradient but adds bias.
- A2C is synchronous with advantage $A=Q-V$; A3C is asynchronous with parallel workers.
- IRL infers the reward from expert behaviour; forward RL infers the policy from the reward.
- IRL is ill-posed: many rewards, including $R=0$, explain the same behaviour.
- Feature matching: $\mu(\pi)=\mu_E$ with $R(s)=\theta^\top\phi(s)$.
- MaxEnt IRL: $P(\tau)\propto\exp(R_\theta(\tau))$, gradient = expert visitation minus learner visitation.
- GAIL: discriminator reward $-\log D(s,a)$ trains the policy; no explicit reward is recovered.
- Recent trends: PPO, SAC, model-based, offline RL, Decision Transformer, RLHF.
Memory hooks
- Actor acts, critic criticises: the TD error is the critic's feedback.
- IRL is RL in reverse: behaviour goes in, reward comes out.
- MaxEnt: assume the least, so be as random as the data allows.
- GAIL is a GAN whose generator is the policy.
- PPO clips, SAC adds entropy, RLHF learns the reward from people.
Coverage checklist
- Actor-Critic Method: actor, critic, TD error, update rule, A2C and A3C.
- Inverse reinforcement learning: definition, contrast with forward RL, methods and applications (Jun 2025, 7 marks: Explain inverse reinforcement learning).
- Maximum Entropy Deep Inverse Reinforcement Learning: maximum-entropy principle, trajectory distribution, gradient, deep reward.
- Generative Adversarial Imitation Learning: generator-discriminator imitation and its objective.
- Recent Trends in RL Architectures: PPO, SAC, model-based, offline RL, transformers, RLHF.