How unit 1 is examined
This unit covers the RL framework and MDP, policies, value functions, exploration versus exploitation, the parts of an agent and RL's problems; the marks sit in the framework topic (RL definition, RL versus supervised learning) and the agent approaches.
Defining RL Framework and Markov Decision Process
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Reinforcement learning is a type of machine learning in which an agent learns by interacting with an environment, taking actions and receiving rewards, so as to maximise its cumulative reward over time.</mark>
Diagram.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-01" viewBox="0 0 338 80" width="338" height="80" role="img" aria-label="RL loop. Ag = agent, En = environment; at each step the agent sends an action and the environment returns the next state and a reward."><style>#dsfig-u1-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-01 .t{fill:#16181D;font-weight:500}#dsfig-u1-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-01 .dot{fill:#16181D}#dsfig-u1-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-01 .ah{fill:#454C5A}#dsfig-u1-01 .ah.hi{fill:#2340B8}#dsfig-u1-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-01 .e{stroke:#B1B7C3}html.dark #dsfig-u1-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-01 .t{fill:#E6E8ED}html.dark #dsfig-u1-01 .t.inv{fill:#0F1115}html.dark #dsfig-u1-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-01 .dot{fill:#E6E8ED}html.dark #dsfig-u1-01 .ann{fill:#8FA3FF}html.dark #dsfig-u1-01 .lbl{fill:#858D9C}html.dark #dsfig-u1-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-01 .ah{fill:#B1B7C3}html.dark #dsfig-u1-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M58.4,44.6 Q169,72 277.6,45.1" marker-end="url(#ah1)"/><path class="e" d="M279.6,35.4 Q169,8 60.4,34.9" marker-end="url(#ah1)"/><g class="wl"><rect x="141.4" y="49.4" width="54.3" height="18" rx="9"/><text class="t" x="168.5" y="58.4" dy=".35em" text-anchor="middle">action</text></g><g class="wl"><rect x="121.2" y="12.6" width="96.6" height="18" rx="9"/><text class="t" x="169.5" y="21.6" dy=".35em" text-anchor="middle">state,reward</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Ag</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">En</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">RL loop. Ag = agent, En = environment; at each step the agent sends an action and the environment returns the next state and a reward.</figcaption></figure>
Key points.
- The agent is the learner and decision maker, while the environment is everything outside the agent that it interacts with and that responds to its actions.
- At each step the agent observes state $s_t$, takes action $a_t$, and receives reward $r_{t+1}$ and next state $s_{t+1}$.
- There is no teacher giving correct answers; the only feedback is a scalar reward, which is often delayed, so the agent must work out which earlier actions deserve credit.
- The goal is to maximise the expected return $G_t = r_{t+1} + \gamma r_{t+2} + \gamma^2 r_{t+3} + \dots$ with discount factor $0 \le \gamma \le 1$.
- Problems are formalised as a Markov Decision Process (MDP), the tuple $(S, A, P, R, \gamma)$: states, actions, transition probabilities $P(s'|s,a)$, reward function and discount.
- The Markov property says the future depends only on the present state and action, not on the earlier history.
- The agent must balance exploring new actions against exploiting known good ones, which is what separates RL from other learning types.
- Example. A self-driving car learns to steer, brake and change lanes from a positive reward for safe progress and a penalty for collisions; a chess program learns from +1 for a win and -1 for a loss.
Comparison.
| Basis | Supervised learning | Reinforcement learning |
|---|---|---|
| Data | Labelled input-output pairs | No labels; experience gathered by acting |
| Feedback | Correct answer for every input | Scalar reward, often delayed |
| Goal | Predict or classify accurately | Maximise cumulative reward |
| Training | Learns from a fixed dataset | Learns by trial and error interacting with the environment |
| Data nature | Independent, identically distributed samples | Sequential; actions change later data |
| Example | Spam detection | Robot walking, game playing |
Answer frame. Definition question: open with the definition, draw the agent-environment loop, develop points 1-7, give the example and close with "the agent learns the best behaviour by maximising cumulative reward". Comparison question: one line defining each, then the 5-6 row table, close with "supervised learning is taught by answers, RL by rewards".
Asked: [7 marks] (Jun 2025) What is Reinforcement learning? State one appropriate real-life example. Asked: [7 marks] (Jun 2025) Differentiate between Reinforcement Learning and Supervised Learning. Pitfall: Do not call the reward a label; it evaluates an action taken, it does not say which action was correct.
Polices (policies)
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A policy $\pi$ is the agent's behaviour: a mapping from states to actions.</mark>
Key points.
- A deterministic policy gives one action per state, $a = \pi(s)$.
- A stochastic policy gives a probability distribution, $\pi(a|s) = P(A_t = a \mid S_t = s)$.
- The aim of RL is to find the optimal policy $\pi^*$ that maximises expected return.
- A policy can be a lookup table for small problems or a function approximator such as a neural network for large ones.
- Example: in a maze, a policy says go left in cell A and go up in cell B, and the agent follows it at every step.
Value Functions and Bellman Equations
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A value function is the expected return from a state (or state-action pair) when following a policy.</mark>
Formula.
$$V^\pi(s) = \mathbb{E}_\pi[G_t \mid S_t = s], \qquad Q^\pi(s,a) = \mathbb{E}_\pi[G_t \mid S_t = s, A_t = a]$$
$$V^\pi(s) = \sum_a \pi(a|s) \sum_{s'} P(s'|s,a)\,[R(s,a,s') + \gamma V^\pi(s')]$$
Key points.
- $V$ measures how good a state is; $Q$ measures how good an action is in a state.
- The Bellman equation writes a state's value recursively as the immediate reward plus the discounted value of the next state.
- It turns long-term return into a one-step relation, which makes computing values possible by iteration.
- The optimal version replaces the policy average by a maximum over actions: $V^*(s) = \max_a \sum_{s'} P(s'|s,a)[R + \gamma V^*(s')]$.
Explorations (exploration)
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Exploration means trying actions not known to be best in order to gather new information about the environment.</mark>
Key points.
- It reduces uncertainty about rewards and transitions and may discover better actions.
- Common methods are $\epsilon$-greedy (random action with probability $\epsilon$), softmax and optimism in the face of uncertainty (UCB).
- Too much exploration wastes reward on poor actions, so it must be limited.
- Example: a restaurant-goer trying a new dish instead of the usual favourite is exploring.
Exploitation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Exploitation means choosing the action currently believed to give the highest reward, using what has already been learned.</mark>
Key points.
- It gives good immediate return based on current estimates.
- Pure exploitation can get stuck on a suboptimal action because better ones are never tried.
- Example: ordering the dish already known to be tasty is exploitation.
- The explore-exploit trade-off is balancing the two: explore early, exploit more as knowledge grows, for example by decaying $\epsilon$.
Inside an RL agent
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>An RL agent may contain up to three components: a policy, a value function and a model of the environment.</mark>
Key points.
- Value-based approach: learn a value function and act greedily on it, with no explicit policy; example Q-learning and DQN.
- Policy-based approach: learn the policy $\pi(a|s)$ directly by gradient ascent on expected return; example REINFORCE (policy gradients).
- Actor-critic approach combines both: the actor is the policy and the critic is the value function that evaluates it.
- Model-based methods learn or use a model of transitions and rewards to plan ahead; model-free methods, such as Q-learning and policy gradients, learn directly from experience without one.
- The model predicts the next state and reward for a given state and action, so a model-based agent can simulate outcomes before acting and needs fewer real interactions.
- Value-based methods suit discrete actions, policy-based methods suit continuous actions and stochastic policies, and actor-critic reduces the variance of pure policy gradients.
Answer frame. Open with "an RL agent can be built in three ways, by learning values, policies or a model"; draw a Venn of policy, value function and model with actor-critic in the overlap; develop points 1-5 in order with one algorithm each; close with a value-based versus policy-based contrast.
Asked: [7 marks] (Jun 2025) Explain various approaches to implement reinforcement learning.
Problems with reinforcement learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>RL problems are the practical difficulties that make agents hard to train.</mark>
Key points.
- Exploration versus exploitation must be balanced.
- Rewards are sparse and delayed, so credit assignment (which action earned the reward) is hard.
- RL needs very many samples, so it is sample-inefficient, and real-world trial and error can be unsafe or costly.
- Training is unstable and reward design is difficult; a badly designed reward gets exploited by the agent in unintended ways.
- Generalising to new situations is hard, and simulation-trained policies often fail on real hardware.
Last-minute revision
- RL: an agent learns by interacting with an environment to maximise cumulative reward.
- Loop: state, action, reward, next state.
- Return $G_t = r_{t+1} + \gamma r_{t+2} + \dots$; $0 \le \gamma \le 1$.
- MDP is $(S, A, P, R, \gamma)$ with the Markov property.
- Supervised learning has labels; RL has only delayed rewards.
- Policy $\pi$ maps states to actions: deterministic or stochastic.
- $V^\pi(s)$ is state value, $Q^\pi(s,a)$ is action value; Bellman equation is recursive.
- Explore to learn, exploit to earn; $\epsilon$-greedy balances them.
- Approaches: value-based (Q-learning), policy-based (policy gradient), actor-critic, model-based, model-free.
- Optimal Bellman: $V^*(s) = \max_a \sum_{s'} P(s'|s,a)[R + \gamma V^*(s')]$.
- $\gamma$ near 0 is short-sighted, near 1 is far-sighted.
- Problems: sparse reward, credit assignment, sample inefficiency, unsafe exploration.
Memory hooks
- SARS: State, Action, Reward, next State.
- Explore = learn, exploit = earn.
- Actor acts (policy), critic criticises (value).
- Supervised has a teacher, RL has a scorekeeper.
- VPM: Value, Policy, Model are the three parts of an agent.
Coverage checklist
- Defining RL Framework and Markov Decision Process: RL definition with example; RL versus supervised learning.
- Polices: policy definition, deterministic and stochastic.
- Value Functions and Bellman Equations: $V$, $Q$, Bellman equation.
- Explorations: definition, $\epsilon$-greedy.
- Exploitation: definition, trade-off.
- Inside an RL agent: approaches to implement RL.
- Problems with reinforcement learning: challenges of RL.