Skip to content
AD-802 (B) · Reinforcement Learning/Quick Revision Short Notes

Reinforcement Learning (AD-802 (B)) - Unit 1 Short Notes

How unit 1 is examined

This unit covers the RL framework and MDP, policies, value functions, exploration versus exploitation, the parts of an agent and RL's problems; the marks sit in the framework topic (RL definition, RL versus supervised learning) and the agent approaches.

Defining RL Framework and Markov Decision Process

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Reinforcement learning is a type of machine learning in which an agent learns by interacting with an environment, taking actions and receiving rewards, so as to maximise its cumulative reward over time.</mark>

Diagram.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-01" viewBox="0 0 338 80" width="338" height="80" role="img" aria-label="RL loop. Ag = agent, En = environment; at each step the agent sends an action and the environment returns the next state and a reward."><style>#dsfig-u1-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-01 .t{fill:#16181D;font-weight:500}#dsfig-u1-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-01 .dot{fill:#16181D}#dsfig-u1-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-01 .ah{fill:#454C5A}#dsfig-u1-01 .ah.hi{fill:#2340B8}#dsfig-u1-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-01 .e{stroke:#B1B7C3}html.dark #dsfig-u1-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-01 .t{fill:#E6E8ED}html.dark #dsfig-u1-01 .t.inv{fill:#0F1115}html.dark #dsfig-u1-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-01 .dot{fill:#E6E8ED}html.dark #dsfig-u1-01 .ann{fill:#8FA3FF}html.dark #dsfig-u1-01 .lbl{fill:#858D9C}html.dark #dsfig-u1-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-01 .ah{fill:#B1B7C3}html.dark #dsfig-u1-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M58.4,44.6 Q169,72 277.6,45.1" marker-end="url(#ah1)"/><path class="e" d="M279.6,35.4 Q169,8 60.4,34.9" marker-end="url(#ah1)"/><g class="wl"><rect x="141.4" y="49.4" width="54.3" height="18" rx="9"/><text class="t" x="168.5" y="58.4" dy=".35em" text-anchor="middle">action</text></g><g class="wl"><rect x="121.2" y="12.6" width="96.6" height="18" rx="9"/><text class="t" x="169.5" y="21.6" dy=".35em" text-anchor="middle">state,reward</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Ag</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">En</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">RL loop. Ag = agent, En = environment; at each step the agent sends an action and the environment returns the next state and a reward.</figcaption></figure>

Key points.

  1. The agent is the learner and decision maker, while the environment is everything outside the agent that it interacts with and that responds to its actions.
  2. At each step the agent observes state $s_t$, takes action $a_t$, and receives reward $r_{t+1}$ and next state $s_{t+1}$.
  3. There is no teacher giving correct answers; the only feedback is a scalar reward, which is often delayed, so the agent must work out which earlier actions deserve credit.
  4. The goal is to maximise the expected return $G_t = r_{t+1} + \gamma r_{t+2} + \gamma^2 r_{t+3} + \dots$ with discount factor $0 \le \gamma \le 1$.
  5. Problems are formalised as a Markov Decision Process (MDP), the tuple $(S, A, P, R, \gamma)$: states, actions, transition probabilities $P(s'|s,a)$, reward function and discount.
  6. The Markov property says the future depends only on the present state and action, not on the earlier history.
  7. The agent must balance exploring new actions against exploiting known good ones, which is what separates RL from other learning types.
  8. Example. A self-driving car learns to steer, brake and change lanes from a positive reward for safe progress and a penalty for collisions; a chess program learns from +1 for a win and -1 for a loss.

Comparison.

Basis Supervised learning Reinforcement learning
Data Labelled input-output pairs No labels; experience gathered by acting
Feedback Correct answer for every input Scalar reward, often delayed
Goal Predict or classify accurately Maximise cumulative reward
Training Learns from a fixed dataset Learns by trial and error interacting with the environment
Data nature Independent, identically distributed samples Sequential; actions change later data
Example Spam detection Robot walking, game playing

Answer frame. Definition question: open with the definition, draw the agent-environment loop, develop points 1-7, give the example and close with "the agent learns the best behaviour by maximising cumulative reward". Comparison question: one line defining each, then the 5-6 row table, close with "supervised learning is taught by answers, RL by rewards".

Asked: [7 marks] (Jun 2025) What is Reinforcement learning? State one appropriate real-life example. Asked: [7 marks] (Jun 2025) Differentiate between Reinforcement Learning and Supervised Learning. Pitfall: Do not call the reward a label; it evaluates an action taken, it does not say which action was correct.

Polices (policies)

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A policy $\pi$ is the agent's behaviour: a mapping from states to actions.</mark>

Key points.

  1. A deterministic policy gives one action per state, $a = \pi(s)$.
  2. A stochastic policy gives a probability distribution, $\pi(a|s) = P(A_t = a \mid S_t = s)$.
  3. The aim of RL is to find the optimal policy $\pi^*$ that maximises expected return.
  4. A policy can be a lookup table for small problems or a function approximator such as a neural network for large ones.
  5. Example: in a maze, a policy says go left in cell A and go up in cell B, and the agent follows it at every step.

Value Functions and Bellman Equations

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A value function is the expected return from a state (or state-action pair) when following a policy.</mark>

Formula.

$$V^\pi(s) = \mathbb{E}_\pi[G_t \mid S_t = s], \qquad Q^\pi(s,a) = \mathbb{E}_\pi[G_t \mid S_t = s, A_t = a]$$

$$V^\pi(s) = \sum_a \pi(a|s) \sum_{s'} P(s'|s,a)\,[R(s,a,s') + \gamma V^\pi(s')]$$

Key points.

  1. $V$ measures how good a state is; $Q$ measures how good an action is in a state.
  2. The Bellman equation writes a state's value recursively as the immediate reward plus the discounted value of the next state.
  3. It turns long-term return into a one-step relation, which makes computing values possible by iteration.
  4. The optimal version replaces the policy average by a maximum over actions: $V^*(s) = \max_a \sum_{s'} P(s'|s,a)[R + \gamma V^*(s')]$.

Explorations (exploration)

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Exploration means trying actions not known to be best in order to gather new information about the environment.</mark>

Key points.

  1. It reduces uncertainty about rewards and transitions and may discover better actions.
  2. Common methods are $\epsilon$-greedy (random action with probability $\epsilon$), softmax and optimism in the face of uncertainty (UCB).
  3. Too much exploration wastes reward on poor actions, so it must be limited.
  4. Example: a restaurant-goer trying a new dish instead of the usual favourite is exploring.

Exploitation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Exploitation means choosing the action currently believed to give the highest reward, using what has already been learned.</mark>

Key points.

  1. It gives good immediate return based on current estimates.
  2. Pure exploitation can get stuck on a suboptimal action because better ones are never tried.
  3. Example: ordering the dish already known to be tasty is exploitation.
  4. The explore-exploit trade-off is balancing the two: explore early, exploit more as knowledge grows, for example by decaying $\epsilon$.

Inside an RL agent

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>An RL agent may contain up to three components: a policy, a value function and a model of the environment.</mark>

Key points.

  1. Value-based approach: learn a value function and act greedily on it, with no explicit policy; example Q-learning and DQN.
  2. Policy-based approach: learn the policy $\pi(a|s)$ directly by gradient ascent on expected return; example REINFORCE (policy gradients).
  3. Actor-critic approach combines both: the actor is the policy and the critic is the value function that evaluates it.
  4. Model-based methods learn or use a model of transitions and rewards to plan ahead; model-free methods, such as Q-learning and policy gradients, learn directly from experience without one.
  5. The model predicts the next state and reward for a given state and action, so a model-based agent can simulate outcomes before acting and needs fewer real interactions.
  6. Value-based methods suit discrete actions, policy-based methods suit continuous actions and stochastic policies, and actor-critic reduces the variance of pure policy gradients.

Answer frame. Open with "an RL agent can be built in three ways, by learning values, policies or a model"; draw a Venn of policy, value function and model with actor-critic in the overlap; develop points 1-5 in order with one algorithm each; close with a value-based versus policy-based contrast.

Asked: [7 marks] (Jun 2025) Explain various approaches to implement reinforcement learning.

Problems with reinforcement learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>RL problems are the practical difficulties that make agents hard to train.</mark>

Key points.

  1. Exploration versus exploitation must be balanced.
  2. Rewards are sparse and delayed, so credit assignment (which action earned the reward) is hard.
  3. RL needs very many samples, so it is sample-inefficient, and real-world trial and error can be unsafe or costly.
  4. Training is unstable and reward design is difficult; a badly designed reward gets exploited by the agent in unintended ways.
  5. Generalising to new situations is hard, and simulation-trained policies often fail on real hardware.

Last-minute revision

  • RL: an agent learns by interacting with an environment to maximise cumulative reward.
  • Loop: state, action, reward, next state.
  • Return $G_t = r_{t+1} + \gamma r_{t+2} + \dots$; $0 \le \gamma \le 1$.
  • MDP is $(S, A, P, R, \gamma)$ with the Markov property.
  • Supervised learning has labels; RL has only delayed rewards.
  • Policy $\pi$ maps states to actions: deterministic or stochastic.
  • $V^\pi(s)$ is state value, $Q^\pi(s,a)$ is action value; Bellman equation is recursive.
  • Explore to learn, exploit to earn; $\epsilon$-greedy balances them.
  • Approaches: value-based (Q-learning), policy-based (policy gradient), actor-critic, model-based, model-free.
  • Optimal Bellman: $V^*(s) = \max_a \sum_{s'} P(s'|s,a)[R + \gamma V^*(s')]$.
  • $\gamma$ near 0 is short-sighted, near 1 is far-sighted.
  • Problems: sparse reward, credit assignment, sample inefficiency, unsafe exploration.

Memory hooks

  • SARS: State, Action, Reward, next State.
  • Explore = learn, exploit = earn.
  • Actor acts (policy), critic criticises (value).
  • Supervised has a teacher, RL has a scorekeeper.
  • VPM: Value, Policy, Model are the three parts of an agent.

Coverage checklist

  • Defining RL Framework and Markov Decision Process: RL definition with example; RL versus supervised learning.
  • Polices: policy definition, deterministic and stochastic.
  • Value Functions and Bellman Equations: $V$, $Q$, Bellman equation.
  • Explorations: definition, $\epsilon$-greedy.
  • Exploitation: definition, trade-off.
  • Inside an RL agent: approaches to implement RL.
  • Problems with reinforcement learning: challenges of RL.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in