Skip to content
CY-802 (B) · Deep & Reinforcement Learning/Important Questions

Deep & Reinforcement Learning (CY-802 (B)) - Important Questions

  1. Unit 414 Marks High Priority

    Derive the Bellman optimality equation for the state-value function $V^*(s)$ in a finite Markov Decision Process (MDP). Write the corresponding Bellman optimality operator and prove that it is a contraction mapping under the sup-norm, hence showing why value iteration converges to $V^*$. In your derivation include the equation

    $$V^*(s)=\max_a\sum_{s'}P(s'\mid s,a)\left(R(s,a,s')+\gamma V^*(s')\right)$$

    and show the contraction inequality

    $$\|T_V V- T_V W\|_\infty \le \gamma\|V-W\|_\infty\;.$$

    Core derivation from Unit 4: Bellman optimality and convergence — fundamental theoretical result frequently examined in exams.

  2. Unit 47 Marks Medium Priority

    Define a Markov Decision Process (MDP). Clearly state and explain its components: state space $S$, action space $A$, transition probabilities $P(s'\mid s,a)$, reward function $R(s,a,s')$, and discount factor $\gamma$. Provide a concise tabular specification (states, actions, transitions, rewards) for a small example MDP (at least 3 states) and explain how an episode would proceed.

    Fundamental definition and components — basic but often asked to test understanding of MDP formalism.

  3. Unit 414 Marks High Priority

    Explain Value Iteration and Policy Iteration algorithms for solving finite MDPs. For Value Iteration give the update rule and policy extraction step. For Policy Iteration describe policy evaluation and policy improvement steps. For a small 3-state MDP perform two iterations of value iteration numerically (showing intermediate value vectors) and demonstrate how the greedy policy is extracted from the current value estimate. Discuss time and space complexity differences between the two methods.

    Algorithmic comparison frequently examined: both algorithm steps and complexity are typical exam requirements.

  4. Unit 410 Marks High Priority

    Derive the Q-learning update rule starting from the Bellman optimality equation for the action-value function $Q^*(s,a)$. State why Q-learning is considered an off-policy temporal-difference control algorithm and list the standard conditions required for its convergence (step-size schedules, sufficient exploration, Markovian environment). Include the Bellman optimality equation:

    $$Q^*(s,a)=\mathbb{E}_{s'}\left[R(s,a,s')+\gamma\max_{a'}Q^*(s',a')\right]\;,$$

    and the incremental update form used in practice.

    Core model-free control question; Q-learning derivation and convergence conditions are commonly asked.

  5. Unit 47 Marks High Priority

    Describe Temporal-Difference (TD(0)) learning for state-value prediction and give its update rule. Present the SARSA algorithm and its update equation, then contrast SARSA (on-policy) with Q-learning (off-policy) in terms of their update targets and behaviour under exploration. Use the SARSA update:

    $$Q(s_t,a_t)\leftarrow Q(s_t,a_t)+\alpha\left(R_{t+1}+\gamma Q(s_{t+1},a_{t+1})-Q(s_t,a_t)\right)\;.$$

    Temporal-Difference family question: SARSA vs Q-learning and TD(0) are standard unit-4 topics.

  6. Unit 414 Marks High Priority

    Explain the Deep Q-Network (DQN) algorithm. Describe the role of the Q-network, experience replay buffer, and target network. Write the DQN squared-error loss used for a minibatch sampled from replay memory and define the target value explicitly. Use the following loss expressions:

    $$L(\theta)=\mathbb{E}_{(s,a,r,s')\sim D}\left[\left(y^{\mathrm{DQN}}-Q(s,a;\theta)\right)^2\right]\;,$$

    where

    $$y^{\mathrm{DQN}}=r+\gamma\max_{a'}Q\left(s',a';\theta^{-}\right)\;,$$

    and discuss how the target network parameters $\theta^{-}$ are maintained and why these components improve training stability.

    Deep RL central practical topic: students must know DQN components and loss formulation; often asked as a long question.

  7. Unit 47 Marks Medium Priority

    Explain experience replay in Deep Q-Learning. State its purpose (breaking temporal correlations and improving sample efficiency), describe how a replay buffer operates, and summarize the idea of prioritized experience replay. Discuss practical trade-offs such as memory cost, staleness of samples, and non-stationarity introduced by changing policy.

    Experience replay is a frequently examined practical element of DQN and its variants; prioritized replay may be asked as extension.

  8. Unit 47 Marks Medium Priority

    Compare and contrast common exploration strategies: $\epsilon$-greedy, Boltzmann (softmax) exploration, and Upper Confidence Bound (UCB). Provide the Boltzmann action selection formula

    $$\pi(a\mid s)=\frac{\exp\left(\dfrac{Q(s,a)}{\tau}\right)}{\sum_b\exp\left(\dfrac{Q(s,b)}{\tau}\right)}\;,$$

    and a standard UCB selection rule:

    $$a_t=\arg\max_a\left(\overline{Q}_a+c\sqrt{\dfrac{\ln t}{N_a}}\right)\;.$$

    Exploration-exploitation strategies are standard short-answer topics in Unit 4 exams.

  9. Unit 410 Marks Medium Priority

    Explain the use of nonlinear function approximation (neural networks) for value and Q-function estimation in reinforcement learning. Describe why combining temporal-difference updates with nonlinear function approximation can cause instability or divergence. List and explain standard remedies (experience replay, target networks, double Q-learning, gradient clipping, batch normalization) that address these issues.

    Function approximation and stability issues are key in deep RL; expected in long-answer or discussion questions.

  10. Unit 414 Marks Medium Priority

    Describe the Policy Gradient framework and derive the REINFORCE (Monte Carlo policy gradient) estimator. Present the policy gradient theorem and the empirical estimator

    $$\nabla_\theta J(\theta)=\mathbb{E}_{\tau}\left[\sum_t\nabla_\theta\log\pi_\theta(a_t\mid s_t)G_t\right]\;,$$

    explain the role of the return $G_t$, and discuss variance reduction using a baseline. Outline how an actor-critic algorithm implements a learned baseline and briefly contrast actor-critic with pure policy-gradient methods.

    Policy-based methods and variance reduction form an important complement to value-based methods in Unit 4.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in