Deep & Reinforcement Learning (CY-802 (B)) - Important Questions
-
Unit 514 Marks High Priority
Derive the Policy Gradient theorem. Starting from the objective $J(\theta)=\mathbb{E}_{\tau\sim\pi_{\theta}}\left[\sum_{t=0}^{T} r_t\right]$, show that the gradient can be written as $\nabla_{\theta} J(\theta)=\mathbb{E}_{\tau\sim\pi_{\theta}}\left[\sum_{t=0}^{T}\nabla_{\theta}\log\pi_{\theta}(a_t\mid s_t)\,G_t\right]$. Explain every step in the derivation and state the assumptions used.
Core derivation of the policy gradient theorem; central concept in Unit 5 and frequently examined.
-
Unit 510 Marks High Priority
Explain the Actor-Critic architecture. Provide the parameter update rules for the actor and the critic using a policy-gradient actor with a value-function critic. Explicitly show the actor update as $\theta\leftarrow\theta+\alpha_{a}\,\delta_{t}\,\nabla_{\theta}\log\pi_{\theta}(a_{t}\mid s_{t})$ and the critic update as $w\leftarrow w+\alpha_{c}\,\delta_{t}\,\nabla_{w}V_{w}(s_{t})$, where $\delta_{t}=r_{t}+\gamma V_{w}(s_{t+1})-V_{w}(s_{t})$. Discuss the roles of learning rates $\alpha_{a}$ and $\alpha_{c}$ and practical stability considerations.
Standard Actor-Critic algorithm formulation and parameter update rules; asked repeatedly in exams as an applied policy-gradient method.
-
Unit 57 Marks High Priority
Show that adding a baseline $b(s)$ to the policy gradient estimator does not introduce bias. Prove that $\mathbb{E}_{a\sim\pi_{\theta}}\left[\nabla_{\theta}\log\pi_{\theta}(a\mid s)\,b(s)\right]=0$ and hence justify why choosing $b(s)=V^{\pi}(s)$ reduces variance. Include all mathematical steps and state any expectations explicitly.
Variance reduction technique that is fundamental to practical policy-gradient methods; tested for understanding of unbiasedness and variance properties.
-
Unit 57 Marks High Priority
Describe the REINFORCE algorithm for episodic tasks and derive its parameter update rule. Starting from $J(\theta)=\mathbb{E}_{\tau\sim\pi_{\theta}}\left[\sum_{t=0}^{T} r_t\right]$, derive the update $\theta\leftarrow\theta+\alpha\sum_{t=0}^{T}\nabla_{\theta}\log\pi_{\theta}(a_t\mid s_t)\,G_t$. Explain how to incorporate a baseline and discuss the effect on learning.
REINFORCE algorithm is the canonical episodic policy-gradient method; derivation and baseline incorporation are commonly examined.
-
Unit 510 Marks Medium Priority
Explain the Maximum Entropy Inverse Reinforcement Learning (MaxEnt IRL) formulation. Given expert feature expectations $\hat{\phi}$ and model expected features $\mathbb{E}_{\pi_{\theta}}\left[\phi(\tau)\right]$, write the MaxEnt IRL objective and derive the gradient with respect to reward parameters $\theta$. Discuss how the partition function (normalizer) influences the gradient and practical ways to approximate it.
Maximum Entropy IRL is a principled IRL method appearing in Unit 5; tests understanding of objective and gradient computation.
-
Unit 57 Marks Medium Priority
Compare Behavioral Cloning and DAGGER (Dataset Aggregation) for imitation learning. Describe the algorithms step-by-step, state their assumptions, and explain situations where DAGGER improves over Behavioral Cloning. Discuss practical limitations such as distribution shift and sample efficiency.
Imitation learning fundamentals comparing direct supervised approaches and interactive methods; frequently asked to test algorithmic trade-offs.
Quick Add to Notes
Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.
Create free accountHave an account? Log in
Notes Panel