Skip to content
CY-802 (B) · Deep & Reinforcement Learning/Important Questions

Deep & Reinforcement Learning (CY-802 (B)) - Important Questions

  1. Unit 114 Marks High Priority

    Derive backpropagation for a multilayer perceptron with one hidden layer using sigmoid activation. Show the chain-rule steps for computing $\frac{\partial L}{\partial W^{(2)}}$ and $\frac{\partial L}{\partial W^{(1)}}$ where $L$ is the loss, $W^{(1)}$ and $W^{(2)}$ are the first-layer and second-layer weight matrices respectively. Obtain the weight update rule for stochastic gradient descent (SGD) and state clearly any assumptions used (batch size, learning rate notation).

    Core derivation from Unit 1; standard RGPV theory question on backpropagation and SGD update.

  2. Unit 17 Marks High Priority

    Explain the Adam optimizer. Present the update formulas for the first and second moment estimates $m_t$ and $v_t$, the bias-corrected estimates $\hat{m}_t$ and $\hat{v}_t$, and the final parameter update rule. Explicitly write the formulas for $m_t$, $v_t$, $\hat{m}_t$, $\hat{v}_t$, and $\theta_{t+1}$ using $\beta_1$, $\beta_2$, learning rate $\alpha$, gradient $g_t$, and small constant $\epsilon$.

    Core optimizer question; frequently examined in exams and interviews; requires formula derivation and discussion of advantages.

  3. Unit 110 Marks High Priority

    Compare and contrast LSTM and GRU. Describe their architectures, gating mechanisms (which gates exist and their roles), forward equations (state update formulas in compact form), relative parameter counts, and discuss situations or tasks where one is preferred over the other. Mention advantages and limitations of each.

    High-frequency recurrent model comparison; expects architecture-level comparison and practical guidance.

  4. Unit 17 Marks High Priority

    Explain the vanishing and exploding gradient problems in training recurrent neural networks. Derive intuitively why repeated multiplication by Jacobians can lead to vanishing or exploding gradients. Describe at least four practical techniques to mitigate these problems (architectural and algorithmic), and explain how each technique helps.

    Core RNN stability topic; commonly paired with mitigation techniques such as advanced architectures and regularization.

  5. Unit 110 Marks High Priority

    Explain Backpropagation Through Time (BPTT) for a simple RNN. Consider the recurrence $h_t=\phi\left(W_h h_{t-1}+W_x x_t + b\right)$ and output $y_t = f\left(U h_t\right)$. Show how to unroll the network for $T$ time steps and derive the expression for the gradient $\frac{\partial L}{\partial W_h}$ in terms of time-indexed partial derivatives. Discuss truncation of BPTT and its practical implications.

    Standard RNN training derivation question; checks understanding of temporal unrolling and gradient accumulation.

  6. Unit 110 Marks High Priority

    Explain attention mechanisms and the encoder–decoder architecture used in sequence-to-sequence models. Derive and write the scaled dot-product attention formula: $$\text{Attention}(Q,K,V)=\text{softmax}\left(\frac{QK^{T}}{\sqrt{d_k}}\right)V\,. $$ Describe the roles of $Q$, $K$, and $V$, explain why the scale factor $\sqrt{d_k}$ is used, and discuss how attention helps with long-range dependencies compared to pure RNNs.

    Modern sequence-to-sequence core concept; includes canonical formula for scaled dot-product attention.

  7. Unit 17 Marks High Priority

    Compare the following optimization algorithms: Momentum, AdaGrad, and RMSProp. For each algorithm give the update rule (mathematical formula), explain the intuition behind it, and state the typical advantages and drawbacks. Explicitly write the update formulas for Momentum, AdaGrad, and RMSProp using standard notation (learning rate $\alpha$, gradient $g_t$, decay/momentum parameters).

    Covers classical optimizer families; expect update rules and comparison of convergence behaviour and use-cases.

  8. Unit 15 Marks Medium Priority

    Explain the differences between batch gradient descent, stochastic gradient descent, and mini-batch gradient descent. Discuss their convergence behaviour, computational trade-offs, and typical scenarios where each is preferred.

    Foundational practical question on gradient-descent variants; standard short-answer type.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in