UNIT 2: DEEP LEARNING - EXAM-FOCUSED SHORT NOTES
1. FOUNDATIONS & HISTORICAL CONTEXT
History and Evolution of Deep Learning
-
Key Milestones:
-
1943: McCulloch-Pitts neuron (first computational model).
-
1958: Perceptron (single-layer, linear).
-
1986: Backpropagation popularized (Rumelhart, Hinton, Williams).
-
2006: Deep Belief Networks (Hinton) - greedy pre-training, start of modern DL.
-
2012: AlexNet (Krizhevsky, Sutskever, Hinton) - ReLU, Dropout, GPU training, won ImageNet.
-
2014-2016: GANs (Goodfellow), ResNet (He et al.), Transformer (Vaswani et al.).
-
-
Eras: Early neural networks (1940s-60s) → AI Winter (1970s-80s) → Modern Resurgence (2006-present) driven by big data, GPU compute, algorithmic advances.
Core Definitions & Paradigms
| Term | Definition | Key Idea |
|---|---|---|
| AI | Broad field of creating intelligent machines. | Encompasses all. |
| ML | Subset of AI; algorithms learn patterns from data. | Feature engineering often manual. |
| DL | Subset of ML; uses deep neural networks with multiple layers to learn hierarchical features automatically. | End-to-end learning. |
| Supervised | Learning from labeled data (input-output pairs). | Classification, regression. |
| Unsupervised | Learning from unlabeled data (find structure). | Clustering, dimensionality reduction. |
| Reinforcement | Learning via interaction (agent, environment, reward). | Policy optimization. |
[!TIP] Exam Distinction: DL ≠ just "more layers". It's about hierarchical feature learning where lower layers learn simple features (edges) and deeper layers learn complex concepts (objects).
Basic Neural Network Building Blocks
-
Single-Layer Perceptron (SLP): $$\displaystyle y = f(\mathbf{w}^T\mathbf{x} + b) $$. Limitation: Can only learn linearly separable problems (XOR fails).
-
Multilayer Perceptron (MLP): Stack of fully connected layers with non-linear activations.
-
Universal Approximation Theorem: An MLP with one hidden layer (sufficiently wide) can approximate any continuous function.
-
Representation Power of Depth: Deep networks can represent certain functions exponentially more efficiently than shallow ones (fewer parameters needed for same complexity). Depth enables compositional representation.
-
-
Biological vs. Artificial: Biological neurons are sparse, asynchronous, plastic; ANNs are dense, synchronous, fixed-structure. Limitation of DL: Lack of causal reasoning, common sense, energy efficiency, and sample efficiency compared to human brain.
2. TRAINING DEEP NEURAL NETWORKS (CORE ALGORITHMS & THEORY)
Forward and Backward Propagation
-
Forward Pass: Compute output layer-by-layer: $$\displaystyle a^{[l]} = g^{[l]}(z^{[l]}) $$, $$\displaystyle z^{[l]} = W^{[l]}a^{[l-1]} + b^{[l]} $$.
-
Backpropagation (BPTT for RNNs):
-
Forward pass to compute loss $L$.
-
Backward pass: Compute gradients $$\displaystyle \frac{\partial L}{\partial W^{[l]}}, \frac{\partial L}{\partial b^{[l]}} $$ using chain rule from output to input.
-
Update weights: $$\displaystyle W^{[l]} := W^{[l]} - \alpha \frac{\partial L}{\partial W^{[l]}} $$.
- Computation Graph: Visualizes operations and dependencies; gradients flow backward along edges.
-
-
Application Areas: Training any differentiable model (CNNs, RNNs, Transformers).
Optimization Algorithms
| Algorithm | Core Idea | Update Rule (for parameter $\theta$) | Pros/Cons |
|---|---|---|---|
| Batch GD | Use entire dataset per update. | $$\displaystyle \theta := \theta - \alpha \nabla_\theta J(\theta) $$ | Stable, slow, memory heavy. |
| SGD | Use one sample per update. | $$\displaystyle \theta := \theta - \alpha \nabla_\theta J(\theta; x^{(i)}, y^{(i)}) $$ | Fast updates, noisy convergence. |
| Mini-batch SGD | Use small batch (e.g., 32, 128). | $$\displaystyle \theta := \theta - \alpha \frac{1}{m} \sum_{i=1}^{m} \nabla_\theta J(\theta; x^{(i)}, y^{(i)}) $$ | Practical standard: balance of speed/stability. |
| Momentum | Add velocity term to smooth updates. | $$\displaystyle v := \beta v + (1-\beta)\nabla_\theta J $$, $$\displaystyle \theta := \theta - \alpha v $$ | Accelerates across shallow regions, dampens oscillations. |
| Nesterov | "Lookahead" momentum. | $$\displaystyle v := \beta v + (1-\beta)\nabla_\theta J(\theta - \alpha\beta v) $$, $$\displaystyle \theta := \theta - \alpha v $$ | More accurate gradient estimate. |
| AdaGrad | Per-parameter adaptive learning rate. Accumulates squared gradients. | $$\displaystyle \theta := \theta - \frac{\alpha}{\sqrt{G_t + \epsilon}} \odot g_t $$, $$\displaystyle G_t = \sum_{\tau=1}^{t} g_\tau^2 $$ | Learning rates decay aggressively; not ideal for non-convex. |
| RMSProp | Fix AdaGrad's aggressive decay via decaying average. | $$\displaystyle G_t := \beta G_{t-1} + (1-\beta) g_t^2 $$, $$\displaystyle \theta := \theta - \frac{\alpha}{\sqrt{G_t + \epsilon}} \odot g_t $$ | Handles non-stationary objectives well. |
| Adam (Adaptive Moment Estimation) | Most popular. Combines Momentum (1st moment) & RMSProp (2nd moment) with bias correction. | $$\displaystyle m_t = \beta_1 m_{t-1} + (1-\beta_1)g_t $$, $$\displaystyle v_t = \beta_2 v_{t-1} + (1-\beta_2)g_t^2 $$, $$\displaystyle \hat{m}_t = \frac{m_t}{1-\beta_1^t} $$, $$\displaystyle \hat{v}_t = \frac{v_t}{1-\beta_2^t} $$, $$\displaystyle \theta := \theta - \alpha \frac{\hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon} $$ | Robust, low memory, good default ($$\displaystyle \alpha=0.001 $$, $$\displaystyle \beta_1=0.9 $$, $$\displaystyle \beta_2=0.999 $$). |
[!TIP] Adam is default. For very large datasets/CNNs, SGD with momentum + learning rate decay can yield better final generalization.
Weight Initialization & Normalization
-
Why Initialization Matters: Break symmetry, control variance of activations/gradients to avoid vanishing/exploding gradients.
-
Methods:
-
Random (Uniform/Normal): Simple but problematic for deep nets.
-
Xavier/Glorot: For tanh/sigmoid. $$\displaystyle W \sim \mathcal{U}\left(-\sqrt{\frac{6}{n_{in}+n_{out}}}, \sqrt{\frac{6}{n_{in}+n_{out}}}\right) $$ or $$\displaystyle \mathcal{N}(0, \sqrt{\frac{2}{n_{in}+n_{out}}}) $$.
-
He Initialization: For ReLU. $$\displaystyle W \sim \mathcal{N}(0, \sqrt{\frac{2}{n_{in}}}) $$. Standard for ReLU CNNs.
-
-
Batch Normalization (BN):
-
Mechanics: For each mini-batch, normalize layer inputs: $$\displaystyle \hat{x}^{(k)} = \frac{x^{(k)} - \mu_B}{\sqrt{\sigma_B^2 + \epsilon}} $$, then scale/shift: $$\displaystyle y^{(k)} = \gamma \hat{x}^{(k)} + \beta $$. $\gamma, \beta$ are learnable.
-
Purpose: Reduce Internal Covariate Shift (distribution change of layer inputs during training). Smooths optimization landscape.
-
Advantages: Allows higher learning rates, reduces need for careful initialization, acts as slight regularizer. Applied before activation (except for RNNs).
-
-
Normalization vs. Standardization (Data Preprocessing):
-
Normalization (Min-Max): $$\displaystyle x' = \frac{x - x_{min}}{x_{max} - x_{min}} $$ → scales to [0,1]. Sensitive to outliers.
-
Standardization (Z-score): $$\displaystyle x' = \frac{x - \mu}{\sigma} $$ → zero mean, unit variance. More robust, preferred for DL.
-
Activation Functions
| Function | Formula | Range | Pros | Cons | Use Case |
|---|---|---|---|---|---|
| Sigmoid | $$\displaystyle \sigma(x) = \frac{1}{1+e^{-x}} $$ | (0,1) | Smooth, probabilistic output. | Vanishing gradients (saturates), not zero-centered. | Output layer for binary classification. |
| Tanh | $$\displaystyle \tanh(x) = \frac{e^x - e^{-x}}{e^x + e^{-x}} $$ | (-1,1) | Zero-centered, steeper than sigmoid. | Vanishing gradients (saturates). | Hidden layers (older nets), RNNs. |
| ReLU | $$\displaystyle \text{ReLU}(x) = \max(0, x) $$ | [0, ∞) | Computationally cheap, mitigates vanishing gradient (positive slope). Sparse activations. | Dying ReLU (neurons stuck at 0 for negative inputs). | Default for hidden layers in CNNs/MLPs. |
| Leaky ReLU | $$\displaystyle \text{LReLU}(x) = \max(\alpha x, x) $$, $\alpha \approx 0.01$ | (-∞, ∞) | Fixes dying ReLU (small gradient for $$\displaystyle x<0 $$). | Results may vary; $\alpha$ is hyperparameter. | Alternative to ReLU. |
| ELU | $$\displaystyle \text{ELU}(x) = \begin{cases} x & x>0 \\ \alpha(e^x -1) & x \le 0 \end{cases} $$ | (-α, ∞) | Smoother near zero, pushes mean to zero. | Computationally heavier (exp). | Alternative to ReLU. |
| Softmax | $$\displaystyle \sigma(\mathbf{z})_j = \frac{e^{z_j}}{\sum_{k=1}^{K} e^{z_k}} $$ | (0,1), sums to 1 | Multi-class probability output. | Sensitive to outliers/naive implementations. | Output layer for multi-class classification. |
[!TIP] ReLU Dominance: ReLU's non-saturating positive region is key for training very deep networks (e.g., ResNet). Always use He initialization with ReLU.
Regularization Techniques
-
Overfitting vs. Underfitting:
-
Overfitting: Low training error, high test error. Model too complex, memorizes noise.
-
Underfitting: High training & test error. Model too simple, cannot capture pattern.
-
-
Dropout:
-
Mechanism: During training, randomly "drop" (set to 0) a fraction $p$ of neurons in a layer per mini-batch. At test time, use all neurons but scale activations by $(1-p)$ or use inverted dropout (scale during training).
-
Why it works: Prevents co-adaptation of neurons, forces network to learn redundant representations, acts as an ensemble of many thinned networks.
-
-
Weight Decay (L1/L2 Regularization):
-
Add penalty term to loss: $$\displaystyle L_{reg} = L + \lambda \|\mathbf{W}\|_p^p $$.
-
L2 (Ridge): $$\displaystyle \lambda \sum W_{ij}^2 $$. Shrinks weights uniformly.
-
L1 (Lasso): $$\displaystyle \lambda \sum |W_{ij}| $$. Encourages sparsity (some weights exactly zero).
-
Benefit: Constrains model complexity, reduces overfitting.
-
-
Regularization in Autoencoders: Techniques like denoising (corrupt input, reconstruct clean) or contractive (penalize Jacobian norm) force the AE to learn robust, meaningful features, not just identity mapping.
Gradient Problems & Solutions
-
Vanishing Gradient Problem:
-
Cause: Gradients multiplied by values $$\displaystyle <1 $$ repeatedly through layers (sigmoid/tanh saturation, deep networks).
-
Impact: Early layers learn extremely slowly or not at all. Severe in BPTT for RNNs (long-term dependencies lost).
-
Solutions: ReLU family, Batch Normalization, Residual Connections (ResNet), proper initialization (He/Xavier), LSTM/GRU (for RNNs).
-
-
Exploding Gradient Problem:
-
Cause: Gradients multiplied by values $$\displaystyle >1 $$ repeatedly (poor initialization, deep nets).
-
Impact: Unstable training, NaN weights.
-
Solutions: Gradient clipping (set max norm), proper initialization, Batch Norm, smaller learning rate.
-
-
Residual Connections (ResNet): $$\displaystyle y = F(x) + x $$. Identity shortcut allows gradient to flow directly backward, mitigates vanishing gradient in very deep (>100 layers) CNNs.
3. CONVOLUTIONAL NEURAL NETWORKS (CNNs)
Fundamental Concepts & Operations
-
Convolution Operation:
-
Purpose: Extract local spatial features (edges, textures) using learnable filters/kernels.
-
Mechanics: Kernel $K$ (size $$\displaystyle k \times k \times c_{in} $$) slides over input volume $I$ (width $w$, height $h$, channels $$\displaystyle c_{in} $$). Compute dot product at each spatial location: $$\displaystyle (I * K)_{i,j} = \sum_m \sum_n I_{i+m, j+n} K_{m,n} $$.
-
Key Parameters:
-
Stride ($s$): Step size of kernel. Larger stride → smaller output volume.
-
Padding ($p$): Add zeros around input border.
-
Valid: No padding ($$\displaystyle p=0 $$). Output shrinks.
-
Same: Output size = input size (with padding).
-
-
-
Output Size: $$\displaystyle \left\lfloor \frac{W - K + 2P}{S} \right\rfloor + 1 $$ (same for height).
-
-
Filters/Kernels: Each filter learns to detect a specific feature (e.g., vertical edge). Output volume depth = number of filters. Weight sharing across spatial locations makes CNNs parameter-efficient and translationally invariant.
Key Architectural Components
-
Pooling Layers (Downsampling):
-
Purpose: Reduce spatial dimensions (width/height), increase receptive field, provide translation invariance, reduce parameters/computation.
-
Max Pooling: Take maximum value in window (most common). Preserves dominant features.
-
Average Pooling: Take average. Used in some modern architectures (e.g., Inception).
-
Typical: $2 \times 2$ window, stride 2 → halves spatial dimensions.
-
-
Fully Connected (FC) Layers: At end of CNN, flatten final feature maps and apply MLP for classification/regression.
Classic CNN Architectures (Detailed Study)
| Architecture | Year | Key Innovations | Significance |
|---|---|---|---|
| LeNet-5 | 1998 | First successful CNN (handwritten digit recognition). Conv → Pool → Conv → Pool → FC → FC → Output. | Pioneered CNN structure for vision. |
| AlexNet | 2012 | ReLU (faster training), Dropout (regularization), GPU training, LRN (now obsolete), 5 conv + 3 FC layers. | Won ImageNet 2012, sparked modern DL revolution. Showed depth + ReLU + Dropout works. |
| ZFNet | 2013 | Improved AlexNet: smaller first filter (7x7→3x3), stride 2 in first conv, more filters. | Visualized learned features; showed importance of architecture tweaks. |
| GoogLeNet (Inception v1) | 2014 | Inception Module: Parallel convs (1x1, 3x3, 5x5) + pooling, concatenated. Auxiliary classifiers for gradient flow. | Won ImageNet 2014. Efficient use of parameters (1M vs AlexNet's 60M). Depth without huge parameter cost. |
| ResNet | 2015 | Residual Block: $$\displaystyle y = F(x) + x $$. Very deep (up to 152 layers). | Won ImageNet 2015 with human-level performance. Solved vanishing gradient for very deep nets via skip connections. |
Specialized CNN Topics
-
Structured Output in CNNs: Output is not a single label but a structure (e.g., pixel-wise segmentation map, bounding box coordinates). Requires fully convolutional architectures (no FC layers) or specialized heads (e.g., YOLO, Mask R-CNN).
-
Data Formats Compatible with CNNs:
| Data Type | Format | Example | | :--- | :--- | :--- | | Images | $(Height, Width, Channels)$ or $(Channels, Height, Width)$ | RGB: (224, 224, 3) | | Video | $(Frames, Height, Width, Channels)$ | (30, 224, 224, 3) | | 1D Signals | $(Timesteps, Channels)$ | Audio waveform: (16000, 1) | | 3D Volumes | $(Depth, Height, Width, Channels)$ | MRI scan: (128, 128, 128, 1) | | Graphs | $(Nodes, Features)$ + Adjacency Matrix | Social network, molecules. |
-
Deep Dream: Technique to visualize learned features. Amplify patterns in input image that activate specific neurons/filters by performing gradient ascent on the input (not weights). Generates surreal, dream-like images. Shows what CNNs "see".
4. RECURRENT NEURAL NETWORKS (RNNs) & SEQUENCE MODELS
Basic RNN Architecture
-
Core Idea: Have hidden state $$\displaystyle h_t $$ that captures information from previous time steps. Process sequences step-by-step.
-
Equations: $$\displaystyle h_t = \tanh(W_{hh} h_{t-1} + W_{xh} x_t + b_h) $$, $$\displaystyle y_t = W_{hy} h_t + b_y $$.
-
Unfolding Through Time: RNN can be seen as a deep network with shared parameters ($$\displaystyle W_{hh}, W_{xh} $$) across time steps. Handles variable-length sequences.
-
vs. Feedforward: RNNs have cycles (memory), process sequences; FFNs are acyclic, process fixed-size inputs independently.
Backpropagation Through Time (BPTT)
-
Process: Unfold RNN for $T$ steps → compute total loss $$\displaystyle L = \sum_{t=1}^{T} L_t $$ → apply standard backpropagation through the unfolded graph.
-
Challenges:
-
Vanishing/Exploding Gradients: Gradients must flow through many time steps. Multiplicative nature leads to exponential growth/shrinkage.
-
Computational Cost: $O(T)$ for forward/backward pass. Truncated BPTT used for long sequences.
-
Long-range dependencies: Hard to learn connections between distant time steps due to gradient issues.
-
Advanced RNN Cells
-
Long Short-Term Memory (LSTM):
-
Purpose: Solve vanishing gradient in standard RNNs for long-range dependencies.
-
Cell Structure (per timestep $t$):
-
Forget Gate: $$\displaystyle f_t = \sigma(W_f \cdot [h_{t-1}, x_t] + b_f) $$ → decides what to discard from cell state $$\displaystyle c_{t-1} $$.
-
Input Gate: $$\displaystyle i_t = \sigma(W_i \cdot [h_{t-1}, x_t] + b_i) $$, $$\displaystyle \tilde{c}_t = \tanh(W_c \cdot [h_{t-1}, x_t] + b_c) $$ → decides what new info to store.
-
Cell State Update: $$\displaystyle c_t = f_t \odot c_{t-1} + i_t \odot \tilde{c}_t $$. Additive update → stable gradient flow.
-
Output Gate: $$\displaystyle o_t = \sigma(W_o \cdot [h_{t-1}, x_t] + b_o) $$, $$\displaystyle h_t = o_t \odot \tanh(c_t) $$ → decides what to output as hidden state.
-
-
Advantages: Explicit memory cell ($$\displaystyle c_t $$) with additive updates; gates control information flow.
-
-
Gated Recurrent Unit (GRU):
-
Simplified LSTM: Merges forget/input gates into single update gate $$\displaystyle z_t $$. No separate cell state; hidden state $$\displaystyle h_t $$ is the memory.
-
Equations: $$\displaystyle z_t = \sigma(W_z \cdot [h_{t-1}, x_t]) $$, $$\displaystyle r_t = \sigma(W_r \cdot [h_{t-1}, x_t]) $$, $$\displaystyle \tilde{h}_t = \tanh(W \cdot [r_t \odot h_{t-1}, x_t]) $$, $$\displaystyle h_t = (1-z_t) \odot h_{t-1} + z_t \odot \tilde{h}_t $$.
-
Comparison: GRU has fewer parameters, often trains faster, performance similar to LSTM. LSTM may be better for longer sequences or more complex dependencies.
-
Sequence-to-Sequence Learning
-
Encoder-Decoder Framework:
-
Encoder (RNN): Reads input sequence $$\displaystyle \mathbf{x} = (x_1, ..., x_T) $$ → produces final hidden state (or context vector) $c$.
-
Decoder (RNN): Generates output sequence $$\displaystyle \mathbf{y} = (y_1, ..., y_{T'}) $$ conditioned on $c$ and previous outputs. Often uses teacher forcing during training.
-
-
Applications: Machine translation, text summarization, image captioning (CNN encoder + RNN decoder), speech recognition.
Specific Advanced Models
-
Deep RNNs: Stack multiple RNN layers. Lower layers learn short-term patterns, higher layers learn long-term abstractions.
-
PixelRNN / PixelCNN: Autoregressive models for image generation. Generate image pixel-by-pixel (row-major for PixelRNN, masked convs for PixelCNN). Condition each pixel on all previous pixels. Captures complex distribution $p(\text{image})$.
-
Recursive Neural Networks (RecNNs): Process tree-structured data (e.g., parse trees). Apply same neural network at each node, combining children's representations. Captures hierarchical compositionality (e.g., in NLP semantics).
5. AUTOENCODERS & REPRESENTATION LEARNING
Autoencoder (AE) Fundamentals
-
Architecture:
-
Encoder: $$\displaystyle z = f_{enc}(x; \theta_{enc}) $$ → maps input to latent code $z$ (bottleneck, low-dimensional).
-
Decoder: $$\displaystyle \hat{x} = f_{dec}(z; \theta_{dec}) $$ → reconstructs input from $z$.
-
Objective: Minimize reconstruction loss $L(x, \hat{x})$ (e.g., MSE for images, cross-entropy for binary).
-
-
Purpose: Unsupervised representation learning, dimensionality reduction, denoising, anomaly detection. Learns data-specific manifold.
Types of Autoencoders
| Type | Key Mechanism | Goal |
|---|---|---|
| Sparse AE | Add sparsity penalty (e.g., KL divergence to target sparsity $\rho$) on activations of hidden layer. | Learn disentangled, interpretable features (like edge detectors). |
| Denoising AE | Corrupt input $x$ (e.g., add Gaussian noise, drop pixels) → train to reconstruct original clean $x$. | Learn robust features invariant to noise. |
| Contractive AE | Add penalty on Jacobian of encoder: $$\displaystyle \lambda \| J_f(x) \|_F^2 $$. | Force encoder to be locally invariant to small input changes. |
Regularization in AEs
-
Problem: Without constraints, AE can learn trivial identity function ($$\displaystyle z=x $$, $$\displaystyle \hat{x}=x $$).
-
Solutions: Bottleneck (low-dimensional $z$), sparsity, denoising, contractive penalty. These force AE to learn meaningful compressed representation.
Comparison with PCA/SVD
| Aspect | PCA/SVD | Autoencoders |
|---|---|---|
| Linearity | Linear transformation. | Non-linear (with non-linear activations). |
| Manifold | Assumes data lies on linear subspace. | Can learn non-linear manifolds. |
| Flexibility | Fixed objective (maximize variance). | Flexible: can be tailored (denoising, sparse) for specific tasks. |
| When to use AE? | When data has non-linear structure (e.g., images, text) or you need task-specific features (e.g., for classification). PCA for quick, linear, global analysis. |
[!TIP] Exam Justification: "Use Autoencoders instead of PCA when the data lies on a non-linear manifold (e.g., images of faces, handwritten digits) or when you need domain-specific features for downstream tasks."
6. GENERATIVE MODELS
Variational Autoencoders (VAEs)
-
Probabilistic Framework: Learn distribution $p(x)$ by approximating posterior $p(z|x)$ with encoder $$\displaystyle q_\phi(z|x) $$.
-
Reparameterization Trick: Sample $$\displaystyle z = \mu_\phi(x) + \sigma_\phi(x) \odot \epsilon $$, $\epsilon \sim \mathcal{N}(0,I)$. Allows gradient flow through sampling.
-
Objective (ELBO - Evidence Lower Bound):
$$\mathcal{L}_{VAE} = \mathbb{E}_{q_\phi(z|x)}[\log p_\theta(x|z)] - D_{KL}(q_\phi(z|x) \| p(z))$$
* **Reconstruction Term:** $$\displaystyle \mathbb{E}[\log p_\theta(x|z)] $$ (e.g., MSE or cross-entropy). Maximize likelihood of data.
* **Regularization Term:** $$\displaystyle D_{KL}(q_\phi(z|x) \| p(z)) $$ (usually $\mathcal{N}(0,I)$). Forces latent space to be **structured** (continuous, smooth).
- Training: Maximize ELBO (equivalently, minimize $$\displaystyle -\mathcal{L}_{VAE} $$) via SGD.
Generative Adversarial Networks (GANs)
-
Architecture: Two networks play minimax game:
-
Generator $G$: Takes random noise $$\displaystyle z \sim p_z $$ → generates fake data $$\displaystyle \tilde{x} = G(z) $$.
-
Discriminator $D$: Takes real $x$ or fake $\tilde{x}$ → outputs probability $D(x)$ of being real.
-
-
Objective:
$$\min_G \max_D V(D,G) = \mathbb{E}_{x \sim p_{data}}[\log D(x)] + \mathbb{E}_{z \sim p_z}[\log(1 - D(G(z)))]$$
* **D's goal:** Maximize $\log D(x) + \log(1-D(G(z)))$ → distinguish real vs fake.
* **G's goal:** Minimize $\log(1-D(G(z)))$ → fool D.
-
Training Dynamics: Alternate updates: 1) Fix $G$, update $D$ (maximize). 2) Fix $D$, update $G$ (minimize).
-
Challenges: Mode collapse (G generates limited modes), training instability, difficult to evaluate.
Comparison & Hybrids
| Aspect | VAE | GAN |
|---|---|---|
| Density | Explicit (learns $p(x)$ via ELBO). | Implicit (no explicit density, only samples). |
| Training | Stable (single optimization). | Unstable (minimax game). |
| Sample Quality | Blurry (due to Gaussian decoder assumption). | Sharp, high-fidelity samples. |
| Latent Space | Structured, continuous (good for interpolation). | Less structured, no direct encoder. |
| Use Case | Representation learning, controlled generation, semi-supervised. | Pure generation (images, audio) where quality > diversity/structure. |
| VAE-GAN Hybrid | Combine VAE encoder + GAN discriminator. Use VAE's reconstruction loss + GAN's adversarial loss for decoder. | Get sharp samples (from GAN) and meaningful latent space (from VAE). Example: PixelCNN + VAE or BEGAN. |
Energy-Based & Autoregressive Models
-
Restricted Boltzmann Machines (RBMs):
-
Bipartite graph: Visible layer $v$ (data), hidden layer $h$. No visible-visible or hidden-hidden connections.
-
Energy Function: $$\displaystyle E(v,h) = -b^T v - c^T h - v^T W h $$.
-
Gibbs Sampling: For training (Contrastive Divergence): 1) Positive phase: $v$ = data. 2) Sample $h \sim p(h|v)$. 3) Negative phase: Sample $\tilde{v} \sim p(v|h)$ (reconstruction). Update $W$ based on $$\displaystyle \langle vh \rangle_{data} - \langle vh \rangle_{recon} $$.
-
-
Autoregressive Models: Factorize joint distribution as product of conditionals: $$\displaystyle p(x) = \prod_{i=1}^{d} p(x_i | x_{<i}) $$.
-
NADE (Neural Autoregressive Distribution Estimator): Uses a neural network with masked weights to ensure autoregressive property. Each unit depends only on previous inputs.
-
MADE (Masked Autoencoder for Distribution Estimation): Applies fixed binary masks to an autoencoder's weights to enforce autoregressive property. Efficient, flexible.
-
-
Markov Networks (Undirected Graphical Models): Represent dependencies via potential functions $\phi(C)$ over cliques $C$. No directed edges. Energy-based. RBMs are a special case.
Deep Belief Networks (DBNs)
-
Architecture: Stack of RBMs. Train greedily, layer-wise: train first RBM on data → use its hidden activations as "data" for next RBM → repeat.
-
Historical Context: Pre-2012, main method for pre-training deep networks (since backprop struggled with deep, randomly initialized nets). Now largely obsolete, replaced by better initialization, ReLU, and supervised training from scratch.
7. ADVANCED TOPICS & SPECIALIZED ARCHITECTURES
Deep Reinforcement Learning (Deep RL)
-
Integration: Use deep neural networks to represent value functions or policies in RL.
-
Markov Decision Process (MDP): $(\mathcal{S}, \mathcal{A}, P, R, \gamma)$.
-
$\mathcal{S}$: States, $\mathcal{A}$: Actions.
-
$P(s'|s,a)$: Transition probability.
-
$R(s,a,s')$: Reward.
-
$\gamma$: Discount factor.
-
-
Value Iteration (Dynamic Programming): Compute optimal state-value $$\displaystyle V^*(s) $$ iteratively:
$$V_{k+1}(s) = \max_a \sum_{s'} P(s'|s,a) [R(s,a,s') + \gamma V_k(s')]$$
Converges to $$\displaystyle V^* $$.
-
Policy Iteration: Alternate: 1) Policy Evaluation: Compute $$\displaystyle V^\pi $$ for current policy $\pi$. 2) Policy Improvement: $$\displaystyle \pi'(s) = \arg\max_a \sum_{s'} P(s'|s,a)[R + \gamma V^\pi(s')] $$. Converges to $$\displaystyle \pi^* $$.
-
Q-Learning (Model-Free): Learn action-value $Q(s,a)$ off-policy:
$$Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma \max_{a'} Q(s',a') - Q(s,a)]$$
-
Deep Q-Networks (DQN): Use deep CNN to approximate $Q(s,a;\theta)$. Key innovations: Experience replay (break correlations), target network (stable targets).
-
Advanced DQN:
-
Double DQN: Decouple action selection (from online net) and evaluation (from target net) to reduce overestimation bias.
-
Dueling DQN: Separate streams for state-value $V(s)$ and advantage $A(s,a)$: $$\displaystyle Q(s,a) = V(s) + A(s,a) $$. Learns better value estimates.
-
-
Least Squares Policy Iteration (LSPI): Policy iteration using LSTD (Least-Squares Temporal Difference) for policy evaluation (linear function approximation). More sample-efficient than TD.
Model Compression & Efficiency
-
Unit/Neuron Pruning:
-
Concept: Remove unimportant neurons (or weights) from trained network to reduce size/inference time.
-
Need: Deploy models on edge devices (mobile, IoT) with limited compute/memory.
-
Methods: 1) Magnitude-based: Prune small-magnitude weights. 2) Activation-based: Prune neurons with low average activation. 3) Second-order: Use Hessian to estimate importance. Iterative: prune → fine-tune.
-
Representation Learning
-
General Principle: Learn meaningful, abstract features from raw data (pixels, words) automatically, without manual feature engineering.
-
Importance: Core of deep learning. Enables models to capture hierarchical, disentangled representations (edges → parts → objects). Foundation for transfer learning, semi-supervised learning, and generative models.
8. NATURAL LANGUAGE PROCESSING (NLP) & APPLICATIONS
NLP Fundamentals
-
Definition: Field at intersection of AI, linguistics, CS focused on enabling computers to understand, interpret, manipulate human language.
-
Core Elements:
-
Syntax: Structure of sentences (grammar, parsing).
-
Semantics: Meaning of words/sentences.
-
Pragmatics: Contextual meaning (intent, discourse).
-
Morphology: Word formation (stems, affixes).
-
DL Applications in NLP
-
Word Embeddings (Word2Vec, GloVe): Dense vector representations ($$\displaystyle \mathbb{R}^{300} $$) capturing semantic similarity ("king" - "man" + "woman" ≈ "queen"). Pre-trained embeddings are standard input for DL models.
-
RNNs/LSTMs for Language Modeling: Predict next word given history. Used in translation, speech recognition.
-
Sequence-to-Sequence with Attention:
-
Encoder: RNN/LSTM/Transformer encodes input sequence into context vectors.
-
Decoder: Generates output token-by-token.
-
Attention Mechanism: Allow decoder to focus on different parts of input sequence at each output step. Crucial for long sequences (e.g., translation). Formula: $$\displaystyle \text{score}(h_i, s_t) = v^T \tanh(W_1 h_i + W_2 s_t) $$, then softmax → weights → context vector.
-
Broad Application Areas of Deep Learning
| Domain | Key DL Models | Example Tasks |
|---|---|---|
| Computer Vision | CNNs (ResNet, EfficientNet), Vision Transformers | Image classification, object detection (YOLO), segmentation (U-Net). |
| Speech Recognition | RNNs, CTC, Transformers (Whisper) | Transcribe audio to text. |
| Generative Modeling | GANs, VAEs, Diffusion Models | Image synthesis (StyleGAN), text generation (GPT), music. |
| Recommendation Systems | Neural Collaborative Filtering, Two-tower models | Product/movie recommendations (YouTube, Netflix). |
| Game Playing | Deep RL (DQN, A3C, AlphaGo) | Learn policies from pixels/rewards. |
PRIORITY-BASED EXAM FOCUS SUMMARY
Highest Frequency (Must Master)
-
Backpropagation: Derivation, computation graph, applications.
-
CNNs: Convolution operation (filters, stride, padding), pooling (max vs avg), classic architectures (AlexNet, ResNet, LeNet).
-
RNNs/LSTMs: Architecture, BPTT challenges, LSTM cell structure (gates, cell state), GRU vs LSTM.
-
Optimization: Adam (derive update), SGD vs Batch GD vs Mini-batch.
-
Autoencoders: Architecture, types (sparse, denoising), PCA vs AE.
-
Generative Models: GANs (minimax game), VAEs (ELBO, reparameterization), GANs vs VAEs comparison.
-
Activation Functions: ReLU (why dominant), sigmoid/tanh issues.
-
Vanishing Gradient: Causes, impact, solutions (ReLU, BatchNorm, ResNet).
-
Batch Normalization: Mechanics, purpose (internal covariate shift), advantages.
High Frequency (Should Know Well)
-
MLP & Universal Approximation.
-
Weight Initialization (Xavier, He).
-
Dropout mechanism.
-
Weight Decay (L1/L2).
-
Structured Output in CNNs.
-
Deep RL basics (MDP, Q-learning, DQN).
-
Encoding/Decoding in RNNs.
-
GoogLeNet (Inception module).
Moderate/Lower Frequency (Be Familiar)
-
History milestones (AlexNet year, etc.).
-
Deep Dream.
-
NADE/MADE.
-
DBNs (greedy pre-training).
-
Unit Pruning.
-
Representation Learning principles.
-
NLP elements (4 elements, word embeddings).
[!TIP] Final Exam Strategy: For 7-mark questions, structure answer as: 1-2 sentence definition → 3-4 step explanation/derivation → 1-2 line application/significance. Use diagrams where possible (e.g., LSTM cell, CNN layers, computation graph). Always highlight key terms in bold.