UNIT 4: Deep Learning
1. Foundations of Deep Learning
Historical Progression & Milestones
-
Key Breakthroughs: Backpropagation (1986), ReLU (2010), AlexNet (2012), Transformers (2017).
-
Evolution: Perceptrons (1950s) → MLPs (1980s) → Deep Networks (Post-2010, fueled by big data & GPUs).
AI vs. ML vs. DL
| Aspect | Artificial Intelligence (AI) | Machine Learning (ML) | Deep Learning (DL) |
|---|---|---|---|
| Goal | Create systems that perform human-like tasks | Learn patterns from data without explicit programming | Subfield of ML using deep neural networks |
| Feature Engineering | Manual & Knowledge-based | Manual | Automatic (hierarchical feature learning) |
| Data Dependency | Varies | Moderate to High | Very High (needs large datasets) |
| Example | Chess-playing system | Spam classifier | Image recognition with CNNs |
Biological vs. Artificial Neural Networks
-
Analogies: Neurons ↔ Artificial neurons, synapses ↔ weights, brain regions ↔ network layers.
-
Key Differences:
-
Scale: Brain has ~86B neurons; ANNs have far fewer.
-
Learning: Biological learning is unsupervised, continual, and energy-efficient. DL is mostly supervised, episodic, and compute-intensive.
-
Architecture: Brain is sparse, dynamically wired, and plastic. ANNs are dense, static during inference.
-
-
Limitations of DL vs. Brain: Lack of common sense, poor data efficiency, catastrophic forgetting, no true understanding, high energy consumption.
Perceptrons and Multilayer Perceptrons (MLPs)
-
Single-Layer Perceptron:
-
Model: $$\displaystyle y = f(\mathbf{w}^T\mathbf{x} + b) $$, where $f$ is a step function.
-
Limitation: Can only learn linearly separable patterns (e.g., AND, OR). Fails on XOR.
-
-
Multilayer Perceptron (MLP):
-
Architecture: Input layer → one or more fully connected (dense) hidden layers → Output layer. Each neuron applies: $$\displaystyle z = \mathbf{w}^T\mathbf{a}^{prev} + b $$, $$\displaystyle a = f(z) $$.
-
Overcoming Limitations: Multiple layers with non-linear activation functions allow learning complex, non-linear decision boundaries.
-
Universal Approximation Theorem: A feedforward network with a single hidden layer containing a finite number of neurons and appropriate non-linear activations can approximate any continuous function on compact subsets of $$\displaystyle \mathbb{R}^n $$ to any desired accuracy.
-
Activation Functions
-
Purpose: Introduce non-linearity, enabling the network to learn complex mappings.
-
Types & Comparison:
| Function | Formula | Range | Pros | Cons |
|---|---|---|---|---|
| Sigmoid | $$\displaystyle \sigma(x) = \frac{1}{1+e^{-x}} $$ | (0,1) | Smooth, outputs probability | Vanishing gradient, not zero-centered, slow |
| Tanh | $$\displaystyle \tanh(x) = \frac{e^x - e^{-x}}{e^x + e^{-x}} $$ | (-1,1) | Zero-centered, steeper than sigmoid | Vanishing gradient |
| ReLU | $$\displaystyle \text{ReLU}(x) = \max(0, x) $$ | [0, ∞) | Computationally cheap, mitigates vanishing gradient in +ve region, induces sparsity | Dying ReLU problem (neurons can get stuck) |
| Leaky ReLU | $$\displaystyle f(x) = \max(\alpha x, x) $$ | (-∞, ∞) | Fixes dying ReLU (small gradient for $$\displaystyle x<0 $$) | Needs tuning of $\alpha$ |
| ELU | $$\displaystyle f(x) = \begin{cases} x & x>0 \\ \alpha(e^x -1) & x \le 0 \end{cases} $$ | (-α, ∞) | Smoother, pushes mean to zero | Computationally heavier |
| Softmax | $$\displaystyle \sigma(\mathbf{z})_j = \frac{e^{z_j}}{\sum_{k=1}^K e^{z_k}} $$ | (0,1), sums to 1 | Used for multi-class classification output |
[!TIP] Exam Focus: ReLU is the default for hidden layers. Use Softmax only for the final layer in multi-class classification. Sigmoid/Tanh are now rarely used in hidden layers due to vanishing gradients.
Weight Initialization
-
Importance: Breaks symmetry between neurons. Poor initialization leads to vanishing/exploding gradients.
-
Methods:
-
Random Uniform/Normal: Simple but problematic for deep nets.
-
Xavier/Glorot Initialization: For Tanh/Sigmoid. $$\displaystyle w \sim \mathcal{U}\left(-\sqrt{\frac{6}{n_{in}+n_{out}}}, \sqrt{\frac{6}{n_{in}+n_{out}}}\right) $$ or $$\displaystyle \mathcal{N}(0, \sqrt{\frac{2}{n_{in}+n_{out}}}) $$. Keeps variance stable across layers.
-
He Initialization: For ReLU and variants. $$\displaystyle w \sim \mathcal{N}(0, \sqrt{\frac{2}{n_{in}}}) $$. Compensates for ReLU's half-rectification.
-
Backpropagation Algorithm
-
Goal: Compute gradient of loss w.r.t. all weights to update them via gradient descent.
-
Steps:
-
Forward Pass: Compute output $\hat{y}$ and loss $L(\hat{y}, y)$.
-
Backward Pass (Chain Rule): Compute $$\displaystyle \frac{\partial L}{\partial w} $$ for each weight, layer by layer from output to input.
- $$\displaystyle \frac{\partial L}{\partial w^{(l)}} = \frac{\partial L}{\partial a^{(l)}} \cdot \frac{\partial a^{(l)}}{\partial z^{(l)}} \cdot \frac{\partial z^{(l)}}{\partial w^{(l)}} $$
-
Weight Update: $$\displaystyle w^{(l)} \leftarrow w^{(l)} - \eta \frac{\partial L}{\partial w^{(l)}} $$ (for SGD).
-
-
Applications: Training all feedforward neural networks (MLPs, CNNs).
Gradient Descent Variants
| Variant | Batch Size | Pros | Cons |
|---|---|---|---|
| Batch GD | Entire dataset | Stable convergence, exact gradient | Very slow, memory heavy |
| Stochastic GD (SGD) | 1 sample | Fast updates, can escape shallow minima | Noisy gradients, unstable |
| Mini-batch GD | $n$ samples (e.g., 32, 64) | Compromise: Faster than Batch, less noisy than SGD. Most common. | Needs batch size tuning |
Optimization Algorithms (Adaptive Learning Rates)
| Algorithm | Core Idea | Key Formula/Update | Best For |
|---|---|---|---|
| Momentum | Accumulate past gradients to dampen oscillations & accelerate. | $$\displaystyle v_t = \gamma v_{t-1} + \eta \nabla L(\theta_t) $$; $$\displaystyle \theta_{t+1} = \theta_t - v_t $$ | Ravines, noisy gradients |
| AdaGrad | Adapt per-parameter LR by accumulating squared gradients. | $$\displaystyle \theta_{t+1} = \theta_t - \frac{\eta}{\sqrt{G_t + \epsilon}} \odot g_t $$, $$\displaystyle G_t = \sum_{\tau=1}^t g_\tau^2 $$ | Sparse data (e.g., NLP) |
| RMSProp | Modify AdaGrad to prevent aggressive LR decay by using exponentially weighted moving average of squared grads. | $$\displaystyle E[g^2]_t = \beta E[g^2]_{t-1} + (1-\beta) g_t^2 $$; $$\displaystyle \theta_{t+1} = \theta_t - \frac{\eta}{\sqrt{E[g^2]_t + \epsilon}} g_t $$ | Non-stationary objectives, RNNs |
| Adam | Combine Momentum & RMSProp + Bias Correction. | $$\displaystyle m_t = \beta_1 m_{t-1} + (1-\beta_1)g_t $$ (1st moment)<br>$$\displaystyle v_t = \beta_2 v_{t-1} + (1-\beta_2)g_t^2 $$ (2nd moment)<br>$$\displaystyle \hat{m}_t = m_t/(1-\beta_1^t) $$, $$\displaystyle \hat{v}_t = v_t/(1-\beta_2^t) $$<br>$$\displaystyle \theta_{t+1} = \theta_t - \eta \frac{\hat{m}_t}{\sqrt{\hat{v}_t + \epsilon}} $$ | Default choice for most problems. Robust. |
Challenges in Training Deep Networks
-
Vanishing Gradient Problem:
-
Cause: Repeated multiplication of small gradients (e.g., from sigmoid/tanh) during backprop through many layers.
-
Impact: Early layers learn very slowly or not at all.
-
Mitigation: Use ReLU, Batch Normalization, Residual Connections (ResNet), proper weight initialization (He), Gradient Clipping.
-
-
Exploding Gradient Problem:
-
Cause: Large gradients (often in RNNs) amplified through many layers.
-
Impact: Unstable training, weight updates become huge.
-
Mitigation: Gradient Clipping (set max norm), Weight Regularization (L2), smaller learning rate.
-
Regularization Techniques
-
Purpose: Prevent overfitting (high training accuracy, low validation accuracy), improve generalization.
-
Methods:
-
L1/L2 Weight Decay: Add penalty $$\displaystyle \lambda \|\mathbf{w}\|_1 $$ or $$\displaystyle \lambda \|\mathbf{w}\|_2^2 $$ to loss. L2 is more common.
-
Dropout: During training, randomly deactivate a fraction $p$ of neurons in a layer with probability $p$. Effectively trains an ensemble of thinned networks. Not used at test time.
-
Early Stopping: Stop training when validation loss stops improving.
-
Data Augmentation: Artificially increase dataset size by applying transformations (rotation, crop, flip for images).
-
Batch Normalization
-
Concept: Normalize the inputs to a layer (pre-activation) to have zero mean and unit variance per mini-batch.
-
$$\displaystyle \hat{x}^{(k)} = \frac{x^{(k)} - \mu_B}{\sqrt{\sigma_B^2 + \epsilon}} $$ (minibatch mean $$\displaystyle \mu_B $$, variance $$\displaystyle \sigma_B^2 $$)
-
Then apply learnable scale and shift: $$\displaystyle y^{(k)} = \gamma \hat{x}^{(k)} + \beta $$.
-
-
Working:
-
Training: Compute batch stats, normalize, apply $\gamma, \beta$. Update running averages of mean/var.
-
Inference: Use stored running averages to normalize. No batch stats.
-
-
Advantages:
-
Faster convergence (allows higher learning rates).
-
Reduces internal covariate shift (distribution changes in layer inputs).
-
Has a slight regularization effect (noise from batch stats).
-
Helps mitigate vanishing gradients.
-
Data Preprocessing
-
Normalization (Min-Max Scaling): $$\displaystyle x' = \frac{x - x_{min}}{x_{max} - x_{min}} $$. Scales to [0, 1]. Sensitive to outliers.
-
Standardization (Z-score): $$\displaystyle x' = \frac{x - \mu}{\sigma} $$. Scales to mean=0, std=1. More robust, preferred for gradient descent.
-
Impact: Essential for stable and fast gradient descent. Prevents features with large scales from dominating.
Overfitting and Underfitting
| Condition | Training Error | Validation Error | Cause | Solution |
|---|---|---|---|---|
| Underfitting | High | High | Model too simple (high bias) | Increase model capacity (more layers/neurons), train longer, better features |
| Overfitting | Low | High | Model too complex (high variance) | More data, regularization (dropout, weight decay), batch norm, early stopping, reduce model size |
2. Convolutional Neural Networks (CNNs)
Convolution Operation
-
Significance: Exploits spatial/temporal locality and translation equivariance. Drastically reduces parameters via parameter sharing.
-
Mechanics:
-
Kernel/Filter $\mathbf{W}$ (e.g., $$\displaystyle 3 \times 3 \times C_{in} $$) slides over input feature map.
-
Stride ($s$): Step size of filter movement.
-
Output Size: $$\displaystyle \left\lfloor \frac{W_{in} - K + 2P}{s} \right\rfloor + 1 $$ (same for height).
-
Produces a feature map highlighting where a particular pattern (learned by filter) is present.
-
-
Capturing Spatial Features: Local receptive fields allow early layers to detect edges/textures. Stacked layers build hierarchical representations (edges → parts → objects).
Filters/Kernels
-
Role: Learnable feature detectors. First layer filters learn simple features (edges, colors). Deeper layers learn complex patterns (object parts).
-
Parameters: Number of filters = depth of output volume. Each filter has spatial dimensions (e.g., $3\times3$) and spans full input depth.
Padding
-
Valid Padding: No padding. Output shrinks: $$\displaystyle (W_{in} - K + 1) $$.
-
Same Padding: Pad input so output size equals input size (if stride=1). $$\displaystyle P = \frac{K-1}{2} $$ (for odd K). Preserves edge information.
Pooling Layers
-
Role: Downsampling (reduces spatial size, parameters, computation), provides translation invariance, abstracts features.
-
Max Pooling: Output = max value in window. Preserves dominant features, most common.
-
Average Pooling: Output = average in window. Smooths features, used in some architectures (e.g., GoogLeNet).
-
Comparison: Max pooling is generally preferred for retaining salient features. Average pooling can lose strong activations.
CNN Architecture
Typical Stack: [Conv → Activation (ReLU) → Pooling] × N → Flatten → [FC → Activation] × M → Output Layer.
Structured Output in CNNs
-
Goal: Produce spatial output (e.g., per-pixel labels for segmentation, bounding boxes for detection).
-
Architectures:
-
Fully Convolutional Networks (FCNs): Replace final FC layers with convolutional layers to produce spatial output. Enable end-to-end training for segmentation.
-
Skip Connections/Upsampling Paths: Combine coarse, high-level features from deep layers with fine, low-level features from shallow layers (e.g., U-Net) to recover spatial details.
-
Key CNN Architectures
| Architecture | Year | Key Innovations | Impact |
|---|---|---|---|
| LeNet-5 | 1998 | First successful CNN (LeCun). Conv → Pool → FC. Designed for MNIST. | Proof-of-concept for digit recognition. |
| AlexNet | 2012 | ReLU, Dropout, GPU training, LRN (later obsolete), 5 conv layers. | Won ImageNet 2012, sparked the deep learning revolution. |
| ZFNet | 2013 | Visualization-guided tweaks to AlexNet: smaller 1st filter ($7\times7 \to 3\times3$), more filters, stride 2 in conv1. | Showed architecture design matters, improved AlexNet baseline. |
| GoogLeNet (Inception v1) | 2014 | Inception Module: Parallel convs (1x1, 3x3, 5x5) + pooling → concatenate. Auxiliary classifiers for gradient flow. Very efficient (22 layers, 12x fewer params than AlexNet). | Won ImageNet 2014. Emphasized computational efficiency. |
| ResNet | 2015 | Residual Block: $$\displaystyle \mathbf{y} = \mathcal{F}(\mathbf{x}, \{W_i\}) + \mathbf{x} $$. Skip connections solve vanishing gradient, enable 100+ layer networks. | Won ImageNet 2015 with 152 layers. Fundamental for very deep nets. |
Applications of CNNs
-
Image Classification (e.g., ResNet).
-
Object Detection (e.g., YOLO, Faster R-CNN).
-
Semantic Segmentation (e.g., FCN, U-Net).
-
Face Recognition.
-
Extensions: Video (3D convolutions), Spectrograms (audio as image), Graph data (GCNs).
3. Recurrent Neural Networks (RNNs) and Variants
RNN Fundamentals
-
Architecture: Has a hidden state $$\displaystyle \mathbf{h}_t $$ that acts as memory. $$\displaystyle \mathbf{h}_t = f(\mathbf{h}_{t-1}, \mathbf{x}_t; W) $$. Parameters are shared across time steps.
-
vs. Feedforward: Can handle variable-length sequences, has temporal memory. Feedforward has no memory, fixed input size.
Backpropagation Through Time (BPTT)
-
Process: Unfold the RNN for $T$ time steps into a deep feedforward network. Apply standard backpropagation through all $T$ steps.
-
Challenges:
-
Computational Cost: $O(T)$ per sequence.
-
Memory: Need to store activations for all $T$ steps for gradient computation.
-
Long-term dependencies: Gradients must flow through many steps → vanishing/exploding gradients.
-
Vanishing/Exploding Gradients in RNNs
-
Cause: Repeated multiplication of the Jacobian matrix $$\displaystyle \frac{\partial \mathbf{h}_t}{\partial \mathbf{h}_{t-k}} $$ during BPTT. If eigenvalues $$\displaystyle <1 $$, gradients vanish. If $$\displaystyle >1 $$, they explode.
-
Impact: Inability to learn dependencies beyond ~10-20 time steps.
-
Mitigation: Gating mechanisms (LSTM, GRU), Gradient Clipping, Proper initialization (orthogonal), Skip connections.
Long Short-Term Memory (LSTM)
-
Architecture: Introduces a cell state $$\displaystyle \mathbf{c}_t $$ (the "memory highway") and three gates to regulate information flow.
-
Forget Gate: $$\displaystyle f_t = \sigma(W_f \cdot [\mathbf{h}_{t-1}, \mathbf{x}_t] + b_f) $$. What to remove from cell state?
-
Input Gate: $$\displaystyle i_t = \sigma(W_i \cdot [\mathbf{h}_{t-1}, \mathbf{x}_t] + b_i) $$. What new info to store?
-
Candidate Cell State: $$\displaystyle \tilde{c}_t = \tanh(W_c \cdot [\mathbf{h}_{t-1}, \mathbf{x}_t] + b_c) $$.
-
Cell State Update: $$\displaystyle \mathbf{c}_t = f_t \odot \mathbf{c}_{t-1} + i_t \odot \tilde{c}_t $$.
-
Output Gate: $$\displaystyle o_t = \sigma(W_o \cdot [\mathbf{h}_{t-1}, \mathbf{x}_t] + b_o) $$. What to output from cell state?
-
Hidden State: $$\displaystyle \mathbf{h}_t = o_t \odot \tanh(\mathbf{c}_t) $$.
-
-
How it Mitigates Vanishing Gradients: Additive nature of cell state update ($$\displaystyle \mathbf{c}_t = f_t \odot \mathbf{c}_{t-1} + ... $$) allows gradients to flow almost unchanged through the $\mathbf{c}$ path if forget gate is near 1.
-
Advantages over Simple RNN: Learns long-term dependencies, robust to vanishing gradients, widely successful in NLP.
Gated Recurrent Unit (GRU)
-
Architecture: Simpler than LSTM. No separate cell state. Uses two gates:
-
Update Gate: $$\displaystyle z_t = \sigma(W_z \cdot [\mathbf{h}_{t-1}, \mathbf{x}_t]) $$. How much past info to keep?
-
Reset Gate: $$\displaystyle r_t = \sigma(W_r \cdot [\mathbf{h}_{t-1}, \mathbf{x}_t]) $$. How much past to forget?
-
Candidate Hidden State: $$\displaystyle \tilde{\mathbf{h}}_t = \tanh(W \cdot [r_t \odot \mathbf{h}_{t-1}, \mathbf{x}_t]) $$.
-
Hidden State: $$\displaystyle \mathbf{h}_t = (1 - z_t) \odot \mathbf{h}_{t-1} + z_t \odot \tilde{\mathbf{h}}_t $$.
-
-
Comparison with LSTM:
-
Fewer parameters (no cell state, 2 vs 3 gates).
-
Computationally faster.
-
Performance often comparable to LSTM on many tasks, but LSTM may be better for longer sequences or complex dependencies.
-
Encoding and Decoding in RNNs
-
Sequence-to-Sequence (Seq2Seq) Model:
-
Encoder RNN: Reads input sequence $$\displaystyle \mathbf{x}_1, ..., \mathbf{x}_T $$ and produces a context vector (final hidden state $$\displaystyle \mathbf{h}_T $$ or a combination). Compresses input into fixed-size representation.
-
Decoder RNN: Initialized with context vector, generates output sequence $$\displaystyle \mathbf{y}_1, ..., \mathbf{y}_{T'} $$ autoregressively ($$\displaystyle \mathbf{y}_t $$ becomes input for $t+1$).
-
-
Applications: Machine translation, text summarization, image captioning (CNN encoder + RNN decoder).
Deep Recurrent Neural Networks
-
Stacked RNNs: Multiple RNN layers. Lower layers learn short-term patterns, higher layers learn long-term, abstract representations.
-
Bidirectional RNNs (BiRNN): Two RNNs process sequence forward and backward. Final output at $t$ is concatenation of both directions' hidden states. Provides full past and future context at each position. Crucial for tasks like POS tagging, NER.
Applications of Deep RNNs
-
NLP: Language modeling, machine translation, sentiment analysis, speech recognition.
-
Time Series: Stock prediction, sensor data analysis.
-
Image Processing: Image captioning (CNN-RNN), video classification (frame-wise RNN), PixelRNN (autoregressive image generation).
4. Autoencoders and Dimensionality Reduction
Autoencoder Architecture
-
Components:
-
Encoder: $$\displaystyle z = f_{enc}(\mathbf{x}; \theta_{enc}) $$ (compresses input $\mathbf{x}$ to latent code $\mathbf{z}$).
-
Bottleneck/Latent Space: Low-dimensional representation $\mathbf{z}$.
-
Decoder: $$\displaystyle \hat{\mathbf{x}} = f_{dec}(\mathbf{z}; \theta_{dec}) $$ (reconstructs $\mathbf{x}$ from $\mathbf{z}$).
-
-
Objective: Minimize reconstruction loss $L(\mathbf{x}, \hat{\mathbf{x}})$ (e.g., MSE for real-valued, binary cross-entropy for binary inputs).
-
Training: Unsupervised (no labels needed). Forces network to learn efficient data encoding.
Purpose and Applications
-
Dimensionality Reduction: Non-linear alternative to PCA.
-
Feature Learning: Latent space $\mathbf{z}$ can be used as features for downstream tasks.
-
Denoising: Train with noisy input $\tilde{\mathbf{x}}$, reconstruct clean $\mathbf{x}$ (Denoising Autoencoder).
-
Anomaly Detection: Train on normal data; high reconstruction error indicates anomaly.
-
Pretraining: Initialize deep network weights (less common now with better initialization & Adam).
Comparison with PCA/SVD
| Aspect | PCA/SVD | Autoencoder (Deep) |
|---|---|---|
| Transformation | Linear, orthogonal projection | Non-linear (if deep/non-linear activations) |
| Objective | Maximize variance / minimize reconstruction error (linear) | Minimize reconstruction error (flexible) |
| Principal Components | Orthogonal, ordered by variance | Latent dimensions not necessarily orthogonal or ordered |
| Global vs Local | Captures global linear structure | Can capture complex, non-linear manifolds |
| Use Autoencoder When: Data lies on a non-linear manifold, need non-linear features, or want to use supervised pretraining (with labels). |
Types of Autoencoders
| Type | Key Idea | Regularization Mechanism |
|---|---|---|
| Sparse Autoencoder | Enforce sparsity in latent representation (few active neurons). | Add KL divergence penalty: $$\displaystyle \beta \sum_j D_{KL}(\rho \| \hat{\rho}_j) $$ where $$\displaystyle \hat{\rho}_j $$ is average activation of neuron $j$. |
| Contractive Autoencoder | Make encoder robust to small perturbations in input. | Penalize squared Frobenius norm of Jacobian of encoder: $$\displaystyle \lambda \|\mathbf{J}_{enc}(\mathbf{x})\|_F^2 $$. |
| Denoising Autoencoder | Learn to reconstruct clean input from corrupted version. | Corruption is input (e.g., Gaussian noise, masking). Loss computed on clean target. |
Regularization in Autoencoders
-
Goal: Prevent the trivial solution where autoencoder learns identity function ($$\displaystyle \mathbf{z}=\mathbf{x} $$, $$\displaystyle \hat{\mathbf{x}}=\mathbf{x} $$).
-
Techniques: Sparsity constraint, contractive penalty, dropout in encoder/decoder, noise injection (denoising AE), bottleneck size smaller than input.
5. Generative Models
Variational Autoencoders (VAEs)
-
Latent Variable Model: Assumes data $\mathbf{x}$ generated from latent variable $\mathbf{z} \sim p(\mathbf{z})$ (usually $\mathcal{N}(0,I)$).
-
Reparameterization Trick: Sample $$\displaystyle \mathbf{z} = \boldsymbol{\mu} + \boldsymbol{\sigma} \odot \boldsymbol{\epsilon} $$, $\boldsymbol{\epsilon} \sim \mathcal{N}(0,I)$. Allows gradient flow through $\boldsymbol{\mu}, \boldsymbol{\sigma}$.
-
Objective (ELBO):
$$\mathcal{L}_{VAE} = \mathbb{E}_{q(\mathbf{z}|\mathbf{x})}[\log p(\mathbf{x}|\mathbf{z})] - D_{KL}(q(\mathbf{z}|\mathbf{x}) \| p(\mathbf{z}))$$
* **Reconstruction Term:** First term (expected log-likelihood).
* **Regularization Term:** KL divergence forces approximate posterior $q(\mathbf{z}|\mathbf{x})$ (encoder output) to be close to prior $p(\mathbf{z})$.
- Training: Maximize ELBO (or minimize negative ELBO). Produces a smooth, structured latent space.
Generative Adversarial Networks (GANs)
-
Framework: Two networks play a minimax game:
-
Generator ($G$): Takes noise $$\displaystyle \mathbf{z} \sim p_z(\mathbf{z}) $$, outputs fake sample $G(\mathbf{z})$.
-
Discriminator ($D$): Takes sample $\mathbf{x}$ (real or fake), outputs probability $D(\mathbf{x})$ of being real.
-
-
Objective:
$$\min_G \max_D V(D,G) = \mathbb{E}_{\mathbf{x}\sim p_{data}}[\log D(\mathbf{x})] + \mathbb{E}_{\mathbf{z}\sim p_z}[\log(1 - D(G(\mathbf{z})))]$$
-
Training: Alternate between:
-
Fix $G$, update $D$: Maximize $\log D(\mathbf{x}) + \log(1-D(G(\mathbf{z})))$.
-
Fix $D$, update $G$: Minimize $\log(1-D(G(\mathbf{z})))$ (or maximize $\log D(G(\mathbf{z}))$).
-
-
Challenges: Mode collapse, training instability, difficult to evaluate.
Comparison: VAEs vs. GANs
| Aspect | VAE | GAN |
|---|---|---|
| Training | Stable (single optimization, ELBO). | Unstable (adversarial, sensitive to hyperparams). |
| Output Quality | Blurry samples (due to MSE/BCE loss). | Sharp, high-fidelity samples. |
| Latent Space | Structured, continuous, meaningful interpolation. | Often disconnected, interpolation may yield unrealistic samples. |
| Mode Coverage | Good (covers all modes due to KL term). | Poor (prone to mode collapse). |
| Inference | Encoder exists → can infer latent code for new data. | No encoder → cannot easily map $\mathbf{x} \to \mathbf{z}$. |
| Choose VAE When: Need structured latent space, stable training, or inference. | ||
| Choose GAN When: Sample quality is paramount, and inference is not needed. |
VAE-GAN Hybrids
-
Idea: Combine VAE's encoder-decoder structure with GAN's discriminator to get sharper outputs.
-
Architecture: Standard VAE encoder-decoder, but reconstruction loss is replaced or supplemented by an adversarial loss from a discriminator that tries to distinguish real $\mathbf{x}$ from reconstructed $$\displaystyle \hat{\mathbf{x}} = G(\mathbf{z}), \mathbf{z} \sim q(\mathbf{z}|\mathbf{x}) $$.
-
How it Works: The generator (decoder) is trained to fool the discriminator, improving sharpness. The encoder still provides a latent code. The KL term from VAE maintains latent structure.
-
Result: Sharper samples than pure VAE, more stable than pure GAN, retains some latent structure.
Autoregressive Models
-
Core Idea: Factorize joint distribution $p(\mathbf{x})$ as product of conditionals: $$\displaystyle p(\mathbf{x}) = \prod_{i=1}^d p(x_i | \mathbf{x}_{<i}) $$. Generate pixel-by-pixel or token-by-token.
-
MADE (Masked Autoencoder for Distribution Estimation):
-
Standard autoencoder, but with masked connections to ensure autoregressive property.
-
Each output unit $i$ only connects to input units $$\displaystyle j < i $$ (via masks). Can be trained efficiently in parallel.
-
-
NADE (Neural Autoregressive Distribution Estimator): Similar concept, different parameterization (using a hidden layer with specific connectivity).
Restricted Boltzmann Machines (RBMs)
-
Bipartite Undirected Graph: Visible units $\mathbf{v}$ (input), Hidden units $\mathbf{h}$. No visible-visible or hidden-hidden connections.
-
Energy Function: $$\displaystyle E(\mathbf{v}, \mathbf{h}) = -\mathbf{v}^T\mathbf{W}\mathbf{h} - \mathbf{b}^T\mathbf{v} - \mathbf{c}^T\mathbf{h} $$.
-
Gibbs Sampling: Alternate sampling $p(\mathbf{h}|\mathbf{v})$ and $p(\mathbf{v}|\mathbf{h})$ to get model samples.
-
Training (Contrastive Divergence): Approximate gradient of log-likelihood using short Gibbs sampling chains (CD-k, usually k=1).
Deep Belief Networks (DBNs)
-
Composition: Stack of RBMs. Train greedily, layer-wise: train first RBM on data, use its hidden activations as "data" for next RBM.
-
Role in History: Pre-2010, primary method for pretraining deep networks (since backprop through many layers failed). Now largely obsolete.
Deep Generative Models: Applications
-
Image Synthesis: Faces (StyleGAN), super-resolution, inpainting.
-
Data Augmentation: Generate synthetic training data.
-
Drug Discovery: Generate novel molecular structures.
-
Text/Audio Generation: GPT (autoregressive), WaveNet (autoregressive audio).
6. Deep Reinforcement Learning
Reinforcement Learning Basics
-
Agent: Learner/decision maker.
-
Environment: World agent interacts with.
-
State ($s$): Current situation.
-
Action ($a$): Agent's move.
-
Reward ($r$): Immediate feedback from environment.
-
Policy ($\pi$): Strategy mapping states to actions ($$\displaystyle a = \pi(s) $$).
-
Value Function: $$\displaystyle V^\pi(s) = \mathbb{E}[\sum_{t=0}^\infty \gamma^t r_t | s_0=s, \pi] $$ (expected cumulative discounted reward).
-
Goal: Maximize expected return $$\displaystyle \mathbb{E}[\sum \gamma^t r_t] $$.
Markov Decision Processes (MDPs)
-
Formalization: $(S, A, P, R, \gamma)$.
-
$P(s'|s,a)$: Transition probability (Markov property: next state depends only on current $s,a$).
-
$R(s,a,s')$: Reward function.
-
$\gamma \in [0,1]$: Discount factor (future rewards less valuable).
-
-
Solution: Find optimal policy $$\displaystyle \pi^* $$ that maximizes $$\displaystyle V^\pi(s) $$ for all $s$.
Dynamic Programming (DP) - For Known MDPs
-
Value Iteration:
-
Initialize $V(s)$ arbitrarily.
-
Iterate until convergence: $$\displaystyle V_{k+1}(s) = \max_a \sum_{s'} P(s'|s,a) [R(s,a,s') + \gamma V_k(s')] $$.
-
Extract greedy policy: $$\displaystyle \pi(s) = \arg\max_a \sum_{s'} P(...) $$.
- Convergence: Guaranteed. Each iteration improves policy.
-
-
Policy Iteration:
-
Policy Evaluation: For current $$\displaystyle \pi_k $$, solve $$\displaystyle V_{\pi_k}(s) = \sum_{s'} P(s'|s,\pi_k(s)) [R(...) + \gamma V_{\pi_k}(s')] $$ (linear system).
-
Policy Improvement: $$\displaystyle \pi_{k+1}(s) = \arg\max_a \sum_{s'} P(s'|s,a) [R(...) + \gamma V_{\pi_k}(s')] $$.
-
Stop if $$\displaystyle \pi_{k+1} = \pi_k $$.
-
-
Comparison: Policy iteration often faster convergence but each step is costly (solving linear system). Value iteration is simpler per step but may need more iterations.
Q-Learning
-
Off-policy TD learning. Learns action-value function $Q(s,a)$.
-
Update Rule (SARSA is on-policy):
$$Q(s_t,a_t) \leftarrow Q(s_t,a_t) + \alpha \left[ r_{t+1} + \gamma \max_{a'} Q(s_{t+1}, a') - Q(s_t,a_t) \right]$$
- Exploration-Exploitation: Use $\epsilon$-greedy: with prob $\epsilon$, choose random action; else choose $$\displaystyle \arg\max_a Q(s,a) $$.
Deep Q-Networks (DQN)
-
Idea: Use a deep neural network to approximate $Q(s,a; \theta)$.
-
Key Techniques:
-
Experience Replay: Store transitions $(s,a,r,s')$ in replay buffer. Sample mini-batches randomly to break correlations between sequential samples.
-
Target Network: Use a separate, slowly updated network $$\displaystyle Q'(s,a; \theta^-) $$ to compute TD target $$\displaystyle r + \gamma \max_{a'} Q'(s',a';\theta^-) $$. $$\displaystyle \theta^- $$ updated periodically (or via Polyak averaging). Stabilizes target.
-
-
Loss: $$\displaystyle L(\theta) = \mathbb{E}_{(s,a,r,s')\sim U(D)} \left[ \left( r + \gamma \max_{a'} Q'(s',a';\theta^-) - Q(s,a;\theta) \right)^2 \right] $$.
Advanced DQN Algorithms
-
Double DQN:
-
Problem: Standard DQN overestimates Q-values because same network selects and evaluates best action.
-
Solution: Decouple selection and evaluation.
-
$$y = r + \gamma Q\left(s', \arg\max_{a'} Q(s',a';\theta); \theta^-\right)$$
* Reduces overestimation bias.
-
Dueling DQN:
-
Architecture: Two streams from shared CNN: Value stream $V(s;\theta, \beta)$ and Advantage stream $A(s,a;\theta, \alpha)$.
-
Combine: $$\displaystyle Q(s,a;\theta, \alpha, \beta) = V(s;\theta, \beta) + (A(s,a;\theta, \alpha) - \frac{1}{|\mathcal{A}|}\sum_{a'} A(s,a';\theta, \alpha)) $$.
-
Benefit: Learns state value independently of specific actions, better generalization, especially when actions don't affect state value much.
-
Applications of Deep RL
-
Game Playing: Atari (DQN), Go (AlphaGo), Dota 2/StarCraft (multi-agent).
-
Robotics: Control, manipulation, locomotion.
-
Autonomous Vehicles: Decision making.
-
Resource Management: Data center cooling, inventory control.
7. Advanced Topics and Applications
Representation Learning
-
Goal: Learn good features from raw data automatically, reducing need for manual feature engineering.
-
How DL Achieves This: Deep networks learn hierarchical representations—early layers learn simple features (edges), later layers learn complex, abstract concepts.
-
Examples: CNN features for images, word embeddings (Word2Vec, GloVe) for text, hidden states of RNNs for sequences.
Deep Dream
-
Process: Start with an image/noise, perform gradient ascent on the activations of a chosen layer in a trained CNN w.r.t. the input image. Maximizes the "presence" of patterns the layer responds to.
-
Applications: Visualizing what neurons/layers "see", generating surreal art, understanding learned features.
Model Compression: Unit Pruning
-
Goal: Reduce model size/inference time for deployment on mobile/edge devices.
-
Unit Pruning: Remove entire neurons/filters/channels deemed unimportant (e.g., based on weight magnitude, activation magnitude, or contribution to loss).
-
Need: Deep networks are often over-parameterized. Pruning removes redundancy with minimal accuracy loss, improving efficiency.
Deep RNNs in Image Processing
-
PixelRNN/PixelCNN: Autoregressive models that generate images pixel-by-pixel, modeling $$\displaystyle p(\mathbf{x}) = \prod_i p(x_i | \mathbf{x}_{<i}) $$. Captures complex dependencies.
-
Image Captioning: CNN (encoder) extracts image features → RNN/LSTM (decoder) generates descriptive sentence.
-
Video Analysis: Apply RNNs frame-wise (or 3D CNNs) to model temporal dynamics.
Natural Language Processing (NLP)
-
Four Key Elements:
-
Tokenization: Split text into words/subwords.
-
Embeddings: Dense vector representations (Word2Vec, GloVe, BERT). Capture semantic meaning.
-
Syntax & Semantics: Grammatical structure and meaning. Handled by RNNs/LSTMs/Transformers.
-
Downstream Tasks: Translation, sentiment analysis, QA, summarization.
-
-
Role of Architectures:
-
RNNs/LSTMs: Early state-of-the-art for sequence modeling.
-
Transformers: Current dominant architecture (self-attention, parallelization).
-
Deep Generative Models: Extended Applications
-
Text: GPT series (autoregressive), BART (seq2seq denoising).
-
Audio: WaveNet (autoregressive raw audio generation), Tacotron (TTS).
-
Molecule Design: Graph-based VAEs/GANs for drug discovery.
Hybrid Architectures
-
CNN + RNN: Standard for image/video captioning, video classification.
-
CNN + Transformer: Vision Transformers (ViT) treat image patches as sequence.
-
RNN + Attention: Allow decoder to focus on relevant encoder states (improves Seq2Seq).
Current Challenges
-
Interpretability: "Black box" nature. Tools: SHAP, LIME, saliency maps.
-
Data Efficiency: DL needs huge data; humans learn from few examples.
-
Robustness: Vulnerable to adversarial examples (small input perturbations cause wrong predictions).
-
Ethical Concerns: Bias in data/models, fairness, misuse (deepfakes), environmental cost (training emissions).