Skip to content
AL-503 (B) · Deep Learning/Quick Revision Short Notes

Deep Learning (AL-503 (B)) - Unit 4 Short Notes

UNIT 4: Deep Learning


1. Foundations of Deep Learning

Historical Progression & Milestones

  • Key Breakthroughs: Backpropagation (1986), ReLU (2010), AlexNet (2012), Transformers (2017).

  • Evolution: Perceptrons (1950s) → MLPs (1980s) → Deep Networks (Post-2010, fueled by big data & GPUs).

AI vs. ML vs. DL

Aspect Artificial Intelligence (AI) Machine Learning (ML) Deep Learning (DL)
Goal Create systems that perform human-like tasks Learn patterns from data without explicit programming Subfield of ML using deep neural networks
Feature Engineering Manual & Knowledge-based Manual Automatic (hierarchical feature learning)
Data Dependency Varies Moderate to High Very High (needs large datasets)
Example Chess-playing system Spam classifier Image recognition with CNNs

Biological vs. Artificial Neural Networks

  • Analogies: Neurons ↔ Artificial neurons, synapses ↔ weights, brain regions ↔ network layers.

  • Key Differences:

    • Scale: Brain has ~86B neurons; ANNs have far fewer.

    • Learning: Biological learning is unsupervised, continual, and energy-efficient. DL is mostly supervised, episodic, and compute-intensive.

    • Architecture: Brain is sparse, dynamically wired, and plastic. ANNs are dense, static during inference.

  • Limitations of DL vs. Brain: Lack of common sense, poor data efficiency, catastrophic forgetting, no true understanding, high energy consumption.

Perceptrons and Multilayer Perceptrons (MLPs)

  • Single-Layer Perceptron:

    • Model: $$\displaystyle y = f(\mathbf{w}^T\mathbf{x} + b) $$, where $f$ is a step function.

    • Limitation: Can only learn linearly separable patterns (e.g., AND, OR). Fails on XOR.

  • Multilayer Perceptron (MLP):

    • Architecture: Input layer → one or more fully connected (dense) hidden layers → Output layer. Each neuron applies: $$\displaystyle z = \mathbf{w}^T\mathbf{a}^{prev} + b $$, $$\displaystyle a = f(z) $$.

    • Overcoming Limitations: Multiple layers with non-linear activation functions allow learning complex, non-linear decision boundaries.

    • Universal Approximation Theorem: A feedforward network with a single hidden layer containing a finite number of neurons and appropriate non-linear activations can approximate any continuous function on compact subsets of $$\displaystyle \mathbb{R}^n $$ to any desired accuracy.

Activation Functions

  • Purpose: Introduce non-linearity, enabling the network to learn complex mappings.

  • Types & Comparison:

Function Formula Range Pros Cons
Sigmoid $$\displaystyle \sigma(x) = \frac{1}{1+e^{-x}} $$ (0,1) Smooth, outputs probability Vanishing gradient, not zero-centered, slow
Tanh $$\displaystyle \tanh(x) = \frac{e^x - e^{-x}}{e^x + e^{-x}} $$ (-1,1) Zero-centered, steeper than sigmoid Vanishing gradient
ReLU $$\displaystyle \text{ReLU}(x) = \max(0, x) $$ [0, ∞) Computationally cheap, mitigates vanishing gradient in +ve region, induces sparsity Dying ReLU problem (neurons can get stuck)
Leaky ReLU $$\displaystyle f(x) = \max(\alpha x, x) $$ (-∞, ∞) Fixes dying ReLU (small gradient for $$\displaystyle x<0 $$) Needs tuning of $\alpha$
ELU $$\displaystyle f(x) = \begin{cases} x & x>0 \\ \alpha(e^x -1) & x \le 0 \end{cases} $$ (-α, ∞) Smoother, pushes mean to zero Computationally heavier
Softmax $$\displaystyle \sigma(\mathbf{z})_j = \frac{e^{z_j}}{\sum_{k=1}^K e^{z_k}} $$ (0,1), sums to 1 Used for multi-class classification output

[!TIP] Exam Focus: ReLU is the default for hidden layers. Use Softmax only for the final layer in multi-class classification. Sigmoid/Tanh are now rarely used in hidden layers due to vanishing gradients.

Weight Initialization

  • Importance: Breaks symmetry between neurons. Poor initialization leads to vanishing/exploding gradients.

  • Methods:

    • Random Uniform/Normal: Simple but problematic for deep nets.

    • Xavier/Glorot Initialization: For Tanh/Sigmoid. $$\displaystyle w \sim \mathcal{U}\left(-\sqrt{\frac{6}{n_{in}+n_{out}}}, \sqrt{\frac{6}{n_{in}+n_{out}}}\right) $$ or $$\displaystyle \mathcal{N}(0, \sqrt{\frac{2}{n_{in}+n_{out}}}) $$. Keeps variance stable across layers.

    • He Initialization: For ReLU and variants. $$\displaystyle w \sim \mathcal{N}(0, \sqrt{\frac{2}{n_{in}}}) $$. Compensates for ReLU's half-rectification.

Backpropagation Algorithm

  • Goal: Compute gradient of loss w.r.t. all weights to update them via gradient descent.

  • Steps:

    1. Forward Pass: Compute output $\hat{y}$ and loss $L(\hat{y}, y)$.

    2. Backward Pass (Chain Rule): Compute $$\displaystyle \frac{\partial L}{\partial w} $$ for each weight, layer by layer from output to input.

      • $$\displaystyle \frac{\partial L}{\partial w^{(l)}} = \frac{\partial L}{\partial a^{(l)}} \cdot \frac{\partial a^{(l)}}{\partial z^{(l)}} \cdot \frac{\partial z^{(l)}}{\partial w^{(l)}} $$
    3. Weight Update: $$\displaystyle w^{(l)} \leftarrow w^{(l)} - \eta \frac{\partial L}{\partial w^{(l)}} $$ (for SGD).

  • Applications: Training all feedforward neural networks (MLPs, CNNs).

Gradient Descent Variants

Variant Batch Size Pros Cons
Batch GD Entire dataset Stable convergence, exact gradient Very slow, memory heavy
Stochastic GD (SGD) 1 sample Fast updates, can escape shallow minima Noisy gradients, unstable
Mini-batch GD $n$ samples (e.g., 32, 64) Compromise: Faster than Batch, less noisy than SGD. Most common. Needs batch size tuning

Optimization Algorithms (Adaptive Learning Rates)

Algorithm Core Idea Key Formula/Update Best For
Momentum Accumulate past gradients to dampen oscillations & accelerate. $$\displaystyle v_t = \gamma v_{t-1} + \eta \nabla L(\theta_t) $$; $$\displaystyle \theta_{t+1} = \theta_t - v_t $$ Ravines, noisy gradients
AdaGrad Adapt per-parameter LR by accumulating squared gradients. $$\displaystyle \theta_{t+1} = \theta_t - \frac{\eta}{\sqrt{G_t + \epsilon}} \odot g_t $$, $$\displaystyle G_t = \sum_{\tau=1}^t g_\tau^2 $$ Sparse data (e.g., NLP)
RMSProp Modify AdaGrad to prevent aggressive LR decay by using exponentially weighted moving average of squared grads. $$\displaystyle E[g^2]_t = \beta E[g^2]_{t-1} + (1-\beta) g_t^2 $$; $$\displaystyle \theta_{t+1} = \theta_t - \frac{\eta}{\sqrt{E[g^2]_t + \epsilon}} g_t $$ Non-stationary objectives, RNNs
Adam Combine Momentum & RMSProp + Bias Correction. $$\displaystyle m_t = \beta_1 m_{t-1} + (1-\beta_1)g_t $$ (1st moment)<br>$$\displaystyle v_t = \beta_2 v_{t-1} + (1-\beta_2)g_t^2 $$ (2nd moment)<br>$$\displaystyle \hat{m}_t = m_t/(1-\beta_1^t) $$, $$\displaystyle \hat{v}_t = v_t/(1-\beta_2^t) $$<br>$$\displaystyle \theta_{t+1} = \theta_t - \eta \frac{\hat{m}_t}{\sqrt{\hat{v}_t + \epsilon}} $$ Default choice for most problems. Robust.

Challenges in Training Deep Networks

  • Vanishing Gradient Problem:

    • Cause: Repeated multiplication of small gradients (e.g., from sigmoid/tanh) during backprop through many layers.

    • Impact: Early layers learn very slowly or not at all.

    • Mitigation: Use ReLU, Batch Normalization, Residual Connections (ResNet), proper weight initialization (He), Gradient Clipping.

  • Exploding Gradient Problem:

    • Cause: Large gradients (often in RNNs) amplified through many layers.

    • Impact: Unstable training, weight updates become huge.

    • Mitigation: Gradient Clipping (set max norm), Weight Regularization (L2), smaller learning rate.

Regularization Techniques

  • Purpose: Prevent overfitting (high training accuracy, low validation accuracy), improve generalization.

  • Methods:

    • L1/L2 Weight Decay: Add penalty $$\displaystyle \lambda \|\mathbf{w}\|_1 $$ or $$\displaystyle \lambda \|\mathbf{w}\|_2^2 $$ to loss. L2 is more common.

    • Dropout: During training, randomly deactivate a fraction $p$ of neurons in a layer with probability $p$. Effectively trains an ensemble of thinned networks. Not used at test time.

    • Early Stopping: Stop training when validation loss stops improving.

    • Data Augmentation: Artificially increase dataset size by applying transformations (rotation, crop, flip for images).

Batch Normalization

  • Concept: Normalize the inputs to a layer (pre-activation) to have zero mean and unit variance per mini-batch.

    • $$\displaystyle \hat{x}^{(k)} = \frac{x^{(k)} - \mu_B}{\sqrt{\sigma_B^2 + \epsilon}} $$ (minibatch mean $$\displaystyle \mu_B $$, variance $$\displaystyle \sigma_B^2 $$)

    • Then apply learnable scale and shift: $$\displaystyle y^{(k)} = \gamma \hat{x}^{(k)} + \beta $$.

  • Working:

    • Training: Compute batch stats, normalize, apply $\gamma, \beta$. Update running averages of mean/var.

    • Inference: Use stored running averages to normalize. No batch stats.

  • Advantages:

    1. Faster convergence (allows higher learning rates).

    2. Reduces internal covariate shift (distribution changes in layer inputs).

    3. Has a slight regularization effect (noise from batch stats).

    4. Helps mitigate vanishing gradients.

Data Preprocessing

  • Normalization (Min-Max Scaling): $$\displaystyle x' = \frac{x - x_{min}}{x_{max} - x_{min}} $$. Scales to [0, 1]. Sensitive to outliers.

  • Standardization (Z-score): $$\displaystyle x' = \frac{x - \mu}{\sigma} $$. Scales to mean=0, std=1. More robust, preferred for gradient descent.

  • Impact: Essential for stable and fast gradient descent. Prevents features with large scales from dominating.

Overfitting and Underfitting

Condition Training Error Validation Error Cause Solution
Underfitting High High Model too simple (high bias) Increase model capacity (more layers/neurons), train longer, better features
Overfitting Low High Model too complex (high variance) More data, regularization (dropout, weight decay), batch norm, early stopping, reduce model size

2. Convolutional Neural Networks (CNNs)

Convolution Operation

  • Significance: Exploits spatial/temporal locality and translation equivariance. Drastically reduces parameters via parameter sharing.

  • Mechanics:

    • Kernel/Filter $\mathbf{W}$ (e.g., $$\displaystyle 3 \times 3 \times C_{in} $$) slides over input feature map.

    • Stride ($s$): Step size of filter movement.

    • Output Size: $$\displaystyle \left\lfloor \frac{W_{in} - K + 2P}{s} \right\rfloor + 1 $$ (same for height).

    • Produces a feature map highlighting where a particular pattern (learned by filter) is present.

  • Capturing Spatial Features: Local receptive fields allow early layers to detect edges/textures. Stacked layers build hierarchical representations (edges → parts → objects).

Filters/Kernels

  • Role: Learnable feature detectors. First layer filters learn simple features (edges, colors). Deeper layers learn complex patterns (object parts).

  • Parameters: Number of filters = depth of output volume. Each filter has spatial dimensions (e.g., $3\times3$) and spans full input depth.

Padding

  • Valid Padding: No padding. Output shrinks: $$\displaystyle (W_{in} - K + 1) $$.

  • Same Padding: Pad input so output size equals input size (if stride=1). $$\displaystyle P = \frac{K-1}{2} $$ (for odd K). Preserves edge information.

Pooling Layers

  • Role: Downsampling (reduces spatial size, parameters, computation), provides translation invariance, abstracts features.

  • Max Pooling: Output = max value in window. Preserves dominant features, most common.

  • Average Pooling: Output = average in window. Smooths features, used in some architectures (e.g., GoogLeNet).

  • Comparison: Max pooling is generally preferred for retaining salient features. Average pooling can lose strong activations.

CNN Architecture

Typical Stack: [Conv → Activation (ReLU) → Pooling] × N → Flatten → [FC → Activation] × M → Output Layer.

Structured Output in CNNs

  • Goal: Produce spatial output (e.g., per-pixel labels for segmentation, bounding boxes for detection).

  • Architectures:

    • Fully Convolutional Networks (FCNs): Replace final FC layers with convolutional layers to produce spatial output. Enable end-to-end training for segmentation.

    • Skip Connections/Upsampling Paths: Combine coarse, high-level features from deep layers with fine, low-level features from shallow layers (e.g., U-Net) to recover spatial details.

Key CNN Architectures

Architecture Year Key Innovations Impact
LeNet-5 1998 First successful CNN (LeCun). Conv → Pool → FC. Designed for MNIST. Proof-of-concept for digit recognition.
AlexNet 2012 ReLU, Dropout, GPU training, LRN (later obsolete), 5 conv layers. Won ImageNet 2012, sparked the deep learning revolution.
ZFNet 2013 Visualization-guided tweaks to AlexNet: smaller 1st filter ($7\times7 \to 3\times3$), more filters, stride 2 in conv1. Showed architecture design matters, improved AlexNet baseline.
GoogLeNet (Inception v1) 2014 Inception Module: Parallel convs (1x1, 3x3, 5x5) + pooling → concatenate. Auxiliary classifiers for gradient flow. Very efficient (22 layers, 12x fewer params than AlexNet). Won ImageNet 2014. Emphasized computational efficiency.
ResNet 2015 Residual Block: $$\displaystyle \mathbf{y} = \mathcal{F}(\mathbf{x}, \{W_i\}) + \mathbf{x} $$. Skip connections solve vanishing gradient, enable 100+ layer networks. Won ImageNet 2015 with 152 layers. Fundamental for very deep nets.

Applications of CNNs

  • Image Classification (e.g., ResNet).

  • Object Detection (e.g., YOLO, Faster R-CNN).

  • Semantic Segmentation (e.g., FCN, U-Net).

  • Face Recognition.

  • Extensions: Video (3D convolutions), Spectrograms (audio as image), Graph data (GCNs).


3. Recurrent Neural Networks (RNNs) and Variants

RNN Fundamentals

  • Architecture: Has a hidden state $$\displaystyle \mathbf{h}_t $$ that acts as memory. $$\displaystyle \mathbf{h}_t = f(\mathbf{h}_{t-1}, \mathbf{x}_t; W) $$. Parameters are shared across time steps.

  • vs. Feedforward: Can handle variable-length sequences, has temporal memory. Feedforward has no memory, fixed input size.

Backpropagation Through Time (BPTT)

  • Process: Unfold the RNN for $T$ time steps into a deep feedforward network. Apply standard backpropagation through all $T$ steps.

  • Challenges:

    • Computational Cost: $O(T)$ per sequence.

    • Memory: Need to store activations for all $T$ steps for gradient computation.

    • Long-term dependencies: Gradients must flow through many steps → vanishing/exploding gradients.

Vanishing/Exploding Gradients in RNNs

  • Cause: Repeated multiplication of the Jacobian matrix $$\displaystyle \frac{\partial \mathbf{h}_t}{\partial \mathbf{h}_{t-k}} $$ during BPTT. If eigenvalues $$\displaystyle <1 $$, gradients vanish. If $$\displaystyle >1 $$, they explode.

  • Impact: Inability to learn dependencies beyond ~10-20 time steps.

  • Mitigation: Gating mechanisms (LSTM, GRU), Gradient Clipping, Proper initialization (orthogonal), Skip connections.

Long Short-Term Memory (LSTM)

  • Architecture: Introduces a cell state $$\displaystyle \mathbf{c}_t $$ (the "memory highway") and three gates to regulate information flow.

    • Forget Gate: $$\displaystyle f_t = \sigma(W_f \cdot [\mathbf{h}_{t-1}, \mathbf{x}_t] + b_f) $$. What to remove from cell state?

    • Input Gate: $$\displaystyle i_t = \sigma(W_i \cdot [\mathbf{h}_{t-1}, \mathbf{x}_t] + b_i) $$. What new info to store?

    • Candidate Cell State: $$\displaystyle \tilde{c}_t = \tanh(W_c \cdot [\mathbf{h}_{t-1}, \mathbf{x}_t] + b_c) $$.

    • Cell State Update: $$\displaystyle \mathbf{c}_t = f_t \odot \mathbf{c}_{t-1} + i_t \odot \tilde{c}_t $$.

    • Output Gate: $$\displaystyle o_t = \sigma(W_o \cdot [\mathbf{h}_{t-1}, \mathbf{x}_t] + b_o) $$. What to output from cell state?

    • Hidden State: $$\displaystyle \mathbf{h}_t = o_t \odot \tanh(\mathbf{c}_t) $$.

  • How it Mitigates Vanishing Gradients: Additive nature of cell state update ($$\displaystyle \mathbf{c}_t = f_t \odot \mathbf{c}_{t-1} + ... $$) allows gradients to flow almost unchanged through the $\mathbf{c}$ path if forget gate is near 1.

  • Advantages over Simple RNN: Learns long-term dependencies, robust to vanishing gradients, widely successful in NLP.

Gated Recurrent Unit (GRU)

  • Architecture: Simpler than LSTM. No separate cell state. Uses two gates:

    • Update Gate: $$\displaystyle z_t = \sigma(W_z \cdot [\mathbf{h}_{t-1}, \mathbf{x}_t]) $$. How much past info to keep?

    • Reset Gate: $$\displaystyle r_t = \sigma(W_r \cdot [\mathbf{h}_{t-1}, \mathbf{x}_t]) $$. How much past to forget?

    • Candidate Hidden State: $$\displaystyle \tilde{\mathbf{h}}_t = \tanh(W \cdot [r_t \odot \mathbf{h}_{t-1}, \mathbf{x}_t]) $$.

    • Hidden State: $$\displaystyle \mathbf{h}_t = (1 - z_t) \odot \mathbf{h}_{t-1} + z_t \odot \tilde{\mathbf{h}}_t $$.

  • Comparison with LSTM:

    • Fewer parameters (no cell state, 2 vs 3 gates).

    • Computationally faster.

    • Performance often comparable to LSTM on many tasks, but LSTM may be better for longer sequences or complex dependencies.

Encoding and Decoding in RNNs

  • Sequence-to-Sequence (Seq2Seq) Model:

    • Encoder RNN: Reads input sequence $$\displaystyle \mathbf{x}_1, ..., \mathbf{x}_T $$ and produces a context vector (final hidden state $$\displaystyle \mathbf{h}_T $$ or a combination). Compresses input into fixed-size representation.

    • Decoder RNN: Initialized with context vector, generates output sequence $$\displaystyle \mathbf{y}_1, ..., \mathbf{y}_{T'} $$ autoregressively ($$\displaystyle \mathbf{y}_t $$ becomes input for $t+1$).

  • Applications: Machine translation, text summarization, image captioning (CNN encoder + RNN decoder).

Deep Recurrent Neural Networks

  • Stacked RNNs: Multiple RNN layers. Lower layers learn short-term patterns, higher layers learn long-term, abstract representations.

  • Bidirectional RNNs (BiRNN): Two RNNs process sequence forward and backward. Final output at $t$ is concatenation of both directions' hidden states. Provides full past and future context at each position. Crucial for tasks like POS tagging, NER.

Applications of Deep RNNs

  • NLP: Language modeling, machine translation, sentiment analysis, speech recognition.

  • Time Series: Stock prediction, sensor data analysis.

  • Image Processing: Image captioning (CNN-RNN), video classification (frame-wise RNN), PixelRNN (autoregressive image generation).


4. Autoencoders and Dimensionality Reduction

Autoencoder Architecture

  • Components:

    1. Encoder: $$\displaystyle z = f_{enc}(\mathbf{x}; \theta_{enc}) $$ (compresses input $\mathbf{x}$ to latent code $\mathbf{z}$).

    2. Bottleneck/Latent Space: Low-dimensional representation $\mathbf{z}$.

    3. Decoder: $$\displaystyle \hat{\mathbf{x}} = f_{dec}(\mathbf{z}; \theta_{dec}) $$ (reconstructs $\mathbf{x}$ from $\mathbf{z}$).

  • Objective: Minimize reconstruction loss $L(\mathbf{x}, \hat{\mathbf{x}})$ (e.g., MSE for real-valued, binary cross-entropy for binary inputs).

  • Training: Unsupervised (no labels needed). Forces network to learn efficient data encoding.

Purpose and Applications

  • Dimensionality Reduction: Non-linear alternative to PCA.

  • Feature Learning: Latent space $\mathbf{z}$ can be used as features for downstream tasks.

  • Denoising: Train with noisy input $\tilde{\mathbf{x}}$, reconstruct clean $\mathbf{x}$ (Denoising Autoencoder).

  • Anomaly Detection: Train on normal data; high reconstruction error indicates anomaly.

  • Pretraining: Initialize deep network weights (less common now with better initialization & Adam).

Comparison with PCA/SVD

Aspect PCA/SVD Autoencoder (Deep)
Transformation Linear, orthogonal projection Non-linear (if deep/non-linear activations)
Objective Maximize variance / minimize reconstruction error (linear) Minimize reconstruction error (flexible)
Principal Components Orthogonal, ordered by variance Latent dimensions not necessarily orthogonal or ordered
Global vs Local Captures global linear structure Can capture complex, non-linear manifolds
Use Autoencoder When: Data lies on a non-linear manifold, need non-linear features, or want to use supervised pretraining (with labels).

Types of Autoencoders

Type Key Idea Regularization Mechanism
Sparse Autoencoder Enforce sparsity in latent representation (few active neurons). Add KL divergence penalty: $$\displaystyle \beta \sum_j D_{KL}(\rho \| \hat{\rho}_j) $$ where $$\displaystyle \hat{\rho}_j $$ is average activation of neuron $j$.
Contractive Autoencoder Make encoder robust to small perturbations in input. Penalize squared Frobenius norm of Jacobian of encoder: $$\displaystyle \lambda \|\mathbf{J}_{enc}(\mathbf{x})\|_F^2 $$.
Denoising Autoencoder Learn to reconstruct clean input from corrupted version. Corruption is input (e.g., Gaussian noise, masking). Loss computed on clean target.

Regularization in Autoencoders

  • Goal: Prevent the trivial solution where autoencoder learns identity function ($$\displaystyle \mathbf{z}=\mathbf{x} $$, $$\displaystyle \hat{\mathbf{x}}=\mathbf{x} $$).

  • Techniques: Sparsity constraint, contractive penalty, dropout in encoder/decoder, noise injection (denoising AE), bottleneck size smaller than input.


5. Generative Models

Variational Autoencoders (VAEs)

  • Latent Variable Model: Assumes data $\mathbf{x}$ generated from latent variable $\mathbf{z} \sim p(\mathbf{z})$ (usually $\mathcal{N}(0,I)$).

  • Reparameterization Trick: Sample $$\displaystyle \mathbf{z} = \boldsymbol{\mu} + \boldsymbol{\sigma} \odot \boldsymbol{\epsilon} $$, $\boldsymbol{\epsilon} \sim \mathcal{N}(0,I)$. Allows gradient flow through $\boldsymbol{\mu}, \boldsymbol{\sigma}$.

  • Objective (ELBO):

$$\mathcal{L}_{VAE} = \mathbb{E}_{q(\mathbf{z}|\mathbf{x})}[\log p(\mathbf{x}|\mathbf{z})] - D_{KL}(q(\mathbf{z}|\mathbf{x}) \| p(\mathbf{z}))$$

*   **Reconstruction Term:** First term (expected log-likelihood).

*   **Regularization Term:** KL divergence forces approximate posterior $q(\mathbf{z}|\mathbf{x})$ (encoder output) to be close to prior $p(\mathbf{z})$.
  • Training: Maximize ELBO (or minimize negative ELBO). Produces a smooth, structured latent space.

Generative Adversarial Networks (GANs)

  • Framework: Two networks play a minimax game:

    • Generator ($G$): Takes noise $$\displaystyle \mathbf{z} \sim p_z(\mathbf{z}) $$, outputs fake sample $G(\mathbf{z})$.

    • Discriminator ($D$): Takes sample $\mathbf{x}$ (real or fake), outputs probability $D(\mathbf{x})$ of being real.

  • Objective:

$$\min_G \max_D V(D,G) = \mathbb{E}_{\mathbf{x}\sim p_{data}}[\log D(\mathbf{x})] + \mathbb{E}_{\mathbf{z}\sim p_z}[\log(1 - D(G(\mathbf{z})))]$$

  • Training: Alternate between:

    1. Fix $G$, update $D$: Maximize $\log D(\mathbf{x}) + \log(1-D(G(\mathbf{z})))$.

    2. Fix $D$, update $G$: Minimize $\log(1-D(G(\mathbf{z})))$ (or maximize $\log D(G(\mathbf{z}))$).

  • Challenges: Mode collapse, training instability, difficult to evaluate.

Comparison: VAEs vs. GANs

Aspect VAE GAN
Training Stable (single optimization, ELBO). Unstable (adversarial, sensitive to hyperparams).
Output Quality Blurry samples (due to MSE/BCE loss). Sharp, high-fidelity samples.
Latent Space Structured, continuous, meaningful interpolation. Often disconnected, interpolation may yield unrealistic samples.
Mode Coverage Good (covers all modes due to KL term). Poor (prone to mode collapse).
Inference Encoder exists → can infer latent code for new data. No encoder → cannot easily map $\mathbf{x} \to \mathbf{z}$.
Choose VAE When: Need structured latent space, stable training, or inference.
Choose GAN When: Sample quality is paramount, and inference is not needed.

VAE-GAN Hybrids

  • Idea: Combine VAE's encoder-decoder structure with GAN's discriminator to get sharper outputs.

  • Architecture: Standard VAE encoder-decoder, but reconstruction loss is replaced or supplemented by an adversarial loss from a discriminator that tries to distinguish real $\mathbf{x}$ from reconstructed $$\displaystyle \hat{\mathbf{x}} = G(\mathbf{z}), \mathbf{z} \sim q(\mathbf{z}|\mathbf{x}) $$.

  • How it Works: The generator (decoder) is trained to fool the discriminator, improving sharpness. The encoder still provides a latent code. The KL term from VAE maintains latent structure.

  • Result: Sharper samples than pure VAE, more stable than pure GAN, retains some latent structure.

Autoregressive Models

  • Core Idea: Factorize joint distribution $p(\mathbf{x})$ as product of conditionals: $$\displaystyle p(\mathbf{x}) = \prod_{i=1}^d p(x_i | \mathbf{x}_{<i}) $$. Generate pixel-by-pixel or token-by-token.

  • MADE (Masked Autoencoder for Distribution Estimation):

    • Standard autoencoder, but with masked connections to ensure autoregressive property.

    • Each output unit $i$ only connects to input units $$\displaystyle j < i $$ (via masks). Can be trained efficiently in parallel.

  • NADE (Neural Autoregressive Distribution Estimator): Similar concept, different parameterization (using a hidden layer with specific connectivity).

Restricted Boltzmann Machines (RBMs)

  • Bipartite Undirected Graph: Visible units $\mathbf{v}$ (input), Hidden units $\mathbf{h}$. No visible-visible or hidden-hidden connections.

  • Energy Function: $$\displaystyle E(\mathbf{v}, \mathbf{h}) = -\mathbf{v}^T\mathbf{W}\mathbf{h} - \mathbf{b}^T\mathbf{v} - \mathbf{c}^T\mathbf{h} $$.

  • Gibbs Sampling: Alternate sampling $p(\mathbf{h}|\mathbf{v})$ and $p(\mathbf{v}|\mathbf{h})$ to get model samples.

  • Training (Contrastive Divergence): Approximate gradient of log-likelihood using short Gibbs sampling chains (CD-k, usually k=1).

Deep Belief Networks (DBNs)

  • Composition: Stack of RBMs. Train greedily, layer-wise: train first RBM on data, use its hidden activations as "data" for next RBM.

  • Role in History: Pre-2010, primary method for pretraining deep networks (since backprop through many layers failed). Now largely obsolete.

Deep Generative Models: Applications

  • Image Synthesis: Faces (StyleGAN), super-resolution, inpainting.

  • Data Augmentation: Generate synthetic training data.

  • Drug Discovery: Generate novel molecular structures.

  • Text/Audio Generation: GPT (autoregressive), WaveNet (autoregressive audio).


6. Deep Reinforcement Learning

Reinforcement Learning Basics

  • Agent: Learner/decision maker.

  • Environment: World agent interacts with.

  • State ($s$): Current situation.

  • Action ($a$): Agent's move.

  • Reward ($r$): Immediate feedback from environment.

  • Policy ($\pi$): Strategy mapping states to actions ($$\displaystyle a = \pi(s) $$).

  • Value Function: $$\displaystyle V^\pi(s) = \mathbb{E}[\sum_{t=0}^\infty \gamma^t r_t | s_0=s, \pi] $$ (expected cumulative discounted reward).

  • Goal: Maximize expected return $$\displaystyle \mathbb{E}[\sum \gamma^t r_t] $$.

Markov Decision Processes (MDPs)

  • Formalization: $(S, A, P, R, \gamma)$.

    • $P(s'|s,a)$: Transition probability (Markov property: next state depends only on current $s,a$).

    • $R(s,a,s')$: Reward function.

    • $\gamma \in [0,1]$: Discount factor (future rewards less valuable).

  • Solution: Find optimal policy $$\displaystyle \pi^* $$ that maximizes $$\displaystyle V^\pi(s) $$ for all $s$.

Dynamic Programming (DP) - For Known MDPs

  • Value Iteration:

    1. Initialize $V(s)$ arbitrarily.

    2. Iterate until convergence: $$\displaystyle V_{k+1}(s) = \max_a \sum_{s'} P(s'|s,a) [R(s,a,s') + \gamma V_k(s')] $$.

    3. Extract greedy policy: $$\displaystyle \pi(s) = \arg\max_a \sum_{s'} P(...) $$.

    • Convergence: Guaranteed. Each iteration improves policy.
  • Policy Iteration:

    1. Policy Evaluation: For current $$\displaystyle \pi_k $$, solve $$\displaystyle V_{\pi_k}(s) = \sum_{s'} P(s'|s,\pi_k(s)) [R(...) + \gamma V_{\pi_k}(s')] $$ (linear system).

    2. Policy Improvement: $$\displaystyle \pi_{k+1}(s) = \arg\max_a \sum_{s'} P(s'|s,a) [R(...) + \gamma V_{\pi_k}(s')] $$.

    3. Stop if $$\displaystyle \pi_{k+1} = \pi_k $$.

  • Comparison: Policy iteration often faster convergence but each step is costly (solving linear system). Value iteration is simpler per step but may need more iterations.

Q-Learning

  • Off-policy TD learning. Learns action-value function $Q(s,a)$.

  • Update Rule (SARSA is on-policy):

$$Q(s_t,a_t) \leftarrow Q(s_t,a_t) + \alpha \left[ r_{t+1} + \gamma \max_{a'} Q(s_{t+1}, a') - Q(s_t,a_t) \right]$$

  • Exploration-Exploitation: Use $\epsilon$-greedy: with prob $\epsilon$, choose random action; else choose $$\displaystyle \arg\max_a Q(s,a) $$.

Deep Q-Networks (DQN)

  • Idea: Use a deep neural network to approximate $Q(s,a; \theta)$.

  • Key Techniques:

    1. Experience Replay: Store transitions $(s,a,r,s')$ in replay buffer. Sample mini-batches randomly to break correlations between sequential samples.

    2. Target Network: Use a separate, slowly updated network $$\displaystyle Q'(s,a; \theta^-) $$ to compute TD target $$\displaystyle r + \gamma \max_{a'} Q'(s',a';\theta^-) $$. $$\displaystyle \theta^- $$ updated periodically (or via Polyak averaging). Stabilizes target.

  • Loss: $$\displaystyle L(\theta) = \mathbb{E}_{(s,a,r,s')\sim U(D)} \left[ \left( r + \gamma \max_{a'} Q'(s',a';\theta^-) - Q(s,a;\theta) \right)^2 \right] $$.

Advanced DQN Algorithms

  • Double DQN:

    • Problem: Standard DQN overestimates Q-values because same network selects and evaluates best action.

    • Solution: Decouple selection and evaluation.

$$y = r + \gamma Q\left(s', \arg\max_{a'} Q(s',a';\theta); \theta^-\right)$$

*   Reduces overestimation bias.
  • Dueling DQN:

    • Architecture: Two streams from shared CNN: Value stream $V(s;\theta, \beta)$ and Advantage stream $A(s,a;\theta, \alpha)$.

    • Combine: $$\displaystyle Q(s,a;\theta, \alpha, \beta) = V(s;\theta, \beta) + (A(s,a;\theta, \alpha) - \frac{1}{|\mathcal{A}|}\sum_{a'} A(s,a';\theta, \alpha)) $$.

    • Benefit: Learns state value independently of specific actions, better generalization, especially when actions don't affect state value much.

Applications of Deep RL

  • Game Playing: Atari (DQN), Go (AlphaGo), Dota 2/StarCraft (multi-agent).

  • Robotics: Control, manipulation, locomotion.

  • Autonomous Vehicles: Decision making.

  • Resource Management: Data center cooling, inventory control.


7. Advanced Topics and Applications

Representation Learning

  • Goal: Learn good features from raw data automatically, reducing need for manual feature engineering.

  • How DL Achieves This: Deep networks learn hierarchical representations—early layers learn simple features (edges), later layers learn complex, abstract concepts.

  • Examples: CNN features for images, word embeddings (Word2Vec, GloVe) for text, hidden states of RNNs for sequences.

Deep Dream

  • Process: Start with an image/noise, perform gradient ascent on the activations of a chosen layer in a trained CNN w.r.t. the input image. Maximizes the "presence" of patterns the layer responds to.

  • Applications: Visualizing what neurons/layers "see", generating surreal art, understanding learned features.

Model Compression: Unit Pruning

  • Goal: Reduce model size/inference time for deployment on mobile/edge devices.

  • Unit Pruning: Remove entire neurons/filters/channels deemed unimportant (e.g., based on weight magnitude, activation magnitude, or contribution to loss).

  • Need: Deep networks are often over-parameterized. Pruning removes redundancy with minimal accuracy loss, improving efficiency.

Deep RNNs in Image Processing

  • PixelRNN/PixelCNN: Autoregressive models that generate images pixel-by-pixel, modeling $$\displaystyle p(\mathbf{x}) = \prod_i p(x_i | \mathbf{x}_{<i}) $$. Captures complex dependencies.

  • Image Captioning: CNN (encoder) extracts image features → RNN/LSTM (decoder) generates descriptive sentence.

  • Video Analysis: Apply RNNs frame-wise (or 3D CNNs) to model temporal dynamics.

Natural Language Processing (NLP)

  • Four Key Elements:

    1. Tokenization: Split text into words/subwords.

    2. Embeddings: Dense vector representations (Word2Vec, GloVe, BERT). Capture semantic meaning.

    3. Syntax & Semantics: Grammatical structure and meaning. Handled by RNNs/LSTMs/Transformers.

    4. Downstream Tasks: Translation, sentiment analysis, QA, summarization.

  • Role of Architectures:

    • RNNs/LSTMs: Early state-of-the-art for sequence modeling.

    • Transformers: Current dominant architecture (self-attention, parallelization).

Deep Generative Models: Extended Applications

  • Text: GPT series (autoregressive), BART (seq2seq denoising).

  • Audio: WaveNet (autoregressive raw audio generation), Tacotron (TTS).

  • Molecule Design: Graph-based VAEs/GANs for drug discovery.

Hybrid Architectures

  • CNN + RNN: Standard for image/video captioning, video classification.

  • CNN + Transformer: Vision Transformers (ViT) treat image patches as sequence.

  • RNN + Attention: Allow decoder to focus on relevant encoder states (improves Seq2Seq).

Current Challenges

  • Interpretability: "Black box" nature. Tools: SHAP, LIME, saliency maps.

  • Data Efficiency: DL needs huge data; humans learn from few examples.

  • Robustness: Vulnerable to adversarial examples (small input perturbations cause wrong predictions).

  • Ethical Concerns: Bias in data/models, fairness, misuse (deepfakes), environmental cost (training emissions).

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in