Skip to content
AL-503 (B) · Deep Learning/Quick Revision Short Notes

Deep Learning (AL-503 (B)) - Unit 2 Short Notes

UNIT 2: DEEP LEARNING - EXAM-FOCUSED SHORT NOTES


1. FOUNDATIONS & HISTORICAL CONTEXT

History and Evolution of Deep Learning

  • Key Milestones:

    • 1943: McCulloch-Pitts neuron (first computational model).

    • 1958: Perceptron (single-layer, linear).

    • 1986: Backpropagation popularized (Rumelhart, Hinton, Williams).

    • 2006: Deep Belief Networks (Hinton) - greedy pre-training, start of modern DL.

    • 2012: AlexNet (Krizhevsky, Sutskever, Hinton) - ReLU, Dropout, GPU training, won ImageNet.

    • 2014-2016: GANs (Goodfellow), ResNet (He et al.), Transformer (Vaswani et al.).

  • Eras: Early neural networks (1940s-60s) → AI Winter (1970s-80s) → Modern Resurgence (2006-present) driven by big data, GPU compute, algorithmic advances.

Core Definitions & Paradigms

Term Definition Key Idea
AI Broad field of creating intelligent machines. Encompasses all.
ML Subset of AI; algorithms learn patterns from data. Feature engineering often manual.
DL Subset of ML; uses deep neural networks with multiple layers to learn hierarchical features automatically. End-to-end learning.
Supervised Learning from labeled data (input-output pairs). Classification, regression.
Unsupervised Learning from unlabeled data (find structure). Clustering, dimensionality reduction.
Reinforcement Learning via interaction (agent, environment, reward). Policy optimization.

[!TIP] Exam Distinction: DL ≠ just "more layers". It's about hierarchical feature learning where lower layers learn simple features (edges) and deeper layers learn complex concepts (objects).

Basic Neural Network Building Blocks

  • Single-Layer Perceptron (SLP): $$\displaystyle y = f(\mathbf{w}^T\mathbf{x} + b) $$. Limitation: Can only learn linearly separable problems (XOR fails).

  • Multilayer Perceptron (MLP): Stack of fully connected layers with non-linear activations.

    • Universal Approximation Theorem: An MLP with one hidden layer (sufficiently wide) can approximate any continuous function.

    • Representation Power of Depth: Deep networks can represent certain functions exponentially more efficiently than shallow ones (fewer parameters needed for same complexity). Depth enables compositional representation.

  • Biological vs. Artificial: Biological neurons are sparse, asynchronous, plastic; ANNs are dense, synchronous, fixed-structure. Limitation of DL: Lack of causal reasoning, common sense, energy efficiency, and sample efficiency compared to human brain.


2. TRAINING DEEP NEURAL NETWORKS (CORE ALGORITHMS & THEORY)

Forward and Backward Propagation

  • Forward Pass: Compute output layer-by-layer: $$\displaystyle a^{[l]} = g^{[l]}(z^{[l]}) $$, $$\displaystyle z^{[l]} = W^{[l]}a^{[l-1]} + b^{[l]} $$.

  • Backpropagation (BPTT for RNNs):

    1. Forward pass to compute loss $L$.

    2. Backward pass: Compute gradients $$\displaystyle \frac{\partial L}{\partial W^{[l]}}, \frac{\partial L}{\partial b^{[l]}} $$ using chain rule from output to input.

    3. Update weights: $$\displaystyle W^{[l]} := W^{[l]} - \alpha \frac{\partial L}{\partial W^{[l]}} $$.

    • Computation Graph: Visualizes operations and dependencies; gradients flow backward along edges.
  • Application Areas: Training any differentiable model (CNNs, RNNs, Transformers).

Optimization Algorithms

Algorithm Core Idea Update Rule (for parameter $\theta$) Pros/Cons
Batch GD Use entire dataset per update. $$\displaystyle \theta := \theta - \alpha \nabla_\theta J(\theta) $$ Stable, slow, memory heavy.
SGD Use one sample per update. $$\displaystyle \theta := \theta - \alpha \nabla_\theta J(\theta; x^{(i)}, y^{(i)}) $$ Fast updates, noisy convergence.
Mini-batch SGD Use small batch (e.g., 32, 128). $$\displaystyle \theta := \theta - \alpha \frac{1}{m} \sum_{i=1}^{m} \nabla_\theta J(\theta; x^{(i)}, y^{(i)}) $$ Practical standard: balance of speed/stability.
Momentum Add velocity term to smooth updates. $$\displaystyle v := \beta v + (1-\beta)\nabla_\theta J $$, $$\displaystyle \theta := \theta - \alpha v $$ Accelerates across shallow regions, dampens oscillations.
Nesterov "Lookahead" momentum. $$\displaystyle v := \beta v + (1-\beta)\nabla_\theta J(\theta - \alpha\beta v) $$, $$\displaystyle \theta := \theta - \alpha v $$ More accurate gradient estimate.
AdaGrad Per-parameter adaptive learning rate. Accumulates squared gradients. $$\displaystyle \theta := \theta - \frac{\alpha}{\sqrt{G_t + \epsilon}} \odot g_t $$, $$\displaystyle G_t = \sum_{\tau=1}^{t} g_\tau^2 $$ Learning rates decay aggressively; not ideal for non-convex.
RMSProp Fix AdaGrad's aggressive decay via decaying average. $$\displaystyle G_t := \beta G_{t-1} + (1-\beta) g_t^2 $$, $$\displaystyle \theta := \theta - \frac{\alpha}{\sqrt{G_t + \epsilon}} \odot g_t $$ Handles non-stationary objectives well.
Adam (Adaptive Moment Estimation) Most popular. Combines Momentum (1st moment) & RMSProp (2nd moment) with bias correction. $$\displaystyle m_t = \beta_1 m_{t-1} + (1-\beta_1)g_t $$, $$\displaystyle v_t = \beta_2 v_{t-1} + (1-\beta_2)g_t^2 $$, $$\displaystyle \hat{m}_t = \frac{m_t}{1-\beta_1^t} $$, $$\displaystyle \hat{v}_t = \frac{v_t}{1-\beta_2^t} $$, $$\displaystyle \theta := \theta - \alpha \frac{\hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon} $$ Robust, low memory, good default ($$\displaystyle \alpha=0.001 $$, $$\displaystyle \beta_1=0.9 $$, $$\displaystyle \beta_2=0.999 $$).

[!TIP] Adam is default. For very large datasets/CNNs, SGD with momentum + learning rate decay can yield better final generalization.

Weight Initialization & Normalization

  • Why Initialization Matters: Break symmetry, control variance of activations/gradients to avoid vanishing/exploding gradients.

  • Methods:

    • Random (Uniform/Normal): Simple but problematic for deep nets.

    • Xavier/Glorot: For tanh/sigmoid. $$\displaystyle W \sim \mathcal{U}\left(-\sqrt{\frac{6}{n_{in}+n_{out}}}, \sqrt{\frac{6}{n_{in}+n_{out}}}\right) $$ or $$\displaystyle \mathcal{N}(0, \sqrt{\frac{2}{n_{in}+n_{out}}}) $$.

    • He Initialization: For ReLU. $$\displaystyle W \sim \mathcal{N}(0, \sqrt{\frac{2}{n_{in}}}) $$. Standard for ReLU CNNs.

  • Batch Normalization (BN):

    • Mechanics: For each mini-batch, normalize layer inputs: $$\displaystyle \hat{x}^{(k)} = \frac{x^{(k)} - \mu_B}{\sqrt{\sigma_B^2 + \epsilon}} $$, then scale/shift: $$\displaystyle y^{(k)} = \gamma \hat{x}^{(k)} + \beta $$. $\gamma, \beta$ are learnable.

    • Purpose: Reduce Internal Covariate Shift (distribution change of layer inputs during training). Smooths optimization landscape.

    • Advantages: Allows higher learning rates, reduces need for careful initialization, acts as slight regularizer. Applied before activation (except for RNNs).

  • Normalization vs. Standardization (Data Preprocessing):

    • Normalization (Min-Max): $$\displaystyle x' = \frac{x - x_{min}}{x_{max} - x_{min}} $$ → scales to [0,1]. Sensitive to outliers.

    • Standardization (Z-score): $$\displaystyle x' = \frac{x - \mu}{\sigma} $$ → zero mean, unit variance. More robust, preferred for DL.

Activation Functions

Function Formula Range Pros Cons Use Case
Sigmoid $$\displaystyle \sigma(x) = \frac{1}{1+e^{-x}} $$ (0,1) Smooth, probabilistic output. Vanishing gradients (saturates), not zero-centered. Output layer for binary classification.
Tanh $$\displaystyle \tanh(x) = \frac{e^x - e^{-x}}{e^x + e^{-x}} $$ (-1,1) Zero-centered, steeper than sigmoid. Vanishing gradients (saturates). Hidden layers (older nets), RNNs.
ReLU $$\displaystyle \text{ReLU}(x) = \max(0, x) $$ [0, ∞) Computationally cheap, mitigates vanishing gradient (positive slope). Sparse activations. Dying ReLU (neurons stuck at 0 for negative inputs). Default for hidden layers in CNNs/MLPs.
Leaky ReLU $$\displaystyle \text{LReLU}(x) = \max(\alpha x, x) $$, $\alpha \approx 0.01$ (-∞, ∞) Fixes dying ReLU (small gradient for $$\displaystyle x<0 $$). Results may vary; $\alpha$ is hyperparameter. Alternative to ReLU.
ELU $$\displaystyle \text{ELU}(x) = \begin{cases} x & x>0 \\ \alpha(e^x -1) & x \le 0 \end{cases} $$ (-α, ∞) Smoother near zero, pushes mean to zero. Computationally heavier (exp). Alternative to ReLU.
Softmax $$\displaystyle \sigma(\mathbf{z})_j = \frac{e^{z_j}}{\sum_{k=1}^{K} e^{z_k}} $$ (0,1), sums to 1 Multi-class probability output. Sensitive to outliers/naive implementations. Output layer for multi-class classification.

[!TIP] ReLU Dominance: ReLU's non-saturating positive region is key for training very deep networks (e.g., ResNet). Always use He initialization with ReLU.

Regularization Techniques

  • Overfitting vs. Underfitting:

    • Overfitting: Low training error, high test error. Model too complex, memorizes noise.

    • Underfitting: High training & test error. Model too simple, cannot capture pattern.

  • Dropout:

    • Mechanism: During training, randomly "drop" (set to 0) a fraction $p$ of neurons in a layer per mini-batch. At test time, use all neurons but scale activations by $(1-p)$ or use inverted dropout (scale during training).

    • Why it works: Prevents co-adaptation of neurons, forces network to learn redundant representations, acts as an ensemble of many thinned networks.

  • Weight Decay (L1/L2 Regularization):

    • Add penalty term to loss: $$\displaystyle L_{reg} = L + \lambda \|\mathbf{W}\|_p^p $$.

    • L2 (Ridge): $$\displaystyle \lambda \sum W_{ij}^2 $$. Shrinks weights uniformly.

    • L1 (Lasso): $$\displaystyle \lambda \sum |W_{ij}| $$. Encourages sparsity (some weights exactly zero).

    • Benefit: Constrains model complexity, reduces overfitting.

  • Regularization in Autoencoders: Techniques like denoising (corrupt input, reconstruct clean) or contractive (penalize Jacobian norm) force the AE to learn robust, meaningful features, not just identity mapping.

Gradient Problems & Solutions

  • Vanishing Gradient Problem:

    • Cause: Gradients multiplied by values $$\displaystyle <1 $$ repeatedly through layers (sigmoid/tanh saturation, deep networks).

    • Impact: Early layers learn extremely slowly or not at all. Severe in BPTT for RNNs (long-term dependencies lost).

    • Solutions: ReLU family, Batch Normalization, Residual Connections (ResNet), proper initialization (He/Xavier), LSTM/GRU (for RNNs).

  • Exploding Gradient Problem:

    • Cause: Gradients multiplied by values $$\displaystyle >1 $$ repeatedly (poor initialization, deep nets).

    • Impact: Unstable training, NaN weights.

    • Solutions: Gradient clipping (set max norm), proper initialization, Batch Norm, smaller learning rate.

  • Residual Connections (ResNet): $$\displaystyle y = F(x) + x $$. Identity shortcut allows gradient to flow directly backward, mitigates vanishing gradient in very deep (>100 layers) CNNs.


3. CONVOLUTIONAL NEURAL NETWORKS (CNNs)

Fundamental Concepts & Operations

  • Convolution Operation:

    • Purpose: Extract local spatial features (edges, textures) using learnable filters/kernels.

    • Mechanics: Kernel $K$ (size $$\displaystyle k \times k \times c_{in} $$) slides over input volume $I$ (width $w$, height $h$, channels $$\displaystyle c_{in} $$). Compute dot product at each spatial location: $$\displaystyle (I * K)_{i,j} = \sum_m \sum_n I_{i+m, j+n} K_{m,n} $$.

    • Key Parameters:

      • Stride ($s$): Step size of kernel. Larger stride → smaller output volume.

      • Padding ($p$): Add zeros around input border.

        • Valid: No padding ($$\displaystyle p=0 $$). Output shrinks.

        • Same: Output size = input size (with padding).

    • Output Size: $$\displaystyle \left\lfloor \frac{W - K + 2P}{S} \right\rfloor + 1 $$ (same for height).

  • Filters/Kernels: Each filter learns to detect a specific feature (e.g., vertical edge). Output volume depth = number of filters. Weight sharing across spatial locations makes CNNs parameter-efficient and translationally invariant.

Key Architectural Components

  • Pooling Layers (Downsampling):

    • Purpose: Reduce spatial dimensions (width/height), increase receptive field, provide translation invariance, reduce parameters/computation.

    • Max Pooling: Take maximum value in window (most common). Preserves dominant features.

    • Average Pooling: Take average. Used in some modern architectures (e.g., Inception).

    • Typical: $2 \times 2$ window, stride 2 → halves spatial dimensions.

  • Fully Connected (FC) Layers: At end of CNN, flatten final feature maps and apply MLP for classification/regression.

Classic CNN Architectures (Detailed Study)

Architecture Year Key Innovations Significance
LeNet-5 1998 First successful CNN (handwritten digit recognition). Conv → Pool → Conv → Pool → FC → FC → Output. Pioneered CNN structure for vision.
AlexNet 2012 ReLU (faster training), Dropout (regularization), GPU training, LRN (now obsolete), 5 conv + 3 FC layers. Won ImageNet 2012, sparked modern DL revolution. Showed depth + ReLU + Dropout works.
ZFNet 2013 Improved AlexNet: smaller first filter (7x7→3x3), stride 2 in first conv, more filters. Visualized learned features; showed importance of architecture tweaks.
GoogLeNet (Inception v1) 2014 Inception Module: Parallel convs (1x1, 3x3, 5x5) + pooling, concatenated. Auxiliary classifiers for gradient flow. Won ImageNet 2014. Efficient use of parameters (1M vs AlexNet's 60M). Depth without huge parameter cost.
ResNet 2015 Residual Block: $$\displaystyle y = F(x) + x $$. Very deep (up to 152 layers). Won ImageNet 2015 with human-level performance. Solved vanishing gradient for very deep nets via skip connections.

Specialized CNN Topics

  • Structured Output in CNNs: Output is not a single label but a structure (e.g., pixel-wise segmentation map, bounding box coordinates). Requires fully convolutional architectures (no FC layers) or specialized heads (e.g., YOLO, Mask R-CNN).

  • Data Formats Compatible with CNNs:

    | Data Type | Format | Example | | :--- | :--- | :--- | | Images | $(Height, Width, Channels)$ or $(Channels, Height, Width)$ | RGB: (224, 224, 3) | | Video | $(Frames, Height, Width, Channels)$ | (30, 224, 224, 3) | | 1D Signals | $(Timesteps, Channels)$ | Audio waveform: (16000, 1) | | 3D Volumes | $(Depth, Height, Width, Channels)$ | MRI scan: (128, 128, 128, 1) | | Graphs | $(Nodes, Features)$ + Adjacency Matrix | Social network, molecules. |

  • Deep Dream: Technique to visualize learned features. Amplify patterns in input image that activate specific neurons/filters by performing gradient ascent on the input (not weights). Generates surreal, dream-like images. Shows what CNNs "see".


4. RECURRENT NEURAL NETWORKS (RNNs) & SEQUENCE MODELS

Basic RNN Architecture

  • Core Idea: Have hidden state $$\displaystyle h_t $$ that captures information from previous time steps. Process sequences step-by-step.

  • Equations: $$\displaystyle h_t = \tanh(W_{hh} h_{t-1} + W_{xh} x_t + b_h) $$, $$\displaystyle y_t = W_{hy} h_t + b_y $$.

  • Unfolding Through Time: RNN can be seen as a deep network with shared parameters ($$\displaystyle W_{hh}, W_{xh} $$) across time steps. Handles variable-length sequences.

  • vs. Feedforward: RNNs have cycles (memory), process sequences; FFNs are acyclic, process fixed-size inputs independently.

Backpropagation Through Time (BPTT)

  • Process: Unfold RNN for $T$ steps → compute total loss $$\displaystyle L = \sum_{t=1}^{T} L_t $$ → apply standard backpropagation through the unfolded graph.

  • Challenges:

    • Vanishing/Exploding Gradients: Gradients must flow through many time steps. Multiplicative nature leads to exponential growth/shrinkage.

    • Computational Cost: $O(T)$ for forward/backward pass. Truncated BPTT used for long sequences.

    • Long-range dependencies: Hard to learn connections between distant time steps due to gradient issues.

Advanced RNN Cells

  • Long Short-Term Memory (LSTM):

    • Purpose: Solve vanishing gradient in standard RNNs for long-range dependencies.

    • Cell Structure (per timestep $t$):

      1. Forget Gate: $$\displaystyle f_t = \sigma(W_f \cdot [h_{t-1}, x_t] + b_f) $$ → decides what to discard from cell state $$\displaystyle c_{t-1} $$.

      2. Input Gate: $$\displaystyle i_t = \sigma(W_i \cdot [h_{t-1}, x_t] + b_i) $$, $$\displaystyle \tilde{c}_t = \tanh(W_c \cdot [h_{t-1}, x_t] + b_c) $$ → decides what new info to store.

      3. Cell State Update: $$\displaystyle c_t = f_t \odot c_{t-1} + i_t \odot \tilde{c}_t $$. Additive update → stable gradient flow.

      4. Output Gate: $$\displaystyle o_t = \sigma(W_o \cdot [h_{t-1}, x_t] + b_o) $$, $$\displaystyle h_t = o_t \odot \tanh(c_t) $$ → decides what to output as hidden state.

    • Advantages: Explicit memory cell ($$\displaystyle c_t $$) with additive updates; gates control information flow.

  • Gated Recurrent Unit (GRU):

    • Simplified LSTM: Merges forget/input gates into single update gate $$\displaystyle z_t $$. No separate cell state; hidden state $$\displaystyle h_t $$ is the memory.

    • Equations: $$\displaystyle z_t = \sigma(W_z \cdot [h_{t-1}, x_t]) $$, $$\displaystyle r_t = \sigma(W_r \cdot [h_{t-1}, x_t]) $$, $$\displaystyle \tilde{h}_t = \tanh(W \cdot [r_t \odot h_{t-1}, x_t]) $$, $$\displaystyle h_t = (1-z_t) \odot h_{t-1} + z_t \odot \tilde{h}_t $$.

    • Comparison: GRU has fewer parameters, often trains faster, performance similar to LSTM. LSTM may be better for longer sequences or more complex dependencies.

Sequence-to-Sequence Learning

  • Encoder-Decoder Framework:

    • Encoder (RNN): Reads input sequence $$\displaystyle \mathbf{x} = (x_1, ..., x_T) $$ → produces final hidden state (or context vector) $c$.

    • Decoder (RNN): Generates output sequence $$\displaystyle \mathbf{y} = (y_1, ..., y_{T'}) $$ conditioned on $c$ and previous outputs. Often uses teacher forcing during training.

  • Applications: Machine translation, text summarization, image captioning (CNN encoder + RNN decoder), speech recognition.

Specific Advanced Models

  • Deep RNNs: Stack multiple RNN layers. Lower layers learn short-term patterns, higher layers learn long-term abstractions.

  • PixelRNN / PixelCNN: Autoregressive models for image generation. Generate image pixel-by-pixel (row-major for PixelRNN, masked convs for PixelCNN). Condition each pixel on all previous pixels. Captures complex distribution $p(\text{image})$.

  • Recursive Neural Networks (RecNNs): Process tree-structured data (e.g., parse trees). Apply same neural network at each node, combining children's representations. Captures hierarchical compositionality (e.g., in NLP semantics).


5. AUTOENCODERS & REPRESENTATION LEARNING

Autoencoder (AE) Fundamentals

  • Architecture:

    • Encoder: $$\displaystyle z = f_{enc}(x; \theta_{enc}) $$ → maps input to latent code $z$ (bottleneck, low-dimensional).

    • Decoder: $$\displaystyle \hat{x} = f_{dec}(z; \theta_{dec}) $$ → reconstructs input from $z$.

    • Objective: Minimize reconstruction loss $L(x, \hat{x})$ (e.g., MSE for images, cross-entropy for binary).

  • Purpose: Unsupervised representation learning, dimensionality reduction, denoising, anomaly detection. Learns data-specific manifold.

Types of Autoencoders

Type Key Mechanism Goal
Sparse AE Add sparsity penalty (e.g., KL divergence to target sparsity $\rho$) on activations of hidden layer. Learn disentangled, interpretable features (like edge detectors).
Denoising AE Corrupt input $x$ (e.g., add Gaussian noise, drop pixels) → train to reconstruct original clean $x$. Learn robust features invariant to noise.
Contractive AE Add penalty on Jacobian of encoder: $$\displaystyle \lambda \| J_f(x) \|_F^2 $$. Force encoder to be locally invariant to small input changes.

Regularization in AEs

  • Problem: Without constraints, AE can learn trivial identity function ($$\displaystyle z=x $$, $$\displaystyle \hat{x}=x $$).

  • Solutions: Bottleneck (low-dimensional $z$), sparsity, denoising, contractive penalty. These force AE to learn meaningful compressed representation.

Comparison with PCA/SVD

Aspect PCA/SVD Autoencoders
Linearity Linear transformation. Non-linear (with non-linear activations).
Manifold Assumes data lies on linear subspace. Can learn non-linear manifolds.
Flexibility Fixed objective (maximize variance). Flexible: can be tailored (denoising, sparse) for specific tasks.
When to use AE? When data has non-linear structure (e.g., images, text) or you need task-specific features (e.g., for classification). PCA for quick, linear, global analysis.

[!TIP] Exam Justification: "Use Autoencoders instead of PCA when the data lies on a non-linear manifold (e.g., images of faces, handwritten digits) or when you need domain-specific features for downstream tasks."


6. GENERATIVE MODELS

Variational Autoencoders (VAEs)

  • Probabilistic Framework: Learn distribution $p(x)$ by approximating posterior $p(z|x)$ with encoder $$\displaystyle q_\phi(z|x) $$.

  • Reparameterization Trick: Sample $$\displaystyle z = \mu_\phi(x) + \sigma_\phi(x) \odot \epsilon $$, $\epsilon \sim \mathcal{N}(0,I)$. Allows gradient flow through sampling.

  • Objective (ELBO - Evidence Lower Bound):

$$\mathcal{L}_{VAE} = \mathbb{E}_{q_\phi(z|x)}[\log p_\theta(x|z)] - D_{KL}(q_\phi(z|x) \| p(z))$$

*   **Reconstruction Term:** $$\displaystyle \mathbb{E}[\log p_\theta(x|z)] $$ (e.g., MSE or cross-entropy). Maximize likelihood of data.

*   **Regularization Term:** $$\displaystyle D_{KL}(q_\phi(z|x) \| p(z)) $$ (usually $\mathcal{N}(0,I)$). Forces latent space to be **structured** (continuous, smooth).
  • Training: Maximize ELBO (equivalently, minimize $$\displaystyle -\mathcal{L}_{VAE} $$) via SGD.

Generative Adversarial Networks (GANs)

  • Architecture: Two networks play minimax game:

    • Generator $G$: Takes random noise $$\displaystyle z \sim p_z $$ → generates fake data $$\displaystyle \tilde{x} = G(z) $$.

    • Discriminator $D$: Takes real $x$ or fake $\tilde{x}$ → outputs probability $D(x)$ of being real.

  • Objective:

$$\min_G \max_D V(D,G) = \mathbb{E}_{x \sim p_{data}}[\log D(x)] + \mathbb{E}_{z \sim p_z}[\log(1 - D(G(z)))]$$

*   **D's goal:** Maximize $\log D(x) + \log(1-D(G(z)))$ → distinguish real vs fake.

*   **G's goal:** Minimize $\log(1-D(G(z)))$ → fool D.
  • Training Dynamics: Alternate updates: 1) Fix $G$, update $D$ (maximize). 2) Fix $D$, update $G$ (minimize).

  • Challenges: Mode collapse (G generates limited modes), training instability, difficult to evaluate.

Comparison & Hybrids

Aspect VAE GAN
Density Explicit (learns $p(x)$ via ELBO). Implicit (no explicit density, only samples).
Training Stable (single optimization). Unstable (minimax game).
Sample Quality Blurry (due to Gaussian decoder assumption). Sharp, high-fidelity samples.
Latent Space Structured, continuous (good for interpolation). Less structured, no direct encoder.
Use Case Representation learning, controlled generation, semi-supervised. Pure generation (images, audio) where quality > diversity/structure.
VAE-GAN Hybrid Combine VAE encoder + GAN discriminator. Use VAE's reconstruction loss + GAN's adversarial loss for decoder. Get sharp samples (from GAN) and meaningful latent space (from VAE). Example: PixelCNN + VAE or BEGAN.

Energy-Based & Autoregressive Models

  • Restricted Boltzmann Machines (RBMs):

    • Bipartite graph: Visible layer $v$ (data), hidden layer $h$. No visible-visible or hidden-hidden connections.

    • Energy Function: $$\displaystyle E(v,h) = -b^T v - c^T h - v^T W h $$.

    • Gibbs Sampling: For training (Contrastive Divergence): 1) Positive phase: $v$ = data. 2) Sample $h \sim p(h|v)$. 3) Negative phase: Sample $\tilde{v} \sim p(v|h)$ (reconstruction). Update $W$ based on $$\displaystyle \langle vh \rangle_{data} - \langle vh \rangle_{recon} $$.

  • Autoregressive Models: Factorize joint distribution as product of conditionals: $$\displaystyle p(x) = \prod_{i=1}^{d} p(x_i | x_{<i}) $$.

    • NADE (Neural Autoregressive Distribution Estimator): Uses a neural network with masked weights to ensure autoregressive property. Each unit depends only on previous inputs.

    • MADE (Masked Autoencoder for Distribution Estimation): Applies fixed binary masks to an autoencoder's weights to enforce autoregressive property. Efficient, flexible.

  • Markov Networks (Undirected Graphical Models): Represent dependencies via potential functions $\phi(C)$ over cliques $C$. No directed edges. Energy-based. RBMs are a special case.

Deep Belief Networks (DBNs)

  • Architecture: Stack of RBMs. Train greedily, layer-wise: train first RBM on data → use its hidden activations as "data" for next RBM → repeat.

  • Historical Context: Pre-2012, main method for pre-training deep networks (since backprop struggled with deep, randomly initialized nets). Now largely obsolete, replaced by better initialization, ReLU, and supervised training from scratch.


7. ADVANCED TOPICS & SPECIALIZED ARCHITECTURES

Deep Reinforcement Learning (Deep RL)

  • Integration: Use deep neural networks to represent value functions or policies in RL.

  • Markov Decision Process (MDP): $(\mathcal{S}, \mathcal{A}, P, R, \gamma)$.

    • $\mathcal{S}$: States, $\mathcal{A}$: Actions.

    • $P(s'|s,a)$: Transition probability.

    • $R(s,a,s')$: Reward.

    • $\gamma$: Discount factor.

  • Value Iteration (Dynamic Programming): Compute optimal state-value $$\displaystyle V^*(s) $$ iteratively:

$$V_{k+1}(s) = \max_a \sum_{s'} P(s'|s,a) [R(s,a,s') + \gamma V_k(s')]$$

Converges to $$\displaystyle V^* $$.
  • Policy Iteration: Alternate: 1) Policy Evaluation: Compute $$\displaystyle V^\pi $$ for current policy $\pi$. 2) Policy Improvement: $$\displaystyle \pi'(s) = \arg\max_a \sum_{s'} P(s'|s,a)[R + \gamma V^\pi(s')] $$. Converges to $$\displaystyle \pi^* $$.

  • Q-Learning (Model-Free): Learn action-value $Q(s,a)$ off-policy:

$$Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma \max_{a'} Q(s',a') - Q(s,a)]$$

  • Deep Q-Networks (DQN): Use deep CNN to approximate $Q(s,a;\theta)$. Key innovations: Experience replay (break correlations), target network (stable targets).

  • Advanced DQN:

    • Double DQN: Decouple action selection (from online net) and evaluation (from target net) to reduce overestimation bias.

    • Dueling DQN: Separate streams for state-value $V(s)$ and advantage $A(s,a)$: $$\displaystyle Q(s,a) = V(s) + A(s,a) $$. Learns better value estimates.

  • Least Squares Policy Iteration (LSPI): Policy iteration using LSTD (Least-Squares Temporal Difference) for policy evaluation (linear function approximation). More sample-efficient than TD.

Model Compression & Efficiency

  • Unit/Neuron Pruning:

    • Concept: Remove unimportant neurons (or weights) from trained network to reduce size/inference time.

    • Need: Deploy models on edge devices (mobile, IoT) with limited compute/memory.

    • Methods: 1) Magnitude-based: Prune small-magnitude weights. 2) Activation-based: Prune neurons with low average activation. 3) Second-order: Use Hessian to estimate importance. Iterative: prune → fine-tune.

Representation Learning

  • General Principle: Learn meaningful, abstract features from raw data (pixels, words) automatically, without manual feature engineering.

  • Importance: Core of deep learning. Enables models to capture hierarchical, disentangled representations (edges → parts → objects). Foundation for transfer learning, semi-supervised learning, and generative models.


8. NATURAL LANGUAGE PROCESSING (NLP) & APPLICATIONS

NLP Fundamentals

  • Definition: Field at intersection of AI, linguistics, CS focused on enabling computers to understand, interpret, manipulate human language.

  • Core Elements:

    1. Syntax: Structure of sentences (grammar, parsing).

    2. Semantics: Meaning of words/sentences.

    3. Pragmatics: Contextual meaning (intent, discourse).

    4. Morphology: Word formation (stems, affixes).

DL Applications in NLP

  • Word Embeddings (Word2Vec, GloVe): Dense vector representations ($$\displaystyle \mathbb{R}^{300} $$) capturing semantic similarity ("king" - "man" + "woman" ≈ "queen"). Pre-trained embeddings are standard input for DL models.

  • RNNs/LSTMs for Language Modeling: Predict next word given history. Used in translation, speech recognition.

  • Sequence-to-Sequence with Attention:

    • Encoder: RNN/LSTM/Transformer encodes input sequence into context vectors.

    • Decoder: Generates output token-by-token.

    • Attention Mechanism: Allow decoder to focus on different parts of input sequence at each output step. Crucial for long sequences (e.g., translation). Formula: $$\displaystyle \text{score}(h_i, s_t) = v^T \tanh(W_1 h_i + W_2 s_t) $$, then softmax → weights → context vector.

Broad Application Areas of Deep Learning

Domain Key DL Models Example Tasks
Computer Vision CNNs (ResNet, EfficientNet), Vision Transformers Image classification, object detection (YOLO), segmentation (U-Net).
Speech Recognition RNNs, CTC, Transformers (Whisper) Transcribe audio to text.
Generative Modeling GANs, VAEs, Diffusion Models Image synthesis (StyleGAN), text generation (GPT), music.
Recommendation Systems Neural Collaborative Filtering, Two-tower models Product/movie recommendations (YouTube, Netflix).
Game Playing Deep RL (DQN, A3C, AlphaGo) Learn policies from pixels/rewards.

PRIORITY-BASED EXAM FOCUS SUMMARY

Highest Frequency (Must Master)

  1. Backpropagation: Derivation, computation graph, applications.

  2. CNNs: Convolution operation (filters, stride, padding), pooling (max vs avg), classic architectures (AlexNet, ResNet, LeNet).

  3. RNNs/LSTMs: Architecture, BPTT challenges, LSTM cell structure (gates, cell state), GRU vs LSTM.

  4. Optimization: Adam (derive update), SGD vs Batch GD vs Mini-batch.

  5. Autoencoders: Architecture, types (sparse, denoising), PCA vs AE.

  6. Generative Models: GANs (minimax game), VAEs (ELBO, reparameterization), GANs vs VAEs comparison.

  7. Activation Functions: ReLU (why dominant), sigmoid/tanh issues.

  8. Vanishing Gradient: Causes, impact, solutions (ReLU, BatchNorm, ResNet).

  9. Batch Normalization: Mechanics, purpose (internal covariate shift), advantages.

High Frequency (Should Know Well)

  • MLP & Universal Approximation.

  • Weight Initialization (Xavier, He).

  • Dropout mechanism.

  • Weight Decay (L1/L2).

  • Structured Output in CNNs.

  • Deep RL basics (MDP, Q-learning, DQN).

  • Encoding/Decoding in RNNs.

  • GoogLeNet (Inception module).

Moderate/Lower Frequency (Be Familiar)

  • History milestones (AlexNet year, etc.).

  • Deep Dream.

  • NADE/MADE.

  • DBNs (greedy pre-training).

  • Unit Pruning.

  • Representation Learning principles.

  • NLP elements (4 elements, word embeddings).

[!TIP] Final Exam Strategy: For 7-mark questions, structure answer as: 1-2 sentence definition → 3-4 step explanation/derivation → 1-2 line application/significance. Use diagrams where possible (e.g., LSTM cell, CNN layers, computation graph). Always highlight key terms in bold.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in