UNIT 1: DEEP LEARNING – FOUNDATIONS & FUNDAMENTAL ARCHITECTURES
I. FOUNDATIONS OF DEEP LEARNING
Introduction and Scope
Deep Learning (DL) is a subset of machine learning that uses deep neural networks with multiple layers to learn hierarchical representations of data.
-
Core paradigm: Automatically extract features from raw data via layered nonlinear transformations.
-
Historical milestones:
-
1957: Perceptron (Rosenblatt) – single-layer linear classifier.
-
1986: Backpropagation popularized (Rumelhart, Hinton, Williams).
-
2012: AlexNet breakthrough (Krizhevsky et al.) – deep CNNs with ReLU, GPU training, won ImageNet.
-
-
AI vs ML vs DL:
| Aspect | AI | ML | DL | |------------------|-------------------------|---------------------------|----------------------------| | Definition | Broad field creating intelligent systems | Subset of AI; algorithms learn from data | Subset of ML; uses deep neural networks | | Feature Engineering | Manual or rule-based | Often manual | Automatic hierarchical learning | | Data Needs | Varies | Moderate to large | Very large | | Hardware | CPU often sufficient | CPU/GPU | GPU essential |
[!TIP]
Exam Focus: Distinguish AI, ML, DL clearly. AlexNet is a frequent 7-mark question – know ReLU, dropout, GPU training.
Biological vs. Artificial Neural Networks
-
Biological neuron: Dendrites (input), soma (processing), axon (output). Spiking signals, analog/digital hybrid, massive parallelism, energy-efficient (~20 watts).
-
Artificial neuron: Mathematical model: \( z = \sum w_i x_i + b \), \( a = g(z) \). Digital, sequential (mostly), high energy consumption.
-
Limitations of DL vs. human brain:
-
Data efficiency: Humans learn from few examples; DL needs thousands/millions.
-
Reasoning & common sense: DL lacks causal reasoning, abstract thought.
-
Energy efficiency: Brain ~20W; large DL models consume megawatts.
-
Continual learning: DL suffers catastrophic forgetting; brain integrates new knowledge seamlessly.
-
[!TIP]
Common Pitfall: Don’t overstate biological plausibility – artificial neurons are crude simplifications.
Representation Learning
-
Concept: Learning multiple layers of increasingly abstract features directly from data.
-
Hierarchical example (image):
-
Layer 1: Edges, corners.
-
Layer 2: Textures, patterns.
-
Layer 3: Object parts.
-
Layer 4: Whole objects.
-
-
Importance: Eliminates manual feature engineering, adapts to data distribution.
Applications of Deep Learning
| Domain | Key Applications |
|---|---|
| Computer Vision | Image classification, object detection, segmentation |
| NLP | Translation, sentiment analysis, chatbots |
| Speech Recognition | Voice assistants, transcription |
| Generative Models | Image synthesis (GANs), data augmentation |
| Reinforcement Learning | Game playing (AlphaGo), robotics |
| Healthcare | Medical image analysis, drug discovery |
II. FUNDAMENTAL NEURAL NETWORK ARCHITECTURES
A. Single-Layer and Logistic Regression
-
Single-layer perceptron:
-
Model: \( y = f(\sum w_i x_i + b) \), where \( f \) is step function.
-
Limitation: Only linearly separable problems (e.g., AND, OR). Cannot learn XOR.
-
-
Logistic regression (single-layer NN for binary classification):
-
Uses sigmoid activation: \( \sigma(z) = \frac{1}{1+e^{-z}} \).
-
Output: probability \( P(y=1|x) \).
-
Loss: Binary cross-entropy: \( \mathcal{L} = -\frac{1}{N}\sum [y \log(\hat{y}) + (1-y)\log(1-\hat{y})] \).
-
Trained via gradient descent.
-
[!TIP]
Exam Alert: "How does MLP overcome single-layer perceptron limitations?" – Answer: Multiple layers + nonlinear activations enable learning complex decision boundaries (e.g., XOR).
B. Multilayer Perceptron (MLP)
-
Architecture:
-
Input layer: Raw features.
-
Hidden layer(s): Nonlinear transformations.
-
Output layer: Task-specific (e.g., softmax for classification).
-
Fully connected: Each neuron in layer \( l \) connected to all in layer \( l-1 \).
-
DiagramCANVAS: MLP with input layer, two hidden layers (ReLU), output layer (softmax)
-
-
Representation power:
-
Universal Approximation Theorem: A single hidden layer with sufficient neurons can approximate any continuous function on compact sets.
-
Depth vs. width:
-
Depth: Enables hierarchical feature learning, exponential reduction in parameters for same function class.
-
Width: Increases capacity but may require exponentially more neurons for complex functions.
-
Deep networks (many layers) learn more abstract features with fewer parameters.
-
-
-
Activation functions:
| Function | Formula | Pros | Cons | |--------------|--------------------------------------|-----------------------------------|-----------------------------------| | Sigmoid | \( \sigma(x) = \frac{1}{1+e^{-x}} \) | Smooth, output in (0,1) | Saturates, vanishing gradient | | Tanh | \( \tanh(x) = \frac{e^x-e^{-x}}{e^x+e^{-x}} \) | Zero-centered, steeper than sigmoid | Saturates, vanishing gradient | | ReLU | \( \text{ReLU}(x) = \max(0,x) \) | Non-saturating, sparse activation, fast | Dying ReLU (negative gradients zero) | | Leaky ReLU | \( \max(\alpha x, x), \alpha \approx 0.01 \) | Fixes dying ReLU | Unproven benefits | | ELU | \( x \) if \( x>0 \), else \( \alpha(e^x-1) \) | Smooth, negative saturation | Computationally heavier | | Softmax | \( \text{softmax}(z_i) = \frac{e^{z_i}}{\sum_j e^{z_j}} \) | Multi-class probability output | Used only in output layer |
-
Forward propagation:
\[ z^{[l]} = W^{[l]} a^{[l-1]} + b^{[l]}, \quad a^{[l]} = g^{[l]}(z^{[l]}) \]
where \( l \) = layer index, \( g^{[l]} \) = activation function.
[!TIP]
Key Formula: ReLU is default for hidden layers due to non-saturation. Softmax for multi-class output.
C. Data Preprocessing and Dimensionality Reduction
-
Normalization vs. Standardization:
| Normalization | Standardization | |----------------------------|-----------------------------| | Scale to [0,1] range | Zero mean, unit variance | | \( x' = \frac{x - \min}{\max - \min} \) | \( x' = \frac{x - \mu}{\sigma} \) | | Sensitive to outliers | Robust to outliers | | Used for image pixels | Used for features with varying scales |
-
PCA and SVD:
-
PCA: Finds orthogonal axes (principal components) maximizing variance.
- Steps: Center data → compute covariance matrix → eigen-decomposition.
-
SVD: Decomposes matrix \( X = U \Sigma V^T \).
-
Relationship: PCA of centered \( X \) is SVD of \( X \); principal components = right singular vectors \( V \).
-
Dimensionality reduction: Keep top \( k \) singular values/vectors → \( X_k = U_k \Sigma_k V_k^T \).
-
-
Autoencoders vs. PCA/SVD:
| Aspect | PCA/SVD | Autoencoders | |------------------|---------------------------|--------------------------------| | Linearity | Linear | Nonlinear (with nonlinear activations) | | Flexibility | Limited to linear subspaces | Can learn complex manifolds | | Data needs | Works with small data | Requires large data | | Interpretability | High (eigenvectors) | Low (latent space opaque) |
- Use autoencoders when: Data lies on nonlinear manifold, sufficient data, need nonlinear features.
-
[!TIP]
Exam Question: "When should autoencoders be used instead of PCA/SVD?" – Answer: For nonlinear data, when you have enough data, and need flexible representations.
III. TRAINING NEURAL NETWORKS
A. Backpropagation
-
Algorithm: Computes gradient of loss w.r.t. all weights via chain rule.
-
Forward pass: compute activations layer by layer.
-
Backward pass: compute error \( \delta^{[l]} = \frac{\partial \mathcal{L}}{\partial z^{[l]}} \) from output to input.
\[ \delta^{[l]} = (W^{[l+1]T} \delta^{[l+1]}) \odot g'^{[l]}(z^{[l]}) \]
-
Gradients: \( \frac{\partial \mathcal{L}}{\partial W^{[l]}} = \delta^{[l]} a^{[l-1]T} \), \( \frac{\partial \mathcal{L}}{\partial b^{[l]}} = \delta^{[l]} \).
-
-
Backpropagation Through Time (BPTT): Unfold RNN in time, apply backpropagation through time steps. Computationally expensive for long sequences.
-
Applications: Weight optimization in all feedforward and recurrent networks.
[!TIP]
Derivation Tip: Start from output layer, propagate errors backward. Remember element-wise multiplication (\(\odot\)) for local gradients.
B. Optimization Algorithms
-
Gradient Descent (GD):
\[ W \leftarrow W - \eta \nabla_W \mathcal{L} \]
Uses full batch → stable but slow for large data.
-
Stochastic Gradient Descent (SGD): Update per sample → noisy but escapes local minima, fast per update.
-
Mini-batch SGD: Compromise – use small batches (e.g., 32, 64). Standard in practice.
-
Momentum:
\[ v \leftarrow \beta v + (1-\beta) \nabla_W \mathcal{L}, \quad W \leftarrow W - \eta v \]
Accelerates along consistent directions, dampens oscillations.
-
Adaptive methods:
-
AdaGrad: Adapts learning rate per parameter.
\[ G_t = G_{t-1} + g_t^2, \quad W \leftarrow W - \frac{\eta}{\sqrt{G_t + \epsilon}} \odot g_t \]
Problem: \( G_t \) accumulates → learning rates vanish.
-
RMSProp: Fixes AdaGrad with decay.
\[ G_t = \beta G_{t-1} + (1-\beta) g_t^2 \]
-
Adam: Combines momentum and RMSProp with bias correction.
\[ m_t = \beta_1 m_{t-1} + (1-\beta_1) g_t, \quad v_t = \beta_2 v_{t-1} + (1-\beta_2) g_t^2 \]
\[ \hat{m}_t = \frac{m_t}{1-\beta_1^t}, \quad \hat{v}_t = \frac{v_t}{1-\beta_2^t}, \quad W \leftarrow W - \eta \frac{\hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon} \]
Default: \( \beta_1=0.9, \beta_2=0.999, \epsilon=10^{-8} \).
-
-
Problems with gradient descent:
-
Local minima/plateaus.
-
Ill-conditioned curvature (ravines).
-
Choice of learning rate.
-
Saddle points in high dimensions.
-
[!TIP]
Adam is default for most DL tasks due to adaptive rates and momentum. SGD with momentum may generalize better for some tasks.
C. Weight Initialization
-
Importance:
-
Break symmetry: identical neurons must learn different features.
-
Avoid vanishing/exploding gradients: maintain variance of activations/gradients across layers.
-
-
Methods:
-
Xavier/Glorot (for tanh, sigmoid):
\[ W \sim \mathcal{U}\left(-\sqrt{\frac{6}{n_{in}+n_{out}}}, \sqrt{\frac{6}{n_{in}+n_{out}}}\right) \]
or normal with std \( \sqrt{\frac{2}{n_{in}+n_{out}}} \).
-
He initialization (for ReLU):
\[ W \sim \mathcal{N}\left(0, \sqrt{\frac{2}{n_{in}}}\right) \]
or uniform \( \pm \sqrt{\frac{6}{n_{in}}} \).
-
[!TIP]
Rule of thumb: ReLU → He init; tanh → Xavier init.
D. Regularization Techniques
-
L1 regularization (lasso): \( \mathcal{L} + \lambda \|W\|_1 \) → sparsity.
-
L2 regularization (weight decay): \( \mathcal{L} + \lambda \|W\|_2^2 \) → small weights.
-
Dropout:
-
Randomly set fraction \( p \) of hidden units to zero during training.
-
At test time, use all units but scale by \( 1-p \) (or inverted dropout: scale at train time).
-
Prevents co-adaptation, acts as ensemble.
-
-
Batch Normalization:
-
Normalize layer inputs: \( \hat{x}^{(k)} = \frac{x^{(k)} - \mu_B}{\sqrt{\sigma_B^2 + \epsilon}} \), then \( y^{(k)} = \gamma \hat{x}^{(k)} + \beta \).
-
\( \mu_B, \sigma_B^2 \) computed over mini-batch.
-
Advantages: Faster convergence, reduces internal covariate shift, slight regularization effect.
-
Difference from Layer Norm: BN uses batch statistics; LN uses per-sample statistics across features.
-
-
Early stopping: Monitor validation loss; stop when it increases to prevent overfitting.
-
Regularization in autoencoders:
-
Sparse: Penalize hidden activations (e.g., KL divergence to target sparsity).
-
Denoising: Train to reconstruct clean input from corrupted version.
-
Contractive: Penalize Frobenius norm of Jacobian \( \| \frac{\partial z}{\partial x} \|_F^2 \) for robustness.
-
E. Common Training Challenges
-
Overfitting vs. Underfitting:
| Overfitting | Underfitting | |----------------------------------|--------------------------------| | High training accuracy, low validation | Low training & validation accuracy | | Model too complex | Model too simple | | Prevention: More data, regularization, dropout, early stopping | Prevention: Increase model capacity, reduce regularization |
-
Vanishing/Exploding Gradients:
-
Causes:
-
Saturating activations (sigmoid/tanh) → small derivatives.
-
Deep networks → repeated multiplication of small/large values.
-
Poor initialization (too large/small weights).
-
-
Impact: Hinders learning long-range dependencies (especially in RNNs).
-
Mitigation:
-
Use ReLU/non-saturating activations.
-
Proper initialization (He/Xavier).
-
Batch normalization.
-
Residual connections (ResNet).
-
Gradient clipping (for exploding).
-
-
-
Strategies for convergence:
-
Learning rate schedules (step decay, cosine annealing).
-
Adaptive optimizers (Adam).
-
Batch normalization.
-
Warm-up: gradually increase learning rate.
-
[!TIP]
Vanishing gradients are critical in RNNs – LSTM/GRU designed to mitigate. In CNNs/MLPs, ReLU + proper init usually suffices.
IV. CONVOLUTIONAL NEURAL NETWORKS (CNN)
A. Core Concepts and Operations
-
Convolution operation:
\[ (I * K)(i,j) = \sum_m \sum_n I(i+m, j+n) K(m,n) \]
-
Filter/kernel: Learnable weights \( K \), detects features (edges, textures).
-
Feature map: Output of convolution; depth = number of filters.
-
Stride \( S \): Step size; larger stride → smaller output.
-
Padding \( P \): Add zeros around input; preserves spatial size.
-
Output size: \( \frac{W - F + 2P}{S} + 1 \) (similarly for height).
-
-
Significance:
-
Spatial hierarchies: Lower layers detect simple features, higher layers detect complex patterns.
-
Translation invariance: Due to weight sharing – same filter applied across space.
-
Parameter efficiency: Fewer parameters than fully connected layers.
-
-
Pooling layers:
-
Max pooling: Take maximum in window → retains prominent features, provides translation invariance.
-
Average pooling: Take average → smoother, used in some architectures (e.g., Inception).
-
Role: Spatial downsampling, reduces computation, increases receptive field.
-
-
DiagramCANVAS: CNN layer showing input, convolution with filters, feature maps, max pooling
B. CNN Architectures
| Architecture | Key Innovations | Impact |
|---|---|---|
| LeNet-5 (1998) | Early CNN for digit recognition (MNIST) | Foundation for CNNs |
| AlexNet (2012) | ReLU, dropout, GPU training, larger depth | ImageNet breakthrough, sparked DL revival |
| ZFNet (2013) | Smaller filters (11×11 → 7×7), visualization | Improved AlexNet |
| GoogleNet/Inception (2014) | Inception modules (multiple filter sizes), depth | Efficient, won ImageNet 2014 |
| ResNet (2015) | Residual connections \( y = F(x) + x \) | Very deep networks (100+ layers), mitigates vanishing gradients |
[!TIP]
AlexNet: Must know ReLU (vs. tanh), dropout (0.5), LRN (local response normalization, now obsolete), 5 conv layers.
ResNet: Residual blocks enable training of very deep networks by allowing gradient flow via identity mapping.
C. Specialized CNN Applications
-
Structured output:
-
Semantic segmentation: Pixel-wise classification (e.g., U-Net, FCN).
-
Object detection: Bounding boxes (e.g., YOLO, Faster R-CNN).
-
CNNs with upsampling (transpose convolutions) or encoder-decoder architectures.
-
-
Deep Dream:
-
Technique: Gradient ascent on input to maximize activations of specific layers/filters.
-
Generates surreal, dream-like images by amplifying patterns the network recognizes.
-
Used for visualization, art.
-
D. Comparison with RNN
| Aspect | CNN | RNN |
|---|---|---|
| Data type | Spatial (images, grids) | Sequential (text, time series) |
| Connectivity | Local receptive fields, weight sharing | Recurrent connections, temporal weight sharing |
| Invariance | Translation invariance | Time invariance (via recurrence) |
| Parallelism | Highly parallelizable | Sequential computation (limited parallelism) |
| Memory | No inherent memory | Hidden state acts as memory |
V. RECURRENT NEURAL NETWORKS (RNN)
A. Basic RNN Architecture
-
Recurrent connection: Hidden state \( h_t \) depends on previous state \( h_{t-1} \) and input \( x_t \).
\[ h_t = f(W_{hh} h_{t-1} + W_{xh} x_t + b_h), \quad y_t = g(W_{hy} h_t + b_y) \]
-
Unfolding in time: Expand recurrence for \( T \) steps → deep network with shared parameters \( W_{hh}, W_{xh} \).
-
Suitability: Sequential data (text, audio, video) where order matters.
-
Comparison with feedforward:
-
Memory: RNN retains history via hidden state; feedforward no memory.
-
Parameter sharing: Same weights across time → handles variable-length sequences.
-
B. Training RNNs: BPTT and Challenges
-
BPTT: Unfold RNN through time, apply backpropagation. Compute gradients by summing over time steps.
-
Vanishing/exploding gradients:
-
Cause: Repeated multiplication of Jacobian matrices \( \frac{\partial h_t}{\partial h_{t-1}} = \text{diag}(f'(z_t)) W_{hh} \). If eigenvalues of \( W_{hh} \) <1 → vanishing; >1 → exploding.
-
Impact: Difficulty learning long-range dependencies (gradients vanish after many steps).
-
Mitigation:
-
Use LSTM/GRU (gating mechanisms).
-
Proper initialization (orthogonal matrices for \( W_{hh} \)).
-
Gradient clipping (for exploding).
-
Skip connections (e.g., residual RNNs).
-
-
C. Advanced RNN Variants
-
Long Short-Term Memory (LSTM):
-
Components:
-
Cell state \( C_t \): "Highway" for information, constant error flow.
-
Forget gate \( f_t = \sigma(W_f \cdot [h_{t-1}, x_t] + b_f) \): What to discard from \( C_{t-1} \).
-
Input gate \( i_t = \sigma(W_i \cdot [h_{t-1}, x_t] + b_i) \): What to store from candidate \( \tilde{C}_t \).
-
Candidate cell \( \tilde{C}_t = \tanh(W_C \cdot [h_{t-1}, x_t] + b_C) \).
-
Cell state update: \( C_t = f_t \odot C_{t-1} + i_t \odot \tilde{C}_t \).
-
Output gate \( o_t = \sigma(W_o \cdot [h_{t-1}, x_t] + b_o) \): What to output from \( C_t \).
-
Hidden state: \( h_t = o_t \odot \tanh(C_t) \).
-
-
Mitigates vanishing gradients: Cell state allows gradients to flow unchanged via additive updates.
-
Advantages over simple RNN: Learns long-term dependencies, robust to vanishing gradients.
-
-
Gated Recurrent Unit (GRU):
-
Simpler than LSTM: no cell state, two gates (reset, update).
-
Reset gate \( r_t = \sigma(W_r \cdot [h_{t-1}, x_t]) \): Controls how much past to forget.
-
Update gate \( z_t = \sigma(W_z \cdot [h_{t-1}, x_t]) \): Trade-off between old and new.
-
Candidate hidden: \( \tilde{h}_t = \tanh(W_h \cdot [r_t \odot h_{t-1}, x_t]) \).
-
Hidden state: \( h_t = (1-z_t) \odot h_{t-1} + z_t \odot \tilde{h}_t \).
-
Comparison: Fewer parameters, often similar performance; LSTM more expressive for complex tasks.
-
D. Encoding and Decoding in RNNs
-
Sequence-to-sequence (seq2seq):
-
Encoder: Processes input sequence \( x_1, ..., x_T \) into context vector \( c \) (often final hidden state).
-
Decoder: Generates output sequence \( y_1, ..., y_{T'} \) conditioned on \( c \) and previous outputs.
-
Challenges:
-
Variable length: Both input/output lengths vary.
-
Information bottleneck: Fixed-size context vector \( c \) may lose information for long sequences.
-
Solutions: Attention mechanisms (allow decoder to attend to all encoder states), bidirectional encoders.
-
-
Applications: Machine translation, text summarization, image captioning.
-
E. Recursive Neural Networks
-
Architecture: Tree-structured; same weights applied to each node in parse tree.
- Each node computes: \( h_{(i,j)} = f(W \cdot [h_i, h_j] + b) \), where \( h_i, h_j \) are children.
-
Comparison with RNN:
-
RNN: Linear sequence (chain).
-
Recursive NN: Arbitrary tree structure → better for hierarchical data (e.g., sentences, parse trees).
-
-
Applications: Natural language parsing, sentiment analysis (capture compositionality).
F. Applications of Deep RNNs
-
NLP:
-
4 elements:
-
Morphology: Word forms (handled by embeddings).
-
Syntax: Grammar (captured by RNNs/transformers).
-
Semantics: Meaning (contextual embeddings).
-
Pragmatics: Contextual use (discourse, intent).
-
-
-
Image processing:
-
Image captioning (CNN encoder + RNN decoder).
-
Video analysis (spatiotemporal features).
-
-
Time series forecasting: Stock prices, weather, sensor data.
VI. AUTOENCODERS
A. Basic Architecture
-
Components:
-
Encoder: \( z = f_\theta(x) \) → latent representation (bottleneck).
-
Decoder: \( \hat{x} = g_\phi(z) \) → reconstruction.
-
-
Training objective: Minimize reconstruction loss.
-
For continuous: \( \mathcal{L} = \|x - \hat{x}\|^2 \).
-
For binary: cross-entropy.
-
-
Latent space: Compressed representation; ideally captures essential features.
B. Types of Autoencoders
-
Sparse autoencoder:
-
Add sparsity penalty on hidden activations (e.g., KL divergence to desired sparsity \( \rho \)).
-
Forces network to learn meaningful features by limiting active neurons.
-
-
Contractive autoencoder:
-
Add penalty on Jacobian: \( \mathcal{L} + \lambda \| \frac{\partial z}{\partial x} \|_F^2 \).
-
Encourages robustness to small input perturbations.
-
-
Denoising autoencoder:
-
Corrupt input \( \tilde{x} \) (e.g., add noise, mask pixels), train to reconstruct clean \( x \).
-
Learns to remove noise → robust features.
-
C. Regularization in Autoencoders
-
Prevents trivial identity mapping (where encoder/decoder just copy input).
-
Encourages learning of meaningful latent representations by constraining capacity or adding noise.
D. Applications
-
Dimensionality reduction: Nonlinear alternative to PCA.
-
Feature extraction: Pretrained encoders for downstream tasks.
-
Anomaly detection: High reconstruction error indicates anomaly.
-
Pretraining (historical): Unsupervised pretraining for deep networks (now less common with better initialization/optimization).
E. Comparison with Linear Methods (PCA/SVD)
| Aspect | PCA/SVD | Autoencoders |
|---|---|---|
| Linearity | Linear transformation | Nonlinear (with nonlinear activations) |
| Capacity | Limited to linear subspace | Can model complex manifolds |
| Data efficiency | Works with small data | Requires large data |
| Interpretability | Principal components are interpretable | Latent space often opaque |
| Reconstruction quality | Optimal linear reconstruction | Can achieve better nonlinear reconstruction |
[!TIP]
Autoencoders vs PCA: Use autoencoders for complex, nonlinear data (e.g., images) with sufficient data; PCA for linear relationships or small data.
VII. GENERATIVE MODELS
A. Variational Autoencoders (VAEs)
-
Probabilistic framework: Latent variable model \( p(x) = \int p(x|z)p(z) dz \).
-
Encoder: Approximate posterior \( q_\phi(z|x) \) (usually Gaussian).
-
Reparameterization trick: Sample \( z = \mu_\phi(x) + \sigma_\phi(x) \odot \epsilon \), \( \epsilon \sim \mathcal{N}(0,I) \) → gradients flow through \( \mu, \sigma \).
-
Loss:
\[ \mathcal{L} = \mathbb{E}_{q_\phi(z|x)}[\log p_\theta(x|z)] - D_{KL}(q_\phi(z|x) \| p(z)) \]
-
Reconstruction term (e.g., MSE or cross-entropy).
-
KL divergence regularizes \( q_\phi(z|x) \) to prior \( p(z) = \mathcal{N}(0,I) \).
-
-
Applications: Generation, interpolation, latent space manipulation.
B. Generative Adversarial Networks (GANs)
-
Adversarial training:
-
Generator \( G(z) \): Maps noise \( z \sim p_z \) to fake data.
-
Discriminator \( D(x) \): Outputs probability real vs fake.
-
-
Min-max game:
\[ \min_G \max_D V(D,G) = \mathbb{E}_{x\sim p_{data}}[\log D(x)] + \mathbb{E}_{z\sim p_z}[\log(1-D(G(z)))] \]
-
Training:
-
Alternate: Train \( D \) (maximize), then train \( G \) (minimize).
-
\( D \) trained on real and fake batches.
-
-
Challenges: Mode collapse, training instability, difficult to evaluate.
-
Applications: High-fidelity image synthesis (StyleGAN), style transfer, data augmentation.
C. VAE-GAN Hybrids
-
Combine VAE’s encoder-decoder structure with GAN’s discriminator.
-
Example (VAE-GAN):
-
Encoder → latent \( z \) → decoder → reconstruction \( \hat{x} \).
-
Discriminator judges realism of \( \hat{x} \) (not just \( x \)).
-
Loss: VAE reconstruction + KL + GAN adversarial loss.
-
-
Benefits: VAE provides stable training, GAN improves sharpness of generated images.
D. Autoregressive Models
-
Principle: Factorize joint distribution as product of conditionals.
\[ p(x) = \prod_{i=1}^D p(x_i | x_{<i}) \]
-
Neural Autoregressive Distribution Estimator (NADE):
-
Uses masked connections to ensure autoregressive property.
-
Each \( p(x_i | x_{<i}) \) modeled by neural network.
-
-
Masked Autoencoder for Distribution Estimation (MADE):
- Extends NADE with arbitrary masking; efficient training.
-
PixelRNN/PixelCNN:
-
Generate images pixel-by-pixel.
-
PixelRNN: Rows/columns with RNNs (slow).
-
PixelCNN: Convolutional with masking (faster, parallelizable).
-
E. Boltzmann Machines
-
Restricted Boltzmann Machines (RBMs):
-
Bipartite graph: visible units \( v \), hidden units \( h \).
-
Energy: \( E(v,h) = -v^T W h - b^T v - c^T h \).
-
Probability: \( p(v,h) = \frac{e^{-E(v,h)}}{Z} \).
-
Training: Contrastive Divergence (CD-k) – approximate gradient via Gibbs sampling.
-
-
Deep Belief Networks (DBNs):
-
Stack RBMs; train layer-wise (unsupervised pretraining).
-
Historical importance for deep learning (pre-2012).
-
F. Markov Networks
-
Undirected graphical models: Cliques with potential functions \( \phi(C) \).
-
Energy-based: \( p(x) = \frac{1}{Z} \prod_C \phi_C(x_C) \).
-
Relationship to RBMs: RBMs are a type of Markov random field with bipartite structure.
G. Applications of Deep Generative Models
-
Image synthesis: Photorealistic faces (StyleGAN), art.
-
Data augmentation: Generate rare class samples.
-
Drug discovery: Generate molecular structures.
-
Anomaly detection: Model normal data, detect outliers via likelihood.
VIII. DEEP REINFORCEMENT LEARNING
A. Markov Decision Processes (MDPs)
-
Components: \( (S, A, P, R, \gamma) \)
-
\( S \): States.
-
\( A \): Actions.
-
\( P(s'|s,a) \): Transition probability.
-
\( R(s,a,s') \): Reward.
-
\( \gamma \in [0,1] \): Discount factor.
-
-
Goal: Maximize expected discounted return \( \mathbb{E}[\sum_{t=0}^\infty \gamma^t r_t] \).
B. Dynamic Programming Methods
-
Value iteration:
\[ V_{k+1}(s) = \max_a \sum_{s'} P(s'|s,a) [R(s,a,s') + \gamma V_k(s')] \]
Iterate until convergence; policy \( \pi(s) = \arg\max_a Q(s,a) \).
-
Policy iteration:
-
Policy evaluation: Compute \( V^\pi \) by solving linear system.
-
Policy improvement: \( \pi' = \arg\max_\pi \sum_a \pi(a|s) Q^\pi(s,a) \).
-
Comparison: Policy iteration often converges faster but each iteration costly (solving linear system); value iteration cheaper per iteration but may converge slowly.
-
C. Q-Learning and Deep Q-Networks (DQN)
-
Q-learning (tabular):
\[ Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma \max_{a'} Q(s',a') - Q(s,a)] \]
Off-policy, model-free.
-
DQN:
-
Use deep network to approximate \( Q(s,a;\theta) \).
-
Experience replay: Store transitions \( (s,a,r,s') \) in buffer; sample random minibatches → breaks correlation, stabilizes training.
-
Target network: Slow-updated target \( Q_{\text{target}}(s,a;\theta^-) \) to stabilize bootstrapping.
-
Loss: \( \mathcal{L} = \mathbb{E}[(r + \gamma \max_{a'} Q_{\text{target}}(s',a';\theta^-) - Q(s,a;\theta))^2] \).
-
D. Advanced DQN Algorithms
-
Double DQN:
- Decouple action selection and evaluation to reduce overestimation bias.
\[ y = r + \gamma Q_{\text{target}}(s', \arg\max_a Q(s',a;\theta); \theta^-) \]
-
Dueling DQN:
- Separate value \( V(s) \) and advantage \( A(s,a) \) streams.
\[ Q(s,a) = V(s) + A(s,a) - \frac{1}{|\mathcal{A}|} \sum_{a'} A(s,a') \]
- Better value estimation, improved learning.
E. Policy Gradient and Actor-Critic Methods (if covered)
-
Policy gradient (REINFORCE):
\[ \nabla_\theta J(\theta) = \mathbb{E}_{\tau \sim \pi_\theta} \left[ \sum_t \nabla_\theta \log \pi_\theta(a_t|s_t) G_t \right] \]
where \( G_t = \sum_{k=t}^T \gamma^{k-t} r_k \).
-
Actor-critic:
-
Actor: Policy \( \pi_\theta \), updated by policy gradient.
-
Critic: Value function \( V_\phi \) or \( Q_\psi \), estimates return.
-
Advantage: Lower variance than pure policy gradient.
-
F. Least Squares Methods
-
Least Squares Policy Iteration (LSPI):
-
Use least-squares to solve Bellman equation from samples.
-
Represent \( Q(s,a) \) with linear features \( \phi(s,a) \).
-
Solve \( (A - \gamma \Phi' \Phi)^{-1} \Phi' R \) where \( A = \Phi' \Phi \), etc.
-
Comparison: More sample-efficient than standard policy iteration; avoids explicit policy evaluation.
-
G. Applications
-
Game playing: Atari (DQN), Go (AlphaGo), chess (AlphaZero).
-
Robotics: Control policies for manipulation, locomotion.
-
Resource management: Data center cooling, inventory control.
IX. ADVANCED TOPICS AND MODEL EFFICIENCY
A. Deep Dream
-
Technique: Gradient ascent on input to maximize activation of specific layer/filter.
\[ x^* = \arg\max_x \mathcal{L}_{\text{activation}}(x) + \text{regularization} \]
-
Role: Visualization of learned features, generating surreal/artistic images.
B. Model Compression and Pruning
-
Unit pruning: Remove entire neurons/filters with low importance (e.g., based on activation magnitude).
-
Weight pruning: Set small weights to zero → sparse networks.
-
Quantization: Reduce precision of weights/activations (e.g., 32-bit → 8-bit).
-
Knowledge distillation: Train small "student" network to mimic large "teacher" network (soft targets).
-
Benefits: Reduced memory, faster inference, deployment on edge devices.
C. Computational Considerations
-
GPU implementation: Parallelize matrix operations (e.g., SVD via cuSOLVER). Randomized SVD for large matrices.
-
Scalability: Distributed training (data/model parallelism), mixed precision training.
D. Directed Graphical Models
-
Bayesian networks: Directed acyclic graph (DAG); nodes = random variables, edges = conditional dependencies.
-
Conditional independence: \( X \perp Y | Z \) means \( p(x|y,z) = p(x|z) \).
-
Inference: Variable elimination, belief propagation.
-
Learning: Parameter estimation (MLE, Bayesian), structure learning.
X. COMPARATIVE AND CRITICAL ANALYSIS
A. Learning Paradigm Comparisons
| Paradigm | Data | Feedback | Examples in DL |
|---|---|---|---|
| Supervised | Labeled \( (x,y) \) | Direct (loss) | Classification, regression |
| Unsupervised | Unlabeled \( x \) | None (intrinsic) | Autoencoders, GANs, clustering |
| Reinforcement | State, action, reward | Delayed (reward) | DQN, policy gradient |
B. Model Comparisons
-
GANs vs VAEs:
| Aspect | GANs | VAEs | |------------------|-----------------------------------|-----------------------------------| | Objective | Min-max game (adversarial) | Variational inference (ELBO) | | Training | Unstable, mode collapse possible | Stable, convergent | | Samples | Sharp, diverse | Blurry, less diverse | | Latent space | Not necessarily continuous | Continuous, structured | | Use case | High-quality synthesis | Interpolation, representation learning |
-
CNN vs RNN vs Transformer:
-
CNN: Spatial data, translation invariance, parallelizable.
-
RNN: Sequential data, memory, but sequential computation.
-
Transformer: Sequential data with self-attention → parallelizable, handles long-range dependencies (now dominant in NLP).
-
-
Autoencoders vs PCA: As above.
C. Limitations and Future Directions
-
Current limitations:
-
Data hunger: Requires massive labeled data.
-
Reasoning: Lacks causal, logical reasoning.
-
Common sense: Struggles with everyday physics, social norms.
-
Explainability: Black-box models.
-
Robustness: Sensitive to adversarial examples.
-
-
Ethical considerations: Bias in data/models, privacy (generative models), misuse (deepfakes), environmental cost.
-
Emerging trends:
-
Self-supervised learning: Pretrain on unlabeled data (e.g., contrastive learning).
-
Neuro-symbolic AI: Combine neural networks with symbolic reasoning.
-
Efficient DL: Model compression, sparse training, federated learning.
-
Foundation models: Large pretrained models (LLMs, vision transformers) fine-tuned for downstream tasks.
-
[!TIP]
Critical analysis: Always mention both strengths and weaknesses. For example, DL excels at pattern recognition but lacks reasoning; ethical concerns are increasingly important.