Skip to content
AL-503 (B) · Deep Learning/Quick Revision Short Notes

Deep Learning (AL-503 (B)) - Unit 1 Short Notes

UNIT 1: DEEP LEARNING – FOUNDATIONS & FUNDAMENTAL ARCHITECTURES


I. FOUNDATIONS OF DEEP LEARNING

Introduction and Scope

Deep Learning (DL) is a subset of machine learning that uses deep neural networks with multiple layers to learn hierarchical representations of data.

  • Core paradigm: Automatically extract features from raw data via layered nonlinear transformations.

  • Historical milestones:

    • 1957: Perceptron (Rosenblatt) – single-layer linear classifier.

    • 1986: Backpropagation popularized (Rumelhart, Hinton, Williams).

    • 2012: AlexNet breakthrough (Krizhevsky et al.) – deep CNNs with ReLU, GPU training, won ImageNet.

  • AI vs ML vs DL:

    | Aspect | AI | ML | DL | |------------------|-------------------------|---------------------------|----------------------------| | Definition | Broad field creating intelligent systems | Subset of AI; algorithms learn from data | Subset of ML; uses deep neural networks | | Feature Engineering | Manual or rule-based | Often manual | Automatic hierarchical learning | | Data Needs | Varies | Moderate to large | Very large | | Hardware | CPU often sufficient | CPU/GPU | GPU essential |

[!TIP]

Exam Focus: Distinguish AI, ML, DL clearly. AlexNet is a frequent 7-mark question – know ReLU, dropout, GPU training.

Biological vs. Artificial Neural Networks
  • Biological neuron: Dendrites (input), soma (processing), axon (output). Spiking signals, analog/digital hybrid, massive parallelism, energy-efficient (~20 watts).

  • Artificial neuron: Mathematical model: \( z = \sum w_i x_i + b \), \( a = g(z) \). Digital, sequential (mostly), high energy consumption.

  • Limitations of DL vs. human brain:

    • Data efficiency: Humans learn from few examples; DL needs thousands/millions.

    • Reasoning & common sense: DL lacks causal reasoning, abstract thought.

    • Energy efficiency: Brain ~20W; large DL models consume megawatts.

    • Continual learning: DL suffers catastrophic forgetting; brain integrates new knowledge seamlessly.

[!TIP]

Common Pitfall: Don’t overstate biological plausibility – artificial neurons are crude simplifications.

Representation Learning
  • Concept: Learning multiple layers of increasingly abstract features directly from data.

  • Hierarchical example (image):

    • Layer 1: Edges, corners.

    • Layer 2: Textures, patterns.

    • Layer 3: Object parts.

    • Layer 4: Whole objects.

  • Importance: Eliminates manual feature engineering, adapts to data distribution.

Applications of Deep Learning
Domain Key Applications
Computer Vision Image classification, object detection, segmentation
NLP Translation, sentiment analysis, chatbots
Speech Recognition Voice assistants, transcription
Generative Models Image synthesis (GANs), data augmentation
Reinforcement Learning Game playing (AlphaGo), robotics
Healthcare Medical image analysis, drug discovery

II. FUNDAMENTAL NEURAL NETWORK ARCHITECTURES

A. Single-Layer and Logistic Regression
  • Single-layer perceptron:

    • Model: \( y = f(\sum w_i x_i + b) \), where \( f \) is step function.

    • Limitation: Only linearly separable problems (e.g., AND, OR). Cannot learn XOR.

  • Logistic regression (single-layer NN for binary classification):

    • Uses sigmoid activation: \( \sigma(z) = \frac{1}{1+e^{-z}} \).

    • Output: probability \( P(y=1|x) \).

    • Loss: Binary cross-entropy: \( \mathcal{L} = -\frac{1}{N}\sum [y \log(\hat{y}) + (1-y)\log(1-\hat{y})] \).

    • Trained via gradient descent.

[!TIP]

Exam Alert: "How does MLP overcome single-layer perceptron limitations?" – Answer: Multiple layers + nonlinear activations enable learning complex decision boundaries (e.g., XOR).

B. Multilayer Perceptron (MLP)
  • Architecture:

    • Input layer: Raw features.

    • Hidden layer(s): Nonlinear transformations.

    • Output layer: Task-specific (e.g., softmax for classification).

    • Fully connected: Each neuron in layer \( l \) connected to all in layer \( l-1 \).

    • DiagramCANVAS: MLP with input layer, two hidden layers (ReLU), output layer (softmax)
  • Representation power:

    • Universal Approximation Theorem: A single hidden layer with sufficient neurons can approximate any continuous function on compact sets.

    • Depth vs. width:

      • Depth: Enables hierarchical feature learning, exponential reduction in parameters for same function class.

      • Width: Increases capacity but may require exponentially more neurons for complex functions.

      • Deep networks (many layers) learn more abstract features with fewer parameters.

  • Activation functions:

    | Function | Formula | Pros | Cons | |--------------|--------------------------------------|-----------------------------------|-----------------------------------| | Sigmoid | \( \sigma(x) = \frac{1}{1+e^{-x}} \) | Smooth, output in (0,1) | Saturates, vanishing gradient | | Tanh | \( \tanh(x) = \frac{e^x-e^{-x}}{e^x+e^{-x}} \) | Zero-centered, steeper than sigmoid | Saturates, vanishing gradient | | ReLU | \( \text{ReLU}(x) = \max(0,x) \) | Non-saturating, sparse activation, fast | Dying ReLU (negative gradients zero) | | Leaky ReLU | \( \max(\alpha x, x), \alpha \approx 0.01 \) | Fixes dying ReLU | Unproven benefits | | ELU | \( x \) if \( x>0 \), else \( \alpha(e^x-1) \) | Smooth, negative saturation | Computationally heavier | | Softmax | \( \text{softmax}(z_i) = \frac{e^{z_i}}{\sum_j e^{z_j}} \) | Multi-class probability output | Used only in output layer |

  • Forward propagation:

    \[ z^{[l]} = W^{[l]} a^{[l-1]} + b^{[l]}, \quad a^{[l]} = g^{[l]}(z^{[l]}) \]

    where \( l \) = layer index, \( g^{[l]} \) = activation function.

[!TIP]

Key Formula: ReLU is default for hidden layers due to non-saturation. Softmax for multi-class output.

C. Data Preprocessing and Dimensionality Reduction
  • Normalization vs. Standardization:

    | Normalization | Standardization | |----------------------------|-----------------------------| | Scale to [0,1] range | Zero mean, unit variance | | \( x' = \frac{x - \min}{\max - \min} \) | \( x' = \frac{x - \mu}{\sigma} \) | | Sensitive to outliers | Robust to outliers | | Used for image pixels | Used for features with varying scales |

  • PCA and SVD:

    • PCA: Finds orthogonal axes (principal components) maximizing variance.

      • Steps: Center data → compute covariance matrix → eigen-decomposition.
    • SVD: Decomposes matrix \( X = U \Sigma V^T \).

      • Relationship: PCA of centered \( X \) is SVD of \( X \); principal components = right singular vectors \( V \).

      • Dimensionality reduction: Keep top \( k \) singular values/vectors → \( X_k = U_k \Sigma_k V_k^T \).

    • Autoencoders vs. PCA/SVD:

      | Aspect | PCA/SVD | Autoencoders | |------------------|---------------------------|--------------------------------| | Linearity | Linear | Nonlinear (with nonlinear activations) | | Flexibility | Limited to linear subspaces | Can learn complex manifolds | | Data needs | Works with small data | Requires large data | | Interpretability | High (eigenvectors) | Low (latent space opaque) |

      • Use autoencoders when: Data lies on nonlinear manifold, sufficient data, need nonlinear features.

[!TIP]

Exam Question: "When should autoencoders be used instead of PCA/SVD?" – Answer: For nonlinear data, when you have enough data, and need flexible representations.


III. TRAINING NEURAL NETWORKS

A. Backpropagation
  • Algorithm: Computes gradient of loss w.r.t. all weights via chain rule.

    • Forward pass: compute activations layer by layer.

    • Backward pass: compute error \( \delta^{[l]} = \frac{\partial \mathcal{L}}{\partial z^{[l]}} \) from output to input.

      \[ \delta^{[l]} = (W^{[l+1]T} \delta^{[l+1]}) \odot g'^{[l]}(z^{[l]}) \]

    • Gradients: \( \frac{\partial \mathcal{L}}{\partial W^{[l]}} = \delta^{[l]} a^{[l-1]T} \), \( \frac{\partial \mathcal{L}}{\partial b^{[l]}} = \delta^{[l]} \).

  • Backpropagation Through Time (BPTT): Unfold RNN in time, apply backpropagation through time steps. Computationally expensive for long sequences.

  • Applications: Weight optimization in all feedforward and recurrent networks.

[!TIP]

Derivation Tip: Start from output layer, propagate errors backward. Remember element-wise multiplication (\(\odot\)) for local gradients.

B. Optimization Algorithms
  • Gradient Descent (GD):

    \[ W \leftarrow W - \eta \nabla_W \mathcal{L} \]

    Uses full batch → stable but slow for large data.

  • Stochastic Gradient Descent (SGD): Update per sample → noisy but escapes local minima, fast per update.

  • Mini-batch SGD: Compromise – use small batches (e.g., 32, 64). Standard in practice.

  • Momentum:

    \[ v \leftarrow \beta v + (1-\beta) \nabla_W \mathcal{L}, \quad W \leftarrow W - \eta v \]

    Accelerates along consistent directions, dampens oscillations.

  • Adaptive methods:

    • AdaGrad: Adapts learning rate per parameter.

      \[ G_t = G_{t-1} + g_t^2, \quad W \leftarrow W - \frac{\eta}{\sqrt{G_t + \epsilon}} \odot g_t \]

      Problem: \( G_t \) accumulates → learning rates vanish.

    • RMSProp: Fixes AdaGrad with decay.

      \[ G_t = \beta G_{t-1} + (1-\beta) g_t^2 \]

    • Adam: Combines momentum and RMSProp with bias correction.

      \[ m_t = \beta_1 m_{t-1} + (1-\beta_1) g_t, \quad v_t = \beta_2 v_{t-1} + (1-\beta_2) g_t^2 \]

      \[ \hat{m}_t = \frac{m_t}{1-\beta_1^t}, \quad \hat{v}_t = \frac{v_t}{1-\beta_2^t}, \quad W \leftarrow W - \eta \frac{\hat{m}_t}{\sqrt{\hat{v}_t} + \epsilon} \]

      Default: \( \beta_1=0.9, \beta_2=0.999, \epsilon=10^{-8} \).

  • Problems with gradient descent:

    • Local minima/plateaus.

    • Ill-conditioned curvature (ravines).

    • Choice of learning rate.

    • Saddle points in high dimensions.

[!TIP]

Adam is default for most DL tasks due to adaptive rates and momentum. SGD with momentum may generalize better for some tasks.

C. Weight Initialization
  • Importance:

    • Break symmetry: identical neurons must learn different features.

    • Avoid vanishing/exploding gradients: maintain variance of activations/gradients across layers.

  • Methods:

    • Xavier/Glorot (for tanh, sigmoid):

      \[ W \sim \mathcal{U}\left(-\sqrt{\frac{6}{n_{in}+n_{out}}}, \sqrt{\frac{6}{n_{in}+n_{out}}}\right) \]

      or normal with std \( \sqrt{\frac{2}{n_{in}+n_{out}}} \).

    • He initialization (for ReLU):

      \[ W \sim \mathcal{N}\left(0, \sqrt{\frac{2}{n_{in}}}\right) \]

      or uniform \( \pm \sqrt{\frac{6}{n_{in}}} \).

[!TIP]

Rule of thumb: ReLU → He init; tanh → Xavier init.

D. Regularization Techniques
  • L1 regularization (lasso): \( \mathcal{L} + \lambda \|W\|_1 \) → sparsity.

  • L2 regularization (weight decay): \( \mathcal{L} + \lambda \|W\|_2^2 \) → small weights.

  • Dropout:

    • Randomly set fraction \( p \) of hidden units to zero during training.

    • At test time, use all units but scale by \( 1-p \) (or inverted dropout: scale at train time).

    • Prevents co-adaptation, acts as ensemble.

  • Batch Normalization:

    • Normalize layer inputs: \( \hat{x}^{(k)} = \frac{x^{(k)} - \mu_B}{\sqrt{\sigma_B^2 + \epsilon}} \), then \( y^{(k)} = \gamma \hat{x}^{(k)} + \beta \).

    • \( \mu_B, \sigma_B^2 \) computed over mini-batch.

    • Advantages: Faster convergence, reduces internal covariate shift, slight regularization effect.

    • Difference from Layer Norm: BN uses batch statistics; LN uses per-sample statistics across features.

  • Early stopping: Monitor validation loss; stop when it increases to prevent overfitting.

  • Regularization in autoencoders:

    • Sparse: Penalize hidden activations (e.g., KL divergence to target sparsity).

    • Denoising: Train to reconstruct clean input from corrupted version.

    • Contractive: Penalize Frobenius norm of Jacobian \( \| \frac{\partial z}{\partial x} \|_F^2 \) for robustness.

E. Common Training Challenges
  • Overfitting vs. Underfitting:

    | Overfitting | Underfitting | |----------------------------------|--------------------------------| | High training accuracy, low validation | Low training & validation accuracy | | Model too complex | Model too simple | | Prevention: More data, regularization, dropout, early stopping | Prevention: Increase model capacity, reduce regularization |

  • Vanishing/Exploding Gradients:

    • Causes:

      • Saturating activations (sigmoid/tanh) → small derivatives.

      • Deep networks → repeated multiplication of small/large values.

      • Poor initialization (too large/small weights).

    • Impact: Hinders learning long-range dependencies (especially in RNNs).

    • Mitigation:

      • Use ReLU/non-saturating activations.

      • Proper initialization (He/Xavier).

      • Batch normalization.

      • Residual connections (ResNet).

      • Gradient clipping (for exploding).

  • Strategies for convergence:

    • Learning rate schedules (step decay, cosine annealing).

    • Adaptive optimizers (Adam).

    • Batch normalization.

    • Warm-up: gradually increase learning rate.

[!TIP]

Vanishing gradients are critical in RNNs – LSTM/GRU designed to mitigate. In CNNs/MLPs, ReLU + proper init usually suffices.


IV. CONVOLUTIONAL NEURAL NETWORKS (CNN)

A. Core Concepts and Operations
  • Convolution operation:

    \[ (I * K)(i,j) = \sum_m \sum_n I(i+m, j+n) K(m,n) \]

    • Filter/kernel: Learnable weights \( K \), detects features (edges, textures).

    • Feature map: Output of convolution; depth = number of filters.

    • Stride \( S \): Step size; larger stride → smaller output.

    • Padding \( P \): Add zeros around input; preserves spatial size.

    • Output size: \( \frac{W - F + 2P}{S} + 1 \) (similarly for height).

  • Significance:

    • Spatial hierarchies: Lower layers detect simple features, higher layers detect complex patterns.

    • Translation invariance: Due to weight sharing – same filter applied across space.

    • Parameter efficiency: Fewer parameters than fully connected layers.

  • Pooling layers:

    • Max pooling: Take maximum in window → retains prominent features, provides translation invariance.

    • Average pooling: Take average → smoother, used in some architectures (e.g., Inception).

    • Role: Spatial downsampling, reduces computation, increases receptive field.

  • DiagramCANVAS: CNN layer showing input, convolution with filters, feature maps, max pooling
B. CNN Architectures
Architecture Key Innovations Impact
LeNet-5 (1998) Early CNN for digit recognition (MNIST) Foundation for CNNs
AlexNet (2012) ReLU, dropout, GPU training, larger depth ImageNet breakthrough, sparked DL revival
ZFNet (2013) Smaller filters (11×11 → 7×7), visualization Improved AlexNet
GoogleNet/Inception (2014) Inception modules (multiple filter sizes), depth Efficient, won ImageNet 2014
ResNet (2015) Residual connections \( y = F(x) + x \) Very deep networks (100+ layers), mitigates vanishing gradients

[!TIP]

AlexNet: Must know ReLU (vs. tanh), dropout (0.5), LRN (local response normalization, now obsolete), 5 conv layers.

ResNet: Residual blocks enable training of very deep networks by allowing gradient flow via identity mapping.

C. Specialized CNN Applications
  • Structured output:

    • Semantic segmentation: Pixel-wise classification (e.g., U-Net, FCN).

    • Object detection: Bounding boxes (e.g., YOLO, Faster R-CNN).

    • CNNs with upsampling (transpose convolutions) or encoder-decoder architectures.

  • Deep Dream:

    • Technique: Gradient ascent on input to maximize activations of specific layers/filters.

    • Generates surreal, dream-like images by amplifying patterns the network recognizes.

    • Used for visualization, art.

D. Comparison with RNN
Aspect CNN RNN
Data type Spatial (images, grids) Sequential (text, time series)
Connectivity Local receptive fields, weight sharing Recurrent connections, temporal weight sharing
Invariance Translation invariance Time invariance (via recurrence)
Parallelism Highly parallelizable Sequential computation (limited parallelism)
Memory No inherent memory Hidden state acts as memory

V. RECURRENT NEURAL NETWORKS (RNN)

A. Basic RNN Architecture
  • Recurrent connection: Hidden state \( h_t \) depends on previous state \( h_{t-1} \) and input \( x_t \).

    \[ h_t = f(W_{hh} h_{t-1} + W_{xh} x_t + b_h), \quad y_t = g(W_{hy} h_t + b_y) \]

  • Unfolding in time: Expand recurrence for \( T \) steps → deep network with shared parameters \( W_{hh}, W_{xh} \).

  • Suitability: Sequential data (text, audio, video) where order matters.

  • Comparison with feedforward:

    • Memory: RNN retains history via hidden state; feedforward no memory.

    • Parameter sharing: Same weights across time → handles variable-length sequences.

B. Training RNNs: BPTT and Challenges
  • BPTT: Unfold RNN through time, apply backpropagation. Compute gradients by summing over time steps.

  • Vanishing/exploding gradients:

    • Cause: Repeated multiplication of Jacobian matrices \( \frac{\partial h_t}{\partial h_{t-1}} = \text{diag}(f'(z_t)) W_{hh} \). If eigenvalues of \( W_{hh} \) <1 → vanishing; >1 → exploding.

    • Impact: Difficulty learning long-range dependencies (gradients vanish after many steps).

    • Mitigation:

      • Use LSTM/GRU (gating mechanisms).

      • Proper initialization (orthogonal matrices for \( W_{hh} \)).

      • Gradient clipping (for exploding).

      • Skip connections (e.g., residual RNNs).

C. Advanced RNN Variants
  • Long Short-Term Memory (LSTM):

    • Components:

      • Cell state \( C_t \): "Highway" for information, constant error flow.

      • Forget gate \( f_t = \sigma(W_f \cdot [h_{t-1}, x_t] + b_f) \): What to discard from \( C_{t-1} \).

      • Input gate \( i_t = \sigma(W_i \cdot [h_{t-1}, x_t] + b_i) \): What to store from candidate \( \tilde{C}_t \).

      • Candidate cell \( \tilde{C}_t = \tanh(W_C \cdot [h_{t-1}, x_t] + b_C) \).

      • Cell state update: \( C_t = f_t \odot C_{t-1} + i_t \odot \tilde{C}_t \).

      • Output gate \( o_t = \sigma(W_o \cdot [h_{t-1}, x_t] + b_o) \): What to output from \( C_t \).

      • Hidden state: \( h_t = o_t \odot \tanh(C_t) \).

    • Mitigates vanishing gradients: Cell state allows gradients to flow unchanged via additive updates.

    • Advantages over simple RNN: Learns long-term dependencies, robust to vanishing gradients.

  • Gated Recurrent Unit (GRU):

    • Simpler than LSTM: no cell state, two gates (reset, update).

    • Reset gate \( r_t = \sigma(W_r \cdot [h_{t-1}, x_t]) \): Controls how much past to forget.

    • Update gate \( z_t = \sigma(W_z \cdot [h_{t-1}, x_t]) \): Trade-off between old and new.

    • Candidate hidden: \( \tilde{h}_t = \tanh(W_h \cdot [r_t \odot h_{t-1}, x_t]) \).

    • Hidden state: \( h_t = (1-z_t) \odot h_{t-1} + z_t \odot \tilde{h}_t \).

    • Comparison: Fewer parameters, often similar performance; LSTM more expressive for complex tasks.

D. Encoding and Decoding in RNNs
  • Sequence-to-sequence (seq2seq):

    • Encoder: Processes input sequence \( x_1, ..., x_T \) into context vector \( c \) (often final hidden state).

    • Decoder: Generates output sequence \( y_1, ..., y_{T'} \) conditioned on \( c \) and previous outputs.

    • Challenges:

      • Variable length: Both input/output lengths vary.

      • Information bottleneck: Fixed-size context vector \( c \) may lose information for long sequences.

      • Solutions: Attention mechanisms (allow decoder to attend to all encoder states), bidirectional encoders.

    • Applications: Machine translation, text summarization, image captioning.

E. Recursive Neural Networks
  • Architecture: Tree-structured; same weights applied to each node in parse tree.

    • Each node computes: \( h_{(i,j)} = f(W \cdot [h_i, h_j] + b) \), where \( h_i, h_j \) are children.
  • Comparison with RNN:

    • RNN: Linear sequence (chain).

    • Recursive NN: Arbitrary tree structure → better for hierarchical data (e.g., sentences, parse trees).

  • Applications: Natural language parsing, sentiment analysis (capture compositionality).

F. Applications of Deep RNNs
  • NLP:

    • 4 elements:

      1. Morphology: Word forms (handled by embeddings).

      2. Syntax: Grammar (captured by RNNs/transformers).

      3. Semantics: Meaning (contextual embeddings).

      4. Pragmatics: Contextual use (discourse, intent).

  • Image processing:

    • Image captioning (CNN encoder + RNN decoder).

    • Video analysis (spatiotemporal features).

  • Time series forecasting: Stock prices, weather, sensor data.


VI. AUTOENCODERS

A. Basic Architecture
  • Components:

    • Encoder: \( z = f_\theta(x) \) → latent representation (bottleneck).

    • Decoder: \( \hat{x} = g_\phi(z) \) → reconstruction.

  • Training objective: Minimize reconstruction loss.

    • For continuous: \( \mathcal{L} = \|x - \hat{x}\|^2 \).

    • For binary: cross-entropy.

  • Latent space: Compressed representation; ideally captures essential features.

B. Types of Autoencoders
  • Sparse autoencoder:

    • Add sparsity penalty on hidden activations (e.g., KL divergence to desired sparsity \( \rho \)).

    • Forces network to learn meaningful features by limiting active neurons.

  • Contractive autoencoder:

    • Add penalty on Jacobian: \( \mathcal{L} + \lambda \| \frac{\partial z}{\partial x} \|_F^2 \).

    • Encourages robustness to small input perturbations.

  • Denoising autoencoder:

    • Corrupt input \( \tilde{x} \) (e.g., add noise, mask pixels), train to reconstruct clean \( x \).

    • Learns to remove noise → robust features.

C. Regularization in Autoencoders
  • Prevents trivial identity mapping (where encoder/decoder just copy input).

  • Encourages learning of meaningful latent representations by constraining capacity or adding noise.

D. Applications
  • Dimensionality reduction: Nonlinear alternative to PCA.

  • Feature extraction: Pretrained encoders for downstream tasks.

  • Anomaly detection: High reconstruction error indicates anomaly.

  • Pretraining (historical): Unsupervised pretraining for deep networks (now less common with better initialization/optimization).

E. Comparison with Linear Methods (PCA/SVD)
Aspect PCA/SVD Autoencoders
Linearity Linear transformation Nonlinear (with nonlinear activations)
Capacity Limited to linear subspace Can model complex manifolds
Data efficiency Works with small data Requires large data
Interpretability Principal components are interpretable Latent space often opaque
Reconstruction quality Optimal linear reconstruction Can achieve better nonlinear reconstruction

[!TIP]

Autoencoders vs PCA: Use autoencoders for complex, nonlinear data (e.g., images) with sufficient data; PCA for linear relationships or small data.


VII. GENERATIVE MODELS

A. Variational Autoencoders (VAEs)
  • Probabilistic framework: Latent variable model \( p(x) = \int p(x|z)p(z) dz \).

  • Encoder: Approximate posterior \( q_\phi(z|x) \) (usually Gaussian).

  • Reparameterization trick: Sample \( z = \mu_\phi(x) + \sigma_\phi(x) \odot \epsilon \), \( \epsilon \sim \mathcal{N}(0,I) \) → gradients flow through \( \mu, \sigma \).

  • Loss:

    \[ \mathcal{L} = \mathbb{E}_{q_\phi(z|x)}[\log p_\theta(x|z)] - D_{KL}(q_\phi(z|x) \| p(z)) \]

    • Reconstruction term (e.g., MSE or cross-entropy).

    • KL divergence regularizes \( q_\phi(z|x) \) to prior \( p(z) = \mathcal{N}(0,I) \).

  • Applications: Generation, interpolation, latent space manipulation.

B. Generative Adversarial Networks (GANs)
  • Adversarial training:

    • Generator \( G(z) \): Maps noise \( z \sim p_z \) to fake data.

    • Discriminator \( D(x) \): Outputs probability real vs fake.

  • Min-max game:

    \[ \min_G \max_D V(D,G) = \mathbb{E}_{x\sim p_{data}}[\log D(x)] + \mathbb{E}_{z\sim p_z}[\log(1-D(G(z)))] \]

  • Training:

    • Alternate: Train \( D \) (maximize), then train \( G \) (minimize).

    • \( D \) trained on real and fake batches.

  • Challenges: Mode collapse, training instability, difficult to evaluate.

  • Applications: High-fidelity image synthesis (StyleGAN), style transfer, data augmentation.

C. VAE-GAN Hybrids
  • Combine VAE’s encoder-decoder structure with GAN’s discriminator.

  • Example (VAE-GAN):

    • Encoder → latent \( z \) → decoder → reconstruction \( \hat{x} \).

    • Discriminator judges realism of \( \hat{x} \) (not just \( x \)).

    • Loss: VAE reconstruction + KL + GAN adversarial loss.

  • Benefits: VAE provides stable training, GAN improves sharpness of generated images.

D. Autoregressive Models
  • Principle: Factorize joint distribution as product of conditionals.

    \[ p(x) = \prod_{i=1}^D p(x_i | x_{<i}) \]

  • Neural Autoregressive Distribution Estimator (NADE):

    • Uses masked connections to ensure autoregressive property.

    • Each \( p(x_i | x_{<i}) \) modeled by neural network.

  • Masked Autoencoder for Distribution Estimation (MADE):

    • Extends NADE with arbitrary masking; efficient training.
  • PixelRNN/PixelCNN:

    • Generate images pixel-by-pixel.

    • PixelRNN: Rows/columns with RNNs (slow).

    • PixelCNN: Convolutional with masking (faster, parallelizable).

E. Boltzmann Machines
  • Restricted Boltzmann Machines (RBMs):

    • Bipartite graph: visible units \( v \), hidden units \( h \).

    • Energy: \( E(v,h) = -v^T W h - b^T v - c^T h \).

    • Probability: \( p(v,h) = \frac{e^{-E(v,h)}}{Z} \).

    • Training: Contrastive Divergence (CD-k) – approximate gradient via Gibbs sampling.

  • Deep Belief Networks (DBNs):

    • Stack RBMs; train layer-wise (unsupervised pretraining).

    • Historical importance for deep learning (pre-2012).

F. Markov Networks
  • Undirected graphical models: Cliques with potential functions \( \phi(C) \).

  • Energy-based: \( p(x) = \frac{1}{Z} \prod_C \phi_C(x_C) \).

  • Relationship to RBMs: RBMs are a type of Markov random field with bipartite structure.

G. Applications of Deep Generative Models
  • Image synthesis: Photorealistic faces (StyleGAN), art.

  • Data augmentation: Generate rare class samples.

  • Drug discovery: Generate molecular structures.

  • Anomaly detection: Model normal data, detect outliers via likelihood.


VIII. DEEP REINFORCEMENT LEARNING

A. Markov Decision Processes (MDPs)
  • Components: \( (S, A, P, R, \gamma) \)

    • \( S \): States.

    • \( A \): Actions.

    • \( P(s'|s,a) \): Transition probability.

    • \( R(s,a,s') \): Reward.

    • \( \gamma \in [0,1] \): Discount factor.

  • Goal: Maximize expected discounted return \( \mathbb{E}[\sum_{t=0}^\infty \gamma^t r_t] \).

B. Dynamic Programming Methods
  • Value iteration:

    \[ V_{k+1}(s) = \max_a \sum_{s'} P(s'|s,a) [R(s,a,s') + \gamma V_k(s')] \]

    Iterate until convergence; policy \( \pi(s) = \arg\max_a Q(s,a) \).

  • Policy iteration:

    • Policy evaluation: Compute \( V^\pi \) by solving linear system.

    • Policy improvement: \( \pi' = \arg\max_\pi \sum_a \pi(a|s) Q^\pi(s,a) \).

    • Comparison: Policy iteration often converges faster but each iteration costly (solving linear system); value iteration cheaper per iteration but may converge slowly.

C. Q-Learning and Deep Q-Networks (DQN)
  • Q-learning (tabular):

    \[ Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma \max_{a'} Q(s',a') - Q(s,a)] \]

    Off-policy, model-free.

  • DQN:

    • Use deep network to approximate \( Q(s,a;\theta) \).

    • Experience replay: Store transitions \( (s,a,r,s') \) in buffer; sample random minibatches → breaks correlation, stabilizes training.

    • Target network: Slow-updated target \( Q_{\text{target}}(s,a;\theta^-) \) to stabilize bootstrapping.

    • Loss: \( \mathcal{L} = \mathbb{E}[(r + \gamma \max_{a'} Q_{\text{target}}(s',a';\theta^-) - Q(s,a;\theta))^2] \).

D. Advanced DQN Algorithms
  • Double DQN:

    • Decouple action selection and evaluation to reduce overestimation bias.

    \[ y = r + \gamma Q_{\text{target}}(s', \arg\max_a Q(s',a;\theta); \theta^-) \]

  • Dueling DQN:

    • Separate value \( V(s) \) and advantage \( A(s,a) \) streams.

    \[ Q(s,a) = V(s) + A(s,a) - \frac{1}{|\mathcal{A}|} \sum_{a'} A(s,a') \]

    • Better value estimation, improved learning.
E. Policy Gradient and Actor-Critic Methods (if covered)
  • Policy gradient (REINFORCE):

    \[ \nabla_\theta J(\theta) = \mathbb{E}_{\tau \sim \pi_\theta} \left[ \sum_t \nabla_\theta \log \pi_\theta(a_t|s_t) G_t \right] \]

    where \( G_t = \sum_{k=t}^T \gamma^{k-t} r_k \).

  • Actor-critic:

    • Actor: Policy \( \pi_\theta \), updated by policy gradient.

    • Critic: Value function \( V_\phi \) or \( Q_\psi \), estimates return.

    • Advantage: Lower variance than pure policy gradient.

F. Least Squares Methods
  • Least Squares Policy Iteration (LSPI):

    • Use least-squares to solve Bellman equation from samples.

    • Represent \( Q(s,a) \) with linear features \( \phi(s,a) \).

    • Solve \( (A - \gamma \Phi' \Phi)^{-1} \Phi' R \) where \( A = \Phi' \Phi \), etc.

    • Comparison: More sample-efficient than standard policy iteration; avoids explicit policy evaluation.

G. Applications
  • Game playing: Atari (DQN), Go (AlphaGo), chess (AlphaZero).

  • Robotics: Control policies for manipulation, locomotion.

  • Resource management: Data center cooling, inventory control.


IX. ADVANCED TOPICS AND MODEL EFFICIENCY

A. Deep Dream
  • Technique: Gradient ascent on input to maximize activation of specific layer/filter.

    \[ x^* = \arg\max_x \mathcal{L}_{\text{activation}}(x) + \text{regularization} \]

  • Role: Visualization of learned features, generating surreal/artistic images.

B. Model Compression and Pruning
  • Unit pruning: Remove entire neurons/filters with low importance (e.g., based on activation magnitude).

  • Weight pruning: Set small weights to zero → sparse networks.

  • Quantization: Reduce precision of weights/activations (e.g., 32-bit → 8-bit).

  • Knowledge distillation: Train small "student" network to mimic large "teacher" network (soft targets).

  • Benefits: Reduced memory, faster inference, deployment on edge devices.

C. Computational Considerations
  • GPU implementation: Parallelize matrix operations (e.g., SVD via cuSOLVER). Randomized SVD for large matrices.

  • Scalability: Distributed training (data/model parallelism), mixed precision training.

D. Directed Graphical Models
  • Bayesian networks: Directed acyclic graph (DAG); nodes = random variables, edges = conditional dependencies.

  • Conditional independence: \( X \perp Y | Z \) means \( p(x|y,z) = p(x|z) \).

  • Inference: Variable elimination, belief propagation.

  • Learning: Parameter estimation (MLE, Bayesian), structure learning.


X. COMPARATIVE AND CRITICAL ANALYSIS

A. Learning Paradigm Comparisons
Paradigm Data Feedback Examples in DL
Supervised Labeled \( (x,y) \) Direct (loss) Classification, regression
Unsupervised Unlabeled \( x \) None (intrinsic) Autoencoders, GANs, clustering
Reinforcement State, action, reward Delayed (reward) DQN, policy gradient
B. Model Comparisons
  • GANs vs VAEs:

    | Aspect | GANs | VAEs | |------------------|-----------------------------------|-----------------------------------| | Objective | Min-max game (adversarial) | Variational inference (ELBO) | | Training | Unstable, mode collapse possible | Stable, convergent | | Samples | Sharp, diverse | Blurry, less diverse | | Latent space | Not necessarily continuous | Continuous, structured | | Use case | High-quality synthesis | Interpolation, representation learning |

  • CNN vs RNN vs Transformer:

    • CNN: Spatial data, translation invariance, parallelizable.

    • RNN: Sequential data, memory, but sequential computation.

    • Transformer: Sequential data with self-attention → parallelizable, handles long-range dependencies (now dominant in NLP).

  • Autoencoders vs PCA: As above.

C. Limitations and Future Directions
  • Current limitations:

    • Data hunger: Requires massive labeled data.

    • Reasoning: Lacks causal, logical reasoning.

    • Common sense: Struggles with everyday physics, social norms.

    • Explainability: Black-box models.

    • Robustness: Sensitive to adversarial examples.

  • Ethical considerations: Bias in data/models, privacy (generative models), misuse (deepfakes), environmental cost.

  • Emerging trends:

    • Self-supervised learning: Pretrain on unlabeled data (e.g., contrastive learning).

    • Neuro-symbolic AI: Combine neural networks with symbolic reasoning.

    • Efficient DL: Model compression, sparse training, federated learning.

    • Foundation models: Large pretrained models (LLMs, vision transformers) fine-tuned for downstream tasks.

[!TIP]

Critical analysis: Always mention both strengths and weaknesses. For example, DL excels at pattern recognition but lacks reasoning; ethical concerns are increasingly important.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in