UNIT 3: Deep Learning for Information Retrieval
I. Foundations of Deep Learning
Historical Progression and Key Milestones
-
1940s-50s: McCulloch-Pitts neuron (binary threshold), Perceptron (Rosenblatt, 1958).
-
1980s: Backpropagation popularized (Rumelhart, Hinton, Williams), Representation Learning concept emerges.
-
2006: "Deep Learning" coined (Hinton et al.) with successful pre-training of Deep Belief Networks (DBNs).
-
2012: AlexNet wins ImageNet, proving deep CNNs' superiority, catalyzing the modern DL era.
-
2014-Present: Generative models (GANs, VAEs), Transformers (2017), and large-scale models dominate.
Learning Paradigms
| Paradigm | Goal | Example in DL |
|---|---|---|
| Supervised | Learn mapping from input x to label y using labeled data. |
Image classification (CNN), Sentiment analysis (RNN/LSTM). |
| Unsupervised | Discover inherent structure/patterns in unlabeled data. | Clustering (Autoencoders), Density estimation (GANs, VAEs). |
| Reinforcement | Learn optimal actions through trial-and-error to maximize cumulative reward. | Game playing (Deep Q-Networks), Robotics. |
Representation Learning
-
Definition: The process where a model automatically discovers the latent representations (features) needed for a task from raw data, rather than relying on manual feature engineering.
-
Significance: Eliminates human bias, handles high-dimensional data (images, text), and enables transfer learning. Deep neural networks are powerful representation learners.
Representation Power of Neural Networks: MLPs and Sigmoid Neurons
-
A Multilayer Perceptron (MLP) with at least one hidden layer and non-linear activation (like sigmoid) is a Universal Function Approximator.
-
It can approximate any continuous function on a compact domain to arbitrary precision given sufficient hidden units.
-
Sigmoid Neurons ($$\displaystyle \sigma(z) = \frac{1}{1+e^{-z}} $$) provide the necessary non-linearity but suffer from vanishing gradients in deep networks.
II. Neural Network Fundamentals
Activation Functions
| Function | Formula | Pros | Cons | Typical Use |
|---|---|---|---|---|
| Sigmoid | $$\displaystyle \sigma(z) = \frac{1}{1+e^{-z}} $$ | Outputs (0,1), smooth gradient. | Vanishing gradients, not zero-centered, slow. | Output layer for binary classification. |
| Tanh | $$\displaystyle \tanh(z) = \frac{e^z - e^{-z}}{e^z + e^{-z}} $$ | Outputs (-1,1), zero-centered. | Vanishing gradients. | Hidden layers (often better than sigmoid). |
| ReLU | $$\displaystyle f(z) = \max(0, z) $$ | Computationally cheap, sparsifies activations, mitigates vanishing gradient (for +ve inputs). | "Dying ReLU" problem (neurons stuck at 0). | Default choice for hidden layers in CNNs/MLPs. |
| Leaky ReLU | $$\displaystyle f(z) = \max(\alpha z, z) $$, $\alpha$ small (e.g., 0.01) | Fixes "dying ReLU", small gradient for -ve inputs. | Results inconsistent, may not converge. | Alternative to ReLU. |
| Softmax | $$\displaystyle S(z)_j = \frac{e^{z_j}}{\sum_{k=1}^K e^{z_k}} $$ | Normalizes outputs to probability distribution (sum=1). | Sensitive to outliers. | Output layer for multi-class classification. |
[!TIP] Exam Focus: ReLU is the most critical activation. Be ready to explain why it mitigates vanishing gradients and its "dying" issue.
Backpropagation Algorithm (Derivation & Weight Updates)
-
Forward Pass: Compute network output and loss $L$ for a given input.
-
Backward Pass (Chain Rule): Compute gradient of loss w.r.t. each weight $$\displaystyle \frac{\partial L}{\partial w_{ij}} $$, starting from output layer and moving backwards.
-
For output layer: $$\displaystyle \delta^{(L)} = \nabla_a L \odot f'(z^{(L)}) $$
-
For hidden layer
l: $$\displaystyle \delta^{(l)} = ((w^{(l+1)})^T \delta^{(l+1)}) \odot f'(z^{(l)}) $$
-
-
Weight Update (SGD): $$\displaystyle w_{ij}^{(l)} \leftarrow w_{ij}^{(l)} - \eta \frac{\partial L}{\partial w_{ij}^{(l)}} = w_{ij}^{(l)} - \eta \delta_j^{(l)} a_i^{(l-1)} $$
- $\eta$: learning rate, $\delta$: error term, $a$: activation.
Optimization Algorithms
| Algorithm | Key Idea | Update Rule (for weight w) |
Pros/Cons |
|---|---|---|---|
| SGD | Fixed learning rate $\eta$. | $$\displaystyle w \leftarrow w - \eta \cdot \nabla L(w) $$ | Simple, but noisy convergence. |
| Momentum | Accumulate past gradients as velocity. | $$\displaystyle v \leftarrow \beta v + (1-\beta) \nabla L(w) $$ <br> $$\displaystyle w \leftarrow w - \eta v $$ | Accelerates convergence, dampens oscillations. |
| AdaGrad | Adapts per-parameter learning rate using sum of squared gradients. | $$\displaystyle g \leftarrow g + \nabla L(w)^2 $$ <br> $$\displaystyle w \leftarrow w - \frac{\eta}{\sqrt{g+\epsilon}} \nabla L(w) $$ | Good for sparse data, but learning rates may decay too fast. |
| RMSProp | Uses exponentially weighted moving average of squared gradients. | $$\displaystyle s \leftarrow \beta s + (1-\beta) \nabla L(w)^2 $$ <br> $$\displaystyle w \leftarrow w - \frac{\eta}{\sqrt{s+\epsilon}} \nabla L(w) $$ | Fixes AdaGrad's aggressive decay. |
| Adam (Most Popular) | Combines Momentum (1st moment) & RMSProp (2nd moment) with bias correction. | $$\displaystyle m \leftarrow \beta_1 m + (1-\beta_1) \nabla L $$ <br> $$\displaystyle v \leftarrow \beta_2 v + (1-\beta_2) \nabla L^2 $$ <br> $$\displaystyle \hat{m} = m/(1-\beta_1^t) $$, $$\displaystyle \hat{v} = v/(1-\beta_2^t) $$ <br> $$\displaystyle w \leftarrow w - \eta \frac{\hat{m}}{\sqrt{\hat{v}}+\epsilon} $$ | Robust, handles noisy/sparse gradients, default choice. |
III. Feedforward Neural Networks
Multilayer Perceptrons (MLPs)
-
Architecture: Fully connected layers. Input layer → one or more hidden layers (with non-linear activation) → output layer.
-
Working: Information flows forward (feedforward). Each neuron computes $$\displaystyle z = w^T x + b $$, then applies activation $$\displaystyle a = f(z) $$.
-
Universal Approximation: A single hidden layer MLP can approximate any continuous function given enough neurons.
Deep Feedforward Networks
-
Architecture: MLP with multiple hidden layers.
-
Advantages:
-
Hierarchical Feature Learning: Early layers learn simple features (edges), later layers compose them into complex concepts (objects).
-
Parameter Efficiency: Can represent some functions exponentially more efficiently than shallow networks.
-
Better Generalization: Often learns more abstract, robust representations.
-
-
Applications: Tabular data classification/regression, as final layers in CNNs/RNNs.
Logistic Regression as a Linear Classifier
-
Model: Single-layer neural network (no hidden layer). $$\displaystyle P(y=1|x) = \sigma(w^T x + b) $$.
-
Decision Boundary: $$\displaystyle w^T x + b = 0 $$ is a linear hyperplane.
-
Training: Minimize binary cross-entropy loss using gradient descent.
-
Limitation: Cannot learn non-linear decision boundaries; requires feature engineering for complex data.
IV. Convolutional Neural Networks (CNNs)
Core Concepts
-
Convolution: Apply a filter/kernel (small weight matrix) across the input (image, feature map) to produce a feature map. Captures local patterns (edges, textures).
- $$\displaystyle Output_{i,j} = \sum_{m}\sum_{n} Input_{i+m, j+n} \cdot Kernel_{m,n} $$
-
Filters: Learnable parameters. Different filters detect different features (vertical edge, color blob).
-
Feature Maps: 2D activation maps resulting from a filter's convolution over the input.
Padding and Pooling
| Concept | Purpose | Types |
|---|---|---|
| Padding | Preserve spatial dimensions of output; avoid information loss at borders. | Valid: No padding (output shrinks). <br> Same: Pad with zeros so output size = input size. |
| Pooling | Downsample feature maps; introduce spatial invariance; reduce parameters/computation. | Max Pooling: Take maximum value in window. (Most common) <br> Average Pooling: Take average value in window. |
ReLU in CNNs
-
Applied element-wise after each convolutional layer.
-
Significance: Introduces non-linearity, is computationally cheap, and helps mitigate vanishing gradients. Standard in modern CNNs (e.g., after every conv layer in ResNet).
CNN Architectures for Image Recognition
-
LeNet-5 (1998): First successful CNN (handwritten digit recognition).
-
AlexNet (2012): Deeper (5 conv layers), used ReLU & Dropout, won ImageNet.
-
VGGNet (2014): Very deep (16-19 layers), used small (3x3) filters throughout.
-
ResNet (2015): Introduced residual connections (skip connections) to train very deep networks (100+ layers) by solving vanishing gradient.
-
Inception/GoogLeNet (2014): Used inception modules with parallel conv layers of different filter sizes.
Data Formats for CNNs
| Data Type | Typical Format | Example | CNN Adaptation |
|---|---|---|---|
| 2D Image | (Height, Width, Channels) |
(224, 224, 3) for RGB |
Standard 2D convolutions. |
| 3D Video | (Frames, Height, Width, Channels) |
(30, 224, 224, 3) |
3D Convolutions (kernel moves in time) or 2D+1D (spatial 2D + temporal 1D). |
| 1D Signal | (Timesteps, Channels) |
Audio waveform (16000,) |
1D Convolutions (kernel moves along time). |
| Text | (Sequence_Length,) (word indices) |
Sentence of 50 words | 1D Convolutions over word embeddings. |
[!TIP] Exam Question: "Develop a table with examples..." – Use the table above. Be ready to explain why 1D convs are used for text and 3D for video.
V. Recurrent Neural Networks (RNNs)
Basic RNN Architecture for Sequential Data
-
Core Idea: Has a hidden state $$\displaystyle h_t $$ that acts as memory, passed from one timestep to the next.
-
Equations:
-
$$\displaystyle h_t = f(W_{xh} x_t + W_{hh} h_{t-1} + b_h) $$
-
$$\displaystyle y_t = g(W_{hy} h_t + b_y) $$
-
$f, g$: Activation functions (typically tanh for $$\displaystyle h_t $$, softmax for $$\displaystyle y_t $$).
-
-
Use: Natural language (words in sentence), time series (stock prices), audio.
Backpropagation Through Time (BPTT)
-
Unfold the RNN for
Ttimesteps, creating a deep feedforward network. -
Apply standard backpropagation through this unfolded graph.
-
Gradient at time
tdepends on gradients and activations from all future timestepst+1...T.- $$\displaystyle \frac{\partial L}{\partial h_t} = \sum_{k=t}^{T} \frac{\partial L}{\partial h_k} \frac{\partial h_k}{\partial h_t} $$
Vanishing and Exploding Gradients in RNNs
-
Cause: Repeated multiplication by the recurrent weight matrix $$\displaystyle W_{hh} $$ during BPTT.
- Gradient involves product: $$\displaystyle \prod_{i=1}^{T} \frac{\partial h_{t+i}}{\partial h_{t+i-1}} \approx (W_{hh})^T $$
-
Vanishing Gradient: If eigenvalues of $$\displaystyle W_{hh} < 1 $$, gradients shrink exponentially → early timesteps learn very slowly. Cannot learn long-range dependencies.
-
Exploding Gradient: If eigenvalues $$\displaystyle > 1 $$, gradients explode → unstable training, NaN weights.
-
Impact: Simple RNNs fail on sequences longer than ~10-20 steps.
Long Short-Term Memory (LSTM)
-
Goal: Design a cell that can maintain a constant error flow, mitigating vanishing gradients.
-
Key Components:
-
Cell State ($$\displaystyle C_t $$): The "conveyor belt" carrying information unchanged across timesteps.
-
Gates: Sigmoid layers controlling information flow.
-
Forget Gate ($$\displaystyle f_t $$): What to remove from cell state? $$\displaystyle f_t = \sigma(W_f \cdot [h_{t-1}, x_t] + b_f) $$
-
Input Gate ($$\displaystyle i_t $$): What new info to store? $$\displaystyle i_t = \sigma(W_i \cdot [h_{t-1}, x_t] + b_i) $$
-
Candidate Cell State ($$\displaystyle \tilde{C}_t $$): New candidate values. $$\displaystyle \tilde{C}_t = \tanh(W_C \cdot [h_{t-1}, x_t] + b_C) $$
-
Output Gate ($$\displaystyle o_t $$): What to output as hidden state? $$\displaystyle o_t = \sigma(W_o \cdot [h_{t-1}, x_t] + b_o) $$
-
-
-
Working (Equations):
-
$$\displaystyle C_t = f_t \odot C_{t-1} + i_t \odot \tilde{C}_t $$ (Element-wise multiplication)
-
$$\displaystyle h_t = o_t \odot \tanh(C_t) $$
-
-
Advantages over Simple RNNs:
-
Explicitly models long-term dependencies via cell state.
-
Gates allow stable gradient flow (additive update to $$\displaystyle C_t $$).
-
Can learn over hundreds of timesteps.
-
Deep RNNs (Stacked Architectures)
-
Stack multiple RNN/LSTM layers on top of each other.
-
Lower layers learn low-level temporal features, higher layers learn high-level abstractions.
-
Increases model capacity and representational power.
Recursive Neural Networks (Tree Structures)
-
Architecture: Process tree-structured data (e.g., parse trees of sentences).
-
Working: Apply the same neural network (same weights) in a bottom-up manner at each node, combining child representations to form parent representation.
-
Applications: Sentiment analysis (on parse trees), Semantic parsing.
VI. Autoencoders and Generative Models
Autoencoders (AEs)
-
Architecture: Encoder ($$\displaystyle z = f(x) $$) → Bottleneck (latent code $z$) → Decoder ($$\displaystyle \hat{x} = g(z) $$). Trained to reconstruct input: minimize $L(x, \hat{x})$.
-
Training: Unsupervised (no labels needed, just input data).
-
Applications: Dimensionality reduction, denoising, anomaly detection, feature learning.
Autoencoder Variants
| Variant | Modification | Purpose |
|---|---|---|
| Sparse AE | Add sparsity penalty (e.g., KL divergence) to activations of hidden layer. | Forces network to learn meaningful, compact features; useful for feature extraction. |
| Contractive AE | Add penalty on the Frobenius norm of the Jacobian of encoder activations. | Encourages encoder to be robust to small input perturbations → learns locally invariant features. |
| Denoising AE | Input is corrupted version $\tilde{x}$ (e.g., added noise, masked), target is original clean $x$. | Forces network to learn robust features and reconstruct from partial info. More powerful than standard AE. |
Autoencoders vs PCA/SVD for Dimensionality Reduction
| Feature | PCA/SVD | Autoencoder |
|---|---|---|
| Linearity | Linear transformation only. | Non-linear (due to activations). |
| Flexibility | Single, fixed transformation. | Can learn complex, multi-stage non-linear mappings. |
| Data Assumption | Assumes Gaussian distribution, linear relationships. | Makes no strong distributional assumptions. |
| When to use AE | When data has non-linear structure (e.g., manifold of images). | PCA fails to capture complex manifolds. |
Regularization in Autoencoders
-
Purpose: Prevent autoencoder from learning trivial identity function ($$\displaystyle x \rightarrow x $$), especially when latent dimension is large.
-
Methods:
-
Sparsity Penalty: Force few active neurons.
-
Denoising: Corrupt input, reconstruct clean.
-
Contractive Penalty: Penalize large Jacobian.
-
Dropout: Applied to inputs/hidden layers.
-
Variational Autoencoders (VAEs)
-
Key Idea: Generative model. Learns a latent variable model $$\displaystyle p(x) = \int p(x|z)p(z)dz $$.
-
Architecture: Encoder outputs parameters of a distribution (usually Gaussian: mean $\mu$ and variance $$\displaystyle \sigma^2 $$) for latent code $z$, not a single point.
-
Reparameterization Trick: Sample $$\displaystyle z = \mu + \sigma \odot \epsilon $$, $\epsilon \sim \mathcal{N}(0,I)$. Allows gradient flow through sampling.
-
Loss: Reconstruction loss (e.g., MSE) + KL Divergence between encoder's distribution and prior (usually $\mathcal{N}(0,I)$). Balances reconstruction quality and latent space structure.
-
Output: Can generate new samples by sampling $z$ from prior and decoding.
Generative Adversarial Networks (GANs)
-
Architecture: Two networks in competition:
-
Generator ($G$): Takes random noise $$\displaystyle z \sim p_z $$, generates fake data $$\displaystyle \hat{x} = G(z) $$.
-
Discriminator ($D$): Classifies input as real ($x$) or fake ($\hat{x}$).
-
-
Training (Adversarial/Minimax Game):
-
$D$ maximizes $$\displaystyle \mathbb{E}_{x\sim p_{data}}[\log D(x)] + \mathbb{E}_{z\sim p_z}[\log(1-D(G(z)))] $$
-
$G$ minimizes $$\displaystyle \mathbb{E}_{z\sim p_z}[\log(1-D(G(z)))] $$
-
-
Result: $G$ generates increasingly realistic data; $D$ becomes a powerful classifier.
GANs vs VAEs
| Aspect | GAN | VAE |
|---|---|---|
| Approach | Adversarial, game-theoretic. | Probabilistic, variational inference. |
| Output Quality | Sharper, more realistic samples. | Often blurry samples. |
| Training Stability | Unstable, mode collapse, hard to evaluate. | Stable, well-defined loss. |
| Latent Space | Not necessarily continuous or meaningful. | Structured, continuous, good for interpolation. |
| Choose GAN when | Ultimate sample quality is critical (e.g., art, super-resolution). | Choose VAE when you need structured latent space for manipulation, interpolation, or semi-supervised learning. |
Auto-regressive Models: NADE & MADE
-
Core Idea: Factorize joint distribution $p(x)$ as product of conditionals: $$\displaystyle p(x) = \prod_{i=1}^d p(x_i | x_{<i}) $$.
-
NADE (Neural Autoregressive Distribution Estimator): Uses a neural network with masked connections to ensure each input $$\displaystyle x_i $$ only depends on previous inputs $$\displaystyle x_{1..i-1} $$. Efficient training via BPTT.
-
MADE (Masked Autoencoder for Distribution Estimation): Applies the masking idea to a feedforward autoencoder. Faster parallel training than NADE. Foundation for PixelCNN (images), WaveNet (audio).
VII. Training Deep Networks
Overfitting and Underfitting
| Underfitting | Overfitting | |
|---|---|---|
| Definition | Model is too simple; fails to capture underlying pattern. | Model learns noise/irrelevant details in training data. |
| Training Error | High | Very Low |
| Validation Error | High | High (gap between train/val is large) |
| Cause | High bias, low model capacity. | High variance, low training data, high model capacity. |
Regularization Techniques
| Technique | Mechanism | Effect |
|---|---|---|
| Dropout | Randomly "drop" (set to 0) a fraction p of neurons during training only. |
Prevents co-adaptation of neurons; acts like an ensemble of many thinned networks. |
| Weight Decay (L2) | Add $$\displaystyle \frac{\lambda}{2} \|w\|^2 $$ to loss. | Penalizes large weights, encourages smaller, more distributed weights. |
| Early Stopping | Monitor validation loss; stop training when it starts to increase. | Prevents memorization of training set noise. |
| Data Augmentation | Artificially enlarge training set by applying transformations (rotate, crop, flip images). | Increases effective dataset size, teaches invariances. |
Batch Normalization
-
Mechanism: For each mini-batch, normalize activations of a layer to have zero mean and unit variance. Then scale ($\gamma$) and shift ($\beta$) with learnable parameters.
-
$$\displaystyle \hat{x}^{(k)} = \frac{x^{(k)} - \mu_B}{\sqrt{\sigma_B^2 + \epsilon}} $$
-
$$\displaystyle y^{(k)} = \gamma \hat{x}^{(k)} + \beta $$
-
-
Benefits:
-
Reduces internal covariate shift (distribution changes in layer inputs during training).
-
Allows use of higher learning rates.
-
Acts as a regularizer (adds noise via mini-batch stats).
-
Reduces need for careful weight initialization.
-
-
Usage: After convolution/linear layer, before non-linearity (ReLU).
Gradient Problems & Mitigation
| Problem | Cause | Mitigation Strategies |
|---|---|---|
| Vanishing Gradients | Repeated multiplication by small values (sigmoid/tanh derivatives, $$\displaystyle |W|<1 $$). | Use ReLU activations, BatchNorm, Residual Connections (ResNet), LSTM/GRU for RNNs. |
| Exploding Gradients | Repeated multiplication by large values ($$\displaystyle |W|>1 $$). | Gradient Clipping (cap gradient norm), Weight Initialization (Xavier/He), BatchNorm. |
Data Preprocessing: Normalization vs Standardization
| Normalization (Min-Max Scaling) | Standardization (Z-score) | |
|---|---|---|
| Formula | $$\displaystyle x' = \frac{x - x_{min}}{x_{max} - x_{min}} $$ | $$\displaystyle x' = \frac{x - \mu}{\sigma} $$ |
| Output Range | [0, 1] (or [-1, 1]) |
Mean=0, Std=1 (unbounded). |
| Sensitivity | Sensitive to outliers. | Less sensitive to outliers. |
| Use Case | When algorithm requires bounded input (e.g., neural nets with sigmoid output). | Default choice for most DL algorithms (works well with zero-centered data). |
VIII. Advanced Topics
Deep Belief Networks (DBNs)
-
Architecture: Stack of Restricted Boltzmann Machines (RBMs). Each RBM is a bipartite graph (visible layer, hidden layer).
-
Pre-training (Greedy Layer-Wise):
-
Train first RBM on raw data.
-
Use its hidden layer activations as "data" for the next RBM.
-
Repeat for all layers.
-
-
Fine-tuning: After pre-training, optionally fine-tune all weights jointly with backpropagation.
-
Significance: One of the first successful methods for deep network pre-training, helping overcome vanishing gradients before ReLU/BatchNorm era.
Deep Reinforcement Learning
-
Markov Decision Process (MDP): $(S, A, P, R, \gamma)$. State $s$, Action $a$, Transition $P(s'|s,a)$, Reward $R(s,a)$, Discount $\gamma$.
-
Value Iteration (Dynamic Programming): Iteratively compute optimal state-value function $$\displaystyle V^*(s) $$.
- $$\displaystyle V_{k+1}(s) = \max_a \sum_{s'} P(s'|s,a) [R(s,a) + \gamma V_k(s')] $$
-
Policy Iteration: Alternate between:
-
Policy Evaluation: Compute $$\displaystyle V^\pi $$ for current policy $\pi$.
-
Policy Improvement: $$\displaystyle \pi'(s) = \arg\max_a \sum_{s'} P(s'|s,a)[R(s,a) + \gamma V^\pi(s')] $$.
-
-
Q-learning (Model-Free): Learn action-value function $Q(s,a)$.
- Update: $$\displaystyle Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma \max_{a'} Q(s',a') - Q(s,a)] $$
-
Deep Q-Network (DQN): Use deep CNN to approximate $Q(s,a;\theta)$. Key tricks: Experience Replay, Target Network.
-
Double DQN: Decouples action selection and evaluation to reduce overestimation bias.
- $$\displaystyle y = r + \gamma Q(s', \arg\max_{a'} Q(s',a'; \theta); \theta^-) $$
-
Dueling DQN: Separately estimates state-value $V(s)$ and advantage $A(s,a)$, then combines: $$\displaystyle Q(s,a) = V(s) + A(s,a) - \frac{1}{|\mathcal{A}|}\sum_{a'} A(s,a') $$.
-
LSPI (Least Squares Policy Iteration): Uses linear function approximation (e.g., RBF features) and solves a least-squares problem for policy evaluation. More sample-efficient but less scalable than DQN.
Deep Dream
-
Goal: Generate surreal, dream-like images by maximizing activations of specific layers/neurons in a trained CNN.
-
Process:
-
Start with noise or input image.
-
Perform forward pass, compute loss = activation of target layer/neuron.
-
Compute gradient of loss w.r.t. input image.
-
Update input image in the direction of the gradient (ascent).
-
Repeat.
-
-
Result: Amplifies patterns the network has learned (e.g., eyes, dog faces), creating hallucinatory images.
Model Compression: Unit Pruning
-
Goal: Reduce model size/inference time by removing unimportant neurons (units) or connections.
-
Unit Pruning: Remove entire neurons (and their incoming/outgoing connections) based on a importance score (e.g., L1 norm of weights, activation magnitude).
-
Need: Deploy models on resource-constrained devices (mobile, IoT), reduce latency, energy consumption.
-
Process: Train full model → prune units based on score → fine-tune remaining weights → iterate.
Directed Graphical Models (Basics)
-
Definition: Probabilistic graphical models where nodes represent random variables and directed edges represent conditional dependencies (causal relationships).
-
Example: Bayesian Network. Encodes joint distribution as $$\displaystyle p(x_1,...,x_n) = \prod_i p(x_i | \text{parents}(x_i)) $$.
-
Use in DL: VAEs have a directed graphical model structure: $$\displaystyle z \rightarrow x $$. Provides probabilistic interpretation.
Computational Efficiency: GPU Implementation, Randomized SVD
-
GPU: Highly parallel hardware. Ideal for matrix/tensor operations (convolutions, matrix multiplies) in DL. Requires data batch processing and careful memory management.
-
Randomized SVD: Approximates SVD of large matrix $A$ using random projections. Faster ($O(mn \log k)$) than deterministic SVD ($O(mn \min(m,n))$). Used for:
-
Fast PCA on large datasets.
-
Pre-conditioning or initializing layers (e.g., in DBNs).
-
Low-rank approximation of weight matrices for compression.
-
IX. Evaluation and Applications (Brief)
Evaluation Metrics
| Task Type | Common Metrics |
|---|---|
| Classification | Accuracy, Precision, Recall, F1-Score, ROC-AUC, Confusion Matrix. |
| Generation | Inception Score (IS), Fréchet Inception Distance (FID), BLEU (text). |
| Sequence Modeling | Perplexity (language models), BLEU (translation). |
Applications
-
Computer Vision: Image classification (ResNet), Object detection (YOLO), Segmentation (U-Net).
-
Natural Language Processing: Machine translation (Transformer), Sentiment analysis (LSTM), Question answering (BERT).
-
Recommendation Systems: Collaborative filtering with neural embeddings (NeuMF), sequential recommendations (RNNs).
Case Studies
-
Image Recognition: AlexNet (2012) demonstrated deep CNNs' superiority, reducing top-5 error from 26% to 15% on ImageNet. Key: ReLU, Dropout, GPU training.
-
Text Sequence Modeling: LSTMs dominated machine translation and language modeling before Transformers. Handled variable-length sequences, captured long-range dependencies, enabled Google's Neural Machine Translation system.
[!TIP] Final Exam Strategy: For 7-mark questions, structure your answer: 1. Clear Definition, 2. Core Mechanism/Architecture (use equations/diagrams if possible), 3. Key Advantages/Disadvantages, 4. One Concrete Example/Application. Always link back to Information Retrieval context where possible (e.g., "RNNs for query suggestion", "CNNs for document image retrieval", "Autoencoders for query/document embedding").