Skip to content
AL-503 (A) · Information Retrieval/Quick Revision Short Notes

Information Retrieval (AL-503 (A)) - Unit 3 Short Notes

UNIT 3: Deep Learning for Information Retrieval


I. Foundations of Deep Learning

Historical Progression and Key Milestones

  • 1940s-50s: McCulloch-Pitts neuron (binary threshold), Perceptron (Rosenblatt, 1958).

  • 1980s: Backpropagation popularized (Rumelhart, Hinton, Williams), Representation Learning concept emerges.

  • 2006: "Deep Learning" coined (Hinton et al.) with successful pre-training of Deep Belief Networks (DBNs).

  • 2012: AlexNet wins ImageNet, proving deep CNNs' superiority, catalyzing the modern DL era.

  • 2014-Present: Generative models (GANs, VAEs), Transformers (2017), and large-scale models dominate.

Learning Paradigms

Paradigm Goal Example in DL
Supervised Learn mapping from input x to label y using labeled data. Image classification (CNN), Sentiment analysis (RNN/LSTM).
Unsupervised Discover inherent structure/patterns in unlabeled data. Clustering (Autoencoders), Density estimation (GANs, VAEs).
Reinforcement Learn optimal actions through trial-and-error to maximize cumulative reward. Game playing (Deep Q-Networks), Robotics.

Representation Learning

  • Definition: The process where a model automatically discovers the latent representations (features) needed for a task from raw data, rather than relying on manual feature engineering.

  • Significance: Eliminates human bias, handles high-dimensional data (images, text), and enables transfer learning. Deep neural networks are powerful representation learners.

Representation Power of Neural Networks: MLPs and Sigmoid Neurons

  • A Multilayer Perceptron (MLP) with at least one hidden layer and non-linear activation (like sigmoid) is a Universal Function Approximator.

  • It can approximate any continuous function on a compact domain to arbitrary precision given sufficient hidden units.

  • Sigmoid Neurons ($$\displaystyle \sigma(z) = \frac{1}{1+e^{-z}} $$) provide the necessary non-linearity but suffer from vanishing gradients in deep networks.


II. Neural Network Fundamentals

Activation Functions

Function Formula Pros Cons Typical Use
Sigmoid $$\displaystyle \sigma(z) = \frac{1}{1+e^{-z}} $$ Outputs (0,1), smooth gradient. Vanishing gradients, not zero-centered, slow. Output layer for binary classification.
Tanh $$\displaystyle \tanh(z) = \frac{e^z - e^{-z}}{e^z + e^{-z}} $$ Outputs (-1,1), zero-centered. Vanishing gradients. Hidden layers (often better than sigmoid).
ReLU $$\displaystyle f(z) = \max(0, z) $$ Computationally cheap, sparsifies activations, mitigates vanishing gradient (for +ve inputs). "Dying ReLU" problem (neurons stuck at 0). Default choice for hidden layers in CNNs/MLPs.
Leaky ReLU $$\displaystyle f(z) = \max(\alpha z, z) $$, $\alpha$ small (e.g., 0.01) Fixes "dying ReLU", small gradient for -ve inputs. Results inconsistent, may not converge. Alternative to ReLU.
Softmax $$\displaystyle S(z)_j = \frac{e^{z_j}}{\sum_{k=1}^K e^{z_k}} $$ Normalizes outputs to probability distribution (sum=1). Sensitive to outliers. Output layer for multi-class classification.

[!TIP] Exam Focus: ReLU is the most critical activation. Be ready to explain why it mitigates vanishing gradients and its "dying" issue.

Backpropagation Algorithm (Derivation & Weight Updates)

  1. Forward Pass: Compute network output and loss $L$ for a given input.

  2. Backward Pass (Chain Rule): Compute gradient of loss w.r.t. each weight $$\displaystyle \frac{\partial L}{\partial w_{ij}} $$, starting from output layer and moving backwards.

    • For output layer: $$\displaystyle \delta^{(L)} = \nabla_a L \odot f'(z^{(L)}) $$

    • For hidden layer l: $$\displaystyle \delta^{(l)} = ((w^{(l+1)})^T \delta^{(l+1)}) \odot f'(z^{(l)}) $$

  3. Weight Update (SGD): $$\displaystyle w_{ij}^{(l)} \leftarrow w_{ij}^{(l)} - \eta \frac{\partial L}{\partial w_{ij}^{(l)}} = w_{ij}^{(l)} - \eta \delta_j^{(l)} a_i^{(l-1)} $$

    • $\eta$: learning rate, $\delta$: error term, $a$: activation.

Optimization Algorithms

Algorithm Key Idea Update Rule (for weight w) Pros/Cons
SGD Fixed learning rate $\eta$. $$\displaystyle w \leftarrow w - \eta \cdot \nabla L(w) $$ Simple, but noisy convergence.
Momentum Accumulate past gradients as velocity. $$\displaystyle v \leftarrow \beta v + (1-\beta) \nabla L(w) $$ <br> $$\displaystyle w \leftarrow w - \eta v $$ Accelerates convergence, dampens oscillations.
AdaGrad Adapts per-parameter learning rate using sum of squared gradients. $$\displaystyle g \leftarrow g + \nabla L(w)^2 $$ <br> $$\displaystyle w \leftarrow w - \frac{\eta}{\sqrt{g+\epsilon}} \nabla L(w) $$ Good for sparse data, but learning rates may decay too fast.
RMSProp Uses exponentially weighted moving average of squared gradients. $$\displaystyle s \leftarrow \beta s + (1-\beta) \nabla L(w)^2 $$ <br> $$\displaystyle w \leftarrow w - \frac{\eta}{\sqrt{s+\epsilon}} \nabla L(w) $$ Fixes AdaGrad's aggressive decay.
Adam (Most Popular) Combines Momentum (1st moment) & RMSProp (2nd moment) with bias correction. $$\displaystyle m \leftarrow \beta_1 m + (1-\beta_1) \nabla L $$ <br> $$\displaystyle v \leftarrow \beta_2 v + (1-\beta_2) \nabla L^2 $$ <br> $$\displaystyle \hat{m} = m/(1-\beta_1^t) $$, $$\displaystyle \hat{v} = v/(1-\beta_2^t) $$ <br> $$\displaystyle w \leftarrow w - \eta \frac{\hat{m}}{\sqrt{\hat{v}}+\epsilon} $$ Robust, handles noisy/sparse gradients, default choice.

III. Feedforward Neural Networks

Multilayer Perceptrons (MLPs)

  • Architecture: Fully connected layers. Input layer → one or more hidden layers (with non-linear activation) → output layer.

  • Working: Information flows forward (feedforward). Each neuron computes $$\displaystyle z = w^T x + b $$, then applies activation $$\displaystyle a = f(z) $$.

  • Universal Approximation: A single hidden layer MLP can approximate any continuous function given enough neurons.

Deep Feedforward Networks

  • Architecture: MLP with multiple hidden layers.

  • Advantages:

    1. Hierarchical Feature Learning: Early layers learn simple features (edges), later layers compose them into complex concepts (objects).

    2. Parameter Efficiency: Can represent some functions exponentially more efficiently than shallow networks.

    3. Better Generalization: Often learns more abstract, robust representations.

  • Applications: Tabular data classification/regression, as final layers in CNNs/RNNs.

Logistic Regression as a Linear Classifier

  • Model: Single-layer neural network (no hidden layer). $$\displaystyle P(y=1|x) = \sigma(w^T x + b) $$.

  • Decision Boundary: $$\displaystyle w^T x + b = 0 $$ is a linear hyperplane.

  • Training: Minimize binary cross-entropy loss using gradient descent.

  • Limitation: Cannot learn non-linear decision boundaries; requires feature engineering for complex data.


IV. Convolutional Neural Networks (CNNs)

Core Concepts

  • Convolution: Apply a filter/kernel (small weight matrix) across the input (image, feature map) to produce a feature map. Captures local patterns (edges, textures).

    • $$\displaystyle Output_{i,j} = \sum_{m}\sum_{n} Input_{i+m, j+n} \cdot Kernel_{m,n} $$
  • Filters: Learnable parameters. Different filters detect different features (vertical edge, color blob).

  • Feature Maps: 2D activation maps resulting from a filter's convolution over the input.

Padding and Pooling

Concept Purpose Types
Padding Preserve spatial dimensions of output; avoid information loss at borders. Valid: No padding (output shrinks). <br> Same: Pad with zeros so output size = input size.
Pooling Downsample feature maps; introduce spatial invariance; reduce parameters/computation. Max Pooling: Take maximum value in window. (Most common) <br> Average Pooling: Take average value in window.

ReLU in CNNs

  • Applied element-wise after each convolutional layer.

  • Significance: Introduces non-linearity, is computationally cheap, and helps mitigate vanishing gradients. Standard in modern CNNs (e.g., after every conv layer in ResNet).

CNN Architectures for Image Recognition

  • LeNet-5 (1998): First successful CNN (handwritten digit recognition).

  • AlexNet (2012): Deeper (5 conv layers), used ReLU & Dropout, won ImageNet.

  • VGGNet (2014): Very deep (16-19 layers), used small (3x3) filters throughout.

  • ResNet (2015): Introduced residual connections (skip connections) to train very deep networks (100+ layers) by solving vanishing gradient.

  • Inception/GoogLeNet (2014): Used inception modules with parallel conv layers of different filter sizes.

Data Formats for CNNs

Data Type Typical Format Example CNN Adaptation
2D Image (Height, Width, Channels) (224, 224, 3) for RGB Standard 2D convolutions.
3D Video (Frames, Height, Width, Channels) (30, 224, 224, 3) 3D Convolutions (kernel moves in time) or 2D+1D (spatial 2D + temporal 1D).
1D Signal (Timesteps, Channels) Audio waveform (16000,) 1D Convolutions (kernel moves along time).
Text (Sequence_Length,) (word indices) Sentence of 50 words 1D Convolutions over word embeddings.

[!TIP] Exam Question: "Develop a table with examples..." – Use the table above. Be ready to explain why 1D convs are used for text and 3D for video.


V. Recurrent Neural Networks (RNNs)

Basic RNN Architecture for Sequential Data

  • Core Idea: Has a hidden state $$\displaystyle h_t $$ that acts as memory, passed from one timestep to the next.

  • Equations:

    • $$\displaystyle h_t = f(W_{xh} x_t + W_{hh} h_{t-1} + b_h) $$

    • $$\displaystyle y_t = g(W_{hy} h_t + b_y) $$

    • $f, g$: Activation functions (typically tanh for $$\displaystyle h_t $$, softmax for $$\displaystyle y_t $$).

  • Use: Natural language (words in sentence), time series (stock prices), audio.

Backpropagation Through Time (BPTT)

  • Unfold the RNN for T timesteps, creating a deep feedforward network.

  • Apply standard backpropagation through this unfolded graph.

  • Gradient at time t depends on gradients and activations from all future timesteps t+1...T.

    • $$\displaystyle \frac{\partial L}{\partial h_t} = \sum_{k=t}^{T} \frac{\partial L}{\partial h_k} \frac{\partial h_k}{\partial h_t} $$

Vanishing and Exploding Gradients in RNNs

  • Cause: Repeated multiplication by the recurrent weight matrix $$\displaystyle W_{hh} $$ during BPTT.

    • Gradient involves product: $$\displaystyle \prod_{i=1}^{T} \frac{\partial h_{t+i}}{\partial h_{t+i-1}} \approx (W_{hh})^T $$
  • Vanishing Gradient: If eigenvalues of $$\displaystyle W_{hh} < 1 $$, gradients shrink exponentially → early timesteps learn very slowly. Cannot learn long-range dependencies.

  • Exploding Gradient: If eigenvalues $$\displaystyle > 1 $$, gradients explode → unstable training, NaN weights.

  • Impact: Simple RNNs fail on sequences longer than ~10-20 steps.

Long Short-Term Memory (LSTM)

  • Goal: Design a cell that can maintain a constant error flow, mitigating vanishing gradients.

  • Key Components:

    1. Cell State ($$\displaystyle C_t $$): The "conveyor belt" carrying information unchanged across timesteps.

    2. Gates: Sigmoid layers controlling information flow.

      • Forget Gate ($$\displaystyle f_t $$): What to remove from cell state? $$\displaystyle f_t = \sigma(W_f \cdot [h_{t-1}, x_t] + b_f) $$

      • Input Gate ($$\displaystyle i_t $$): What new info to store? $$\displaystyle i_t = \sigma(W_i \cdot [h_{t-1}, x_t] + b_i) $$

      • Candidate Cell State ($$\displaystyle \tilde{C}_t $$): New candidate values. $$\displaystyle \tilde{C}_t = \tanh(W_C \cdot [h_{t-1}, x_t] + b_C) $$

      • Output Gate ($$\displaystyle o_t $$): What to output as hidden state? $$\displaystyle o_t = \sigma(W_o \cdot [h_{t-1}, x_t] + b_o) $$

  • Working (Equations):

    • $$\displaystyle C_t = f_t \odot C_{t-1} + i_t \odot \tilde{C}_t $$ (Element-wise multiplication)

    • $$\displaystyle h_t = o_t \odot \tanh(C_t) $$

  • Advantages over Simple RNNs:

    • Explicitly models long-term dependencies via cell state.

    • Gates allow stable gradient flow (additive update to $$\displaystyle C_t $$).

    • Can learn over hundreds of timesteps.

Deep RNNs (Stacked Architectures)

  • Stack multiple RNN/LSTM layers on top of each other.

  • Lower layers learn low-level temporal features, higher layers learn high-level abstractions.

  • Increases model capacity and representational power.

Recursive Neural Networks (Tree Structures)

  • Architecture: Process tree-structured data (e.g., parse trees of sentences).

  • Working: Apply the same neural network (same weights) in a bottom-up manner at each node, combining child representations to form parent representation.

  • Applications: Sentiment analysis (on parse trees), Semantic parsing.


VI. Autoencoders and Generative Models

Autoencoders (AEs)

  • Architecture: Encoder ($$\displaystyle z = f(x) $$) → Bottleneck (latent code $z$) → Decoder ($$\displaystyle \hat{x} = g(z) $$). Trained to reconstruct input: minimize $L(x, \hat{x})$.

  • Training: Unsupervised (no labels needed, just input data).

  • Applications: Dimensionality reduction, denoising, anomaly detection, feature learning.

Autoencoder Variants

Variant Modification Purpose
Sparse AE Add sparsity penalty (e.g., KL divergence) to activations of hidden layer. Forces network to learn meaningful, compact features; useful for feature extraction.
Contractive AE Add penalty on the Frobenius norm of the Jacobian of encoder activations. Encourages encoder to be robust to small input perturbations → learns locally invariant features.
Denoising AE Input is corrupted version $\tilde{x}$ (e.g., added noise, masked), target is original clean $x$. Forces network to learn robust features and reconstruct from partial info. More powerful than standard AE.

Autoencoders vs PCA/SVD for Dimensionality Reduction

Feature PCA/SVD Autoencoder
Linearity Linear transformation only. Non-linear (due to activations).
Flexibility Single, fixed transformation. Can learn complex, multi-stage non-linear mappings.
Data Assumption Assumes Gaussian distribution, linear relationships. Makes no strong distributional assumptions.
When to use AE When data has non-linear structure (e.g., manifold of images). PCA fails to capture complex manifolds.

Regularization in Autoencoders

  • Purpose: Prevent autoencoder from learning trivial identity function ($$\displaystyle x \rightarrow x $$), especially when latent dimension is large.

  • Methods:

    1. Sparsity Penalty: Force few active neurons.

    2. Denoising: Corrupt input, reconstruct clean.

    3. Contractive Penalty: Penalize large Jacobian.

    4. Dropout: Applied to inputs/hidden layers.

Variational Autoencoders (VAEs)

  • Key Idea: Generative model. Learns a latent variable model $$\displaystyle p(x) = \int p(x|z)p(z)dz $$.

  • Architecture: Encoder outputs parameters of a distribution (usually Gaussian: mean $\mu$ and variance $$\displaystyle \sigma^2 $$) for latent code $z$, not a single point.

  • Reparameterization Trick: Sample $$\displaystyle z = \mu + \sigma \odot \epsilon $$, $\epsilon \sim \mathcal{N}(0,I)$. Allows gradient flow through sampling.

  • Loss: Reconstruction loss (e.g., MSE) + KL Divergence between encoder's distribution and prior (usually $\mathcal{N}(0,I)$). Balances reconstruction quality and latent space structure.

  • Output: Can generate new samples by sampling $z$ from prior and decoding.

Generative Adversarial Networks (GANs)

  • Architecture: Two networks in competition:

    • Generator ($G$): Takes random noise $$\displaystyle z \sim p_z $$, generates fake data $$\displaystyle \hat{x} = G(z) $$.

    • Discriminator ($D$): Classifies input as real ($x$) or fake ($\hat{x}$).

  • Training (Adversarial/Minimax Game):

    • $D$ maximizes $$\displaystyle \mathbb{E}_{x\sim p_{data}}[\log D(x)] + \mathbb{E}_{z\sim p_z}[\log(1-D(G(z)))] $$

    • $G$ minimizes $$\displaystyle \mathbb{E}_{z\sim p_z}[\log(1-D(G(z)))] $$

  • Result: $G$ generates increasingly realistic data; $D$ becomes a powerful classifier.

GANs vs VAEs

Aspect GAN VAE
Approach Adversarial, game-theoretic. Probabilistic, variational inference.
Output Quality Sharper, more realistic samples. Often blurry samples.
Training Stability Unstable, mode collapse, hard to evaluate. Stable, well-defined loss.
Latent Space Not necessarily continuous or meaningful. Structured, continuous, good for interpolation.
Choose GAN when Ultimate sample quality is critical (e.g., art, super-resolution). Choose VAE when you need structured latent space for manipulation, interpolation, or semi-supervised learning.

Auto-regressive Models: NADE & MADE

  • Core Idea: Factorize joint distribution $p(x)$ as product of conditionals: $$\displaystyle p(x) = \prod_{i=1}^d p(x_i | x_{<i}) $$.

  • NADE (Neural Autoregressive Distribution Estimator): Uses a neural network with masked connections to ensure each input $$\displaystyle x_i $$ only depends on previous inputs $$\displaystyle x_{1..i-1} $$. Efficient training via BPTT.

  • MADE (Masked Autoencoder for Distribution Estimation): Applies the masking idea to a feedforward autoencoder. Faster parallel training than NADE. Foundation for PixelCNN (images), WaveNet (audio).


VII. Training Deep Networks

Overfitting and Underfitting

Underfitting Overfitting
Definition Model is too simple; fails to capture underlying pattern. Model learns noise/irrelevant details in training data.
Training Error High Very Low
Validation Error High High (gap between train/val is large)
Cause High bias, low model capacity. High variance, low training data, high model capacity.

Regularization Techniques

Technique Mechanism Effect
Dropout Randomly "drop" (set to 0) a fraction p of neurons during training only. Prevents co-adaptation of neurons; acts like an ensemble of many thinned networks.
Weight Decay (L2) Add $$\displaystyle \frac{\lambda}{2} \|w\|^2 $$ to loss. Penalizes large weights, encourages smaller, more distributed weights.
Early Stopping Monitor validation loss; stop training when it starts to increase. Prevents memorization of training set noise.
Data Augmentation Artificially enlarge training set by applying transformations (rotate, crop, flip images). Increases effective dataset size, teaches invariances.

Batch Normalization

  • Mechanism: For each mini-batch, normalize activations of a layer to have zero mean and unit variance. Then scale ($\gamma$) and shift ($\beta$) with learnable parameters.

    • $$\displaystyle \hat{x}^{(k)} = \frac{x^{(k)} - \mu_B}{\sqrt{\sigma_B^2 + \epsilon}} $$

    • $$\displaystyle y^{(k)} = \gamma \hat{x}^{(k)} + \beta $$

  • Benefits:

    1. Reduces internal covariate shift (distribution changes in layer inputs during training).

    2. Allows use of higher learning rates.

    3. Acts as a regularizer (adds noise via mini-batch stats).

    4. Reduces need for careful weight initialization.

  • Usage: After convolution/linear layer, before non-linearity (ReLU).

Gradient Problems & Mitigation

Problem Cause Mitigation Strategies
Vanishing Gradients Repeated multiplication by small values (sigmoid/tanh derivatives, $$\displaystyle |W|<1 $$). Use ReLU activations, BatchNorm, Residual Connections (ResNet), LSTM/GRU for RNNs.
Exploding Gradients Repeated multiplication by large values ($$\displaystyle |W|>1 $$). Gradient Clipping (cap gradient norm), Weight Initialization (Xavier/He), BatchNorm.

Data Preprocessing: Normalization vs Standardization

Normalization (Min-Max Scaling) Standardization (Z-score)
Formula $$\displaystyle x' = \frac{x - x_{min}}{x_{max} - x_{min}} $$ $$\displaystyle x' = \frac{x - \mu}{\sigma} $$
Output Range [0, 1] (or [-1, 1]) Mean=0, Std=1 (unbounded).
Sensitivity Sensitive to outliers. Less sensitive to outliers.
Use Case When algorithm requires bounded input (e.g., neural nets with sigmoid output). Default choice for most DL algorithms (works well with zero-centered data).

VIII. Advanced Topics

Deep Belief Networks (DBNs)

  • Architecture: Stack of Restricted Boltzmann Machines (RBMs). Each RBM is a bipartite graph (visible layer, hidden layer).

  • Pre-training (Greedy Layer-Wise):

    1. Train first RBM on raw data.

    2. Use its hidden layer activations as "data" for the next RBM.

    3. Repeat for all layers.

  • Fine-tuning: After pre-training, optionally fine-tune all weights jointly with backpropagation.

  • Significance: One of the first successful methods for deep network pre-training, helping overcome vanishing gradients before ReLU/BatchNorm era.

Deep Reinforcement Learning

  • Markov Decision Process (MDP): $(S, A, P, R, \gamma)$. State $s$, Action $a$, Transition $P(s'|s,a)$, Reward $R(s,a)$, Discount $\gamma$.

  • Value Iteration (Dynamic Programming): Iteratively compute optimal state-value function $$\displaystyle V^*(s) $$.

    • $$\displaystyle V_{k+1}(s) = \max_a \sum_{s'} P(s'|s,a) [R(s,a) + \gamma V_k(s')] $$
  • Policy Iteration: Alternate between:

    1. Policy Evaluation: Compute $$\displaystyle V^\pi $$ for current policy $\pi$.

    2. Policy Improvement: $$\displaystyle \pi'(s) = \arg\max_a \sum_{s'} P(s'|s,a)[R(s,a) + \gamma V^\pi(s')] $$.

  • Q-learning (Model-Free): Learn action-value function $Q(s,a)$.

    • Update: $$\displaystyle Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma \max_{a'} Q(s',a') - Q(s,a)] $$
  • Deep Q-Network (DQN): Use deep CNN to approximate $Q(s,a;\theta)$. Key tricks: Experience Replay, Target Network.

  • Double DQN: Decouples action selection and evaluation to reduce overestimation bias.

    • $$\displaystyle y = r + \gamma Q(s', \arg\max_{a'} Q(s',a'; \theta); \theta^-) $$
  • Dueling DQN: Separately estimates state-value $V(s)$ and advantage $A(s,a)$, then combines: $$\displaystyle Q(s,a) = V(s) + A(s,a) - \frac{1}{|\mathcal{A}|}\sum_{a'} A(s,a') $$.

  • LSPI (Least Squares Policy Iteration): Uses linear function approximation (e.g., RBF features) and solves a least-squares problem for policy evaluation. More sample-efficient but less scalable than DQN.

Deep Dream

  • Goal: Generate surreal, dream-like images by maximizing activations of specific layers/neurons in a trained CNN.

  • Process:

    1. Start with noise or input image.

    2. Perform forward pass, compute loss = activation of target layer/neuron.

    3. Compute gradient of loss w.r.t. input image.

    4. Update input image in the direction of the gradient (ascent).

    5. Repeat.

  • Result: Amplifies patterns the network has learned (e.g., eyes, dog faces), creating hallucinatory images.

Model Compression: Unit Pruning

  • Goal: Reduce model size/inference time by removing unimportant neurons (units) or connections.

  • Unit Pruning: Remove entire neurons (and their incoming/outgoing connections) based on a importance score (e.g., L1 norm of weights, activation magnitude).

  • Need: Deploy models on resource-constrained devices (mobile, IoT), reduce latency, energy consumption.

  • Process: Train full model → prune units based on score → fine-tune remaining weights → iterate.

Directed Graphical Models (Basics)

  • Definition: Probabilistic graphical models where nodes represent random variables and directed edges represent conditional dependencies (causal relationships).

  • Example: Bayesian Network. Encodes joint distribution as $$\displaystyle p(x_1,...,x_n) = \prod_i p(x_i | \text{parents}(x_i)) $$.

  • Use in DL: VAEs have a directed graphical model structure: $$\displaystyle z \rightarrow x $$. Provides probabilistic interpretation.

Computational Efficiency: GPU Implementation, Randomized SVD

  • GPU: Highly parallel hardware. Ideal for matrix/tensor operations (convolutions, matrix multiplies) in DL. Requires data batch processing and careful memory management.

  • Randomized SVD: Approximates SVD of large matrix $A$ using random projections. Faster ($O(mn \log k)$) than deterministic SVD ($O(mn \min(m,n))$). Used for:

    • Fast PCA on large datasets.

    • Pre-conditioning or initializing layers (e.g., in DBNs).

    • Low-rank approximation of weight matrices for compression.


IX. Evaluation and Applications (Brief)

Evaluation Metrics

Task Type Common Metrics
Classification Accuracy, Precision, Recall, F1-Score, ROC-AUC, Confusion Matrix.
Generation Inception Score (IS), Fréchet Inception Distance (FID), BLEU (text).
Sequence Modeling Perplexity (language models), BLEU (translation).

Applications

  • Computer Vision: Image classification (ResNet), Object detection (YOLO), Segmentation (U-Net).

  • Natural Language Processing: Machine translation (Transformer), Sentiment analysis (LSTM), Question answering (BERT).

  • Recommendation Systems: Collaborative filtering with neural embeddings (NeuMF), sequential recommendations (RNNs).

Case Studies

  • Image Recognition: AlexNet (2012) demonstrated deep CNNs' superiority, reducing top-5 error from 26% to 15% on ImageNet. Key: ReLU, Dropout, GPU training.

  • Text Sequence Modeling: LSTMs dominated machine translation and language modeling before Transformers. Handled variable-length sequences, captured long-range dependencies, enabled Google's Neural Machine Translation system.

[!TIP] Final Exam Strategy: For 7-mark questions, structure your answer: 1. Clear Definition, 2. Core Mechanism/Architecture (use equations/diagrams if possible), 3. Key Advantages/Disadvantages, 4. One Concrete Example/Application. Always link back to Information Retrieval context where possible (e.g., "RNNs for query suggestion", "CNNs for document image retrieval", "Autoencoders for query/document embedding").

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in