Skip to content
IT-701 · Soft Computing/Quick Revision Short Notes

Soft Computing (IT-701) - Unit 2 Short Notes

UNIT 2: NEURAL NETWORKS & DEEP LEARNING

1. INTRODUCTION & BIOLOGICAL INSPIRATION

1.1. What is a Neural Network?

  • Definition: A computing system inspired by the biological brain, consisting of interconnected processing units (neurons) that learn patterns from data through adjustment of connection strengths (weights).

  • Historical Milestones:

    • McCulloch-Pitts Neuron (1943): First mathematical model of a neuron. Binary threshold unit.

    • Perceptron Convergence Theorem (Rosenblatt, 1950s): Proved that the perceptron learning rule converges to a solution for linearly separable problems.

1.2. Biological Neuron vs. Artificial Neuron

Feature Biological Neuron Artificial Neuron (Model)
Structure Soma (cell body), Dendrites (input), Axon (output), Synapse (connection) Inputs ($$\displaystyle x_1, x_2, ..., x_n $$), Weights ($$\displaystyle w_1, w_2, ..., w_n $$), Bias ($b$), Activation Function ($f$)
Signal Electro-chemical impulses (all-or-none action potential) Continuous or discrete numerical values
Processing Summation of inputs at soma; threshold firing Net Input: $$\displaystyle z = \sum_{i=1}^{n} w_i x_i + b $$ <br> Output: $$\displaystyle y = f(z) $$
Learning Synaptic plasticity (strength change) Weight Update: $$\displaystyle w_i^{new} = w_i^{old} + \Delta w_i $$

1.3. Key Terminology & Components

  • Weight ($w$): Strength of a connection. Positive = excitatory, Negative = inhibitory.

  • Bias ($b$): Offset term added to net input; allows shifting the decision boundary.

  • Activation Function ($f$): Introduces non-linearity. Determines if and to what extent a neuron fires.

  • Net Input ($z$): Weighted sum of inputs plus bias.

  • Network Topology:

    • Layer: Input, Hidden (1+), Output.

    • Fully Connected (Dense): Every neuron in layer $l$ connected to every neuron in layer $l+1$.

    • Feedforward: Signals flow one direction (input → output).

[!TIP] Exam Focus: Be prepared to draw and label the artificial neuron model and contrast it with the biological counterpart.


2. SINGLE-LAYER PERCEPTRON (SLP)

2.1. Architecture & Model

  • Structure: One layer of weights connecting input layer directly to output layer. No hidden layers.

  • Decision Boundary: For a single output neuron, the decision surface is a hyperplane in the input space.

    • Equation: $$\displaystyle w_1 x_1 + w_2 x_2 + ... + w_n x_n + b = 0 $$.
  • Linear Separability: A problem is solvable by an SLP if the training data can be perfectly separated by a straight line (2D), plane (3D), or hyperplane (n-D).

2.2. Learning Algorithm: Perceptron Learning Rule

  • Goal: Find weight vector $\mathbf{w}$ and bias $b$ that correctly classify all training patterns.

  • Weight Update Rule (for one sample $(\mathbf{x}, t)$):

$$\Delta w_i = \eta (t - y) x_i$$

$$\Delta b = \eta (t - y)$$

Where:

*   $\eta$ = learning rate (small positive constant, 0 < $\eta$ ≤ 1)

*   $t$ = target (true) label

*   $y$ = predicted output (usually 0 or 1 after step function)

*   $$\displaystyle x_i $$ = $i$-th input feature
  • Algorithm: Iterate over training set. For each misclassified pattern, adjust weights & bias as above. Repeat until convergence.

  • Convergence Theorem: If training data is linearly separable, the perceptron learning rule is guaranteed to find a solution in a finite number of steps.

  • Limitation (XOR Problem): Cannot learn the non-linearly separable XOR function. This limitation necessitated multi-layer networks.

2.3. Applications & Limitations

  • Applications: Simple binary classification tasks (e.g., AND, OR gates, linearly separable datasets).

  • Fundamental Limitations:

    1. Linear Decision Boundary: Cannot solve non-linear problems.

    2. No Hidden Layers: No feature learning; only works on pre-defined features.

    3. Convergence Guarantee Only for Linearly Separable Data.

[!TIP] Common Pitfall: The perceptron can learn the OR and AND functions (linearly separable) but fails on XOR. This is the classic exam question.


3. MULTI-LAYER PERCEPTRON (MLP) & BACKPROPAGATION

3.1. MLP Architecture

  • Structure: At least one input layer, one or more hidden layers, and one output layer. Fully connected between successive layers.

  • Universal Approximation Theorem: A feedforward network with a single hidden layer containing a finite number of neurons and appropriate non-linear activation functions can approximate any continuous function on compact subsets of $$\displaystyle \mathbb{R}^n $$ to any desired accuracy. This justifies the use of hidden layers.

  • Notation:

    • $l$ = layer index (0=input, L=output).

    • $$\displaystyle w_{jk}^{(l)} $$ = weight from neuron $k$ in layer $l-1$ to neuron $j$ in layer $l$.

    • $$\displaystyle b_j^{(l)} $$ = bias for neuron $j$ in layer $l$.

    • $$\displaystyle z_j^{(l)} $$ = net input to neuron $j$ in layer $l$.

    • $$\displaystyle a_j^{(l)} $$ = output/activation of neuron $j$ in layer $l$.

3.2. Activation Functions

Function Formula Range Pros Cons
Linear (Identity) $$\displaystyle f(z) = z $$ $(-\infty, \infty)$ No saturation, gradient constant. Cannot introduce non-linearity. Only for output layer in regression.
Sigmoid (Logistic) $$\displaystyle f(z) = \frac{1}{1 + e^{-z}} $$ (0, 1) Smooth, output interpretable as probability. Saturates (gradients → 0 for large |z|), not zero-centered, slow convergence.
Hyperbolic Tangent (tanh) $$\displaystyle f(z) = \frac{e^z - e^{-z}}{e^z + e^{-z}} $$ (-1, 1) Zero-centered, steeper than sigmoid. Still saturates for large |z|.
ReLU $$\displaystyle f(z) = \max(0, z) $$ $[0, \infty)$ Computationally cheap, non-saturating for $$\displaystyle z>0 $$, sparsifies network. Dying ReLU: Neurons can get stuck in inactive state for $$\displaystyle z<0 $$ (gradient=0).
Leaky ReLU $$\displaystyle f(z) = \max(\alpha z, z) $$, $\alpha$ small (e.g., 0.01) $(-\infty, \infty)$ Fixes dying ReLU (small gradient for $$\displaystyle z<0 $$). Results may be inconsistent.
ELU $$\displaystyle f(z) = \begin{cases} z & z>0 \\ \alpha(e^z - 1) & z \leq 0 \end{cases} $$ $(-\alpha, \infty)$ Smoother near zero, negative saturation helps push mean unit activations closer to zero. Computationally more expensive (exp).

3.3. Backpropagation of Error (Core Algorithm)

Goal: Compute gradient of error $E$ w.r.t. all weights & biases to perform gradient descent. Error Functions:

  • Sum of Squared Errors (SSE): $$\displaystyle E = \frac{1}{2} \sum_{j \in \text{output}} (t_j - a_j^{(L)})^2 $$

  • Cross-Entropy Loss (for classification): $$\displaystyle E = -\sum_{j} t_j \log(a_j^{(L)}) $$

Process:

  1. Forward Pass:

    • For each layer $$\displaystyle l = 1 $$ to $L$:

$$z_j^{(l)} = \sum_{k} w_{jk}^{(l)} a_k^{(l-1)} + b_j^{(l)}$$

$$a_j^{(l)} = f(z_j^{(l)})$$

*   Compute error $E$ using final outputs $$\displaystyle a_j^{(L)} $$ and targets $$\displaystyle t_j $$.
  1. Backward Pass (Compute Gradients):

    • Output Layer Error ($$\displaystyle \delta_j^{(L)} $$):

$$\delta_j^{(L)} = \frac{\partial E}{\partial z_j^{(L)}} = \frac{\partial E}{\partial a_j^{(L)}} \cdot f'(z_j^{(L)})$$

    *   For SSE & sigmoid output: $$\displaystyle \delta_j^{(L)} = (a_j^{(L)} - t_j) a_j^{(L)} (1 - a_j^{(L)}) $$

    *   For Cross-Entropy & sigmoid: $$\displaystyle \delta_j^{(L)} = a_j^{(L)} - t_j $$ (simpler!)

*   **Backpropagate Error to Previous Layers ($$\displaystyle l = L-1, L-2, ..., 1 $$):**

$$\delta_j^{(l)} = \left( \sum_{k} w_{kj}^{(l+1)} \delta_k^{(l+1)} \right) f'(z_j^{(l)})$$

    *   This is the **chain rule** in action. Error at layer $l+1$ is distributed back to layer $l$ via weights $$\displaystyle w_{kj}^{(l+1)} $$ and local gradient $$\displaystyle f'(z_j^{(l)}) $$.
  1. Weight & Bias Gradients:

$$\frac{\partial E}{\partial w_{jk}^{(l)}} = a_k^{(l-1)} \delta_j^{(l)}$$

$$\frac{\partial E}{\partial b_j^{(l)}} = \delta_j^{(l)}$$

  1. Weight Update (Gradient Descent):

$$w_{jk}^{(l)} \leftarrow w_{jk}^{(l)} - \eta \frac{\partial E}{\partial w_{jk}^{(l)}}$$

$$b_j^{(l)} \leftarrow b_j^{(l)} - \eta \frac{\partial E}{\partial b_j^{(l)}}$$

Where $\eta$ is the learning rate.

3.4. Practical Considerations in Training

  • Learning Rate ($\eta$): Critical hyperparameter. Too large → divergence; too small → slow convergence. Often use learning rate schedules (step decay, exponential decay) or adaptive methods (Adam).

  • Momentum: Adds a fraction $\gamma$ (typically 0.9) of the previous weight update to the current one to accelerate convergence and dampen oscillations.

$$v_{jk}^{(l)} \leftarrow \gamma v_{jk}^{(l)} + \eta \frac{\partial E}{\partial w_{jk}^{(l)}}$$

$$w_{jk}^{(l)} \leftarrow w_{jk}^{(l)} - v_{jk}^{(l)}$$

  • Weight Initialization:

    • Random (Small): Can lead to vanishing/exploding gradients.

    • Xavier/Glorot (for tanh/sigmoid): $$\displaystyle w \sim U\left[-\frac{\sqrt{6}}{\sqrt{n_{in}+n_{out}}}, \frac{\sqrt{6}}{\sqrt{n_{in}+n_{out}}}\right] $$

    • He Initialization (for ReLU): $$\displaystyle w \sim \mathcal{N}\left(0, \sqrt{\frac{2}{n_{in}}}\right) $$

  • Stopping Criteria:

    • Fixed number of epochs.

    • Early Stopping: Monitor validation error; stop when it starts increasing (prevents overfitting).

  • Overfitting & Underfitting (Bias-Variance Tradeoff):

    • Underfitting (High Bias): Model too simple, poor performance on train & test.

    • Overfitting (High Variance): Model too complex, excellent train performance, poor test performance.

  • Regularization Techniques:

    • L1/L2 Weight Decay: Add penalty $$\displaystyle \lambda \|\mathbf{w}\|_1 $$ or $$\displaystyle \lambda \|\mathbf{w}\|_2^2 $$ to error $E$.

    • Dropout: During training, randomly "drop" (set to 0) a fraction $p$ of neurons in a layer. Forces network to learn redundant representations.

      • Forward Pass: Apply mask $\mathbf{m} \sim \text{Bernoulli}(1-p)$. Output: $$\displaystyle a' = \mathbf{m} \odot a / (1-p) $$ (scale to maintain expectation).

      • Backward Pass: Gradient only flows through non-dropped neurons ($$\displaystyle \delta' = \mathbf{m} \odot \delta $$).

[!TIP] Exam Focus: Be able to derive the backpropagation equations for a simple 2-layer MLP. Know the difference between L1 and L2 regularization. Understand why dropout works and how it modifies the forward/backward pass.


4. SPECIALIZED NEURAL NETWORK ARCHITECTURES

4.1. Radial Basis Function (RBF) Networks

  • Structure:

    • Input Layer: Passes inputs to hidden layer.

    • Hidden Layer: Uses Radial Basis Functions (typically Gaussian) as activation. Each neuron computes distance from a center $$\displaystyle \mathbf{c}_j $$.

$$\phi_j(\mathbf{x}) = \exp\left(-\frac{\|\mathbf{x} - \mathbf{c}_j\|^2}{2\sigma_j^2}\right)$$

    Where $$\displaystyle \sigma_j $$ is the **width** (spread).

*   **Output Layer:** Linear combination of hidden layer outputs: $$\displaystyle y_k = \sum_j w_{kj} \phi_j(\mathbf{x}) $$.
  • Training (Two-Stage):

    1. Centers & Widths: Use clustering (e.g., K-Means) on training inputs to set $$\displaystyle \mathbf{c}_j $$. Set $$\displaystyle \sigma_j $$ based on distances to nearest centers.

    2. Output Weights: Since output layer is linear, solve using pseudo-inverse or linear regression: $$\displaystyle \mathbf{W} = \mathbf{\Phi}^+ \mathbf{T} $$, where $\mathbf{\Phi}$ is the matrix of RBF outputs.

  • Comparison with MLP:

    • RBF: Faster training (two-stage), hidden layer has interpretable meaning (distance to cluster centers), prone to curse of dimensionality.

    • MLP: Slower (backprop), more flexible universal approximator, handles high-dim data better.

4.2. Self-Organizing Maps (SOM) / Kohonen Networks

  • Purpose: Unsupervised learning for topology-preserving dimensionality reduction and clustering.

  • Architecture: Typically 2D grid of neurons. Each neuron has a weight vector $$\displaystyle \mathbf{w}_j $$ of same dimension as input.

  • Phases:

    1. Competition: Given input $\mathbf{x}$, find Best Matching Unit (BMU) $c$ with weight vector closest to $\mathbf{x}$ (using Euclidean distance).

    2. Cooperation: BMU's neighbors on the grid are selected for adaptation. Neighborhood function $$\displaystyle h_{cj}(t) $$ (e.g., Gaussian) decreases with distance from BMU and time.

    3. Adaptation: Update weights of BMU and its neighbors:

$$\Delta w_j = \eta(t) h_{cj}(t) (\mathbf{x} - \mathbf{w}_j)$$

  • Result: Similar inputs map to nearby neurons on the grid, preserving topological relationships. Used for visualization, clustering, feature mapping.

4.3. Convolutional Neural Networks (CNNs)

  • Motivation: Exploit spatial/temporal invariance in grid-like data (images, audio). Use parameter sharing and local connectivity to drastically reduce parameters vs. fully connected MLP.

  • Core Layers:

    • Convolutional Layer:

      • Applies filters/kernels $\mathbf{K}$ to input. Each filter detects a specific local feature (edge, texture).

      • Output Feature Map: $$\displaystyle a_{i,j}^{(l)} = f\left( \sum_m \sum_n w_{m,n}^{(l)} x_{i+m, j+n} + b^{(l)} \right) $$

      • Hyperparameters:

        • Kernel Size ($k \times k$): Receptive field size.

        • Stride ($s$): Step size of filter movement.

        • Padding ($p$): Zero-padding input to control output size. "same" padding keeps size; "valid" (no padding) reduces size.

      • Depth: Number of filters = number of feature maps in output volume.

    • Pooling (Subsampling) Layer:

      • Purpose: Downsample feature maps, introduce translation invariance, reduce spatial size/computation.

      • Operation: Non-overlapping windows (e.g., $2 \times 2$). Common types:

        • Max Pooling: Output = max value in window.

        • Average Pooling: Output = average value in window.

      • Hyperparameters: Pool size, stride (usually = pool size).

    • Fully Connected (FC) Layer: At the end of network, flattens final feature maps and applies MLP for classification/regression.

  • Classic Architectures:

    • LeNet-5 (1998): First successful CNN for digit recognition. Conv → Pool → Conv → Pool → FC → FC.

    • AlexNet (2012): Deeper (5 conv layers), used ReLU, Dropout, GPU training. Won ImageNet.

    • VGGNet (2014): Demonstrated depth matters. Used only $3 \times 3$ conv filters stacked, increasing depth (16/19 layers). Simple, uniform architecture.

4.4. Recurrent Neural Networks (RNNs)

  • Motivation: Process sequential data (time series, text, speech) where current output depends on previous inputs/hidden state.

  • Architecture: Has a hidden state $$\displaystyle \mathbf{h}_t $$ that acts as memory. At each time step $t$:

$$\mathbf{h}_t = f_h(\mathbf{W}_{xh} \mathbf{x}_t + \mathbf{W}_{hh} \mathbf{h}_{t-1} + \mathbf{b}_h)$$

$$\mathbf{y}_t = f_y(\mathbf{W}_{hy} \mathbf{h}_t + \mathbf{b}_y)$$

*   Parameters are **shared across time steps**.
  • Standard RNN Problems:

    • Vanishing Gradients: Gradients shrink exponentially during backprop through time (BPTT) for long sequences, preventing learning of long-range dependencies.

    • Exploding Gradients: Gradients grow exponentially (can be clipped).

  • Long Short-Term Memory (LSTM):

    • Core Idea: Introduce a cell state $$\displaystyle \mathbf{C}_t $$ that runs through the entire sequence with minimal interference, acting as a "conveyor belt".

    • Gates (Sigmoid activations + pointwise multiplication):

      1. Forget Gate ($$\displaystyle f_t $$): Decides what information to discard from $$\displaystyle \mathbf{C}_{t-1} $$.

        $$\displaystyle f_t = \sigma(\mathbf{W}_f [\mathbf{h}_{t-1}, \mathbf{x}_t] + \mathbf{b}_f) $$

      2. Input Gate ($$\displaystyle i_t $$) & Candidate Cell State ($$\displaystyle \tilde{\mathbf{C}}_t $$): Decides what new information to store.

        $$\displaystyle i_t = \sigma(\mathbf{W}_i [\mathbf{h}_{t-1}, \mathbf{x}_t] + \mathbf{b}_i) $$

        $$\displaystyle \tilde{\mathbf{C}}_t = \tanh(\mathbf{W}_C [\mathbf{h}_{t-1}, \mathbf{x}_t] + \mathbf{b}_C) $$

      3. Update Cell State:

        $$\displaystyle \mathbf{C}_t = f_t \odot \mathbf{C}_{t-1} + i_t \odot \tilde{\mathbf{C}}_t $$

      4. Output Gate ($$\displaystyle o_t $$) & Hidden State:

        $$\displaystyle o_t = \sigma(\mathbf{W}_o [\mathbf{h}_{t-1}, \mathbf{x}_t] + \mathbf{b}_o) $$

        $$\displaystyle \mathbf{h}_t = o_t \odot \tanh(\mathbf{C}_t) $$

    • Mechanism: Additive nature of cell state update ($$\displaystyle \mathbf{C}_t = ... + ... $$) allows gradients to flow relatively unchanged, mitigating vanishing gradient problem.

  • Gated Recurrent Unit (GRU):

    • Simplified LSTM with two gates:

      1. Update Gate ($$\displaystyle z_t $$): Combines forget & input gate. Decides how much past information to keep.

      2. Reset Gate ($$\displaystyle r_t $$): Controls how much past information to forget when computing new candidate.

    • No separate cell state; hidden state $$\displaystyle \mathbf{h}_t $$ is the only state. Often faster to train, performance comparable to LSTM.

[!TIP] Diagram Needed: Draw the LSTM cell showing the cell state $$\displaystyle \mathbf{C}_t $$, hidden state $$\displaystyle \mathbf{h}_t $$, and the three gates (forget, input, output) with their equations.


5. ADVANCED DEEP LEARNING CONCEPTS (Modern Context)

5.1. Deep Neural Networks (DNNs)

  • Definition: MLPs with many hidden layers (e.g., >5). Depth allows learning hierarchical feature representations.

  • Challenges:

    • Vanishing/Exploding Gradients: More severe in deep, narrow networks.

    • Solutions: Careful weight initialization (He/Xavier), non-saturating activations (ReLU), batch normalization, skip connections (ResNet), gradient clipping.

5.2. Attention Mechanism & Transformers

  • Motivation: RNNs/CNNs struggle with long-range dependencies and are sequential (slow). Attention allows model to focus on relevant parts of input sequence when producing each output.

  • Core Idea: Query, Key, Value (Q, K, V):

    • For each element in sequence, compute:

      • Query ($\mathbf{Q}$): What am I looking for?

      • Key ($\mathbf{K}$): What do I contain?

      • Value ($\mathbf{V}$): What do I return if matched?

    • Attention Score: Compute similarity between Query and all Keys (scaled dot-product):

$$\text{Attention}(\mathbf{Q}, \mathbf{K}, \mathbf{V}) = \text{softmax}\left(\frac{\mathbf{Q}\mathbf{K}^T}{\sqrt{d_k}}\right)\mathbf{V}$$

*   Output is weighted sum of Values, weights from softmax.
  • Self-Attention: Q, K, V all come from the same sequence. Allows each position to attend to all others.

  • Transformer Architecture (Vaswani et al., 2017):

    • Encoder: Stack of identical layers. Each layer has:

      1. Multi-Head Self-Attention: Multiple attention heads learn different relationship subspaces.

      2. Feed-Forward Network (MLP): Applied position-wise.

      3. Residual Connections & Layer Normalization around each sub-layer.

    • Decoder: Similar stack, but with:

      1. Masked Multi-Head Self-Attention: Prevents attending to future positions.

      2. Multi-Head Encoder-Decoder Attention: Queries from decoder, Keys/Values from encoder.

      3. Feed-Forward + Norm/Residuals.

    • Positional Encoding: Since no recurrence/convolutions, add fixed or learned positional encodings to input embeddings to give sequence order information.

  • Significance: Foundation for BERT (encoder-only), GPT (decoder-only), T5 (encoder-decoder). Dominates NLP and is expanding to vision (ViT), audio, etc.

5.3. Transfer Learning & Fine-Tuning

  • Concept: Leverage knowledge from a pre-trained model (trained on large dataset like ImageNet) for a new, related task with limited data.

  • Strategies:

    1. Feature Extraction: Freeze all pre-trained layers. Use them as a fixed feature extractor. Only train a new classifier (FC layer) on top.

    2. Fine-Tuning: Unfreeze some/all of the pre-trained layers and continue training (with a lower learning rate) on the new dataset. Allows adaptation of generic features to the specific task.

  • Why it works: Early layers learn generic features (edges, textures) that are useful across many vision tasks; later layers learn more task-specific features.


6. TRAINING, OPTIMIZATION & EVALUATION

6.1. Optimization Algorithms (Beyond SGD)

Algorithm Key Idea Update Rule (simplified)
SGD $$\displaystyle \Delta w = -\eta \nabla E $$ $$\displaystyle w \leftarrow w - \eta \nabla E $$
SGD with Momentum Accumulate velocity to dampen oscillations. $$\displaystyle v \leftarrow \gamma v + \eta \nabla E $$ <br> $$\displaystyle w \leftarrow w - v $$
AdaGrad Adapt $\eta$ per parameter based on historical sum of squared gradients. $$\displaystyle g \leftarrow g + (\nabla E)^2 $$ <br> $$\displaystyle w \leftarrow w - \frac{\eta}{\sqrt{g + \epsilon}} \nabla E $$ <br> (Problem: $g$ accumulates, $\eta$ decays to 0)
RMSProp Fix AdaGrad's aggressive decay using exponential moving average of squared grads. $$\displaystyle g \leftarrow \beta g + (1-\beta) (\nabla E)^2 $$ <br> $$\displaystyle w \leftarrow w - \frac{\eta}{\sqrt{g + \epsilon}} \nabla E $$
Adam (Most Popular) Combines Momentum (1st moment) & RMSProp (2nd moment) with bias correction. $$\displaystyle m \leftarrow \beta_1 m + (1-\beta_1) \nabla E $$ <br> $$\displaystyle v \leftarrow \beta_2 v + (1-\beta_2) (\nabla E)^2 $$ <br> $$\displaystyle \hat{m} = m/(1-\beta_1^t) $$, $$\displaystyle \hat{v} = v/(1-\beta_2^t) $$ <br> $$\displaystyle w \leftarrow w - \frac{\eta}{\sqrt{\hat{v}} + \epsilon} \hat{m} $$ <br> Typical: $$\displaystyle \beta_1=0.9, \beta_2=0.999, \epsilon=10^{-8} $$

6.2. Batch Normalization

  • Purpose: Reduce internal covariate shift (distribution of layer inputs changes during training), stabilize training, allow higher learning rates, act as regularizer.

  • Operation (during training, per mini-batch):

    1. Compute batch mean $$\displaystyle \mu_B = \frac{1}{m} \sum_{i=1}^m x_i $$ and variance $$\displaystyle \sigma_B^2 = \frac{1}{m} \sum_{i=1}^m (x_i - \mu_B)^2 $$.

    2. Normalize: $$\displaystyle \hat{x}_i = \frac{x_i - \mu_B}{\sqrt{\sigma_B^2 + \epsilon}} $$.

    3. Scale and Shift (learnable parameters): $$\displaystyle y_i = \gamma \hat{x}_i + \beta $$.

  • During Inference: Use exponentially averaged (running) mean and variance from training. No gradient computation.

6.3. Evaluation Metrics for Classification

  • Confusion Matrix:

    | | Predicted + | Predicted - | | :--- | :--- | :--- | | Actual + | TP (True Pos) | FN (False Neg) | | Actual - | FP (False Pos) | TN (True Neg) |

  • Key Metrics:

    • Accuracy: $$\displaystyle \frac{TP + TN}{TP + TN + FP + FN} $$ (Good for balanced classes).

    • Precision (Positive Predictive Value): $$\displaystyle \frac{TP}{TP + FP} $$ (How many selected are relevant?).

    • Recall (Sensitivity, True Positive Rate): $$\displaystyle \frac{TP}{TP + FN} $$ (How many relevant are selected?).

    • F1-Score: Harmonic mean: $$\displaystyle F1 = 2 \cdot \frac{\text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}} $$.

    • ROC-AUC: Area under Receiver Operating Characteristic curve (TPR vs FPR). Threshold-invariant measure of separability.


7. APPLICATIONS & CASE STUDIES

7.1. Computer Vision

  • Image Classification: CNNs (ResNet, EfficientNet). Output: single label.

  • Object Detection: Localize & classify multiple objects.

    • Two-stage: R-CNN family (Faster R-CNN). Propose regions → classify.

    • One-stage: YOLO (You Only Look Once), SSD. Predict bounding boxes & classes in single forward pass (faster).

  • Semantic Segmentation: Classify every pixel.

    • U-Net: Encoder-decoder with skip connections (preserves spatial details). Dominant in medical imaging.

7.2. Natural Language Processing (NLP)

  • Sentiment Analysis: Classify text as positive/negative. Use RNNs/LSTMs or Transformers (BERT fine-tuning).

  • Machine Translation: Sequence-to-sequence. LSTM/GRU encoder-decoder with attention, now dominated by Transformer (e.g., Google's Transformer, T5).

  • Named Entity Recognition (NER): Tag tokens with entity types (person, location). Use BiLSTM-CRF or Transformer models.

7.3. Time Series Forecasting

  • Univariate/Multivariate Forecasting: Use LSTMs/GRUs to model temporal dependencies. Can also use Temporal Convolutional Networks (TCNs) or Transformers (Informer, Autoformer).

7.4. Other Domains

  • Recommender Systems: Neural Collaborative Filtering (NCF), using MLPs to learn user-item interactions.

  • Generative Models (Intro):

    • Autoencoders (AE): Learn efficient data encoding (bottleneck). Used for denoising, anomaly detection.

    • Generative Adversarial Networks (GANs): Two networks (Generator $G$, Discriminator $D$) in minimax game. $G$ generates fake data, $D$ tries to distinguish real vs fake. Used for image generation, style transfer.

[!TIP] Exam Focus: Know one key application for each major architecture (CNN for CV, LSTM for sequences, Transformer for NLP). Be able to name one representative model from each category (e.g., VGG for CNN, LSTM for RNN, BERT for Transformer).


Final Note to Student: This document distills the core of Unit 2 based on global standards. Prioritize understanding the derivations (backpropagation, LSTM gates) and comparisons (activation functions, architectures, optimizers). Always cross-check with your RGPV lecture notes and prescribed textbooks (Haykin, Goodfellow). Good luck!

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in