Skip to content
AL-405 ยท Machine Learning/Quick Revision Short Notes

Machine Learning (AL-405) - Unit 1 Short Notes

UNIT 1: Foundations & Core Concepts

I. Foundations of Machine Learning

Definition & Importance

Machine Learning (ML) is a subset of AI that enables systems to learn patterns from data without explicit programming.

Importance: Automation (e.g., self-driving cars), prediction (e.g., stock trends), personalization (e.g., recommender systems), handling complex, high-dimensional data.

Types of Machine Learning

Type Description Examples
Supervised learns from labeled data (input-output pairs) Classification (spam detection), Regression (price prediction)
Unsupervised finds hidden patterns in unlabeled data Clustering (customer segmentation), Dimensionality reduction (PCA)
Reinforcement agent learns via rewards/penalties from environment Game playing (AlphaGo), Robotics control
Semi-supervised mix of labeled & unlabeled data Web page classification (few labeled docs)
One-shot learns from very few examples (often 1) Face recognition (new person from single photo)

AI vs. ML vs. Deep Learning

  • AI: Broad fieldโ€”creating intelligent systems.

  • ML: Subset of AIโ€”algorithms that learn from data.

  • Deep Learning: Subset of MLโ€”uses deep neural networks with many layers.

Perspectives & Issues

  • Design Issues: Choosing hypothesis space, inductive bias (algorithm's assumptions), avoiding overfitting.

  • Limitations: Data quality dependency, model interpretability ("black box"), computational scalability.

  • Performance Factors: Data quality/quantity, algorithm choice, hyperparameter tuning.

  • Statistical Theory: Probability & statistics underpin generalization bounds, confidence intervals, hypothesis testing.

Machine Learning Workflow

  1. Problem definition & data collection

  2. Data preprocessing & feature engineering

  3. Model selection (choose algorithm & hypothesis space)

  4. Training on training set

  5. Validation (tune hyperparameters via validation set)

  6. Testing on unseen test set

  7. Deployment & monitoring


II. Data Preprocessing and Feature Engineering

Need for Preprocessing

Ensures model convergence, stability, and performance. Raw data often contains noise, missing values, and scale inconsistencies.

Common Techniques

Technique Purpose Example
Normalization (Min-Max) Scales features to [0,1] $$\displaystyle X_{\text{norm}} = \frac{X - X_{\min}}{X_{\max} - X_{\min}} $$
Standardization (Z-score) Centers to mean 0, std 1 $$\displaystyle X_{\text{std}} = \frac{X - \mu}{\sigma} $$
One-hot Encoding Converts categorical to binary vectors (increases dimensionality) "Color" โ†’ [Red:1,0,0], [Green:0,1,0], [Blue:0,0,1]
Label Encoding Assigns integer to categories (may imply false ordering) "Size" โ†’ Small:0, Medium:1, Large:2
Handling Missing Data Imputation (mean/median), deletion, or model-based Replace missing age with median age

High-Dimensional Data Challenges

  • Curse of Dimensionality: Data becomes sparse; distance metrics lose meaning; computational cost explodes.

  • Impact: Overfitting risk increases, model performance degrades ("dimensionality curse").

Dimensionality Reduction

Method Type Key Idea
PCA Unsupervised Projects data onto orthogonal axes of max variance
LDA Supervised Maximizes class separability (uses labels)
PLS Supervised Finds components that maximize covariance with target
Feature Selection โ€” Forward selection (add features), backward elimination (remove features)

Data Augmentation

Artificially increases training data by applying transformations (rotation, flipping, cropping) to improve generalizationโ€”critical in computer vision.


III. Supervised Learning: Fundamentals

Regression vs. Classification

Aspect Regression Classification
Output Continuous value (e.g., price) Discrete class label (e.g., spam/ham)
Example Predicting house price Image recognition (cat/dog)

Hypothesis Space & Inductive Bias

  • Hypothesis Space: Set of all possible models the algorithm can consider (e.g., all linear functions).

  • Inductive Bias: Algorithm's assumptions to generalize (e.g., "nearest neighbors are similar"). Guides learning from examples to unseen cases.

Overfitting vs. Underfitting

Overfitting Underfitting
Cause Model too complex (high variance) Model too simple (high bias)
Symptoms Low training error, high validation error High training & validation error
Solutions Regularization, cross-validation, early stopping, more data Increase model complexity, feature engineering

Cross-Validation

  • Purpose: Estimate model performance robustly, tune hyperparameters.

  • k-fold: Split data into k subsets; train on k-1, validate on 1; repeat k times.

  • Leave-One-Out (LOO): k = N (extreme case); high variance, computationally expensive.

Evaluation Metrics

  • Regression:

    • MSE (Mean Squared Error): $$\displaystyle \text{MSE} = \frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2 $$

    • MAE (Mean Absolute Error): $$\displaystyle \text{MAE} = \frac{1}{n}\sum_{i=1}^{n}\|y_i - \hat{y}_i\| $$

    • Rยฒ (Coefficient of Determination): $$\displaystyle R^2 = 1 - \frac{\sum(y_i - \hat{y}_i)^2}{\sum(y_i - \bar{y})^2} $$

  • Classification:

    • Confusion Matrix:

      | | Predicted + | Predicted - | |---|---|---| | Actual + | TP | FN | | Actual - | FP | TN |

    • Derived Metrics:

      • Accuracy: $$\displaystyle \frac{TP+TN}{TP+TN+FP+FN} $$

      • Precision: $$\displaystyle \frac{TP}{TP+FP} $$ (positive predictive value)

      • Recall (TPR): $$\displaystyle \frac{TP}{TP+FN} $$ (sensitivity)

      • F1-score: $$\displaystyle 2 \times \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}} $$

      • ROC-AUC: Area under ROC curve (TPR vs FPR trade-off).


IV. Supervised Learning: Algorithms

A. Linear Models

Linear Regression

  • Assumptions: Linearity, independence, homoscedasticity (constant variance), normality of residuals.

  • Cost Function: Minimize MSE โ†’ closed-form solution: $$\displaystyle \hat{\beta} = (X^TX)^{-1}X^Ty $$ (or gradient descent).

  • Interpretation: $$\displaystyle \hat{y} = \beta_0 + \beta_1 x_1 + ... + \beta_p x_p $$.

Logistic Regression

  • Sigmoid Function: $$\displaystyle \sigma(z) = \frac{1}{1+e^{-z}} $$, maps $z \in \mathbb{R}$ to $(0,1)$.

  • Use: Binary classification; predicts $$\displaystyle P(y=1|x) $$.

  • Difference from Linear Regression: Uses log-loss (cross-entropy) cost; output is probability.

Locally Weighted Linear Regression

  • Non-parametric: Fits a separate linear model for each query point $$\displaystyle x_q $$, weighting training points by distance to $$\displaystyle x_q $$.

  • Weight: $$\displaystyle w_i = \exp\left(-\frac{(x_i - x_q)^2}{2\tau^2}\right) $$ ($\tau$: bandwidth).

B. Decision Trees
  • Splitting Criteria:

    • Entropy: $$\displaystyle H(S) = -\sum p_i \log_2 p_i $$ (impurity measure).

    • Gini Impurity: $$\displaystyle G(S) = 1 - \sum p_i^2 $$.

    • Information Gain: $$\displaystyle \text{IG}(S,A) = H(S) - \sum \frac{|S_v|}{|S|} H(S_v) $$.

  • Pruning: Remove branches to reduce overfitting (pre-pruning vs. post-pruning).

  • Issues: Instability (small data changes โ†’ different tree), bias-variance tradeoff (deep trees low bias, high variance).

C. Support Vector Machines (SVM)
  • Goal: Find optimal hyperplane that maximizes margin (distance to nearest points of each class).

  • Support Vectors: Training points lying on margin boundaries; define the decision boundary.

  • Linear SVM: For linearly separable data, solve constrained optimization: maximize margin subject to $$\displaystyle y_i(w^Tx_i + b) \geq 1 $$.

  • Kernel Trick: Map data to high-dimensional space via kernel (e.g., RBF: $$\displaystyle K(x_i,x_j) = \exp(-\gamma\|x_i-x_j\|^2) $$) to handle non-linear separation.

  • High-Dimensional Performance: Effective when features >> samples (e.g., bioinformatics, text classification) due to margin maximization.

D. Instance-Based Learning: k-NN
  • Algorithm: For a query point, find k nearest training points (by distance metric), predict majority class (classification) or average (regression).

  • Distance Metrics: Euclidean ($$\displaystyle \sqrt{\sum (x_i-y_i)^2} $$), Manhattan ($$\displaystyle \sum |x_i-y_i| $$).

  • Supervised: Uses labeled training data directly (lazy learning).

  • Example: Predict if player drafted given speed=6.75, agility=3 using k=3 (compute distances to all labeled players, pick 3 nearest, majority vote).

E. Ensemble Methods
Method Mechanism Bias/Variance Effect
Bagging (Bootstrap Aggregating) Train multiple models on bootstrapped samples, average/vote Reduces variance (e.g., Random Forest)
Boosting Sequentially train models, focus on misclassified points Reduces bias (e.g., AdaBoost, Gradient Boosting)
Stacking Train meta-model on outputs of base models Combines strengths, can reduce both

Bagging vs. Boosting

  • Bagging: Parallel, reduces overfitting (variance), models independent.

  • Boosting: Sequential, reduces underfitting (bias), models dependent (correct errors).


V. Unsupervised Learning

Goals & Requirements

  • Goal: Discover hidden structure (clusters, low-dim representation) without labels.

  • Requirements: Scalability, ability to handle noise, minimal prior knowledge (e.g., number of clusters).

Clustering Algorithms

Algorithm Key Steps Limitations
K-means 1. Initialize k centroids. 2. Assign points to nearest centroid. 3. Update centroids. Repeat until convergence. Sensitive to outliers, initial centroids; assumes spherical clusters.
Hierarchical (DIANA) Divisive: Start with all points in one cluster; recursively split. Computationally expensive, irreversible splits.
EM (GMM) E-step: Estimate cluster responsibilities given current params. M-step: Update Gaussian parameters (mean, covariance, mixing coeff). Can converge to local optimum; sensitive to initialization.
BIRCH Builds Clustering Feature (CF) Tree; incrementally clusters large datasets. Assumes clusters are spherical/convex; not for arbitrary shapes.

VI. Neural Networks & Deep Learning

A. Basics
  • ANN Structure: Input layer โ†’ hidden layers (with activation functions) โ†’ output layer. Connections have weights.

  • Biological Inspiration: Neurons (nodes) with dendrites (inputs) and axons (outputs); parallel processing.

  • MLP (Multi-Layer Perceptron): Feedforward network with โ‰ฅ1 hidden layer; universal approximator.

Perceptron Learning Algorithm

  1. Initialize weights $w$ randomly.

  2. For each training sample $$\displaystyle (x_i, y_i) $$:

    • Compute output: $$\displaystyle \hat{y} = f(w^T x_i) $$ (step/activation).

    • Update: $$\displaystyle w \leftarrow w + \eta (y_i - \hat{y}) x_i $$.

  3. Repeat until convergence.

B. Training Neural Networks

Backpropagation

  1. Forward pass: compute output & loss $L$.

  2. Backward pass: compute gradients $$\displaystyle \frac{\partial L}{\partial w} $$ via chain rule.

  3. Update weights: $$\displaystyle w \leftarrow w - \eta \frac{\partial L}{\partial w} $$.

  • Characteristics: Gradient-based, supervised, computes error derivatives layer-by-layer.

Gradient Descent Optimizers

Optimizer Mechanism Use Case
SGD Updates using single sample gradient Simple, noisy convergence
Mini-batch GD Updates using small batch (common) Balance speed/stability
Adam Adaptive learning rate, momentum, bias correction Default choice, fast convergence
RMSprop Adapts learning rate per parameter RNNs, non-stationary objectives

Loss Functions

  • MSE: For regression.

  • Cross-Entropy: For classification (binary: $-[y\log\hat{y}+(1-y)\log(1-\hat{y})]$; multi-class: $$\displaystyle -\sum y_i\log\hat{y}_i $$).

  • Role: Guides backpropagation by providing gradient signal to minimize prediction error.

C. Key Components

Activation Functions

Function Formula Pros Cons
Sigmoid $$\displaystyle \sigma(x) = \frac{1}{1+e^{-x}} $$ Output (0,1), smooth Vanishing gradient, not zero-centered
tanh $$\displaystyle \tanh(x) = \frac{e^x-e^{-x}}{e^x+e^{-x}} $$ Output (-1,1), zero-centered Vanishing gradient
ReLU $$\displaystyle \text{ReLU}(x) = \max(0,x) $$ Computationally cheap,็ผ“่งฃ vanishing gradient Dying ReLU (negative outputs)
Softmax $$\displaystyle \sigma(z)_j = \frac{e^{z_j}}{\sum_{k=1}^K e^{z_k}} $$ Multi-class probability output โ€”

Regularization Techniques

  • L1 (Lasso): Adds $$\displaystyle \lambda \sum |w| $$ to loss โ†’ sparsity (some weights become 0).

  • L2 (Ridge): Adds $$\displaystyle \lambda \sum w^2 $$ โ†’ weight decay (weights shrink uniformly).

  • Dropout: Randomly deactivate neurons during training (prevents co-adaptation).

  • Batch Normalization: Normalizes layer inputs to zero mean/unit variance โ†’ stabilizes training, reduces internal covariate shift.

Overfitting/Underfitting in NN

  • Detection: Training loss << validation loss (overfitting); both high (underfitting).

  • Solutions: Dropout, data augmentation, L1/L2, early stopping, reduce network size.

D. Convolutional Neural Networks (CNN)

Architecture: [Input] โ†’ [Conv โ†’ Activation โ†’ Pooling]ร—n โ†’ [Flatten] โ†’ [FC] โ†’ [Output]

  • Convolution Operation: Filter (kernel) slides over input, computes dot product โ†’ feature map.

    • Parameters: filter size (e.g., 3ร—3), stride, number of filters.
  • Padding:

    • Same: Output size = input size (pad with zeros); preserves spatial info.

    • Valid: No padding; output shrinks.

  • Pooling (Subsampling):

    • Max Pooling: Takes max in window โ†’ translation invariance, reduces dimensions.

    • Average Pooling: Takes average โ†’ smoother downsampling.

  • 1ร—1 Convolution:

    • Purpose: Feature reduction (change depth), model efficiency ("network-in-network"), add non-linearity without spatial change.
  • Flattening: Converts multi-dimensional feature maps to 1D vector for fully connected layers.

  • Inception Module:

    • Parallel convolutions with multiple filter sizes (1ร—1, 3ร—3, 5ร—5) + max pooling.

    • Concatenate outputs โ†’ captures multi-scale features efficiently.

  • Transfer Learning:

    • Feature Extraction: Freeze pre-trained CNN layers, train new classifier on top.

    • Fine-tuning: Unfreeze some top layers, train with low learning rate.

  • CNN in TensorFlow (brief):

    
    model = Sequential([
    
        Conv2D(32, (3,3), activation='relu', input_shape=(224,224,3)),
    
        MaxPooling2D(2,2),
    
        Flatten(),
    
        Dense(128, activation='relu'),
    
        Dense(10, activation='softmax')
    
    ])
    
    
E. Recurrent Neural Networks (RNN)
  • Architecture: Recurrent connections; hidden state $$\displaystyle h_t $$ depends on $$\displaystyle h_{t-1} $$ and input $$\displaystyle x_t $$.

    • Unfolding in Time: Expands recurrent connections to a deep network across time steps.
  • Types:

    • Vanilla RNN: $$\displaystyle h_t = \tanh(W_{hh}h_{t-1} + W_{xh}x_t + b) $$; suffers from vanishing/exploding gradients.

    • LSTM (Long Short-Term Memory):

      • Cells: Memory cell $$\displaystyle c_t $$ with gates:

        • Forget gate: $$\displaystyle f_t = \sigma(W_f \cdot [h_{t-1}, x_t] + b_f) $$ โ†’ what to discard from cell.

        • Input gate: $$\displaystyle i_t = \sigma(W_i \cdot [h_{t-1}, x_t] + b_i) $$; $$\displaystyle \tilde{c}_t = \tanh(W_c \cdot [h_{t-1}, x_t] + b_c) $$ โ†’ new candidate.

        • Output gate: $$\displaystyle o_t = \sigma(W_o \cdot [h_{t-1}, x_t] + b_o) $$; $$\displaystyle h_t = o_t * \tanh(c_t) $$.

      • Handles long-term dependencies via cell state.

    • GRU (Gated Recurrent Unit):

      • Merges forget & input gates into update gate $$\displaystyle z_t $$; has reset gate $$\displaystyle r_t $$.

      • Fewer parameters than LSTM, often similar performance.

  • LSTM vs. GRU: LSTM has separate cell state & hidden state; GRU merges them. LSTM more expressive, GRU faster.

  • Applications in NLP: Language modeling, machine translation, text generation (e.g., ChatGPT uses transformer, but LSTMs were foundational).

F. Autoencoders
  • Architecture: [Input] โ†’ [Encoder (bottleneck)] โ†’ [Decoder] โ†’ [Reconstructed Input].

  • Unsupervised Use: Learns efficient data encoding (dimensionality reduction), denoising (train with corrupted input), feature learning.

  • Bottleneck: Forces network to learn compressed representation.

G. Advanced Topics
  • Attention Models: Allows decoder to "focus" on relevant parts of input sequence (e.g., in translation). Key in transformers.

  • Self-Supervised Learning:

    • Pretext Tasks: Solve auxiliary task on unlabeled data (e.g., predict missing image patch, next sentence).

    • Contrastive Learning: Learn representations by pulling similar samples together, pushing dissimilar apart (e.g., SimCLR).


VII. Reinforcement Learning

A. Fundamentals
  • Difference from Supervised/Unsupervised:

    • Supervised: Fixed labeled dataset.

    • Unsupervised: Find structure in unlabeled data.

    • RL: Agent interacts with environment, learns from reward signal (no explicit labels).

  • Markov Decision Process (MDP):

    • Components: States $S$, Actions $A$, Transition probabilities $P(s'|s,a)$, Rewards $R(s,a,s')$, Policy $\pi(a|s)$.

    • Value Function: $$\displaystyle V^\pi(s) = \mathbb{E}[\sum \gamma^t R_t | s_0=s, \pi] $$ (expected return under policy $\pi$).

    • Policy Iteration: Alternate between policy evaluation (compute $$\displaystyle V^\pi $$) and policy improvement (greedy update).

    • Value Iteration: Directly update $$\displaystyle V(s) \leftarrow \max_a \sum_{s'} P(s'|s,a)[R(s,a,s') + \gamma V(s')] $$ until convergence.

  • Exploration vs. Exploitation:

    • Dilemma: Explore (try new actions to discover better rewards) vs. Exploit (choose known best action).

    • Strategies: $\epsilon$-greedy, softmax, Upper Confidence Bound (UCB).

B. Core Algorithms

Q-Learning (Off-policy TD control)

  • Q-value: $Q(s,a)$ = expected future reward taking action $a$ in state $s$.

  • Update Rule: $$\displaystyle Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma \max_{a'} Q(s',a') - Q(s,a)] $$.

  • Algorithm (deterministic rewards/actions):

    1. Initialize $Q(s,a)$ arbitrarily.

    2. For each episode:

      • Initialize state $s$.

      • While $s$ not terminal:

        • Choose $a$ from $s$ using policy derived from $Q$ (e.g., $\epsilon$-greedy).

        • Take $a$, observe $r, s'$.

        • $$\displaystyle Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma \max_{a'} Q(s',a') - Q(s,a)] $$.

        • $$\displaystyle s \leftarrow s' $$.

  • Guides Actions: Agent selects action with highest $Q$-value (greedy) after learning.

SARSA (On-policy TD control)

  • Update Rule: $$\displaystyle Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma Q(s',a') - Q(s,a)] $$.

  • Difference from Q-learning: Uses actual next action $a'$ (from same policy) vs. $$\displaystyle \max_{a'} Q(s',a') $$ (off-policy). More conservative, follows current policy.

Value Iteration vs. Policy Iteration

Value Iteration Policy Iteration
Process Directly update $V(s)$; no explicit policy evaluation Alternate: evaluate $$\displaystyle V^\pi $$, then improve $\pi$
Convergence Often faster per iteration, but each step computationally heavier Slower per iteration (requires full policy evaluation), but stable
Policy Implicit (greedy w.r.t $V$) Explicit
C. Advanced Methods
  • Actor-Critic:

    • Actor: Policy network $\pi(a|s)$ (selects actions).

    • Critic: Value network $V(s)$ or $Q(s,a)$ (evaluates actions).

    • Interaction: Actor proposes action; Critic estimates value; Actor updated via policy gradient using Critic's feedback.

  • Advanced Actor-Critic Models:

    • A2C (Advantage Actor-Critic): Uses advantage function $$\displaystyle A(s,a) = Q(s,a)-V(s) $$.

    • A3C (Asynchronous A2C): Parallel actors with global network.

    • PPO (Proximal Policy Optimization): Constrains policy updates to avoid large destructive steps.

  • RL Frameworks:

    • OpenAI Gym: Standardized environments (Atari, robotics).

    • TensorFlow Agents: Library for RL in TensorFlow.


VIII. Probabilistic & Bayesian Methods

Bayes' Theorem

$$P(A|B) = \frac{P(B|A) P(A)}{P(B)}$$

  • Interpretation: Updates belief about hypothesis $A$ given evidence $B$.

  • Example: Medical test: $$\displaystyle P(\text{disease}|\text{positive}) = \frac{P(\text{positive}|\text{disease})P(\text{disease})}{P(\text{positive})} $$.

  • Fundamental to Probabilistic Models: Forms basis for Naive Bayes, Bayesian networks, Bayesian inference.

Bayesian Networks

  • Structure: Directed acyclic graph (DAG); nodes = random variables, edges = conditional dependencies.

  • Conditional Independence: Node independent of non-descendants given parents.

  • Probability Tables: Each node has Conditional Probability Table (CPT) $P(X|\text{Parents}(X))$.

  • Inference: Compute posterior probabilities given evidence (e.g., $$\displaystyle P(\text{car value}|\text{mileage=Lo, engine=Bad}) $$ via variable elimination or sampling).

Bayesian Learning

  • Definition: Treat parameters as random variables; use prior knowledge $P(\theta)$ + likelihood $P(D|\theta)$ to compute posterior $P(\theta|D) \propto P(D|\theta)P(\theta)$.

  • Impact: Incorporates prior beliefs, quantifies uncertainty, avoids overfitting (regularization via prior).


IX. Applications & Advanced Topics

Natural Language Processing (NLP)

  • Pipeline: Tokenization โ†’ Stemming/Lemmatization โ†’ Stop-word removal โ†’ Vectorization (Bag-of-Words, TF-IDF, embeddings).

  • BLEU Score (for machine translation):

    • n-gram precision: Fraction of n-grams in candidate that appear in reference.

    • Brevity Penalty: Penalizes short candidates: $$\displaystyle BP = \min(1, \exp(1 - \frac{\text{ref len}}{\text{cand len}})) $$.

    • BLEU: $$\displaystyle BLEU = BP \cdot \exp(\sum_{n=1}^N w_n \log p_n) $$ (typically $$\displaystyle N=4 $$, $$\displaystyle w_n=1/4 $$).

  • Applications: Machine translation, text generation, sentiment analysis.

  • LSTMs in ChatGPT: Early language models (pre-transformers) used LSTMs to capture long-term dependencies in sequences; now largely superseded by transformers.

Computer Vision

  • Applications: Image classification (ResNet), object detection (YOLO), segmentation (U-Net).

  • Role of CNNs: Hierarchical feature learning (edges โ†’ textures โ†’ objects); translation invariance via convolution/pooling.

  • ImageNet Competition: Catalyst for deep learning boom (2012: AlexNet with CNNs drastically reduced error).

Speech Processing

  • Applications: Speech-to-text (ASR), speaker identification, emotion recognition.

  • ML Utilization:

    • ASR: CNN for spectrogram feature extraction, RNN/LSTM for temporal modeling, CTC loss.

    • Speaker ID: x-vector systems (DNN embeddings), cosine similarity scoring.

Advanced ML Paradigms

  • One-Shot Learning: Learn from few examples; uses metric learning (Siamese networks), data augmentation, or prior knowledge.

  • Self-Supervised Learning: Create pretext tasks from unlabeled data (e.g., predict rotation angle, masked language modeling in BERT).

  • Generative AI:

    • GANs (Generative Adversarial Networks): Generator vs. Discriminator game; generates realistic images.

    • VAEs (Variational Autoencoders): Learn latent distribution; generates samples via decoder.

Theoretical Aspects

  • Convex Optimization: Ensures gradient descent finds global optimum (loss surface has single minimum). Many ML problems (linear regression, SVM) are convex.

  • Linearity vs. Non-Linearity:

    • Linear models: Simple, interpretable, limited capacity.

    • Non-linearity (via activation functions, kernels): Increases model capacity, but gradient descent may converge to local minima; requires careful initialization.

  • Statistical Hypothesis Testing: Compare models using tests (e.g., t-test for paired results, McNemar's test for classification accuracy differences).


X. Model Evaluation & Optimization

Measuring Classifier Performance

  • Beyond accuracy: use precision-recall tradeoff (especially imbalanced data), ROC-AUC (threshold-independent).

  • Confusion Matrix: Foundation for all derived metrics (TPR, FPR, etc.).

Hyperparameter Tuning

  • Impact: Critical for performance (e.g., learning rate, regularization strength, network depth).

  • Methods:

    • Grid Search: Exhaustive over predefined grid.

    • Random Search: Sample random combinations (often more efficient).

    • Bayesian Optimization: Builds surrogate model to guide search.

Resampling Methods

  • Cross-Validation: As above (k-fold).

  • Bootstrapping: Sample with replacement to estimate statistic distribution (e.g., confidence intervals).

Confusion Matrix Interpretation

Metric Formula When to Use
Accuracy $(TP+TN)/Total$ Balanced classes
Precision $TP/(TP+FP)$ Minimize false positives (e.g., spam detection)
Recall $TP/(TP+FN)$ Minimize false negatives (e.g., cancer screening)
F1-score $$\displaystyle 2 \cdot \frac{Precision \cdot Recall}{Precision+Recall} $$ Balance precision/recall
Specificity $TN/(TN+FP)$ True negative rate

[!TIP] Exam Focus: Past papers frequently ask for definitions (ML, one-shot, Bayes' theorem), algorithm explanations (PCA, SVM, k-NN, EM, backpropagation), comparisons (L1 vs L2, bagging vs boosting, Q-learning vs SARSA, LSTM vs GRU), and applications (CNN in CV, RNN in NLP, RL frameworks). Always include key formulas (e.g., sigmoid, MSE, cross-entropy, Q-update) and diagrams where possible (CNN architecture, LSTM gates). For calculations, show step-by-step (e.g., k-NN distance, PCA eigenvectors, EM E/M steps).

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in