UNIT 1: Foundations & Core Concepts
I. Foundations of Machine Learning
Definition & Importance
Machine Learning (ML) is a subset of AI that enables systems to learn patterns from data without explicit programming.
Importance: Automation (e.g., self-driving cars), prediction (e.g., stock trends), personalization (e.g., recommender systems), handling complex, high-dimensional data.
Types of Machine Learning
| Type | Description | Examples |
|---|---|---|
| Supervised | learns from labeled data (input-output pairs) | Classification (spam detection), Regression (price prediction) |
| Unsupervised | finds hidden patterns in unlabeled data | Clustering (customer segmentation), Dimensionality reduction (PCA) |
| Reinforcement | agent learns via rewards/penalties from environment | Game playing (AlphaGo), Robotics control |
| Semi-supervised | mix of labeled & unlabeled data | Web page classification (few labeled docs) |
| One-shot | learns from very few examples (often 1) | Face recognition (new person from single photo) |
AI vs. ML vs. Deep Learning
-
AI: Broad fieldโcreating intelligent systems.
-
ML: Subset of AIโalgorithms that learn from data.
-
Deep Learning: Subset of MLโuses deep neural networks with many layers.
Perspectives & Issues
-
Design Issues: Choosing hypothesis space, inductive bias (algorithm's assumptions), avoiding overfitting.
-
Limitations: Data quality dependency, model interpretability ("black box"), computational scalability.
-
Performance Factors: Data quality/quantity, algorithm choice, hyperparameter tuning.
-
Statistical Theory: Probability & statistics underpin generalization bounds, confidence intervals, hypothesis testing.
Machine Learning Workflow
-
Problem definition & data collection
-
Data preprocessing & feature engineering
-
Model selection (choose algorithm & hypothesis space)
-
Training on training set
-
Validation (tune hyperparameters via validation set)
-
Testing on unseen test set
-
Deployment & monitoring
II. Data Preprocessing and Feature Engineering
Need for Preprocessing
Ensures model convergence, stability, and performance. Raw data often contains noise, missing values, and scale inconsistencies.
Common Techniques
| Technique | Purpose | Example |
|---|---|---|
| Normalization (Min-Max) | Scales features to [0,1] | $$\displaystyle X_{\text{norm}} = \frac{X - X_{\min}}{X_{\max} - X_{\min}} $$ |
| Standardization (Z-score) | Centers to mean 0, std 1 | $$\displaystyle X_{\text{std}} = \frac{X - \mu}{\sigma} $$ |
| One-hot Encoding | Converts categorical to binary vectors (increases dimensionality) | "Color" โ [Red:1,0,0], [Green:0,1,0], [Blue:0,0,1] |
| Label Encoding | Assigns integer to categories (may imply false ordering) | "Size" โ Small:0, Medium:1, Large:2 |
| Handling Missing Data | Imputation (mean/median), deletion, or model-based | Replace missing age with median age |
High-Dimensional Data Challenges
-
Curse of Dimensionality: Data becomes sparse; distance metrics lose meaning; computational cost explodes.
-
Impact: Overfitting risk increases, model performance degrades ("dimensionality curse").
Dimensionality Reduction
| Method | Type | Key Idea |
|---|---|---|
| PCA | Unsupervised | Projects data onto orthogonal axes of max variance |
| LDA | Supervised | Maximizes class separability (uses labels) |
| PLS | Supervised | Finds components that maximize covariance with target |
| Feature Selection | โ | Forward selection (add features), backward elimination (remove features) |
Data Augmentation
Artificially increases training data by applying transformations (rotation, flipping, cropping) to improve generalizationโcritical in computer vision.
III. Supervised Learning: Fundamentals
Regression vs. Classification
| Aspect | Regression | Classification |
|---|---|---|
| Output | Continuous value (e.g., price) | Discrete class label (e.g., spam/ham) |
| Example | Predicting house price | Image recognition (cat/dog) |
Hypothesis Space & Inductive Bias
-
Hypothesis Space: Set of all possible models the algorithm can consider (e.g., all linear functions).
-
Inductive Bias: Algorithm's assumptions to generalize (e.g., "nearest neighbors are similar"). Guides learning from examples to unseen cases.
Overfitting vs. Underfitting
| Overfitting | Underfitting | |
|---|---|---|
| Cause | Model too complex (high variance) | Model too simple (high bias) |
| Symptoms | Low training error, high validation error | High training & validation error |
| Solutions | Regularization, cross-validation, early stopping, more data | Increase model complexity, feature engineering |
Cross-Validation
-
Purpose: Estimate model performance robustly, tune hyperparameters.
-
k-fold: Split data into k subsets; train on k-1, validate on 1; repeat k times.
-
Leave-One-Out (LOO): k = N (extreme case); high variance, computationally expensive.
Evaluation Metrics
-
Regression:
-
MSE (Mean Squared Error): $$\displaystyle \text{MSE} = \frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2 $$
-
MAE (Mean Absolute Error): $$\displaystyle \text{MAE} = \frac{1}{n}\sum_{i=1}^{n}\|y_i - \hat{y}_i\| $$
-
Rยฒ (Coefficient of Determination): $$\displaystyle R^2 = 1 - \frac{\sum(y_i - \hat{y}_i)^2}{\sum(y_i - \bar{y})^2} $$
-
-
Classification:
-
Confusion Matrix:
| | Predicted + | Predicted - | |---|---|---| | Actual + | TP | FN | | Actual - | FP | TN |
-
Derived Metrics:
-
Accuracy: $$\displaystyle \frac{TP+TN}{TP+TN+FP+FN} $$
-
Precision: $$\displaystyle \frac{TP}{TP+FP} $$ (positive predictive value)
-
Recall (TPR): $$\displaystyle \frac{TP}{TP+FN} $$ (sensitivity)
-
F1-score: $$\displaystyle 2 \times \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}} $$
-
ROC-AUC: Area under ROC curve (TPR vs FPR trade-off).
-
-
IV. Supervised Learning: Algorithms
A. Linear Models
Linear Regression
-
Assumptions: Linearity, independence, homoscedasticity (constant variance), normality of residuals.
-
Cost Function: Minimize MSE โ closed-form solution: $$\displaystyle \hat{\beta} = (X^TX)^{-1}X^Ty $$ (or gradient descent).
-
Interpretation: $$\displaystyle \hat{y} = \beta_0 + \beta_1 x_1 + ... + \beta_p x_p $$.
Logistic Regression
-
Sigmoid Function: $$\displaystyle \sigma(z) = \frac{1}{1+e^{-z}} $$, maps $z \in \mathbb{R}$ to $(0,1)$.
-
Use: Binary classification; predicts $$\displaystyle P(y=1|x) $$.
-
Difference from Linear Regression: Uses log-loss (cross-entropy) cost; output is probability.
Locally Weighted Linear Regression
-
Non-parametric: Fits a separate linear model for each query point $$\displaystyle x_q $$, weighting training points by distance to $$\displaystyle x_q $$.
-
Weight: $$\displaystyle w_i = \exp\left(-\frac{(x_i - x_q)^2}{2\tau^2}\right) $$ ($\tau$: bandwidth).
B. Decision Trees
-
Splitting Criteria:
-
Entropy: $$\displaystyle H(S) = -\sum p_i \log_2 p_i $$ (impurity measure).
-
Gini Impurity: $$\displaystyle G(S) = 1 - \sum p_i^2 $$.
-
Information Gain: $$\displaystyle \text{IG}(S,A) = H(S) - \sum \frac{|S_v|}{|S|} H(S_v) $$.
-
-
Pruning: Remove branches to reduce overfitting (pre-pruning vs. post-pruning).
-
Issues: Instability (small data changes โ different tree), bias-variance tradeoff (deep trees low bias, high variance).
C. Support Vector Machines (SVM)
-
Goal: Find optimal hyperplane that maximizes margin (distance to nearest points of each class).
-
Support Vectors: Training points lying on margin boundaries; define the decision boundary.
-
Linear SVM: For linearly separable data, solve constrained optimization: maximize margin subject to $$\displaystyle y_i(w^Tx_i + b) \geq 1 $$.
-
Kernel Trick: Map data to high-dimensional space via kernel (e.g., RBF: $$\displaystyle K(x_i,x_j) = \exp(-\gamma\|x_i-x_j\|^2) $$) to handle non-linear separation.
-
High-Dimensional Performance: Effective when features >> samples (e.g., bioinformatics, text classification) due to margin maximization.
D. Instance-Based Learning: k-NN
-
Algorithm: For a query point, find k nearest training points (by distance metric), predict majority class (classification) or average (regression).
-
Distance Metrics: Euclidean ($$\displaystyle \sqrt{\sum (x_i-y_i)^2} $$), Manhattan ($$\displaystyle \sum |x_i-y_i| $$).
-
Supervised: Uses labeled training data directly (lazy learning).
-
Example: Predict if player drafted given speed=6.75, agility=3 using k=3 (compute distances to all labeled players, pick 3 nearest, majority vote).
E. Ensemble Methods
| Method | Mechanism | Bias/Variance Effect |
|---|---|---|
| Bagging (Bootstrap Aggregating) | Train multiple models on bootstrapped samples, average/vote | Reduces variance (e.g., Random Forest) |
| Boosting | Sequentially train models, focus on misclassified points | Reduces bias (e.g., AdaBoost, Gradient Boosting) |
| Stacking | Train meta-model on outputs of base models | Combines strengths, can reduce both |
Bagging vs. Boosting
-
Bagging: Parallel, reduces overfitting (variance), models independent.
-
Boosting: Sequential, reduces underfitting (bias), models dependent (correct errors).
V. Unsupervised Learning
Goals & Requirements
-
Goal: Discover hidden structure (clusters, low-dim representation) without labels.
-
Requirements: Scalability, ability to handle noise, minimal prior knowledge (e.g., number of clusters).
Clustering Algorithms
| Algorithm | Key Steps | Limitations |
|---|---|---|
| K-means | 1. Initialize k centroids. 2. Assign points to nearest centroid. 3. Update centroids. Repeat until convergence. | Sensitive to outliers, initial centroids; assumes spherical clusters. |
| Hierarchical (DIANA) | Divisive: Start with all points in one cluster; recursively split. | Computationally expensive, irreversible splits. |
| EM (GMM) | E-step: Estimate cluster responsibilities given current params. M-step: Update Gaussian parameters (mean, covariance, mixing coeff). | Can converge to local optimum; sensitive to initialization. |
| BIRCH | Builds Clustering Feature (CF) Tree; incrementally clusters large datasets. | Assumes clusters are spherical/convex; not for arbitrary shapes. |
VI. Neural Networks & Deep Learning
A. Basics
-
ANN Structure: Input layer โ hidden layers (with activation functions) โ output layer. Connections have weights.
-
Biological Inspiration: Neurons (nodes) with dendrites (inputs) and axons (outputs); parallel processing.
-
MLP (Multi-Layer Perceptron): Feedforward network with โฅ1 hidden layer; universal approximator.
Perceptron Learning Algorithm
-
Initialize weights $w$ randomly.
-
For each training sample $$\displaystyle (x_i, y_i) $$:
-
Compute output: $$\displaystyle \hat{y} = f(w^T x_i) $$ (step/activation).
-
Update: $$\displaystyle w \leftarrow w + \eta (y_i - \hat{y}) x_i $$.
-
-
Repeat until convergence.
B. Training Neural Networks
Backpropagation
-
Forward pass: compute output & loss $L$.
-
Backward pass: compute gradients $$\displaystyle \frac{\partial L}{\partial w} $$ via chain rule.
-
Update weights: $$\displaystyle w \leftarrow w - \eta \frac{\partial L}{\partial w} $$.
- Characteristics: Gradient-based, supervised, computes error derivatives layer-by-layer.
Gradient Descent Optimizers
| Optimizer | Mechanism | Use Case |
|---|---|---|
| SGD | Updates using single sample gradient | Simple, noisy convergence |
| Mini-batch GD | Updates using small batch (common) | Balance speed/stability |
| Adam | Adaptive learning rate, momentum, bias correction | Default choice, fast convergence |
| RMSprop | Adapts learning rate per parameter | RNNs, non-stationary objectives |
Loss Functions
-
MSE: For regression.
-
Cross-Entropy: For classification (binary: $-[y\log\hat{y}+(1-y)\log(1-\hat{y})]$; multi-class: $$\displaystyle -\sum y_i\log\hat{y}_i $$).
-
Role: Guides backpropagation by providing gradient signal to minimize prediction error.
C. Key Components
Activation Functions
| Function | Formula | Pros | Cons |
|---|---|---|---|
| Sigmoid | $$\displaystyle \sigma(x) = \frac{1}{1+e^{-x}} $$ | Output (0,1), smooth | Vanishing gradient, not zero-centered |
| tanh | $$\displaystyle \tanh(x) = \frac{e^x-e^{-x}}{e^x+e^{-x}} $$ | Output (-1,1), zero-centered | Vanishing gradient |
| ReLU | $$\displaystyle \text{ReLU}(x) = \max(0,x) $$ | Computationally cheap,็ผ่งฃ vanishing gradient | Dying ReLU (negative outputs) |
| Softmax | $$\displaystyle \sigma(z)_j = \frac{e^{z_j}}{\sum_{k=1}^K e^{z_k}} $$ | Multi-class probability output | โ |
Regularization Techniques
-
L1 (Lasso): Adds $$\displaystyle \lambda \sum |w| $$ to loss โ sparsity (some weights become 0).
-
L2 (Ridge): Adds $$\displaystyle \lambda \sum w^2 $$ โ weight decay (weights shrink uniformly).
-
Dropout: Randomly deactivate neurons during training (prevents co-adaptation).
-
Batch Normalization: Normalizes layer inputs to zero mean/unit variance โ stabilizes training, reduces internal covariate shift.
Overfitting/Underfitting in NN
-
Detection: Training loss << validation loss (overfitting); both high (underfitting).
-
Solutions: Dropout, data augmentation, L1/L2, early stopping, reduce network size.
D. Convolutional Neural Networks (CNN)
Architecture: [Input] โ [Conv โ Activation โ Pooling]รn โ [Flatten] โ [FC] โ [Output]
-
Convolution Operation: Filter (kernel) slides over input, computes dot product โ feature map.
- Parameters: filter size (e.g., 3ร3), stride, number of filters.
-
Padding:
-
Same: Output size = input size (pad with zeros); preserves spatial info.
-
Valid: No padding; output shrinks.
-
-
Pooling (Subsampling):
-
Max Pooling: Takes max in window โ translation invariance, reduces dimensions.
-
Average Pooling: Takes average โ smoother downsampling.
-
-
1ร1 Convolution:
- Purpose: Feature reduction (change depth), model efficiency ("network-in-network"), add non-linearity without spatial change.
-
Flattening: Converts multi-dimensional feature maps to 1D vector for fully connected layers.
-
Inception Module:
-
Parallel convolutions with multiple filter sizes (1ร1, 3ร3, 5ร5) + max pooling.
-
Concatenate outputs โ captures multi-scale features efficiently.
-
-
Transfer Learning:
-
Feature Extraction: Freeze pre-trained CNN layers, train new classifier on top.
-
Fine-tuning: Unfreeze some top layers, train with low learning rate.
-
-
CNN in TensorFlow (brief):
model = Sequential([ Conv2D(32, (3,3), activation='relu', input_shape=(224,224,3)), MaxPooling2D(2,2), Flatten(), Dense(128, activation='relu'), Dense(10, activation='softmax') ])
E. Recurrent Neural Networks (RNN)
-
Architecture: Recurrent connections; hidden state $$\displaystyle h_t $$ depends on $$\displaystyle h_{t-1} $$ and input $$\displaystyle x_t $$.
- Unfolding in Time: Expands recurrent connections to a deep network across time steps.
-
Types:
-
Vanilla RNN: $$\displaystyle h_t = \tanh(W_{hh}h_{t-1} + W_{xh}x_t + b) $$; suffers from vanishing/exploding gradients.
-
LSTM (Long Short-Term Memory):
-
Cells: Memory cell $$\displaystyle c_t $$ with gates:
-
Forget gate: $$\displaystyle f_t = \sigma(W_f \cdot [h_{t-1}, x_t] + b_f) $$ โ what to discard from cell.
-
Input gate: $$\displaystyle i_t = \sigma(W_i \cdot [h_{t-1}, x_t] + b_i) $$; $$\displaystyle \tilde{c}_t = \tanh(W_c \cdot [h_{t-1}, x_t] + b_c) $$ โ new candidate.
-
Output gate: $$\displaystyle o_t = \sigma(W_o \cdot [h_{t-1}, x_t] + b_o) $$; $$\displaystyle h_t = o_t * \tanh(c_t) $$.
-
-
Handles long-term dependencies via cell state.
-
-
GRU (Gated Recurrent Unit):
-
Merges forget & input gates into update gate $$\displaystyle z_t $$; has reset gate $$\displaystyle r_t $$.
-
Fewer parameters than LSTM, often similar performance.
-
-
-
LSTM vs. GRU: LSTM has separate cell state & hidden state; GRU merges them. LSTM more expressive, GRU faster.
-
Applications in NLP: Language modeling, machine translation, text generation (e.g., ChatGPT uses transformer, but LSTMs were foundational).
F. Autoencoders
-
Architecture:
[Input] โ [Encoder (bottleneck)] โ [Decoder] โ [Reconstructed Input]. -
Unsupervised Use: Learns efficient data encoding (dimensionality reduction), denoising (train with corrupted input), feature learning.
-
Bottleneck: Forces network to learn compressed representation.
G. Advanced Topics
-
Attention Models: Allows decoder to "focus" on relevant parts of input sequence (e.g., in translation). Key in transformers.
-
Self-Supervised Learning:
-
Pretext Tasks: Solve auxiliary task on unlabeled data (e.g., predict missing image patch, next sentence).
-
Contrastive Learning: Learn representations by pulling similar samples together, pushing dissimilar apart (e.g., SimCLR).
-
VII. Reinforcement Learning
A. Fundamentals
-
Difference from Supervised/Unsupervised:
-
Supervised: Fixed labeled dataset.
-
Unsupervised: Find structure in unlabeled data.
-
RL: Agent interacts with environment, learns from reward signal (no explicit labels).
-
-
Markov Decision Process (MDP):
-
Components: States $S$, Actions $A$, Transition probabilities $P(s'|s,a)$, Rewards $R(s,a,s')$, Policy $\pi(a|s)$.
-
Value Function: $$\displaystyle V^\pi(s) = \mathbb{E}[\sum \gamma^t R_t | s_0=s, \pi] $$ (expected return under policy $\pi$).
-
Policy Iteration: Alternate between policy evaluation (compute $$\displaystyle V^\pi $$) and policy improvement (greedy update).
-
Value Iteration: Directly update $$\displaystyle V(s) \leftarrow \max_a \sum_{s'} P(s'|s,a)[R(s,a,s') + \gamma V(s')] $$ until convergence.
-
-
Exploration vs. Exploitation:
-
Dilemma: Explore (try new actions to discover better rewards) vs. Exploit (choose known best action).
-
Strategies: $\epsilon$-greedy, softmax, Upper Confidence Bound (UCB).
-
B. Core Algorithms
Q-Learning (Off-policy TD control)
-
Q-value: $Q(s,a)$ = expected future reward taking action $a$ in state $s$.
-
Update Rule: $$\displaystyle Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma \max_{a'} Q(s',a') - Q(s,a)] $$.
-
Algorithm (deterministic rewards/actions):
-
Initialize $Q(s,a)$ arbitrarily.
-
For each episode:
-
Initialize state $s$.
-
While $s$ not terminal:
-
Choose $a$ from $s$ using policy derived from $Q$ (e.g., $\epsilon$-greedy).
-
Take $a$, observe $r, s'$.
-
$$\displaystyle Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma \max_{a'} Q(s',a') - Q(s,a)] $$.
-
$$\displaystyle s \leftarrow s' $$.
-
-
-
-
Guides Actions: Agent selects action with highest $Q$-value (greedy) after learning.
SARSA (On-policy TD control)
-
Update Rule: $$\displaystyle Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma Q(s',a') - Q(s,a)] $$.
-
Difference from Q-learning: Uses actual next action $a'$ (from same policy) vs. $$\displaystyle \max_{a'} Q(s',a') $$ (off-policy). More conservative, follows current policy.
Value Iteration vs. Policy Iteration
| Value Iteration | Policy Iteration | |
|---|---|---|
| Process | Directly update $V(s)$; no explicit policy evaluation | Alternate: evaluate $$\displaystyle V^\pi $$, then improve $\pi$ |
| Convergence | Often faster per iteration, but each step computationally heavier | Slower per iteration (requires full policy evaluation), but stable |
| Policy | Implicit (greedy w.r.t $V$) | Explicit |
C. Advanced Methods
-
Actor-Critic:
-
Actor: Policy network $\pi(a|s)$ (selects actions).
-
Critic: Value network $V(s)$ or $Q(s,a)$ (evaluates actions).
-
Interaction: Actor proposes action; Critic estimates value; Actor updated via policy gradient using Critic's feedback.
-
-
Advanced Actor-Critic Models:
-
A2C (Advantage Actor-Critic): Uses advantage function $$\displaystyle A(s,a) = Q(s,a)-V(s) $$.
-
A3C (Asynchronous A2C): Parallel actors with global network.
-
PPO (Proximal Policy Optimization): Constrains policy updates to avoid large destructive steps.
-
-
RL Frameworks:
-
OpenAI Gym: Standardized environments (Atari, robotics).
-
TensorFlow Agents: Library for RL in TensorFlow.
-
VIII. Probabilistic & Bayesian Methods
Bayes' Theorem
$$P(A|B) = \frac{P(B|A) P(A)}{P(B)}$$
-
Interpretation: Updates belief about hypothesis $A$ given evidence $B$.
-
Example: Medical test: $$\displaystyle P(\text{disease}|\text{positive}) = \frac{P(\text{positive}|\text{disease})P(\text{disease})}{P(\text{positive})} $$.
-
Fundamental to Probabilistic Models: Forms basis for Naive Bayes, Bayesian networks, Bayesian inference.
Bayesian Networks
-
Structure: Directed acyclic graph (DAG); nodes = random variables, edges = conditional dependencies.
-
Conditional Independence: Node independent of non-descendants given parents.
-
Probability Tables: Each node has Conditional Probability Table (CPT) $P(X|\text{Parents}(X))$.
-
Inference: Compute posterior probabilities given evidence (e.g., $$\displaystyle P(\text{car value}|\text{mileage=Lo, engine=Bad}) $$ via variable elimination or sampling).
Bayesian Learning
-
Definition: Treat parameters as random variables; use prior knowledge $P(\theta)$ + likelihood $P(D|\theta)$ to compute posterior $P(\theta|D) \propto P(D|\theta)P(\theta)$.
-
Impact: Incorporates prior beliefs, quantifies uncertainty, avoids overfitting (regularization via prior).
IX. Applications & Advanced Topics
Natural Language Processing (NLP)
-
Pipeline: Tokenization โ Stemming/Lemmatization โ Stop-word removal โ Vectorization (Bag-of-Words, TF-IDF, embeddings).
-
BLEU Score (for machine translation):
-
n-gram precision: Fraction of n-grams in candidate that appear in reference.
-
Brevity Penalty: Penalizes short candidates: $$\displaystyle BP = \min(1, \exp(1 - \frac{\text{ref len}}{\text{cand len}})) $$.
-
BLEU: $$\displaystyle BLEU = BP \cdot \exp(\sum_{n=1}^N w_n \log p_n) $$ (typically $$\displaystyle N=4 $$, $$\displaystyle w_n=1/4 $$).
-
-
Applications: Machine translation, text generation, sentiment analysis.
-
LSTMs in ChatGPT: Early language models (pre-transformers) used LSTMs to capture long-term dependencies in sequences; now largely superseded by transformers.
Computer Vision
-
Applications: Image classification (ResNet), object detection (YOLO), segmentation (U-Net).
-
Role of CNNs: Hierarchical feature learning (edges โ textures โ objects); translation invariance via convolution/pooling.
-
ImageNet Competition: Catalyst for deep learning boom (2012: AlexNet with CNNs drastically reduced error).
Speech Processing
-
Applications: Speech-to-text (ASR), speaker identification, emotion recognition.
-
ML Utilization:
-
ASR: CNN for spectrogram feature extraction, RNN/LSTM for temporal modeling, CTC loss.
-
Speaker ID: x-vector systems (DNN embeddings), cosine similarity scoring.
-
Advanced ML Paradigms
-
One-Shot Learning: Learn from few examples; uses metric learning (Siamese networks), data augmentation, or prior knowledge.
-
Self-Supervised Learning: Create pretext tasks from unlabeled data (e.g., predict rotation angle, masked language modeling in BERT).
-
Generative AI:
-
GANs (Generative Adversarial Networks): Generator vs. Discriminator game; generates realistic images.
-
VAEs (Variational Autoencoders): Learn latent distribution; generates samples via decoder.
-
Theoretical Aspects
-
Convex Optimization: Ensures gradient descent finds global optimum (loss surface has single minimum). Many ML problems (linear regression, SVM) are convex.
-
Linearity vs. Non-Linearity:
-
Linear models: Simple, interpretable, limited capacity.
-
Non-linearity (via activation functions, kernels): Increases model capacity, but gradient descent may converge to local minima; requires careful initialization.
-
-
Statistical Hypothesis Testing: Compare models using tests (e.g., t-test for paired results, McNemar's test for classification accuracy differences).
X. Model Evaluation & Optimization
Measuring Classifier Performance
-
Beyond accuracy: use precision-recall tradeoff (especially imbalanced data), ROC-AUC (threshold-independent).
-
Confusion Matrix: Foundation for all derived metrics (TPR, FPR, etc.).
Hyperparameter Tuning
-
Impact: Critical for performance (e.g., learning rate, regularization strength, network depth).
-
Methods:
-
Grid Search: Exhaustive over predefined grid.
-
Random Search: Sample random combinations (often more efficient).
-
Bayesian Optimization: Builds surrogate model to guide search.
-
Resampling Methods
-
Cross-Validation: As above (k-fold).
-
Bootstrapping: Sample with replacement to estimate statistic distribution (e.g., confidence intervals).
Confusion Matrix Interpretation
| Metric | Formula | When to Use |
|---|---|---|
| Accuracy | $(TP+TN)/Total$ | Balanced classes |
| Precision | $TP/(TP+FP)$ | Minimize false positives (e.g., spam detection) |
| Recall | $TP/(TP+FN)$ | Minimize false negatives (e.g., cancer screening) |
| F1-score | $$\displaystyle 2 \cdot \frac{Precision \cdot Recall}{Precision+Recall} $$ | Balance precision/recall |
| Specificity | $TN/(TN+FP)$ | True negative rate |
[!TIP] Exam Focus: Past papers frequently ask for definitions (ML, one-shot, Bayes' theorem), algorithm explanations (PCA, SVM, k-NN, EM, backpropagation), comparisons (L1 vs L2, bagging vs boosting, Q-learning vs SARSA, LSTM vs GRU), and applications (CNN in CV, RNN in NLP, RL frameworks). Always include key formulas (e.g., sigmoid, MSE, cross-entropy, Q-update) and diagrams where possible (CNN architecture, LSTM gates). For calculations, show step-by-step (e.g., k-NN distance, PCA eigenvectors, EM E/M steps).