1. Introduction to Machine Learning
Definition: Machine Learning (ML) is a subset of Artificial Intelligence that enables systems to learn patterns from data without explicit programming.
Significance: Solves real-world problems in healthcare (diagnosis), finance (fraud detection), NLP (chatbots), CV (autonomous vehicles).
Types:
-
Supervised Learning: Labeled data (e.g., classification, regression).
-
Unsupervised Learning: Unlabeled data (e.g., clustering, dimensionality reduction).
-
Reinforcement Learning: Agent learns via rewards/penalties (e.g., game playing).
Hypothesis Space & Inductive Bias:
-
Hypothesis Space: Set of all possible models/algorithms considered.
-
Inductive Bias: Assumptions made to generalize from training data (e.g., smoothness in linear regression).
Bayesian Learning & Bayes' Theorem:
- Bayes' Theorem:
$$P(A|B) = \frac{P(B|A) P(A)}{P(B)}$$
-
Used in Naïve Bayes classifier: Predicts class by maximizing posterior probability $P(\text{class}|\text{features})$.
-
Fundamental: Handles uncertainty, updates beliefs with new evidence.
[!TIP]
Exam Focus: Distinguish ML types with examples. Bayes' theorem is frequently asked—memorize formula and its role in probabilistic classification.
2. Data Preprocessing and Management
Normalization vs Standardization:
-
Normalization: Scales features to $[0,1]$.
Formula: $$\displaystyle x' = \frac{x - x_{\min}}{x_{\max} - x_{\min}} $$
-
Standardization: Zero mean, unit variance.
Formula: $$\displaystyle x' = \frac{x - \mu}{\sigma} $$
-
Why needed: Improves convergence (gradient descent stability), prevents features with larger scales from dominating.
Categorical Encoding:
-
One-Hot Encoding: Creates binary column for each category. Increases dimensionality (curse of dimensionality risk).
-
Label Encoding: Assigns integer to each category. May imply false ordinality (e.g., "France"=1, "Germany"=2 suggests Germany > France).
Curse of Dimensionality:
-
High-dimensional data becomes sparse.
-
Distance metrics lose meaning.
-
Solutions: Dimensionality reduction (PCA), feature selection.
Missing Data Handling:
-
Imputation: Fill with mean/median/mode, or use ML models (e.g., KNN imputer).
-
Deletion: Remove rows/columns if missingness is random and minimal.
Data Augmentation:
-
Artificially increases dataset size (e.g., rotations, flips for images; synonym replacement for text).
-
Reduces overfitting, especially in CNNs.
[!TIP]
Common Pitfall: Using label encoding for nominal data—always use one-hot for non-ordinal categories. Normalization is critical for distance-based algorithms (KNN, K-means) and gradient descent.
3. Model Evaluation and Validation
Regression Metrics:
| Metric | Formula | Interpretation |
|---|---|---|
| MSE | $$\displaystyle \frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2 $$ | Punishes large errors; sensitive to outliers. |
| MAE | $$\displaystyle \frac{1}{n}\sum_{i=1}^{n}\|y_i - \hat{y}_i\| $$ | Robust to outliers. |
| R² | $$\displaystyle 1 - \frac{\sum(y_i - \hat{y}_i)^2}{\sum(y_i - \bar{y})^2} $$ | Proportion of variance explained; max = 1. |
Classification Metrics (from Confusion Matrix):
-
Accuracy: $$\displaystyle \frac{TP+TN}{TP+TN+FP+FN} $$
-
Precision: $$\displaystyle \frac{TP}{TP+FP} $$ (False positives costly, e.g., spam detection).
-
Recall: $$\displaystyle \frac{TP}{TP+FN} $$ (False negatives costly, e.g., cancer diagnosis).
-
F1-Score: $$\displaystyle 2 \times \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}} $$ (Harmonic mean).
-
ROC-AUC: Area under ROC curve; measures trade-off between TPR and FPR.
Cross-Validation:
-
k-fold: Split data into k folds; train on k-1, test on 1; repeat.
-
Stratified k-fold: Maintains class distribution in each fold (imbalanced data).
-
Purpose: Robust performance estimate, reduces variance.
Overfitting vs Underfitting:
-
Overfitting: High training accuracy, low test accuracy. Causes: complex model, small data. Mitigation: regularization, dropout, more data.
-
Underfitting: Low training/test accuracy. Causes: overly simple model. Mitigation: increase features, reduce regularization.
Hyperparameter Tuning:
-
Grid search, random search, Bayesian optimization.
-
Use validation set (not test set) to tune.
[!TIP]
Exam Trick: For imbalanced datasets, accuracy is misleading—always use precision, recall, F1. ROC-AUC is threshold-independent. Cross-validation prevents lucky splits.
4. Core Machine Learning Algorithms
4.1 Supervised Learning
Linear Regression:
-
Assumptions: Linearity, independence, homoscedasticity, normal residuals, no multicollinearity.
-
Model: $$\displaystyle \hat{y} = \beta_0 + \beta_1 x_1 + ... + \beta_p x_p $$
-
Cost: MSE; solved via normal equation or gradient descent.
Logistic Regression:
-
Used for classification; outputs probability via sigmoid: $$\displaystyle P(y=1|x) = \frac{1}{1+e^{-(\beta^T x)}} $$
-
vs Linear Regression: Logistic models probability (bounded [0,1]), uses log-loss.
K-Nearest Neighbors (KNN):
-
Algorithm: For a new point, find k closest training points (Euclidean distance), predict majority class (classification) or average (regression).
-
Supervised: Uses labeled data.
-
Example Calculation: Given points with speed/agility, compute distances to new point, pick k=3 nearest, vote.
Decision Trees:
-
Splitting Criteria:
-
Entropy: $$\displaystyle H(S) = -\sum p_i \log_2 p_i $$ (measures impurity).
-
Gini Impurity: $$\displaystyle G(S) = 1 - \sum p_i^2 $$ (faster computation).
-
-
Choose split that maximizes information gain: $$\displaystyle \text{IG} = H(\text{parent}) - \sum \frac{n_{\text{child}}}{n_{\text{parent}}} H(\text{child}) $$.
-
Pruning: Remove branches to prevent overfitting (pre-pruning vs post-pruning).
-
Issues: Overfitting, instability, bias towards features with more levels.
Support Vector Machines (SVM):
-
Goal: Find optimal hyperplane maximizing margin (distance to nearest points).
-
Support Vectors: Training points closest to hyperplane; define decision boundary.
-
Linear SVM: Solve constrained optimization: maximize margin subject to $$\displaystyle y_i(w^T x_i + b) \geq 1 $$.
-
Kernel Trick: Map data to high-dimensional space via kernel (e.g., RBF: $$\displaystyle K(x_i,x_j) = e^{-\gamma \|x_i-x_j\|^2} $$) without explicit transformation.
-
High-Dimensional Performance: Effective due to margin maximization; kernel handles non-linearity.
-
Applications: Bioinformatics (protein classification), image recognition (handwritten digits).
[!TIP]
KNN Example: Always standardize features before distance calculation. For k, odd numbers avoid ties in binary classification. SVM's kernel choice is critical—RBF for non-linear, linear for large features.
4.2 Unsupervised Learning
K-Means Clustering:
-
Algorithm:
-
Initialize k centroids randomly.
-
Assign each point to nearest centroid.
-
Update centroids as mean of assigned points.
-
Repeat until convergence.
-
-
Goal: Minimize within-cluster sum of squares (WCSS): $$\displaystyle \sum_{i=1}^{k} \sum_{x \in C_i} \|x - \mu_i\|^2 $$
-
Requirements: Pre-specify k, spherical clusters, similar size.
-
Example: Cluster data into 2 groups with initial centroids $$\displaystyle m_1=2, m_2=4 $$.
Hierarchical Clustering:
-
DIANA (Divisive): Top-down; start with all points in one cluster, split recursively.
-
AGNES (Agglomerative): Bottom-up; start with each point as cluster, merge nearest clusters.
-
Linkage Criteria: Single (min distance), complete (max distance), average (mean distance).
-
Adaptive: Determines number of clusters from dendrogram (cut at desired distance).
Expectation-Maximization (EM):
-
E-step: Estimate missing data (e.g., cluster assignments) given current parameters.
-
M-step: Maximize likelihood to update parameters (e.g., means, variances in GMM).
-
Handles Missing Data: Iteratively fills missing values and updates model.
-
Gaussian Mixture Models (GMM): Assumes data from mixture of Gaussians; EM estimates parameters.
BIRCH Algorithm:
-
Builds Clustering Feature (CF) Tree for large datasets.
-
Uses clustering feature summary: $(N, LS, SS)$ where $N$=count, $LS$=linear sum, $SS$=squared sum.
-
Incremental, memory-efficient; used as preprocessing for other clustering.
Clustering Evaluation:
-
Internal: Silhouette score (cohesion vs separation).
-
External: Adjusted Rand Index (if ground truth available).
-
Applications: Market segmentation, image compression, anomaly detection.
[!TIP]
K-Means Pitfall: Sensitive to initial centroids—use k-means++ initialization. EM may converge to local optimum—run multiple times. BIRCH is ideal for huge datasets but assumes spherical clusters.
4.3 Ensemble Methods
Bagging vs Boosting:
| Bagging | Boosting |
|---|---|
| Parallel training of base learners (e.g., Decision Trees). | Sequential training; each learner corrects previous errors. |
| Reduces variance (e.g., Random Forest). | Reduces bias (e.g., AdaBoost, Gradient Boosting). |
| Bootstrap sampling (with replacement). | Re-weighting of misclassified instances. |
| Example: Random Forest. | Example: XGBoost, AdaBoost. |
Random Forest:
-
Ensemble of decision trees on bootstrapped samples.
-
Regression Output: Average predictions from all trees: $$\displaystyle \hat{y} = \frac{1}{T} \sum_{t=1}^{T} \hat{y}_t $$
-
Feature randomness: Each split considers subset of features.
Stacking:
-
Combines multiple models (level-0) using a meta-learner (level-1).
-
Example: Use SVM, KNN, RF as base; logistic regression as meta to combine predictions.
[!TIP]
Ensemble Rule: Bagging for unstable learners (high variance like trees); boosting for weak learners. Random Forest reduces overfitting via feature randomness and averaging.
4.4 Other Algorithms
Locally Weighted Linear Regression (LWLR):
-
Non-parametric; fits linear model locally around each query point.
-
Weight $$\displaystyle w_i = \exp\left(-\frac{\|x_i - x\|^2}{2\tau^2}\right) $$ (Gaussian kernel).
-
Minimizes weighted MSE: $$\displaystyle \sum w_i (y_i - \beta^T x_i)^2 $$
-
Use Case: Captures local patterns; no global model assumptions.
-
Downside: Computationally expensive; no explicit model for prediction.
5. Neural Networks and Deep Learning
5.1 Fundamentals
Perceptron:
-
Single neuron: $$\displaystyle y = f(w^T x + b) $$
-
Learning Algorithm:
-
Initialize weights randomly.
-
For each sample: compute output $$\displaystyle \hat{y} = f(w^T x) $$.
-
Update weights: $$\displaystyle w_j \leftarrow w_j + \eta (y - \hat{y}) x_j $$.
-
-
Limitation: Only linearly separable problems.
Multi-Layer Perceptron (MLP):
-
Input, hidden, output layers.
-
Universal approximator with non-linear activations.
Backpropagation:
-
Algorithm:
-
Forward pass: compute output and loss $L$.
-
Backward pass: compute gradients $$\displaystyle \frac{\partial L}{\partial w} $$ via chain rule.
-
Update weights: $$\displaystyle w \leftarrow w - \eta \frac{\partial L}{\partial w} $$.
-
-
Weight Adjustment: Gradients flow from output to input layers; adjusts weights to minimize loss.
Activation Functions:
| Function | Formula | Properties | Vanishing Gradient? |
|---|---|---|---|
| Sigmoid | $$\displaystyle \sigma(x) = \frac{1}{1+e^{-x}} $$ | Output [0,1]; smooth. | Yes (saturates). |
| Tanh | $$\displaystyle \tanh(x) = \frac{e^x - e^{-x}}{e^x + e^{-x}} $$ | Output [-1,1]; zero-centered. | Yes (saturates). |
| ReLU | $$\displaystyle \text{ReLU}(x) = \max(0,x) $$ | Sparse activation; fast compute. | No (for $$\displaystyle x>0 $$). |
| Leaky ReLU | $\max(\alpha x, x)$, $\alpha \approx 0.01$ | Fixes "dying ReLU". | No. |
| Softmax | $$\displaystyle \sigma(z)_j = \frac{e^{z_j}}{\sum_{k=1}^{K} e^{z_k}} $$ | Multi-class probability; sum=1. | Yes (in combination). |
Loss Functions:
-
MSE: $$\displaystyle \frac{1}{n}\sum (y_i - \hat{y}_i)^2 $$ (regression).
-
Cross-Entropy: $$\displaystyle -\sum y_i \log(\hat{y}_i) $$ (classification).
-
Hinge Loss: $$\displaystyle \max(0, 1 - y_i \hat{y}_i) $$ (SVM).
Optimization Algorithms:
-
Batch GD: Full dataset per update (stable, slow).
-
Stochastic GD (SGD): One sample per update (noisy, fast).
-
Mini-batch GD: Compromise (common in practice).
-
Momentum: $$\displaystyle v \leftarrow \beta v + (1-\beta) \nabla L $$; accelerates convergence.
-
Adam: Adaptive learning rate; combines momentum and RMSprop.
-
RMSprop: Scales learning rate by moving average of squared gradients.
Regularization:
-
L1 (Lasso): Adds $$\displaystyle \lambda \sum |w_j| $$ to loss. Drives some weights to zero (feature selection).
-
L2 (Ridge): Adds $$\displaystyle \lambda \sum w_j^2 $$ to loss. Shrinks weights toward zero but not zero.
-
Dropout: Randomly deactivate neurons during training; prevents co-adaptation.
-
Batch Normalization: Normalizes layer inputs (mean=0, var=1); stabilizes training, reduces internal covariate shift.
Linearity vs Non-linearity:
-
Linear Models: Limited expressiveness; gradient descent converges to global optimum (convex).
-
Non-linear Models (with activations): Can approximate any function; loss surface non-convex (local minima risk).
-
Impact on GD: Non-linearities enable complex patterns but require careful initialization and learning rates.
Parallel Processing:
-
Matrix operations (e.g., layer computations) are highly parallelizable on GPUs/TPUs.
-
Batch processing leverages hardware parallelism.
[!TIP]
Activation Choice: ReLU default for hidden layers; avoid sigmoid/tanh in deep nets (vanishing gradient). Use softmax for multi-class output. Adam is robust default optimizer. L1 for sparse solutions, L2 for generalization.
5.2 Convolutional Neural Networks (CNN)
Architecture:
-
Convolutional Layers: Apply filters to extract features.
-
Pooling Layers: Downsample (reduce spatial size, increase depth).
-
Fully Connected (FC) Layers: Final classification/regression.
Convolution Operation:
-
Filter/Kernel: $$\displaystyle K \in \mathbb{R}^{k \times k \times C_{\text{in}}} $$ (e.g., 3x3).
-
Stride: Step size of filter movement.
-
Padding:
-
Same: Output size = input size; pads with zeros.
-
Valid: No padding; output shrinks.
-
-
Output Size: $$\displaystyle \frac{W - K + 2P}{S} + 1 $$
-
1×1 Convolution:
-
Purpose: Feature reduction/increase (change depth), add non-linearity, efficient computation.
-
Example: Inception modules use 1x1 convs before 3x3/5x5 to reduce channels.
-
Pooling:
-
Max Pooling: Takes maximum in window; preserves dominant features.
-
Average Pooling: Takes average; smooths features.
-
Sub-sampling: Reduces spatial dimensions, increases invariance.
Flattening: Converts 3D feature maps to 1D vector for FC layers.
Advanced Architectures:
-
Inception Module: Parallel convolutions (1x1, 3x3, 5x5) + pooling; concatenate outputs. Increases width, captures multi-scale features efficiently.
-
Transfer Learning:
-
Feature Extraction: Freeze pretrained CNN (e.g., ResNet), use as fixed feature extractor; train new classifier.
-
Fine-Tuning: Unfreeze some top layers, train with low learning rate.
-
TensorFlow Implementation (Brief):
model = Sequential([
Conv2D(32, (3,3), activation='relu', input_shape=(224,224,3)),
MaxPooling2D(2,2),
Conv2D(64, (3,3), activation='relu'),
Flatten(),
Dense(128, activation='relu'),
Dense(10, activation='softmax')
])
Overfitting in CNNs:
-
Detection: Training accuracy >> validation accuracy.
-
Solutions: Data augmentation (rotations, crops), dropout (after FC layers), L2 regularization, early stopping.
Applications: Image classification (ImageNet), object detection (YOLO), segmentation (U-Net).
[!TIP]
1×1 Convolution: Often used to reduce computational cost before expensive 3×3/5×5 convs (e.g., Inception). Padding choice affects output size: "same" preserves spatial dimensions, "valid" reduces. Max pooling is most common; average pooling used in global pooling for fully convolutional networks.
5.3 Recurrent Neural Networks (RNN) and Sequence Models
RNN Architecture:
-
Recurrent Connection: Hidden state $$\displaystyle h_t $$ depends on previous $$\displaystyle h_{t-1} $$ and input $$\displaystyle x_t $$.
-
Unfolding in Time: $$\displaystyle h_t = f(W_{hh} h_{t-1} + W_{xh} x_t + b_h) $$; output $$\displaystyle y_t = g(W_{hy} h_t + b_y) $$.
-
Vanilla RNN: Simple but suffers from vanishing/exploding gradients.
LSTM (Long Short-Term Memory):
-
Structure: Cell state $$\displaystyle C_t $$ (memory) + hidden state $$\displaystyle h_t $$; gates control information flow.
-
Forget Gate: Decides what to discard from $$\displaystyle C_{t-1} $$.
-
Input Gate: Updates $$\displaystyle C_t $$ with new info.
-
Output Gate: Produces $$\displaystyle h_t $$ based on $$\displaystyle C_t $$.
-
-
Advantage: Handles long-term dependencies via constant error flow through cell state.
GRU (Gated Recurrent Unit):
-
Simplified LSTM: No separate cell state; combines forget/input into update gate and reset gate.
-
Fewer parameters; faster training; similar performance.
Handling Long-Term Dependencies:
-
LSTM/GRU mitigate vanishing gradient via gating mechanisms.
-
Attention Mechanisms: Allow direct access to past states (e.g., in Transformers).
Applications in NLP/Speech:
-
NLP: Machine translation, text generation (e.g., LSTMs in early ChatGPT).
-
Speech: Speech-to-text (transcription), speaker identification (voice patterns).
Attention Mechanisms (Brief):
-
Computes weighted sum of all hidden states; focuses on relevant parts.
-
Foundation for Transformers (self-attention).
[!TIP]
LSTM vs GRU: GRU is faster, less data-hungry; LSTM more expressive for long sequences. Vanilla RNNs rarely used in practice due to gradient issues. For very long sequences, combine with attention or use Transformers.
6. Reinforcement Learning
Markov Decision Processes (MDP):
-
Components: States $S$, Actions $A$, Rewards $R$, Transition probabilities $P(s'|s,a)$, Discount factor $\gamma$.
-
Goal: Maximize cumulative discounted reward: $$\displaystyle G_t = R_{t+1} + \gamma R_{t+2} + ... $$
-
Policy $\pi(a|s)$: Strategy mapping states to actions.
Value Iteration vs Policy Iteration:
-
Value Iteration:
-
Initialize $$\displaystyle V(s)=0 $$.
-
Update: $$\displaystyle V_{k+1}(s) = \max_a \sum_{s'} P(s'|s,a)[R(s,a,s') + \gamma V_k(s')] $$
-
Stop when $$\displaystyle \Delta < \theta $$.
-
-
Policy Iteration:
-
Policy evaluation: Compute $$\displaystyle V^\pi $$.
-
Policy improvement: $$\displaystyle \pi' = \arg\max_a \sum_{s'} P(s'|s,a)[R(s,a,s') + \gamma V^\pi(s')] $$.
-
Repeat until $\pi$ stable.
-
-
Difference: Value iteration updates values directly; policy iteration alternates evaluation/improvement.
Q-Learning:
-
Q-value: $Q(s,a)$ = expected cumulative reward taking action $a$ in state $s$.
-
Update Rule:
$$Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r + \gamma \max_{a'} Q(s',a') - Q(s,a) \right]$$
-
Exploration vs Exploitation: Use $\epsilon$-greedy: with probability $\epsilon$, random action; else $$\displaystyle \arg\max_a Q(s,a) $$.
-
Off-policy: Learns optimal policy independent of current policy.
SARSA:
- Update Rule:
$$Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r + \gamma Q(s',a') - Q(s,a) \right]$$
where $a'$ is action actually taken (from current policy).
-
On-policy: Updates based on current policy's actions.
-
Difference: SARSA is more conservative (accounts for exploration); Q-learning assumes optimal actions.
Actor-Critic Models:
-
Actor: Policy network $\pi(a|s)$; selects actions.
-
Critic: Value network $V(s)$ or $Q(s,a)$; evaluates actions.
-
Interaction: Actor proposes action; critic computes TD error $$\displaystyle \delta = r + \gamma V(s') - V(s) $$; both networks update using $\delta$.
-
Advanced Variants:
-
A2C (Advantage Actor-Critic): Uses advantage $$\displaystyle A(s,a) = Q(s,a) - V(s) $$.
-
A3C (Asynchronous A2C): Parallel actors for faster training.
-
PPO (Proximal Policy Optimization): Constrains policy updates for stability.
-
Frameworks:
-
OpenAI Gym: Standardized environments (Atari, robotics).
-
TensorFlow Agents: RL library built on TensorFlow.
Applications: Game playing (AlphaGo), robotics, resource management, autonomous driving.
[!TIP]
MDP Assumption: Current state captures all past info (Markov property). Q-learning converges to optimal policy with sufficient exploration. SARSA learns safer policy in risky environments (e.g., cliff walking). Actor-Critic combines policy gradient stability with value-based efficiency.
7. Natural Language Processing (NLP)
NLP Pipeline:
-
Tokenization: Split text into words/subwords.
-
Stemming: Crudely chop suffixes (e.g., "running" → "run").
-
Lemmatization: Vocabulary-aware base form (e.g., "better" → "good").
-
Stop Words Removal: Remove common words (the, is).
-
n-grams: Contiguous sequences (unigram, bigram, trigram); capture context.
BLEU Score (for machine translation):
-
n-gram Precision: Fraction of candidate n-grams found in reference(s).
-
Brevity Penalty: Penalizes short candidates: $$\displaystyle \text{BP} = \min\left(1, e^{1 - \frac{r}{c}}\right) $$ where $r$=reference length, $c$=candidate length.
-
BLEU: $$\displaystyle \text{BLEU} = \text{BP} \cdot \exp\left(\sum_{n=1}^{N} w_n \log p_n\right) $$ (geometric mean of n-gram precisions).
Language Models:
-
n-gram: Probabilistic based on previous $n-1$ words; sparse data problem.
-
RNN/LSTM: Sequential processing; captures long-range dependencies.
-
Transformers (brief): Self-attention; parallel processing; state-of-the-art (BERT, GPT).
LSTMs in NLP Generation:
-
Used in early ChatGPT (GPT-1/2) for next-word prediction.
-
Handles long-term dependencies via cell state.
-
Generates coherent text by sampling from predicted probability distribution.
Speech Processing:
-
Speech-to-Text: Convert audio to text (e.g., ASR systems using CTC loss with RNNs/Transformers).
-
Speaker Identification: Extract voice features (MFCCs), classify speaker (embedding-based models like x-vector).
Attention in NLP:
-
Allows decoder to focus on relevant encoder states (e.g., in translation).
-
Self-Attention: Relates all positions in a sequence (Transformer core).
[!TIP]
BLEU Pitfall: High n-gram precision but poor fluency may still yield low BLEU due to brevity penalty. Stemming/lemmatization can improve recall but may lose nuance. For speech, MFCCs are hand-crafted features; modern systems use raw waveforms or filterbanks.
8. Computer Vision (CV)
CNN Applications:
-
Image Classification: Assign label to entire image (e.g., ResNet).
-
Object Detection: Locate and classify objects (e.g., YOLO, Faster R-CNN).
-
Segmentation: Pixel-wise labeling (e.g., U-Net for medical imaging).
ImageNet Competition:
-
History: Started 2010; large-scale image dataset (14M images, 22K categories).
-
Impact:
-
AlexNet (2012): First deep CNN (5 conv layers); dropped top-5 error from 26% to 15%.
-
VGGNet (2014): Deeper (16-19 layers), small 3x3 filters.
-
ResNet (2015): Residual connections; >100 layers; human-level performance.
-
-
Significance: Catalyzed deep learning revolution; demonstrated depth matters; spurred architecture innovations.
Other CV Tasks:
- Image captioning (CNN + LSTM), face recognition (Siamese networks), style transfer (GANs), video analysis (3D CNNs).
[!TIP]
ImageNet Legacy: Proved GPUs + deep CNNs could surpass traditional methods. Residual connections solved vanishing gradient in very deep nets. Transfer learning from ImageNet pretrained models is standard practice.
9. Advanced Topics and Emerging Trends
Autoencoders:
-
Architecture: Encoder (compress input to latent code), Decoder (reconstruct input).
-
Loss: Reconstruction loss (MSE for images, cross-entropy for binary).
-
Types:
-
Denoising Autoencoder: Train to reconstruct clean input from noisy version.
-
Variational Autoencoder (VAE): Learns probabilistic latent space; used for generation.
-
-
Unsupervised Use: Feature learning, dimensionality reduction, anomaly detection.
One-Shot Learning:
-
Definition: Learn from very few examples (often one per class).
-
Difference from Supervised: Traditional supervised requires many examples; one-shot uses prior knowledge/metric learning.
-
Applications: Face recognition (new person with one photo), rare disease diagnosis.
Self-Supervised Learning:
-
Pretext Tasks: Create supervised task from unlabeled data (e.g., predict rotation angle, masked image modeling).
-
Contrastive Learning: Learn representations by pulling similar samples together, pushing dissimilar apart (e.g., SimCLR).
-
Benefit: Reduces need for labeled data; learns robust features.
Generative AI:
-
GANs (Generative Adversarial Networks): Generator vs Discriminator; generates realistic images/videos.
-
VAEs: Probabilistic generative model; smooth latent space.
-
Diffusion Models: Iteratively denoise from random noise; state-of-the-art in image synthesis (DALL-E 2, Stable Diffusion).
Transfer Learning:
-
Concept: Reuse knowledge from source task to target task.
-
Strategies:
-
Feature Extraction: Freeze pretrained layers, train new head.
-
Fine-Tuning: Unfreeze some/all layers, train with small learning rate.
-
-
When to Use: Target dataset small, similar to source (e.g., ImageNet → medical imaging).
Convex Optimization in ML:
-
Convex Loss: Single global minimum (e.g., linear regression MSE, SVM with linear kernel).
-
Importance: Guarantees convergence to global optimum; used in traditional ML (logistic regression, SVM).
-
Non-convex: Deep learning; multiple local minima.
Bayesian Networks:
-
Structure: Directed acyclic graph; nodes = random variables, edges = conditional dependencies.
-
Inference: Compute posterior probabilities given evidence (e.g., using variable elimination).
-
Example: Car diagnosis network—predict car value given mileage, engine condition, AC status.
-
Prediction: $$\displaystyle P(\text{Car Value}=\text{High} \mid \text{Mileage}=\text{Lo}, ...) $$ via conditional probability tables.
[!TIP]
Autoencoders: Latent dimension controls compression; too small loses info. One-shot learning often uses Siamese networks with contrastive loss. Diffusion models currently dominate image generation. Bayesian networks require known dependencies; inference is NP-hard in general (approximate methods used).
10. Historical Milestones and Frameworks
ImageNet Competition:
-
Significance: Benchmark for deep learning progress; spurred architectures (AlexNet, VGG, ResNet); demonstrated depth and GPU training efficacy.
-
Impact: Drop in top-5 error from ~26% (2010) to <3% (2017); catalyzed AI boom.
Popular Frameworks:
-
TensorFlow: Production-ready, static graph (eager mode now), extensive deployment tools.
-
PyTorch: Dynamic graph, research favorite, intuitive debugging.
-
scikit-learn: Traditional ML (SVM, trees, clustering); easy API.
-
OpenAI Gym: RL environments (Atari, MuJoCo).
-
TensorFlow Agents: RL library with agents, networks, environments.
[!TIP]
Framework Choice: PyTorch for research/prototyping; TensorFlow for production/mobile. scikit-learn for classical ML (not deep learning). Gym for RL algorithm testing.