1. FOUNDATIONS & OVERVIEW
Definition & Importance
-
Machine Learning (ML): A field of AI where algorithms learn patterns from data without explicit programming.
Formal Definition (Tom Mitchell): "A computer program is said to learn from experience E with respect to some task T and performance measure P, if its performance on T, as measured by P, improves with E."
-
Importance: Solves complex, real-world problems (image recognition, NLP, recommendation systems) where rule-based systems fail. Enables automation, prediction, and adaptation from data.
-
Perspectives & Issues:
-
Design Issues: Choice of hypothesis space, overfitting vs. underfitting, evaluation metrics, computational cost.
-
Basic Approaches: Symbolic (rule-based), connectionist (neural networks), statistical (probabilistic models).
-
Types of Machine Learning
| Type | Goal | Examples |
|---|---|---|
| Supervised Learning | Learn mapping from input to known output | Classification (spam detection), Regression (price prediction) |
| Unsupervised Learning | Discover hidden patterns in unlabeled data | Clustering (customer segmentation), Dimensionality Reduction (PCA) |
| Reinforcement Learning (RL) | Agent learns via rewards/penalties from environment | Game playing (AlphaGo), Robotics control |
| Semi-supervised | Mix of labeled & unlabeled data | Web page classification (few labeled docs) |
| Self-supervised | Create labels from data itself | Pre-training on image rotations, masked language modeling |
Core Concepts & Terminology
-
Hypothesis Space (H): Set of all possible models/algorithms considered. Inductive Bias: Assumptions used to choose one hypothesis over others (e.g., preference for simpler models).
-
Data Splits:
-
Training Set: Used to fit model parameters.
-
Validation Set: Used for hyperparameter tuning & model selection.
-
Testing Set: Used for final, unbiased performance evaluation.
-
-
Overfitting vs. Underfitting:
-
Overfitting: Model learns noise in training data; high training accuracy, low test accuracy. Solutions: More data, regularization, simpler model, early stopping.
-
Underfitting: Model fails to capture underlying pattern; low training & test accuracy. Solutions: More features, complex model, less regularization.
-
-
Model Selection & Hyperparameter Tuning: Choosing the best model from hypothesis space (e.g., SVM vs. Decision Tree) and optimizing settings (e.g., learning rate, regularization strength) using validation set or cross-validation.
2. DATA PREPROCESSING & FEATURE ENGINEERING
Importance & Need
-
Improves model convergence speed, stability, and final performance.
-
Ensures features are on comparable scales, handles noise, and reduces computational cost.
Common Techniques
| Technique | Purpose | Common Methods |
|---|---|---|
| Normalization/Standardization | Scale features to similar range | Min-Max Scaling ($$\displaystyle x' = \frac{x - \min}{\max - \min} $$), Z-score ($$\displaystyle x' = \frac{x - \mu}{\sigma} $$) |
| Categorical Variables | Convert non-numeric data to numeric | One-Hot Encoding (creates binary columns, increases dimensionality), Label Encoding (assigns integers, may imply ordinal relationship) |
| Missing Data | Handle incomplete records | Deletion, Imputation (mean, median, model-based) |
| Data Augmentation | Artificially increase dataset size (esp. for images) | Rotation, flipping, cropping, adding noise |
Dimensionality Reduction
-
Curse of Dimensionality: As dimensions increase, data becomes sparse, distance metrics lose meaning, and computational cost rises.
-
Principal Component Analysis (PCA):
-
Goal: Transform data to a lower-dimensional space while preserving maximum variance.
-
Algorithm:
-
Standardize data.
-
Compute covariance matrix.
-
Compute eigenvectors & eigenvalues of covariance matrix.
-
Select top k eigenvectors (principal components) with largest eigenvalues.
-
Project data onto these components.
-
-
Key Formula: Projection $$\displaystyle Y = XW $$, where $W$ contains top k eigenvectors.
-
-
Feature Extraction vs. Feature Selection:
-
Extraction: Create new features from original ones (e.g., PCA, LDA).
-
Selection: Choose a subset of original features (e.g., mutual information, recursive feature elimination).
-
-
Partial Least Squares (PLS): Similar to PCA but maximizes covariance between features and target variable (supervised).
3. SUPERVISED LEARNING: CLASSIC ALGORITHMS
Linear Models
-
Linear Regression:
-
Assumptions: Linearity, independence, homoscedasticity, normality of errors, no multicollinearity.
-
Cost Function: Mean Squared Error (MSE) $$\displaystyle J(\theta) = \frac{1}{2m} \sum_{i=1}^{m} (h_\theta(x^{(i)}) - y^{(i)})^2 $$.
-
Optimization: Gradient Descent: $$\displaystyle \theta_j := \theta_j - \alpha \frac{\partial}{\partial \theta_j} J(\theta) $$.
-
-
Logistic Regression:
-
Used for binary classification.
-
Sigmoid Function: $$\displaystyle \sigma(z) = \frac{1}{1 + e^{-z}} $$, maps any real number to (0,1).
-
Decision Boundary: $$\displaystyle \theta^T x = 0 $$.
-
-
Locally Weighted Linear Regression: Fits linear models weighted by proximity to query point; captures non-linear trends.
Support Vector Machines (SVM)
-
Goal: Find optimal hyperplane that maximizes the margin (distance between hyperplane and nearest data points of any class).
-
Support Vectors: Training points that lie on the margin boundaries; solely define the hyperplane.
-
Linear SVM: For linearly separable data, solves constrained optimization to maximize margin.
-
Kernel Trick: Implicitly maps data to higher-dimensional space using kernel functions (e.g., polynomial, RBF) to handle non-linearity.
-
Performance in High-Dimensional Spaces: Effective when features >> samples (e.g., bioinformatics, text classification) due to margin maximization.
Instance-Based Learning: K-Nearest Neighbors (K-NN)
-
Algorithm:
-
Store all training examples.
-
For a new query point, compute distance (Euclidean, Manhattan, etc.) to all training points.
-
Select k nearest neighbors.
-
Classification: Majority vote of neighbors' labels.
-
Regression: Average of neighbors' target values.
-
-
Key Parameters: k (number of neighbors), distance metric.
Decision Trees
-
Structure: Hierarchical tree of nodes (root, internal, leaf). Splits based on feature values.
-
Algorithms: ID3 (uses Information Gain), C4.5 (uses Gain Ratio), CART (uses Gini Impurity).
-
Entropy & Information Gain:
-
Entropy (measure of impurity): $$\displaystyle Entropy(t) = -\sum_{i=1}^{c} p_i(t) \log_2 p_i(t) $$.
-
Information Gain: $$\displaystyle IG(T, a) = Entropy(T) - \sum_{v \in Values(a)} \frac{|T_v|}{|T|} Entropy(T_v) $$. Split on feature with highest IG.
-
-
Issues: Overfitting (deep trees), instability (small data changes alter tree), bias towards features with many levels.
-
Random Forests:
-
Ensemble of decision trees using bagging (bootstrap aggregating) and feature randomness.
-
Reduces variance, improves generalization vs. single tree.
-
Bagging vs. Boosting: Bagging (parallel, reduces variance), Boosting (sequential, reduces bias, e.g., AdaBoost, Gradient Boosting).
-
Evaluation Metrics
-
Regression:
-
MSE: $$\displaystyle \frac{1}{n} \sum (y_i - \hat{y}_i)^2 $$ (sensitive to outliers).
-
MAE: $$\displaystyle \frac{1}{n} \sum |y_i - \hat{y}_i| $$ (robust).
-
Rยฒ (Coefficient of Determination): $$\displaystyle 1 - \frac{\sum (y_i - \hat{y}_i)^2}{\sum (y_i - \bar{y})^2} $$ (proportion of variance explained).
-
-
Classification:
-
Confusion Matrix:
| | Predicted + | Predicted - | |----------------|-------------|-------------| | Actual + | TP | FN | | Actual - | FP | TN |
-
Derived Metrics:
-
Accuracy: $$\displaystyle \frac{TP+TN}{Total} $$
-
Precision: $$\displaystyle \frac{TP}{TP+FP} $$ (of predicted positives, how many correct?)
-
Recall (Sensitivity): $$\displaystyle \frac{TP}{TP+FN} $$ (of actual positives, how many found?)
-
F1-Score: $$\displaystyle 2 \times \frac{Precision \times Recall}{Precision + Recall} $$ (harmonic mean).
-
-
-
Resampling Methods:
-
k-fold Cross-Validation: Split data into k folds; train on k-1, test on 1; repeat k times; average performance.
-
Bootstrap: Sample with replacement; estimate performance on out-of-bag samples.
-
4. UNSUPERVISED LEARNING
Clustering
-
Goals: Group similar instances together, discover inherent structure, reduce data complexity.
-
K-Means Clustering:
-
Algorithm:
-
Initialize k centroids (randomly or k-means++).
-
Repeat until convergence:
-
Assignment: Assign each point to nearest centroid (Euclidean distance).
-
Update: Recompute centroids as mean of assigned points.
-
-
-
Properties: Sensitive to initialization, assumes spherical clusters, requires k.
-
-
Expectation-Maximization (EM) Algorithm:
-
E-step: Estimate "missing" data (e.g., cluster assignments) given current parameters.
-
M-step: Update model parameters to maximize likelihood given estimated data.
-
Application to Gaussian Mixture Models (GMM): Assumes data from mixture of Gaussians; EM estimates means, covariances, mixing coefficients.
-
Handling Missing Data: EM iteratively fills missing values (E-step) and re-trains model (M-step).
-
-
Hierarchical Clustering:
-
DIANA (Divisive): Top-down; start with all points in one cluster, recursively split.
-
AGNES (Agglomerative): Bottom-up; start with each point as cluster, merge closest clusters.
-
Linkage Criteria: Single (min distance), Complete (max distance), Average.
-
-
Density-Based Clustering (DBSCAN): Groups dense regions separated by sparse areas; handles arbitrary shapes, noise.
-
BIRCH Algorithm: Builds CF (Clustering Feature) tree for incremental clustering; efficient for large datasets.
5. NEURAL NETWORKS & DEEP LEARNING FUNDAMENTALS
Artificial Neural Network (ANN) Basics
-
Biological Inspiration: Neurons with dendrites (inputs), soma (processing), axon (output).
-
Artificial Neuron (Perceptron): Computes weighted sum + bias, applies activation function.
$$z = w^T x + b, \quad a = f(z)$$
- Multi-Layer Perceptron (MLP): Input layer, one or more hidden layers, output layer. Fully connected.
Activation Functions
| Function | Formula | Pros | Cons |
|---|---|---|---|
| Sigmoid | $$\displaystyle \sigma(z) = \frac{1}{1+e^{-z}} $$ | Smooth, output (0,1) | Vanishing gradient, not zero-centered |
| Tanh | $$\displaystyle \tanh(z) = \frac{e^z - e^{-z}}{e^z + e^{-z}} $$ | Zero-centered, steeper than sigmoid | Vanishing gradient |
| ReLU | $$\displaystyle f(z) = \max(0, z) $$ | Computationally cheap, alleviates vanishing gradient | Dying ReLU (neurons stuck at 0) |
| Leaky ReLU | $$\displaystyle f(z) = \max(\alpha z, z) $$ | Fixes dying ReLU | Needs tuning of $\alpha$ |
Vanishing Gradient Problem: Gradients become extremely small in deep networks during backpropagation, especially with sigmoid/tanh, hindering learning in early layers.
Training Neural Networks
-
Gradient Descent:
-
Batch GD: Use entire dataset per update (stable, slow).
-
Stochastic GD (SGD): Use one sample per update (noisy, fast).
-
Mini-batch GD: Use small batch (compromise, common in practice).
-
-
Backpropagation Algorithm:
-
Forward Pass: Compute output and loss for a batch.
-
Backward Pass:
-
Compute gradient of loss w.r.t. output layer weights (using chain rule).
-
Propagate gradients backward through layers.
-
Update weights: $$\displaystyle w_{ij} := w_{ij} - \alpha \frac{\partial Loss}{\partial w_{ij}} $$.
-
- Chain Rule Application: $$\displaystyle \frac{\partial Loss}{\partial w_{ij}} = \frac{\partial Loss}{\partial a_j} \cdot \frac{\partial a_j}{\partial z_j} \cdot \frac{\partial z_j}{\partial w_{ij}} $$, where $$\displaystyle a_j $$ is activation, $$\displaystyle z_j $$ is weighted sum.
-
-
Loss Functions:
-
Regression: Mean Squared Error (MSE).
-
Classification: Cross-Entropy Loss (Binary: $$\displaystyle -\frac{1}{m}\sum [y^{(i)}\log(\hat{y}^{(i)}) + (1-y^{(i)})\log(1-\hat{y}^{(i)})] $$; Categorical: $$\displaystyle -\sum y_i \log(\hat{y}_i) $$).
-
-
Optimizers:
-
SGD: Basic update.
-
Adam: Adaptive learning rate, momentum, RMSprop combination; widely used.
-
RMSprop: Adapts learning rate per parameter, divides by root of squared gradients.
-
Regularization Techniques
-
L1 (Lasso) vs. L2 (Ridge):
-
L1: Adds $$\displaystyle \lambda \sum |w_j| $$ to loss. Drives some weights to exactly zero โ sparsity (feature selection).
-
L2: Adds $$\displaystyle \lambda \sum w_j^2 $$ to loss. Shrinks weights toward zero but rarely to zero โ weight decay.
-
Both: Prevent overfitting by penalizing large weights.
-
-
Dropout: Randomly "drop" (set to zero) a fraction of neurons during training; prevents co-adaptation, acts as ensemble.
-
Batch Normalization: Normalizes layer inputs to zero mean, unit variance; stabilizes training, allows higher learning rates, slight regularization effect.
6. CONVOLUTIONAL NEURAL NETWORKS (CNNs)
Core Architecture & Motivation
-
Why for Images?: Exploits spatial locality (nearby pixels related) and translation invariance (object identity independent of position).
-
Operations: Convolution (feature extraction) โ Activation (non-linearity) โ Pooling/Sub-sampling (invariance, dimensionality reduction).
Key Layers & Concepts
-
Convolutional Layer:
-
Applies filters/kernels (small weight matrices) to input, producing feature maps.
-
Stride: Step size of filter movement.
-
Output size: $$\displaystyle \frac{W - F + 2P}{S} + 1 $$, where $W$=input width, $F$=filter size, $P$=padding, $S$=stride.
-
-
$1 \times 1$ Convolution:
-
Purpose: Feature reduction/channel mixing without spatial change. Used in architectures like Inception to reduce computational cost and increase non-linearity.
-
Acts as a per-pixel fully connected layer across channels.
-
-
Pooling/Sub-sampling Layer:
-
Max Pooling: Takes maximum value in window โ retains dominant feature, provides translation invariance.
-
Average Pooling: Takes average โ smooths features.
-
Reduces spatial dimensions, number of parameters, and computation.
-
-
Fully Connected (FC) Layer: At end; flattens features for final classification/regression.
Architectural Design Choices
-
Padding:
-
Valid: No padding; output shrinks.
-
Same: Padding such that output size equals input size (preserves spatial dimensions).
-
-
Flattening Layer: Converts multi-dimensional feature maps to 1D vector for FC layers.
Famous Architectures & Concepts
-
Inception Module:
-
Uses parallel convolutions of different sizes ($1\times1$, $3\times3$, $5\times5$) and pooling, concatenates outputs.
-
Benefits: Multi-scale feature capture, efficient use of parameters (via $1\times1$ bottlenecks).
-
-
Transfer Learning:
-
Feature Extraction: Use pre-trained CNN (e.g., ImageNet) as fixed feature extractor; train only new classifier.
-
Fine-tuning: Unfreeze and train some top layers of pre-trained network along with new classifier.
-
-
ImageNet Competition: Drove CNN innovation (AlexNet 2012, VGG, GoogLeNet/Inception, ResNet). Demonstrated deep CNNs' power.
Practical Considerations
-
Overfitting in CNNs: Use data augmentation, dropout, weight decay, early stopping.
-
Underfitting: Deeper/wider network, more filters, reduce regularization.
-
Implementation (TensorFlow/Keras): Sequential/Functional API, layers (
Conv2D,MaxPooling2D,Dense).
7. RECURRENT NEURAL NETWORKS (RNNs) & SEQUENCE MODELS
Vanilla RNN
-
Architecture: Has hidden state $$\displaystyle h_t $$ that captures past information. $$\displaystyle h_t = f(W_{xh} x_t + W_{hh} h_{t-1} + b_h) $$.
-
Limitations:
-
Vanishing/Exploding Gradients: Gradients shrink/explode over long sequences due to repeated multiplication.
-
Short-term Memory: Struggles with long-range dependencies.
-
Long Short-Term Memory (LSTM)
-
Architecture:
-
Cell State ($$\displaystyle C_t $$): "Conveyor belt" carrying information with minimal interference.
-
Gates (sigmoid + pointwise multiplication):
-
Forget Gate: Decides what to drop from cell state.
-
Input Gate: Decides what new information to store.
-
Output Gate: Decides what to output based on cell state.
-
-
-
Handles Long-Term Dependencies: Cell state allows gradients to flow easily over many time steps.
-
Role in NLP: Predecessor to Transformers; used in language modeling, machine translation (e.g., seq2seq with LSTM encoder-decoder).
Gated Recurrent Unit (GRU)
-
Simplified LSTM: Combines forget and input gates into update gate; no separate cell state (hidden state serves both).
-
Comparison: Fewer parameters, often similar performance; faster training.
Applications in NLP & Speech
-
Speech Processing:
-
Speech-to-Text (ASR): Convert audio waveform to text (RNN/CNN + CTC loss).
-
Speaker Identification: Classify speaker from voice features.
-
-
NLP Pipeline: Tokenization, stemming/lemmatization, stop-word removal, n-grams.
-
BLEU Score (for machine translation):
-
n-gram Precision: Modified precision for n-grams (clips counts to avoid overcounting).
-
Brevity Penalty: Penalizes translations shorter than reference. $$\displaystyle BP = \min(1, e^{1 - r/c}) $$, where $r$=reference length, $c$=candidate length.
-
Final: $$\displaystyle BLEU = BP \cdot \exp(\sum_{n=1}^{N} w_n \log p_n) $$.
-
8. REINFORCEMENT LEARNING (RL)
Fundamental Framework: Markov Decision Process (MDP)
-
Defined by $(S, A, P, R, \gamma)$:
-
$S$: Set of states.
-
$A$: Set of actions.
-
$P(s' | s, a)$: Transition probability to state $s'$ from $s$ taking $a$.
-
$R(s, a, s')$: Reward received.
-
$\gamma$: Discount factor ($0 \leq \gamma \leq 1$).
-
-
Policy ($\pi$): Strategy mapping states to actions ($$\displaystyle a = \pi(s) $$).
-
Difference from Supervised/Unsupervised: No labeled data; agent learns via interaction and delayed rewards.
Core Algorithms
-
Value Iteration:
-
Iteratively updates state values: $$\displaystyle V_{k+1}(s) = \max_a \sum_{s'} P(s'|s,a)[R(s,a,s') + \gamma V_k(s')] $$.
-
Directly computes optimal value function; no explicit policy.
-
-
Policy Iteration:
-
Alternates between Policy Evaluation (compute $$\displaystyle V^\pi $$) and Policy Improvement ($$\displaystyle \pi' = \arg\max_a \sum_{s'} P(s'|s,\pi(s))[R + \gamma V^\pi(s')] $$).
-
Guarantees convergence to optimal policy.
-
-
Q-Learning (Off-policy TD Control):
-
Learns Q-value $Q(s,a)$: expected cumulative reward taking $a$ in $s$ then following optimal policy.
-
Update Rule: $$\displaystyle Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma \max_{a'} Q(s',a') - Q(s,a)] $$.
-
Uses greedy actions for update regardless of behavior policy.
-
-
SARSA (On-policy TD Control):
-
Updates $Q$ based on actual next action taken: $$\displaystyle Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma Q(s',a') - Q(s,a)] $$, where $a' \sim \pi(s')$.
-
More conservative; learns policy being executed.
-
Advanced Actor-Critic Models
-
Actor-Critic:
-
Actor: Policy network ($$\displaystyle \pi_\theta(a|s) $$) selects actions.
-
Critic: Value network ($$\displaystyle V_w(s) $$ or $$\displaystyle Q_w(s,a) $$) evaluates actions.
-
Interaction: Actor proposes action, Critic computes TD error $$\displaystyle \delta = r + \gamma V_w(s') - V_w(s) $$, which guides Actor update (policy gradient) and Critic update (value loss).
-
-
Rise of Advanced Variants:
-
A2C/A3C: Asynchronous Advantage Actor-Critic (parallel actors).
-
PPO (Proximal Policy Optimization): Constrains policy updates for stability; state-of-the-art for many tasks.
-
Practical Frameworks & Applications
-
Frameworks: OpenAI Gym (environments), Stable Baselines3, TensorFlow Agents, Ray RLlib.
-
Applications: Game playing (Atari, Go), robotics control, resource management, autonomous driving.
9. ADVANCED TOPICS & APPLICATIONS
Autoencoders
-
Architecture: Encoder (compresses input to latent code), Decoder (reconstructs input from code).
-
Loss: Reconstruction error (MSE, cross-entropy).
-
Uses in Unsupervised Learning:
-
Dimensionality Reduction: Latent space is lower-dimensional representation.
-
Denoising: Train with corrupted inputs, reconstruct clean outputs.
-
Anomaly Detection: High reconstruction error indicates anomaly.
-
One-Shot Learning
-
Definition: Learning from very few examples (often one or few per class).
-
Difference from Traditional Supervised: Requires massive data; one-shot learns efficiently with prior knowledge (e.g., via meta-learning, metric learning).
-
Beneficial Applications: Rare event classification (medical diagnosis), new object recognition, personalization.
Self-Supervised Learning
-
Concept: Automatically generate supervisory signals from data itself (pretext tasks).
-
Examples: Predict image rotation, mask parts of image/text and recover (BERT, MAE).
-
Rise: Reduces need for labeled data; pre-trains powerful representations.
Generative AI
-
Overview: Models that generate new data samples resembling training data.
-
Connections:
-
Autoencoders: Variational Autoencoders (VAEs) generate samples.
-
GANs (Generative Adversarial Networks): Generator vs. Discriminator game.
-
LLMs (Large Language Models): Transformer-based, trained on vast text corpora (GPT, BERT).
-
Application Domains
-
Computer Vision:
-
Object Detection: Locate & classify objects (YOLO, Faster R-CNN).
-
Image Segmentation: Pixel-level labeling (U-Net, Mask R-CNN).
-
-
Natural Language Processing (NLP):
-
Tasks: Translation, sentiment analysis, summarization, question answering.
-
Transformers: Attention mechanism (context for LSTM/RNN); BERT, GPT.
-
-
Speech Processing:
-
ASR (Automatic Speech Recognition): Audio โ Text (DeepSpeech, Wav2Vec).
-
Synthesis (TTS): Text โ Speech (Tacotron, WaveNet).
-
Speaker Identification: Verify/identify speaker from voice.
-
-
Bioinformatics: SVM for gene classification, protein structure prediction; CNNs for medical image analysis.
10. OPTIMIZATION & THEORETICAL UNDERPINNINGS
Convex Optimization
-
Importance: If loss function is convex, gradient descent is guaranteed to find global optimum (no local minima).
-
Many ML problems (linear regression, logistic regression) have convex loss; deep learning often non-convex.
Linearity vs. Non-linearity
-
Linearity: Model is linear in parameters (e.g., linear regression). Can be solved analytically; limited expressiveness.
-
Non-linearity: Introduced via activation functions or feature transformations. Enables modeling complex relationships but requires iterative optimization (gradient descent); prone to vanishing/exploding gradients.
Probability & Statistics in ML
-
Role: Provides framework for uncertainty, decision making, and probabilistic models (Naive Bayes, HMMs, Bayesian Networks).
-
Bayes' Theorem:
$$P(A|B) = \frac{P(B|A) P(A)}{P(B)}$$
* **Use in Classification (Naive Bayes)**: Assumes feature independence given class. $$\displaystyle P(\text{class}|\text{features}) \propto P(\text{class}) \prod P(\text{feature}|\text{class}) $$.
* **Fundamental**: Underlies Bayesian inference, updating beliefs with evidence.
-
Bayesian Learning:
-
Treats parameters as random variables with prior distributions.
-
Posterior $P(\theta|D) \propto P(D|\theta) P(\theta)$.
-
Allows incorporation of prior knowledge, quantifies uncertainty.
-
-
Bayesian Networks: Directed acyclic graph representing conditional dependencies among variables. Nodes = random variables, edges = conditional dependencies. Used for probabilistic inference.
Loss Calculation
-
Critical Role: Defines objective function to minimize. Guides backpropagation by providing gradient signal.
-
Choice Depends on Task: MSE for regression, cross-entropy for classification, hinge loss for SVM.
-
Directly influences model's learned representation and generalization.
Exam Tips & Common Pitfalls:
- Confusion Matrix: Always label rows (Actual) and columns (Predicted). Precision = PPV, Recall = Sensitivity.
- PCA vs. Feature Selection: PCA creates new features; selection picks existing ones.
- L1 vs. L2: L1 gives sparsity (zero weights), L2 gives small weights.
- Q-learning vs. SARSA: Off-policy vs. on-policy; Q uses $$\displaystyle \max_{a'} Q(s',a') $$, SARSA uses $Q(s',a')$ from current policy.
- Backpropagation: Chain rule applied layer-by-layer from output to input; compute gradients, then update weights.
- CNN Padding: "Same" preserves spatial size; "Valid" reduces size.
- Activation Functions: ReLU most common today; sigmoid/tanh in output layers for probability.
- MDP: Must satisfy Markov property (future depends only on current state, not history).