Skip to content
CY-701 ยท Machine Learning/Quick Revision Short Notes

Machine Learning (CY-701) - Unit 1 Short Notes

1. FOUNDATIONS & OVERVIEW

Definition & Importance

  • Machine Learning (ML): A field of AI where algorithms learn patterns from data without explicit programming.

    Formal Definition (Tom Mitchell): "A computer program is said to learn from experience E with respect to some task T and performance measure P, if its performance on T, as measured by P, improves with E."

  • Importance: Solves complex, real-world problems (image recognition, NLP, recommendation systems) where rule-based systems fail. Enables automation, prediction, and adaptation from data.

  • Perspectives & Issues:

    • Design Issues: Choice of hypothesis space, overfitting vs. underfitting, evaluation metrics, computational cost.

    • Basic Approaches: Symbolic (rule-based), connectionist (neural networks), statistical (probabilistic models).

Types of Machine Learning

Type Goal Examples
Supervised Learning Learn mapping from input to known output Classification (spam detection), Regression (price prediction)
Unsupervised Learning Discover hidden patterns in unlabeled data Clustering (customer segmentation), Dimensionality Reduction (PCA)
Reinforcement Learning (RL) Agent learns via rewards/penalties from environment Game playing (AlphaGo), Robotics control
Semi-supervised Mix of labeled & unlabeled data Web page classification (few labeled docs)
Self-supervised Create labels from data itself Pre-training on image rotations, masked language modeling

Core Concepts & Terminology

  • Hypothesis Space (H): Set of all possible models/algorithms considered. Inductive Bias: Assumptions used to choose one hypothesis over others (e.g., preference for simpler models).

  • Data Splits:

    • Training Set: Used to fit model parameters.

    • Validation Set: Used for hyperparameter tuning & model selection.

    • Testing Set: Used for final, unbiased performance evaluation.

  • Overfitting vs. Underfitting:

    • Overfitting: Model learns noise in training data; high training accuracy, low test accuracy. Solutions: More data, regularization, simpler model, early stopping.

    • Underfitting: Model fails to capture underlying pattern; low training & test accuracy. Solutions: More features, complex model, less regularization.

  • Model Selection & Hyperparameter Tuning: Choosing the best model from hypothesis space (e.g., SVM vs. Decision Tree) and optimizing settings (e.g., learning rate, regularization strength) using validation set or cross-validation.


2. DATA PREPROCESSING & FEATURE ENGINEERING

Importance & Need

  • Improves model convergence speed, stability, and final performance.

  • Ensures features are on comparable scales, handles noise, and reduces computational cost.

Common Techniques

Technique Purpose Common Methods
Normalization/Standardization Scale features to similar range Min-Max Scaling ($$\displaystyle x' = \frac{x - \min}{\max - \min} $$), Z-score ($$\displaystyle x' = \frac{x - \mu}{\sigma} $$)
Categorical Variables Convert non-numeric data to numeric One-Hot Encoding (creates binary columns, increases dimensionality), Label Encoding (assigns integers, may imply ordinal relationship)
Missing Data Handle incomplete records Deletion, Imputation (mean, median, model-based)
Data Augmentation Artificially increase dataset size (esp. for images) Rotation, flipping, cropping, adding noise

Dimensionality Reduction

  • Curse of Dimensionality: As dimensions increase, data becomes sparse, distance metrics lose meaning, and computational cost rises.

  • Principal Component Analysis (PCA):

    • Goal: Transform data to a lower-dimensional space while preserving maximum variance.

    • Algorithm:

      1. Standardize data.

      2. Compute covariance matrix.

      3. Compute eigenvectors & eigenvalues of covariance matrix.

      4. Select top k eigenvectors (principal components) with largest eigenvalues.

      5. Project data onto these components.

    • Key Formula: Projection $$\displaystyle Y = XW $$, where $W$ contains top k eigenvectors.

  • Feature Extraction vs. Feature Selection:

    • Extraction: Create new features from original ones (e.g., PCA, LDA).

    • Selection: Choose a subset of original features (e.g., mutual information, recursive feature elimination).

  • Partial Least Squares (PLS): Similar to PCA but maximizes covariance between features and target variable (supervised).


3. SUPERVISED LEARNING: CLASSIC ALGORITHMS

Linear Models

  • Linear Regression:

    • Assumptions: Linearity, independence, homoscedasticity, normality of errors, no multicollinearity.

    • Cost Function: Mean Squared Error (MSE) $$\displaystyle J(\theta) = \frac{1}{2m} \sum_{i=1}^{m} (h_\theta(x^{(i)}) - y^{(i)})^2 $$.

    • Optimization: Gradient Descent: $$\displaystyle \theta_j := \theta_j - \alpha \frac{\partial}{\partial \theta_j} J(\theta) $$.

  • Logistic Regression:

    • Used for binary classification.

    • Sigmoid Function: $$\displaystyle \sigma(z) = \frac{1}{1 + e^{-z}} $$, maps any real number to (0,1).

    • Decision Boundary: $$\displaystyle \theta^T x = 0 $$.

  • Locally Weighted Linear Regression: Fits linear models weighted by proximity to query point; captures non-linear trends.

Support Vector Machines (SVM)

  • Goal: Find optimal hyperplane that maximizes the margin (distance between hyperplane and nearest data points of any class).

  • Support Vectors: Training points that lie on the margin boundaries; solely define the hyperplane.

  • Linear SVM: For linearly separable data, solves constrained optimization to maximize margin.

  • Kernel Trick: Implicitly maps data to higher-dimensional space using kernel functions (e.g., polynomial, RBF) to handle non-linearity.

  • Performance in High-Dimensional Spaces: Effective when features >> samples (e.g., bioinformatics, text classification) due to margin maximization.

Instance-Based Learning: K-Nearest Neighbors (K-NN)

  • Algorithm:

    1. Store all training examples.

    2. For a new query point, compute distance (Euclidean, Manhattan, etc.) to all training points.

    3. Select k nearest neighbors.

    4. Classification: Majority vote of neighbors' labels.

    5. Regression: Average of neighbors' target values.

  • Key Parameters: k (number of neighbors), distance metric.

Decision Trees

  • Structure: Hierarchical tree of nodes (root, internal, leaf). Splits based on feature values.

  • Algorithms: ID3 (uses Information Gain), C4.5 (uses Gain Ratio), CART (uses Gini Impurity).

  • Entropy & Information Gain:

    • Entropy (measure of impurity): $$\displaystyle Entropy(t) = -\sum_{i=1}^{c} p_i(t) \log_2 p_i(t) $$.

    • Information Gain: $$\displaystyle IG(T, a) = Entropy(T) - \sum_{v \in Values(a)} \frac{|T_v|}{|T|} Entropy(T_v) $$. Split on feature with highest IG.

  • Issues: Overfitting (deep trees), instability (small data changes alter tree), bias towards features with many levels.

  • Random Forests:

    • Ensemble of decision trees using bagging (bootstrap aggregating) and feature randomness.

    • Reduces variance, improves generalization vs. single tree.

    • Bagging vs. Boosting: Bagging (parallel, reduces variance), Boosting (sequential, reduces bias, e.g., AdaBoost, Gradient Boosting).

Evaluation Metrics

  • Regression:

    • MSE: $$\displaystyle \frac{1}{n} \sum (y_i - \hat{y}_i)^2 $$ (sensitive to outliers).

    • MAE: $$\displaystyle \frac{1}{n} \sum |y_i - \hat{y}_i| $$ (robust).

    • Rยฒ (Coefficient of Determination): $$\displaystyle 1 - \frac{\sum (y_i - \hat{y}_i)^2}{\sum (y_i - \bar{y})^2} $$ (proportion of variance explained).

  • Classification:

    • Confusion Matrix:

      | | Predicted + | Predicted - | |----------------|-------------|-------------| | Actual + | TP | FN | | Actual - | FP | TN |

    • Derived Metrics:

      • Accuracy: $$\displaystyle \frac{TP+TN}{Total} $$

      • Precision: $$\displaystyle \frac{TP}{TP+FP} $$ (of predicted positives, how many correct?)

      • Recall (Sensitivity): $$\displaystyle \frac{TP}{TP+FN} $$ (of actual positives, how many found?)

      • F1-Score: $$\displaystyle 2 \times \frac{Precision \times Recall}{Precision + Recall} $$ (harmonic mean).

  • Resampling Methods:

    • k-fold Cross-Validation: Split data into k folds; train on k-1, test on 1; repeat k times; average performance.

    • Bootstrap: Sample with replacement; estimate performance on out-of-bag samples.


4. UNSUPERVISED LEARNING

Clustering

  • Goals: Group similar instances together, discover inherent structure, reduce data complexity.

  • K-Means Clustering:

    • Algorithm:

      1. Initialize k centroids (randomly or k-means++).

      2. Repeat until convergence:

        • Assignment: Assign each point to nearest centroid (Euclidean distance).

        • Update: Recompute centroids as mean of assigned points.

    • Properties: Sensitive to initialization, assumes spherical clusters, requires k.

  • Expectation-Maximization (EM) Algorithm:

    • E-step: Estimate "missing" data (e.g., cluster assignments) given current parameters.

    • M-step: Update model parameters to maximize likelihood given estimated data.

    • Application to Gaussian Mixture Models (GMM): Assumes data from mixture of Gaussians; EM estimates means, covariances, mixing coefficients.

    • Handling Missing Data: EM iteratively fills missing values (E-step) and re-trains model (M-step).

  • Hierarchical Clustering:

    • DIANA (Divisive): Top-down; start with all points in one cluster, recursively split.

    • AGNES (Agglomerative): Bottom-up; start with each point as cluster, merge closest clusters.

    • Linkage Criteria: Single (min distance), Complete (max distance), Average.

  • Density-Based Clustering (DBSCAN): Groups dense regions separated by sparse areas; handles arbitrary shapes, noise.

  • BIRCH Algorithm: Builds CF (Clustering Feature) tree for incremental clustering; efficient for large datasets.


5. NEURAL NETWORKS & DEEP LEARNING FUNDAMENTALS

Artificial Neural Network (ANN) Basics

  • Biological Inspiration: Neurons with dendrites (inputs), soma (processing), axon (output).

  • Artificial Neuron (Perceptron): Computes weighted sum + bias, applies activation function.

$$z = w^T x + b, \quad a = f(z)$$

  • Multi-Layer Perceptron (MLP): Input layer, one or more hidden layers, output layer. Fully connected.

Activation Functions

Function Formula Pros Cons
Sigmoid $$\displaystyle \sigma(z) = \frac{1}{1+e^{-z}} $$ Smooth, output (0,1) Vanishing gradient, not zero-centered
Tanh $$\displaystyle \tanh(z) = \frac{e^z - e^{-z}}{e^z + e^{-z}} $$ Zero-centered, steeper than sigmoid Vanishing gradient
ReLU $$\displaystyle f(z) = \max(0, z) $$ Computationally cheap, alleviates vanishing gradient Dying ReLU (neurons stuck at 0)
Leaky ReLU $$\displaystyle f(z) = \max(\alpha z, z) $$ Fixes dying ReLU Needs tuning of $\alpha$

Vanishing Gradient Problem: Gradients become extremely small in deep networks during backpropagation, especially with sigmoid/tanh, hindering learning in early layers.

Training Neural Networks

  • Gradient Descent:

    • Batch GD: Use entire dataset per update (stable, slow).

    • Stochastic GD (SGD): Use one sample per update (noisy, fast).

    • Mini-batch GD: Use small batch (compromise, common in practice).

  • Backpropagation Algorithm:

    1. Forward Pass: Compute output and loss for a batch.

    2. Backward Pass:

      • Compute gradient of loss w.r.t. output layer weights (using chain rule).

      • Propagate gradients backward through layers.

      • Update weights: $$\displaystyle w_{ij} := w_{ij} - \alpha \frac{\partial Loss}{\partial w_{ij}} $$.

    • Chain Rule Application: $$\displaystyle \frac{\partial Loss}{\partial w_{ij}} = \frac{\partial Loss}{\partial a_j} \cdot \frac{\partial a_j}{\partial z_j} \cdot \frac{\partial z_j}{\partial w_{ij}} $$, where $$\displaystyle a_j $$ is activation, $$\displaystyle z_j $$ is weighted sum.
  • Loss Functions:

    • Regression: Mean Squared Error (MSE).

    • Classification: Cross-Entropy Loss (Binary: $$\displaystyle -\frac{1}{m}\sum [y^{(i)}\log(\hat{y}^{(i)}) + (1-y^{(i)})\log(1-\hat{y}^{(i)})] $$; Categorical: $$\displaystyle -\sum y_i \log(\hat{y}_i) $$).

  • Optimizers:

    • SGD: Basic update.

    • Adam: Adaptive learning rate, momentum, RMSprop combination; widely used.

    • RMSprop: Adapts learning rate per parameter, divides by root of squared gradients.

Regularization Techniques

  • L1 (Lasso) vs. L2 (Ridge):

    • L1: Adds $$\displaystyle \lambda \sum |w_j| $$ to loss. Drives some weights to exactly zero โ†’ sparsity (feature selection).

    • L2: Adds $$\displaystyle \lambda \sum w_j^2 $$ to loss. Shrinks weights toward zero but rarely to zero โ†’ weight decay.

    • Both: Prevent overfitting by penalizing large weights.

  • Dropout: Randomly "drop" (set to zero) a fraction of neurons during training; prevents co-adaptation, acts as ensemble.

  • Batch Normalization: Normalizes layer inputs to zero mean, unit variance; stabilizes training, allows higher learning rates, slight regularization effect.


6. CONVOLUTIONAL NEURAL NETWORKS (CNNs)

Core Architecture & Motivation

  • Why for Images?: Exploits spatial locality (nearby pixels related) and translation invariance (object identity independent of position).

  • Operations: Convolution (feature extraction) โ†’ Activation (non-linearity) โ†’ Pooling/Sub-sampling (invariance, dimensionality reduction).

Key Layers & Concepts

  • Convolutional Layer:

    • Applies filters/kernels (small weight matrices) to input, producing feature maps.

    • Stride: Step size of filter movement.

    • Output size: $$\displaystyle \frac{W - F + 2P}{S} + 1 $$, where $W$=input width, $F$=filter size, $P$=padding, $S$=stride.

  • $1 \times 1$ Convolution:

    • Purpose: Feature reduction/channel mixing without spatial change. Used in architectures like Inception to reduce computational cost and increase non-linearity.

    • Acts as a per-pixel fully connected layer across channels.

  • Pooling/Sub-sampling Layer:

    • Max Pooling: Takes maximum value in window โ†’ retains dominant feature, provides translation invariance.

    • Average Pooling: Takes average โ†’ smooths features.

    • Reduces spatial dimensions, number of parameters, and computation.

  • Fully Connected (FC) Layer: At end; flattens features for final classification/regression.

Architectural Design Choices

  • Padding:

    • Valid: No padding; output shrinks.

    • Same: Padding such that output size equals input size (preserves spatial dimensions).

  • Flattening Layer: Converts multi-dimensional feature maps to 1D vector for FC layers.

Famous Architectures & Concepts

  • Inception Module:

    • Uses parallel convolutions of different sizes ($1\times1$, $3\times3$, $5\times5$) and pooling, concatenates outputs.

    • Benefits: Multi-scale feature capture, efficient use of parameters (via $1\times1$ bottlenecks).

  • Transfer Learning:

    • Feature Extraction: Use pre-trained CNN (e.g., ImageNet) as fixed feature extractor; train only new classifier.

    • Fine-tuning: Unfreeze and train some top layers of pre-trained network along with new classifier.

  • ImageNet Competition: Drove CNN innovation (AlexNet 2012, VGG, GoogLeNet/Inception, ResNet). Demonstrated deep CNNs' power.

Practical Considerations

  • Overfitting in CNNs: Use data augmentation, dropout, weight decay, early stopping.

  • Underfitting: Deeper/wider network, more filters, reduce regularization.

  • Implementation (TensorFlow/Keras): Sequential/Functional API, layers (Conv2D, MaxPooling2D, Dense).


7. RECURRENT NEURAL NETWORKS (RNNs) & SEQUENCE MODELS

Vanilla RNN

  • Architecture: Has hidden state $$\displaystyle h_t $$ that captures past information. $$\displaystyle h_t = f(W_{xh} x_t + W_{hh} h_{t-1} + b_h) $$.

  • Limitations:

    • Vanishing/Exploding Gradients: Gradients shrink/explode over long sequences due to repeated multiplication.

    • Short-term Memory: Struggles with long-range dependencies.

Long Short-Term Memory (LSTM)

  • Architecture:

    • Cell State ($$\displaystyle C_t $$): "Conveyor belt" carrying information with minimal interference.

    • Gates (sigmoid + pointwise multiplication):

      • Forget Gate: Decides what to drop from cell state.

      • Input Gate: Decides what new information to store.

      • Output Gate: Decides what to output based on cell state.

  • Handles Long-Term Dependencies: Cell state allows gradients to flow easily over many time steps.

  • Role in NLP: Predecessor to Transformers; used in language modeling, machine translation (e.g., seq2seq with LSTM encoder-decoder).

Gated Recurrent Unit (GRU)

  • Simplified LSTM: Combines forget and input gates into update gate; no separate cell state (hidden state serves both).

  • Comparison: Fewer parameters, often similar performance; faster training.

Applications in NLP & Speech

  • Speech Processing:

    • Speech-to-Text (ASR): Convert audio waveform to text (RNN/CNN + CTC loss).

    • Speaker Identification: Classify speaker from voice features.

  • NLP Pipeline: Tokenization, stemming/lemmatization, stop-word removal, n-grams.

  • BLEU Score (for machine translation):

    • n-gram Precision: Modified precision for n-grams (clips counts to avoid overcounting).

    • Brevity Penalty: Penalizes translations shorter than reference. $$\displaystyle BP = \min(1, e^{1 - r/c}) $$, where $r$=reference length, $c$=candidate length.

    • Final: $$\displaystyle BLEU = BP \cdot \exp(\sum_{n=1}^{N} w_n \log p_n) $$.


8. REINFORCEMENT LEARNING (RL)

Fundamental Framework: Markov Decision Process (MDP)

  • Defined by $(S, A, P, R, \gamma)$:

    • $S$: Set of states.

    • $A$: Set of actions.

    • $P(s' | s, a)$: Transition probability to state $s'$ from $s$ taking $a$.

    • $R(s, a, s')$: Reward received.

    • $\gamma$: Discount factor ($0 \leq \gamma \leq 1$).

  • Policy ($\pi$): Strategy mapping states to actions ($$\displaystyle a = \pi(s) $$).

  • Difference from Supervised/Unsupervised: No labeled data; agent learns via interaction and delayed rewards.

Core Algorithms

  • Value Iteration:

    • Iteratively updates state values: $$\displaystyle V_{k+1}(s) = \max_a \sum_{s'} P(s'|s,a)[R(s,a,s') + \gamma V_k(s')] $$.

    • Directly computes optimal value function; no explicit policy.

  • Policy Iteration:

    • Alternates between Policy Evaluation (compute $$\displaystyle V^\pi $$) and Policy Improvement ($$\displaystyle \pi' = \arg\max_a \sum_{s'} P(s'|s,\pi(s))[R + \gamma V^\pi(s')] $$).

    • Guarantees convergence to optimal policy.

  • Q-Learning (Off-policy TD Control):

    • Learns Q-value $Q(s,a)$: expected cumulative reward taking $a$ in $s$ then following optimal policy.

    • Update Rule: $$\displaystyle Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma \max_{a'} Q(s',a') - Q(s,a)] $$.

    • Uses greedy actions for update regardless of behavior policy.

  • SARSA (On-policy TD Control):

    • Updates $Q$ based on actual next action taken: $$\displaystyle Q(s,a) \leftarrow Q(s,a) + \alpha [r + \gamma Q(s',a') - Q(s,a)] $$, where $a' \sim \pi(s')$.

    • More conservative; learns policy being executed.

Advanced Actor-Critic Models

  • Actor-Critic:

    • Actor: Policy network ($$\displaystyle \pi_\theta(a|s) $$) selects actions.

    • Critic: Value network ($$\displaystyle V_w(s) $$ or $$\displaystyle Q_w(s,a) $$) evaluates actions.

    • Interaction: Actor proposes action, Critic computes TD error $$\displaystyle \delta = r + \gamma V_w(s') - V_w(s) $$, which guides Actor update (policy gradient) and Critic update (value loss).

  • Rise of Advanced Variants:

    • A2C/A3C: Asynchronous Advantage Actor-Critic (parallel actors).

    • PPO (Proximal Policy Optimization): Constrains policy updates for stability; state-of-the-art for many tasks.

Practical Frameworks & Applications

  • Frameworks: OpenAI Gym (environments), Stable Baselines3, TensorFlow Agents, Ray RLlib.

  • Applications: Game playing (Atari, Go), robotics control, resource management, autonomous driving.


9. ADVANCED TOPICS & APPLICATIONS

Autoencoders

  • Architecture: Encoder (compresses input to latent code), Decoder (reconstructs input from code).

  • Loss: Reconstruction error (MSE, cross-entropy).

  • Uses in Unsupervised Learning:

    • Dimensionality Reduction: Latent space is lower-dimensional representation.

    • Denoising: Train with corrupted inputs, reconstruct clean outputs.

    • Anomaly Detection: High reconstruction error indicates anomaly.

One-Shot Learning

  • Definition: Learning from very few examples (often one or few per class).

  • Difference from Traditional Supervised: Requires massive data; one-shot learns efficiently with prior knowledge (e.g., via meta-learning, metric learning).

  • Beneficial Applications: Rare event classification (medical diagnosis), new object recognition, personalization.

Self-Supervised Learning

  • Concept: Automatically generate supervisory signals from data itself (pretext tasks).

  • Examples: Predict image rotation, mask parts of image/text and recover (BERT, MAE).

  • Rise: Reduces need for labeled data; pre-trains powerful representations.

Generative AI

  • Overview: Models that generate new data samples resembling training data.

  • Connections:

    • Autoencoders: Variational Autoencoders (VAEs) generate samples.

    • GANs (Generative Adversarial Networks): Generator vs. Discriminator game.

    • LLMs (Large Language Models): Transformer-based, trained on vast text corpora (GPT, BERT).

Application Domains

  • Computer Vision:

    • Object Detection: Locate & classify objects (YOLO, Faster R-CNN).

    • Image Segmentation: Pixel-level labeling (U-Net, Mask R-CNN).

  • Natural Language Processing (NLP):

    • Tasks: Translation, sentiment analysis, summarization, question answering.

    • Transformers: Attention mechanism (context for LSTM/RNN); BERT, GPT.

  • Speech Processing:

    • ASR (Automatic Speech Recognition): Audio โ†’ Text (DeepSpeech, Wav2Vec).

    • Synthesis (TTS): Text โ†’ Speech (Tacotron, WaveNet).

    • Speaker Identification: Verify/identify speaker from voice.

  • Bioinformatics: SVM for gene classification, protein structure prediction; CNNs for medical image analysis.


10. OPTIMIZATION & THEORETICAL UNDERPINNINGS

Convex Optimization

  • Importance: If loss function is convex, gradient descent is guaranteed to find global optimum (no local minima).

  • Many ML problems (linear regression, logistic regression) have convex loss; deep learning often non-convex.

Linearity vs. Non-linearity

  • Linearity: Model is linear in parameters (e.g., linear regression). Can be solved analytically; limited expressiveness.

  • Non-linearity: Introduced via activation functions or feature transformations. Enables modeling complex relationships but requires iterative optimization (gradient descent); prone to vanishing/exploding gradients.

Probability & Statistics in ML

  • Role: Provides framework for uncertainty, decision making, and probabilistic models (Naive Bayes, HMMs, Bayesian Networks).

  • Bayes' Theorem:

$$P(A|B) = \frac{P(B|A) P(A)}{P(B)}$$

*   **Use in Classification (Naive Bayes)**: Assumes feature independence given class. $$\displaystyle P(\text{class}|\text{features}) \propto P(\text{class}) \prod P(\text{feature}|\text{class}) $$.

*   **Fundamental**: Underlies Bayesian inference, updating beliefs with evidence.
  • Bayesian Learning:

    • Treats parameters as random variables with prior distributions.

    • Posterior $P(\theta|D) \propto P(D|\theta) P(\theta)$.

    • Allows incorporation of prior knowledge, quantifies uncertainty.

  • Bayesian Networks: Directed acyclic graph representing conditional dependencies among variables. Nodes = random variables, edges = conditional dependencies. Used for probabilistic inference.

Loss Calculation

  • Critical Role: Defines objective function to minimize. Guides backpropagation by providing gradient signal.

  • Choice Depends on Task: MSE for regression, cross-entropy for classification, hinge loss for SVM.

  • Directly influences model's learned representation and generalization.

Exam Tips & Common Pitfalls:

  • Confusion Matrix: Always label rows (Actual) and columns (Predicted). Precision = PPV, Recall = Sensitivity.
  • PCA vs. Feature Selection: PCA creates new features; selection picks existing ones.
  • L1 vs. L2: L1 gives sparsity (zero weights), L2 gives small weights.
  • Q-learning vs. SARSA: Off-policy vs. on-policy; Q uses $$\displaystyle \max_{a'} Q(s',a') $$, SARSA uses $Q(s',a')$ from current policy.
  • Backpropagation: Chain rule applied layer-by-layer from output to input; compute gradients, then update weights.
  • CNN Padding: "Same" preserves spatial size; "Valid" reduces size.
  • Activation Functions: ReLU most common today; sigmoid/tanh in output layers for probability.
  • MDP: Must satisfy Markov property (future depends only on current state, not history).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in