Skip to content
CS-601 · Machine Learning/Quick Revision Short Notes

Machine Learning (CS-601) - Unit 1 Short Notes

How unit 1 is examined

This unit covers what machine learning is, the maths behind it, data preparation and the main learning types; regression, convex optimization, preprocessing, models and supervised learning carry the marks.

Introduction to machine learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Machine learning is the study of algorithms that improve their performance P at a task T with experience E (Tom Mitchell), learning rules from data instead of being explicitly programmed.</mark>

Key points.

  1. Example: a spam filter learns from thousands of emails labelled spam or not spam, so T is classifying mail, E is the labelled emails and P is the percentage classified correctly.
  2. Traditional programming takes data and hand-written rules and produces output; machine learning takes data and output (labels) and produces the rules (the model).
  3. Key components are the dataset, features, model, loss function, optimizer and evaluation.
  4. Perspectives: the statistical view treats learning as estimating a function from samples, the computational view as searching a hypothesis space efficiently, and the AI view as a machine improving with experience.
  5. Design issues: choice of training data and experience, choice of target function, its representation (linear, tree, network) and the function approximation algorithm.
  6. Issues: poor data quality, overfitting and underfitting, high dimensionality, and the bias-variance trade-off.
Point Traditional programming Machine learning
Input Data + rules Data + expected output
Output Answers Rules (model)
Change of behaviour Programmer rewrites code Retrain on new data
Best for Well-defined logic Patterns too complex to code

Answer frame. Open with Mitchell's definition and the spam example; develop the comparison table, then components, perspectives and issues; close with generalization as the aim.

Asked: [7 marks] (Dec 2020, May 2022) Explain the machine learning concept by taking an example. Describe the perspective and issues in machine learning. / What are the basic design issues and approaches to machine learning? Asked: [7 marks] (Dec 2024) Define machine learning and differentiate it from traditional programming approaches. Identify the key components of machine learning.

Scope and limitations

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. Machine learning is the ability of a system to learn patterns from data and improve without explicit programming; its scope is the set of problems it can solve, its limitations the conditions under which it fails.

Key points.

  1. Scope covers the learning types (supervised, unsupervised, semi-supervised, reinforcement) and tasks such as classification, regression, clustering and decision making.
  2. Applications: spam filtering, recommendation systems, medical diagnosis, fraud detection, speech and image recognition, machine translation and self-driving cars.
  3. Limitation: it needs large, good-quality data, and poor or biased data gives poor or biased models.
  4. Limitation: models may overfit and fail to generalize to data unlike the training data.
  5. Limitation: complex models such as deep networks are black boxes, so decisions are hard to explain.

Answer frame. Open with the definition; list scope and applications, then limitations in the order data, generalization, interpretability; close that ML is powerful but needs good data and care.

Asked: [8 marks] (Jun 2025, Jun 2026) Define Machine Learning? Discuss the scope and limitations of Machine Learning. / Explain the scope, limitations and applications of Machine Learning.

Regression

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Regression is a supervised learning method that predicts a continuous output $y$ from input features $x$ by fitting a hypothesis function $h_\theta(x)$ that minimizes the error between prediction and actual value.</mark>

Formula.

$$h_\theta(x)=\theta_0+\theta_1x,\qquad J(\theta)=\frac{1}{2m}\sum_{i=1}^{m}\big(h_\theta(x^{(i)})-y^{(i)}\big)^2$$

Parameter update by gradient descent: $\theta_j := \theta_j-\alpha\,\dfrac{\partial J}{\partial\theta_j}$. For one feature the least-squares fit is $\theta_1=\dfrac{\sum(x-\bar x)(y-\bar y)}{\sum(x-\bar x)^2}$, $\theta_0=\bar y-\theta_1\bar x$.

Key points (linear regression).

  1. Linear regression assumes the output is a linear combination of the inputs, so the model is a straight line (or hyperplane) $h_\theta(x)=\theta^Tx$.
  2. The hypothesis function $h_\theta(x)$ maps an input to a prediction, and training means choosing the parameters $\theta$.
  3. The cost function $J(\theta)$ measures the average squared error over $m$ training examples; it is convex, so it has a single global minimum.
  4. Parameters are learned by gradient descent (repeat the update until $J$ stops falling) or directly by the normal equation.

Example. Data $(x,y)$: (1,2), (2,3), (3,5), (4,6). Here $\bar x=2.5$, $\bar y=4$, $\sum(x-\bar x)(y-\bar y)=7$, $\sum(x-\bar x)^2=5$, so $\theta_1=1.4$ and $\theta_0=4-1.4(2.5)=0.5$. Fitted line: $h(x)=0.5+1.4x$, so the prediction at $x=5$ is 7.5.

Locally weighted linear regression (LWR). LWR is a non-parametric method: it keeps the training data and fits a fresh line around each query point $x$.

  1. Each training example gets a weight from a kernel, $w^{(i)}=\exp\!\big(-\dfrac{(x^{(i)}-x)^2}{2\tau^2}\big)$, so nearby points weigh near 1 and distant points near 0.
  2. The bandwidth $\tau$ controls how quickly the weight falls: small $\tau$ gives a wiggly fit, large $\tau$ approaches ordinary regression.
  3. The cost is $J(\theta)=\sum_i w^{(i)}\big(\theta^Tx^{(i)}-y^{(i)}\big)^2$, and $\theta$ is refit for every query; the update is $\theta_j:=\theta_j-\alpha\sum_i w^{(i)}(h_\theta(x^{(i)})-y^{(i)})x_j^{(i)}$.

Evaluation metrics.

Metric Formula Property
MSE $\frac1n\sum(y-\hat y)^2$ Squares errors, so sensitive to outliers
MAE $\frac1n\sum\lvert y-\hat y\rvert$ Robust to outliers
RMSE $\sqrt{\text{MSE}}$ Same units as the target
$R^2$ $1-\frac{\sum(y-\hat y)^2}{\sum(y-\bar y)^2}$ Fraction of variance explained, 0 to 1
MAPE $\frac{100}{n}\sum\lvert\frac{y-\hat y}{y}\rvert$ Percentage error, comparable across scales

Role in ML. Regression predicts numeric values; probability models uncertainty and noise; statistics supplies estimation, hypothesis tests and the bias-variance view used to judge models.

Answer frame. Linear regression: define with $h_\theta(x)$, sketch scatter points with the fitted line, give cost, update and the example, close with the role of the hypothesis. LWR: non-parametric, kernel weight, cost, update. Metrics: one line each. Role question: one paragraph each for regression, probability, statistics, then a summary.

Asked: [7 marks] (Dec 2020) Discuss linear regression with an example. Explain the role of hypothesis function in machine learning models. Asked: [7 marks] (May 2022) Explain locally weighted linear regression. Asked: [7 marks] (May 2024) Describe common evaluation metrics used for regression models of machine learning. Asked: [6 marks] (Jun 2025) Explain the role of regression, probability and statistics in Machine Learning.

Probability

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Probability measures how likely an event is, between 0 and 1, and gives machine learning a language for uncertainty.

Key points.

  1. Conditional probability is $P(A\mid B)=P(A\cap B)/P(B)$, and Bayes' theorem is $P(A\mid B)=\dfrac{P(B\mid A)P(A)}{P(B)}$.
  2. A random variable has a distribution, and its expectation $E[X]=\sum xP(x)$ is the average value.

Statistics

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. Statistical learning theory studies how a model learned from a finite sample generalizes to unseen data.

Key points.

  1. Empirical risk is the average loss on the training sample; empirical risk minimization picks the hypothesis with the smallest one.
  2. True (expected) risk is the loss over the whole distribution, and the gap between the two is the generalization error.
  3. Structural risk minimization adds a complexity penalty to empirical risk, which guards against overfitting and guides model selection.

Asked: [7 marks] (May 2022) Define statistical theory and how it is performed in machine learning.

Linear algebra for machine learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Linear algebra is the maths of vectors and matrices, which is how data and model parameters are stored and transformed.

Key points.

  1. A data point is a vector, a dataset is a matrix $X$ of $m$ rows (samples) by $n$ columns (features), and predictions are $\hat y=X\theta$.
  2. Key operations are the dot product, matrix multiplication, transpose and inverse; the normal equation is $\theta=(X^TX)^{-1}X^Ty$.
  3. Eigenvalues and eigenvectors ($Av=\lambda v$) are the basis of PCA for dimensionality reduction.

Convex optimization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>A function is convex if the line segment between any two points on its graph lies on or above the graph, so every local minimum is the global minimum; convex optimization minimizes such a function.</mark>

Key points.

  1. The condition is $f(\lambda a+(1-\lambda)b)\le\lambda f(a)+(1-\lambda)f(b)$ for $0\le\lambda\le1$; for a twice-differentiable function, $f''(x)\ge0$.
  2. Gradient descent on a convex function is guaranteed to reach the global minimum, with no bad local minima.
  3. In ML, the squared-error cost of linear regression, the logistic loss and the SVM hinge loss are convex, so they are solved reliably.
  4. Deep network losses are non-convex, so training there depends on initialization and finds only good local minima.
  5. The update is $\theta:=\theta-\alpha\nabla J(\theta)$ with a suitable learning rate $\alpha$.

The other terms in the question (one example each).

  • Multilayer network: a neural network with an input layer, one or more hidden layers and an output layer; hidden layers learn non-linear features, for example digit recognition. It is detailed in Unit 2.
  • Attention model: a mechanism that lets a network weight the parts of the input most relevant to each output step, for example the source words when translating a sentence. It is detailed in Unit 4.
  • Natural Language Processing: the use of ML to let computers understand and generate human language, for example chatbots, translation and sentiment analysis. It is detailed in Unit 5.

Answer frame. Give each part about 3.5 marks: define convex optimization, points 1-4, gradient update; for each other term a definition, working, example and application.

Asked: [14 marks] (May 2022) Explain the following terms with example: i) Convex optimization in machine learning ii) Multilayer network iii) Attention model iv) Natural Language Processing

Data visualization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. Data visualization is the graphical presentation of data to reveal patterns, trends, outliers and relationships before and after modelling.

Key points.

  1. It is needed to understand the data, spot outliers and missing values, choose features and communicate results.
  2. A bar chart compares categories; a line chart shows a trend over time; a scatter plot shows the relationship between two variables.
  3. A histogram shows the distribution of one variable; a box plot shows median, spread and outliers; a heatmap shows a correlation matrix or a confusion matrix.
  4. Common tools are Matplotlib, Seaborn, Plotly, Tableau and Power BI.

Asked: [9 marks] (Jun 2025) Explain the different Data Visualization methods in detail.

Hypothesis function and testing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>A hypothesis function $h(x)$ is the model's mapping from an input feature vector $x$ to a predicted output $\hat y$; the hypothesis space is the set of all functions the algorithm can choose from.</mark>

Key points.

  1. The input is a feature vector $x=(x_1,\dots,x_n)$, and the hypothesis computes a prediction, for example $h_\theta(x)=\theta_0+\theta_1x_1+\dots+\theta_nx_n$ for regression.
  2. For classification the hypothesis outputs a class or probability, for example $h_\theta(x)=\sigma(\theta^Tx)=1/(1+e^{-\theta^Tx})$ in logistic regression.
  3. Training searches the hypothesis space for the parameters $\theta$ that minimize the loss between $h(x)$ and the true $y$.
  4. Hypothesis testing is the statistical procedure: state the null $H_0$ and alternative $H_1$, choose the significance level $\alpha$, compute the test statistic and p-value, and reject $H_0$ if $p<\alpha$.
  5. A Type I error rejects a true $H_0$ (false positive); a Type II error accepts a false $H_0$ (false negative).

Answer frame. Define $h(x)$ and hypothesis space; draw input, h(x), prediction; regression and classification examples, then training. Three-part question: add testing steps, two errors and data distributions.

Asked: [7 marks] (Dec 2024) Explain how a hypothesis function maps input features to output predictions in a machine learning model. Asked: [7 marks] (Jun 2026) Discuss hypothesis functions, hypothesis testing and data distributions.

Data distributions

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. A data distribution describes how the values of a variable are spread, that is, which values occur and how often.

Key points.

  1. The normal (Gaussian) distribution is bell-shaped with mean $\mu$ and standard deviation $\sigma$: $f(x)=\frac{1}{\sigma\sqrt{2\pi}}e^{-(x-\mu)^2/2\sigma^2}$.
  2. Models assume the training and test data come from the same distribution; when it shifts, generalization fails.

Data preprocessing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Data preprocessing is the process of cleaning and transforming raw data into a form a machine learning algorithm can learn from, since raw data is noisy, incomplete and inconsistent.</mark>

Steps.

Step 1: Cleaning - fill or drop missing values, remove duplicates and fix outliers.
Step 2: Integration - merge data from several sources into one dataset.
Step 3: Transformation - encode categories, scale features, build new features.
Step 4: Reduction - drop irrelevant features or apply PCA to reduce dimensions.
Step 5: Split - divide into training, validation and test sets.

Key points.

  1. Missing values are handled by deleting rows or by filling with the mean, median, mode or a model prediction.
  2. Categorical text cannot enter a model, so it is encoded into numbers.
  3. Label encoding assigns each category an integer, for example Red=0, Green=1, Blue=2; it adds no columns, but it implies a false order.
  4. One-hot encoding makes one binary column per category, for example Red=(1,0,0), Green=(0,1,0), Blue=(0,0,1); it implies no order but adds columns.
Point Label encoding One-hot encoding
Output One integer per category One binary vector per category
Dimensionality Unchanged, one column Rises from 1 column to $k$ columns for $k$ categories
Order implied Yes, wrongly No
Suits Ordinal data, tree models Nominal data, linear and distance models

Answer frame. Encoding: define both with the colour example, table, close with the dimensionality effect. General: define, five steps, missing values, encoding, importance. The Jun 2026 question adds the next three topics, one paragraph each.

Asked: [7 marks] (May 2023) Explain One-hot encoding and Label encoding. In what ways do they change the dimensionality of the data? Asked: [5 marks] (Jun 2025) Explain data preprocessing in detail. Asked: [7 marks] (Jun 2026) Explain data preprocessing, data augmentation, normalization and data visualization in Machine Learning.

Data augmentation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. Data augmentation artificially enlarges and diversifies a training set by applying label-preserving transformations to existing samples.

Key points.

  1. Its purpose is to increase the size and diversity of data, which reduces overfitting when data is scarce.
  2. For images the common transforms are rotation, flipping, cropping, scaling, translation, brightness and colour jitter, and noise; for text, synonym replacement and back-translation.
  3. Steps: select suitable transforms, apply them to the training set only, then validate that the labels are still correct.

Asked: [7 marks] (Dec 2024) Describe the process of applying data augmentation techniques to expand the size and diversity of a dataset.

Normalizing data sets

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. Normalization rescales features to a common range or distribution so no feature dominates because of its units.

Formula. Min-max: $x'=\dfrac{x-x_{min}}{x_{max}-x_{min}}$ (range 0 to 1). Z-score: $x'=\dfrac{x-\mu}{\sigma}$ (mean 0, standard deviation 1).

Key points.

  1. Features on large scales, such as salary against age, would otherwise dominate the model.
  2. Gradient descent converges faster because the loss surface becomes more spherical, and larger learning rates become safe.
  3. Scaling avoids numerical overflow and underflow and keeps training stable.
  4. Distance-based algorithms (KNN, K-means, SVM) need it so every feature contributes fairly.
  5. Fit the scaler on training data only, then apply it to the test data.

Asked: [7 marks] (May 2024) Explain the importance of data normalization in machine learning for improving model convergence, stability and performance.

Machine learning models

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. A machine learning model is the function, learned from data, that maps inputs to outputs; the algorithm is the procedure that learns it.

Diagram.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-01" viewBox="0 0 886 194" width="886" height="194" role="img" aria-label="Types of machine learning algorithms"><style>#dsfig-u1-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-01 .t{fill:#16181D;font-weight:500}#dsfig-u1-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-01 .dot{fill:#16181D}#dsfig-u1-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-01 .ah{fill:#454C5A}#dsfig-u1-01 .ah.hi{fill:#2340B8}#dsfig-u1-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-01 .e{stroke:#B1B7C3}html.dark #dsfig-u1-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-01 .t{fill:#E6E8ED}html.dark #dsfig-u1-01 .t.inv{fill:#0F1115}html.dark #dsfig-u1-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-01 .dot{fill:#E6E8ED}html.dark #dsfig-u1-01 .ann{fill:#8FA3FF}html.dark #dsfig-u1-01 .lbl{fill:#858D9C}html.dark #dsfig-u1-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-01 .ah{fill:#B1B7C3}html.dark #dsfig-u1-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="457.5" y1="37" x2="128" y2="101"/><line class="e" x1="457.5" y1="37" x2="397.8" y2="101"/><line class="e" x1="457.5" y1="37" x2="641.5" y2="101"/><line class="e" x1="457.5" y1="37" x2="787" y2="101"/><line class="e" x1="128" y1="101" x2="63" y2="165"/><line class="e" x1="128" y1="101" x2="193" y2="165"/><line class="e" x1="397.8" y1="101" x2="323" y2="165"/><line class="e" x1="397.8" y1="101" x2="472.5" y2="165"/><rect class="n" x="385" y="22" width="145" height="30" rx="8"/><text class="t" x="457.5" y="37" dy=".35em" text-anchor="middle">Machine Learning</text><rect class="n" x="79" y="86" width="98" height="30" rx="8"/><text class="t" x="128" y="101" dy=".35em" text-anchor="middle">Supervised</text><rect class="n" x="14" y="150" width="98" height="30" rx="8"/><text class="t" x="63" y="165" dy=".35em" text-anchor="middle">Regression</text><rect class="n" x="128" y="150" width="130" height="30" rx="8"/><text class="t" x="193" y="165" dy=".35em" text-anchor="middle">Classification</text><rect class="n" x="340.8" y="86" width="114" height="30" rx="8"/><text class="t" x="397.8" y="101" dy=".35em" text-anchor="middle">Unsupervised</text><rect class="n" x="274" y="150" width="98" height="30" rx="8"/><text class="t" x="323" y="165" dy=".35em" text-anchor="middle">Clustering</text><rect class="n" x="388" y="150" width="169" height="30" rx="8"/><text class="t" x="472.5" y="165" dy=".35em" text-anchor="middle">Dimension reduction</text><rect class="n" x="573" y="86" width="137" height="30" rx="8"/><text class="t" x="641.5" y="101" dy=".35em" text-anchor="middle">Semi-supervised</text><rect class="n" x="726" y="86" width="122" height="30" rx="8"/><text class="t" x="787" y="101" dy=".35em" text-anchor="middle">Reinforcement</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Types of machine learning algorithms</figcaption></figure>

Key points.

  1. Supervised: linear regression (house price), classification (email spam) and decision trees (loan approval).
  2. Unsupervised: clustering (customer segmentation with K-means) and PCA (feature extraction).
  3. Semi-supervised: a few labelled and many unlabelled samples, for example text classification with limited labels.
  4. Reinforcement: an agent learns by reward, for example game playing (AlphaGo), robotics control and autonomous vehicles.
  5. Model selection is choosing the best model and hyperparameters for the problem.
  6. Data is split into training (fit), validation (tune and select) and test (final unbiased estimate) sets, and k-fold cross-validation rotates the validation fold to use the data better.
  7. Selection balances bias (underfitting) against variance (overfitting), using criteria such as validation error, AIC and BIC.
Point Traditional ML Deep learning
Features Hand-crafted by experts Learned automatically (representation learning)
Data needed Works with small data Needs large data
Compute Low, CPU is enough High, GPU needed
Interpretability Higher (trees, regression) Low, a black box
Examples SVM, decision tree, random forest CNN for images, RNN for speech

Answer frame. Model selection: define, split, cross-validation, bias-variance. Types: one line each with application. Comparison: the table, then examples (spam filter versus face recognition).

Asked: [7 marks] (May 2022) What is model selection in Machine Learning? Asked: [7 marks] (May 2024) List the main types of machine learning algorithms and provide examples of applications for each type. Asked: [7 marks] (Jun 2026) Compare traditional Machine Learning techniques with deep learning for real-world applications.

Supervised learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Supervised learning learns a mapping from inputs to outputs using labelled training data, then predicts the label of new inputs.</mark>

Key points.

  1. For continuous data the task is regression (linear regression, time-series forecasting), for example predicting house price or temperature.
  2. For discrete or categorical data the task is classification (KNN, logistic regression, decision trees, SVM), for example spam or not spam.
  3. Unsupervised learning, with no labels, handles unlabelled data through clustering; semi-supervised mixes few labels with many unlabelled samples; reinforcement learning learns by reward.
  4. KNN is a lazy learner that stores the data and classifies a new point by the majority label of its $k$ nearest neighbours.

KNN steps.

Step 1: Choose k and a distance metric (Euclidean).
Step 2: Compute the distance from the query to every training point.
Step 3: Pick the k points with the smallest distance.
Step 4: Predict by majority vote (classification) or average (regression).

Example (query speed 6.75, agility 3, k = 3). $d=\sqrt{(s-6.75)^2+(a-3)^2}$.

ID Distance Draft
18 1.275 yes
12 1.820 no
20 2.795 yes
15 3.816 yes
16 3.953 yes

The other IDs are farther (11: 4.854, 19: 5.056, 13: 5.701, 14: 5.836, 17: 6.671). The three nearest are IDs 18, 12 and 20, with votes yes, no, yes. Prediction: Draft = yes (2 of 3 neighbours, likelihood about 67%).

Answer frame. Types: define the four, split continuous (regression) from categorical (classification, clustering), one algorithm each. KNN: steps, formula, table, vote, answer.

Asked: [7 marks] (May 2023) Explain various types of machine learning used for continuous data and non-continuous data. Asked: [7 marks] (May 2023) Explain how the KNN method is implemented. Predict whether a player with speed = 6.75 and agility = 3 makes the team using KNN with k = 3 (10-player table given).

Unsupervised learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. Unsupervised learning finds hidden structure in data that has no labels.

Key points.

  1. The two main tasks are clustering (group similar points) and association (find items that occur together, such as market-basket rules); dimensionality reduction (PCA) is also common.
  2. Example, K-means: pick $k$ centroids, assign every point to the nearest centroid, recompute each centroid as the mean of its cluster, and repeat until assignments stop changing; it can segment customers by spending.
  3. Preprocessing is needed first: cleaning missing values, normalizing so no feature dominates the distances, and encoding categories; without it distances and clusters are wrong.

Asked: [7 marks] (Dec 2020) What is the role of preprocessing of data in machine learning? Why is it needed? Explain the unsupervised model of machine learning in detail with an example.

Last-minute revision

  • Mitchell: a program learns from experience E for task T if performance P improves with E.
  • Traditional programming: data + rules -> output; ML: data + output -> rules.
  • Linear hypothesis $h_\theta(x)=\theta_0+\theta_1x$; cost $J=\frac1{2m}\sum(h-y)^2$; update $\theta_j:=\theta_j-\alpha\,\partial J/\partial\theta_j$.
  • The fit (1,2), (2,3), (3,5), (4,6) gives $h(x)=0.5+1.4x$.
  • LWR weight $w=\exp(-(x_i-x)^2/2\tau^2)$; it is non-parametric and refits at every query.
  • Type I error rejects a true $H_0$; Type II accepts a false $H_0$.
  • Label encoding keeps 1 column; one-hot gives $k$ columns for $k$ categories.
  • KNN with k = 3 on speed 6.75, agility 3: neighbours IDs 18, 12, 20, so Draft = yes.

Memory hooks

  • Mitchell's TPE: Task, Performance, Experience.
  • Convex is a bowl: a ball always rolls to the one bottom.
  • Type I = false alarm; Type II = missed detection.
  • Train to fit, validate to tune, test to judge.

Coverage checklist

  • Introduction to machine learning: Dec 2020/May 2022, Dec 2024.
  • scope and limitations: Jun 2025, Jun 2026.
  • regression: Dec 2020, May 2022, May 2024, Jun 2025.
  • probability: not asked; covered.
  • statistics: May 2022.
  • linear algebra for machine learning: not asked; covered.
  • convex optimization: May 2022.
  • data visualization: Jun 2025.
  • hypothesis function and testing: Dec 2024, Jun 2026.
  • data distributions: not asked; covered.
  • data preprocessing: May 2023, Jun 2025, Jun 2026.
  • data augmentation: Dec 2024.
  • normalizing data sets: May 2024.
  • machine learning models: May 2022, May 2024, Jun 2026.
  • supervised learning: May 2023 (types), May 2023 (KNN).
  • unsupervised learning: Dec 2020.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in