Skip to content
AD-502 · Machine Learning/Quick Revision Short Notes

Machine Learning (AD-502) - Unit 1 Short Notes

How unit 1 is examined

This unit covers what machine learning is, its life cycle, its types, its limits and data preparation basics; the life cycle, the types of ML and the AI-ML-DL-DS relation carry the marks.

Introduction to machine learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Machine learning is the field of study that gives computers the ability to learn from data and improve with experience, without being explicitly programmed.</mark>

Key points.

  1. Tom Mitchell's formal form: a program learns from experience E for task T with performance measure P if its performance at T, measured by P, improves with E.
  2. Instead of hand-written rules, an algorithm finds patterns in training data and builds a model.
  3. The trained model is then used to predict on new, unseen data.
  4. Examples are spam filtering, recommendation, fraud detection and language translation, where writing fixed rules by hand is impractical.

Machine learning life cycle

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>The ML life cycle is the iterative sequence of stages that takes a problem from data collection to a deployed model that is monitored and retrained.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-01" viewBox="0 0 510 217.6" width="510" height="217.6" role="img" aria-label="Col = data collection, Pre = preprocessing, Fea = feature engineering, Trn = training, Evl = evaluation and validation, Dep = deployment, Mon = monitoring and retraining"><style>#dsfig-u1-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-01 .t{fill:#16181D;font-weight:500}#dsfig-u1-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-01 .dot{fill:#16181D}#dsfig-u1-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-01 .ah{fill:#454C5A}#dsfig-u1-01 .ah.hi{fill:#2340B8}#dsfig-u1-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-01 .e{stroke:#B1B7C3}html.dark #dsfig-u1-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-01 .t{fill:#E6E8ED}html.dark #dsfig-u1-01 .t.inv{fill:#0F1115}html.dark #dsfig-u1-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-01 .dot{fill:#E6E8ED}html.dark #dsfig-u1-01 .ann{fill:#8FA3FF}html.dark #dsfig-u1-01 .lbl{fill:#858D9C}html.dark #dsfig-u1-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-01 .ah{fill:#B1B7C3}html.dark #dsfig-u1-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L105,40" marker-end="url(#ah1)"/><path class="e" d="M145,40 L191,40" marker-end="url(#ah1)"/><path class="e" d="M231,40 L277,40" marker-end="url(#ah1)"/><path class="e" d="M317,40 L363,40" marker-end="url(#ah1)"/><path class="e" d="M403,40 L449,40" marker-end="url(#ah1)"/><path class="e" d="M470,59 L470,156.6" marker-end="url(#ah1)"/><path class="e hi" d="M451.9,171.8 L60,46.4" marker-end="url(#ahh1)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Col</text><circle class="n" cx="126" cy="40" r="18"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">Pre</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">Fea</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Trn</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">Evl</text><circle class="n" cx="470" cy="40" r="18"/><text class="t" x="470" y="40" dy=".35em" text-anchor="middle">Dep</text><circle class="n" cx="470" cy="177.6" r="18"/><text class="t" x="470" y="177.6" dy=".35em" text-anchor="middle">Mon</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Col = data collection, Pre = preprocessing, Fea = feature engineering, Trn = training, Evl = evaluation and validation, Dep = deployment, Mon = monitoring and retraining</figcaption></figure>

Key points.

  1. Data collection gathers relevant raw data from databases, sensors, logs or the web; its quality limits the model.
  2. Preprocessing cleans the data by handling missing values, duplicates and outliers, then scales and encodes it.
  3. Feature engineering selects, transforms or creates the input variables that best expose the pattern.
  4. Training fits the chosen algorithm to the training set by minimising a loss function.
  5. Evaluation and validation test the model on unseen validation and test data using metrics such as accuracy or RMSE, and tune hyperparameters.
  6. Deployment puts the model into production, for example as an API or app, so it serves real predictions.
  7. Monitoring tracks live accuracy, latency and input data; retraining is triggered when data drift (inputs change), concept drift (input-output relation changes) or performance degradation appears.
  8. Continuous monitoring and retraining sustain accuracy, reliability and adaptability; the cycle is iterative, not one-way.

Answer frame. For stages: open with the definition; draw the flow diagram with a feedback arrow; develop points 1-6 in order; close that the cycle is iterative. For monitoring: open by defining monitoring and retraining; give the need (drift, degradation), then metrics and triggers (point 7); close with the benefits (point 8).

Asked: [7 marks] (Nov 2023) Explain the importance of continuous model monitoring and retraining in the life cycle. Asked: [7 marks] (Nov 2023) Describe the stages of the machine learning life cycle, from data collection to model deployment.

Types of Machine Learning System

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Machine learning systems are classified by the supervision they get (supervised, unsupervised, semi-supervised, reinforcement), by whether they learn incrementally (batch or online), and by how they generalise (instance-based or model-based).</mark>

Key points.

  1. Supervised learning trains on labelled input-output pairs and learns a mapping by minimising loss; tasks are classification (spam detection, disease diagnosis) and regression (house price prediction).
  2. Unsupervised learning finds structure in unlabelled data; tasks are clustering (customer segments), dimensionality reduction and anomaly detection.
  3. Semi-supervised learning uses a few labelled and many unlabelled examples, as in photo tagging.
  4. Reinforcement learning has an agent act in an environment and learn a policy from rewards and penalties, as in game playing and robotics.
  5. Batch learning trains offline on all data and must be retrained from scratch to update; online learning learns incrementally from small batches or single instances and adapts to changing data.
  6. Instance-based learning memorises examples and predicts by similarity to them (k-NN); model-based learning builds a model with parameters from the data and predicts with it (linear regression).
Type Data Goal Example
Supervised Labelled Predict output Spam detection
Unsupervised Unlabelled Find structure Customer clustering
Reinforcement Rewards Maximise reward Chess agent

Answer frame. Open by defining ML; list the types; give point-wise working and an example for each; draw the table; close with batch/online and instance/model-based. For supervised: define with labelled pairs, loss, examples, then contrast with unsupervised.

Asked: [7 marks] (Nov 2022) Briefly explain the types of Machine Learnings. Asked: [7 marks] (Nov 2023) Define supervised learning and give examples of tasks that can be solved using this approach.

Scope and limitations

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>The scope of ML is the range of problems where learning from data replaces explicit rules; its limits come from data and interpretability.</mark>

Key points.

  1. Scope covers healthcare diagnosis, finance, recommendation, speech, vision and self-driving.
  2. ML needs large, good-quality data and much computation.
  3. Complex models are hard to interpret and can inherit bias from data.
  4. ML finds correlation, not causation, and fails outside the data it saw.

Challenges of Machine learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Challenges are the data and model problems that reduce how well a model generalises.</mark>

Key points.

  1. Insufficient or unrepresentative training data gives poor and biased models.
  2. Poor-quality data (noise, errors, outliers, missing values) gives poor output.
  3. Irrelevant features hurt the model, so feature selection is needed.
  4. Overfitting means memorising training data and failing on new data; underfitting means the model is too simple to capture the pattern.

Data visualization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data visualization is presenting data graphically to reveal patterns, trends, outliers and relations before and after modelling.</mark>

Key points.

  1. A histogram shows a variable's distribution and a scatter plot shows the relation between two variables.
  2. Box plots show spread and outliers; bar and line charts compare categories and trends.
  3. A heat map of correlations shows which features relate.
  4. Matplotlib and Seaborn are the common Python libraries; good plots carry a title, labelled axes and units.

Hypothesis function and testing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A hypothesis function $h(x)$ is the model's proposed mapping from input $x$ to predicted output; hypothesis testing checks whether an observed result is statistically significant.</mark>

Key points.

  1. For linear regression, $h_\theta(x)=\theta_0+\theta_1 x$, and learning searches for the parameters $\theta$ giving the best fit.
  2. The set of all functions an algorithm can represent is its hypothesis space.
  3. Testing starts with a null hypothesis $H_0$ (no effect) against an alternative $H_1$.
  4. If the p-value is below the significance level (usually 0.05), $H_0$ is rejected.

Data pre-processing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data pre-processing converts raw data into a clean, consistent form suitable for a learning algorithm.</mark>

Key points.

  1. Cleaning fills missing values by mean, median or mode imputation, or drops such rows, and removes duplicates.
  2. Outliers are detected, for example by the IQR rule, and treated or removed.
  3. Categorical variables are encoded numerically by label or one-hot encoding.
  4. Data is scaled and split into training, validation and test sets.

Data augmentation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data augmentation artificially enlarges the training set by creating modified copies of existing data or synthetic samples.</mark>

Key points.

  1. For images, flipping, rotating, cropping, scaling and changing brightness create new samples.
  2. For text, synonym replacement and back-translation are used; for tabular data, SMOTE creates synthetic minority samples.
  3. It reduces overfitting and improves generalisation when data is scarce.
  4. Augmented samples must keep the original label.

Normalizing data sets

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Normalization rescales features to a common range or distribution so that no feature dominates because of its units.</mark>

Formula. Min-max: $x'=\dfrac{x-x_{min}}{x_{max}-x_{min}}$ gives values in $[0,1]$. Z-score: $z=\dfrac{x-\mu}{\sigma}$ gives mean 0 and standard deviation 1.

Key points.

  1. Scaling matters for distance-based and gradient-based methods such as k-NN, SVM and neural networks.
  2. For data 10, 20, 30, 50, min-max gives 0, 0.25, 0.5, 1.
  3. Compute the scaling parameters on training data only and reuse them on test data.

Bias-Variance tradeoff

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>The bias-variance tradeoff is the balance between bias (error from over-simple assumptions) and variance (sensitivity to training data), which together decide generalisation error.</mark>

Key points.

  1. Expected error = bias$^2$ + variance + irreducible noise.
  2. High bias means underfitting: high error on training and test data.
  3. High variance means overfitting: low training error, high test error.
  4. Lowering one usually raises the other, so choose model complexity where total error is minimum, using regularisation or cross-validation.

Relation between AI (Artificial Intelligence), ML (Machine Learning), DL (Deep Learning) and DS (Data Science)

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Artificial Intelligence is the broad field of making machines act intelligently; Machine Learning is a subset of AI that learns from data; Deep Learning is a subset of ML that uses multi-layer neural networks; Data Science extracts knowledge and insight from data using statistics, programming and domain knowledge.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-02" viewBox="0 0 184 198" width="184" height="198" role="img" aria-label="AI contains ML, which contains DL; Data Science overlaps all three and uses them as tools"><style>#dsfig-u1-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-02 .t{fill:#16181D;font-weight:500}#dsfig-u1-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-02 .dot{fill:#16181D}#dsfig-u1-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-02 .ah{fill:#454C5A}#dsfig-u1-02 .ah.hi{fill:#2340B8}#dsfig-u1-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-02 .e{stroke:#B1B7C3}html.dark #dsfig-u1-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-02 .t{fill:#E6E8ED}html.dark #dsfig-u1-02 .t.inv{fill:#0F1115}html.dark #dsfig-u1-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-02 .dot{fill:#E6E8ED}html.dark #dsfig-u1-02 .ann{fill:#8FA3FF}html.dark #dsfig-u1-02 .lbl{fill:#858D9C}html.dark #dsfig-u1-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-02 .ah{fill:#B1B7C3}html.dark #dsfig-u1-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="124" y1="39" x2="80" y2="103"/><line class="e" x1="80" y1="103" x2="36" y2="167"/><circle class="n" cx="124" cy="39" r="17"/><text class="t" x="124" y="39" dy=".35em" text-anchor="middle">AI</text><circle class="n" cx="80" cy="103" r="17"/><text class="t" x="80" y="103" dy=".35em" text-anchor="middle">ML</text><circle class="n" cx="36" cy="167" r="17"/><text class="t" x="36" y="167" dy=".35em" text-anchor="middle">DL</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">AI contains ML, which contains DL; Data Science overlaps all three and uses them as tools</figcaption></figure>

Key points.

  1. AI covers any technique that imitates human intelligence, including rule-based expert systems, search and planning.
  2. ML is the part of AI that learns patterns from data, for example spam filters and price prediction.
  3. DL uses deep neural networks and learns features automatically, for example image recognition and speech assistants; it needs large data and GPUs.
  4. Data science is interdisciplinary: it collects, cleans, analyses and visualises data, and uses ML as one tool to support decisions.
  5. Data science applications: healthcare (disease prediction), finance (fraud detection, credit scoring), e-commerce (recommendation), social media (sentiment analysis), transport (route optimisation).

Answer frame. For AI, ML and DL: define each in order broad to narrow, draw the nested diagram, give one example each, and close that DL is a subset of ML which is a subset of AI. For data science: define it, list the applications with one line each on 3-4 of them, and close that it drives data-based decisions.

Asked: [7 marks] (Nov 2022) Explain in detail about of Artificial Intelligence, Machine Learning and Deep Learning. Asked: [7 marks] (Nov 2022) Define data science. List and explain different application areas of data science.

Last-minute revision

  • ML learns from data without explicit programming (Mitchell: T, P, E).
  • Life cycle: collect, preprocess, feature engineering, train, evaluate, deploy, monitor and retrain.
  • Retrain on data drift, concept drift or performance drop.
  • Supervised uses labels; unsupervised finds structure; reinforcement uses rewards.
  • Batch learns offline; online learns incrementally.
  • Instance-based memorises; model-based fits parameters.
  • Overfitting is high variance; underfitting is high bias.
  • Min-max: $(x-x_{min})/(x_{max}-x_{min})$; z-score: $(x-\mu)/\sigma$.
  • AI contains ML contains DL; data science uses all.
  • Reject $H_0$ when p-value is below 0.05.

Memory hooks

  • "Can People Find Truth, Every Day, Maybe": Collect, Preprocess, Feature, Train, Evaluate, Deploy, Monitor.
  • AI > ML > DL is like Russian dolls.
  • Bias = too simple (underfit); variance = too jumpy (overfit).
  • Supervised has a teacher, unsupervised has none, reinforcement has a reward.

Coverage checklist

  • Introduction to machine learning: covered.
  • Machine learning life cycle: Nov 2023 (2 questions).
  • Types of Machine Learning System (supervised and unsupervised learning, Batch and online learning, Instance-Based and Model based Learning): Nov 2022, Nov 2023.
  • scope and limitations: covered.
  • Challenges of Machine learning: covered.
  • data visualization: covered.
  • hypothesis function and testing: covered.
  • data pre-processing: covered.
  • data augmentation: covered.
  • normalizing data sets: covered.
  • Bias-Variance tradeoff: covered.
  • Relation between AI (Artificial Intelligence), ML (Machine Learning), DL (Deep Learning) and DS (Data Science): Nov 2022 (2 questions).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in