How unit 1 is examined
This unit covers what machine learning is, its life cycle, its types, its limits and data preparation basics; the life cycle, the types of ML and the AI-ML-DL-DS relation carry the marks.
Introduction to machine learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Machine learning is the field of study that gives computers the ability to learn from data and improve with experience, without being explicitly programmed.</mark>
Key points.
- Tom Mitchell's formal form: a program learns from experience E for task T with performance measure P if its performance at T, measured by P, improves with E.
- Instead of hand-written rules, an algorithm finds patterns in training data and builds a model.
- The trained model is then used to predict on new, unseen data.
- Examples are spam filtering, recommendation, fraud detection and language translation, where writing fixed rules by hand is impractical.
Machine learning life cycle
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>The ML life cycle is the iterative sequence of stages that takes a problem from data collection to a deployed model that is monitored and retrained.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-01" viewBox="0 0 510 217.6" width="510" height="217.6" role="img" aria-label="Col = data collection, Pre = preprocessing, Fea = feature engineering, Trn = training, Evl = evaluation and validation, Dep = deployment, Mon = monitoring and retraining"><style>#dsfig-u1-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-01 .t{fill:#16181D;font-weight:500}#dsfig-u1-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-01 .dot{fill:#16181D}#dsfig-u1-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-01 .ah{fill:#454C5A}#dsfig-u1-01 .ah.hi{fill:#2340B8}#dsfig-u1-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-01 .e{stroke:#B1B7C3}html.dark #dsfig-u1-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-01 .t{fill:#E6E8ED}html.dark #dsfig-u1-01 .t.inv{fill:#0F1115}html.dark #dsfig-u1-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-01 .dot{fill:#E6E8ED}html.dark #dsfig-u1-01 .ann{fill:#8FA3FF}html.dark #dsfig-u1-01 .lbl{fill:#858D9C}html.dark #dsfig-u1-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-01 .ah{fill:#B1B7C3}html.dark #dsfig-u1-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L105,40" marker-end="url(#ah1)"/><path class="e" d="M145,40 L191,40" marker-end="url(#ah1)"/><path class="e" d="M231,40 L277,40" marker-end="url(#ah1)"/><path class="e" d="M317,40 L363,40" marker-end="url(#ah1)"/><path class="e" d="M403,40 L449,40" marker-end="url(#ah1)"/><path class="e" d="M470,59 L470,156.6" marker-end="url(#ah1)"/><path class="e hi" d="M451.9,171.8 L60,46.4" marker-end="url(#ahh1)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Col</text><circle class="n" cx="126" cy="40" r="18"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">Pre</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">Fea</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Trn</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">Evl</text><circle class="n" cx="470" cy="40" r="18"/><text class="t" x="470" y="40" dy=".35em" text-anchor="middle">Dep</text><circle class="n" cx="470" cy="177.6" r="18"/><text class="t" x="470" y="177.6" dy=".35em" text-anchor="middle">Mon</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Col = data collection, Pre = preprocessing, Fea = feature engineering, Trn = training, Evl = evaluation and validation, Dep = deployment, Mon = monitoring and retraining</figcaption></figure>
Key points.
- Data collection gathers relevant raw data from databases, sensors, logs or the web; its quality limits the model.
- Preprocessing cleans the data by handling missing values, duplicates and outliers, then scales and encodes it.
- Feature engineering selects, transforms or creates the input variables that best expose the pattern.
- Training fits the chosen algorithm to the training set by minimising a loss function.
- Evaluation and validation test the model on unseen validation and test data using metrics such as accuracy or RMSE, and tune hyperparameters.
- Deployment puts the model into production, for example as an API or app, so it serves real predictions.
- Monitoring tracks live accuracy, latency and input data; retraining is triggered when data drift (inputs change), concept drift (input-output relation changes) or performance degradation appears.
- Continuous monitoring and retraining sustain accuracy, reliability and adaptability; the cycle is iterative, not one-way.
Answer frame. For stages: open with the definition; draw the flow diagram with a feedback arrow; develop points 1-6 in order; close that the cycle is iterative. For monitoring: open by defining monitoring and retraining; give the need (drift, degradation), then metrics and triggers (point 7); close with the benefits (point 8).
Asked: [7 marks] (Nov 2023) Explain the importance of continuous model monitoring and retraining in the life cycle. Asked: [7 marks] (Nov 2023) Describe the stages of the machine learning life cycle, from data collection to model deployment.
Types of Machine Learning System
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Machine learning systems are classified by the supervision they get (supervised, unsupervised, semi-supervised, reinforcement), by whether they learn incrementally (batch or online), and by how they generalise (instance-based or model-based).</mark>
Key points.
- Supervised learning trains on labelled input-output pairs and learns a mapping by minimising loss; tasks are classification (spam detection, disease diagnosis) and regression (house price prediction).
- Unsupervised learning finds structure in unlabelled data; tasks are clustering (customer segments), dimensionality reduction and anomaly detection.
- Semi-supervised learning uses a few labelled and many unlabelled examples, as in photo tagging.
- Reinforcement learning has an agent act in an environment and learn a policy from rewards and penalties, as in game playing and robotics.
- Batch learning trains offline on all data and must be retrained from scratch to update; online learning learns incrementally from small batches or single instances and adapts to changing data.
- Instance-based learning memorises examples and predicts by similarity to them (k-NN); model-based learning builds a model with parameters from the data and predicts with it (linear regression).
| Type | Data | Goal | Example |
|---|---|---|---|
| Supervised | Labelled | Predict output | Spam detection |
| Unsupervised | Unlabelled | Find structure | Customer clustering |
| Reinforcement | Rewards | Maximise reward | Chess agent |
Answer frame. Open by defining ML; list the types; give point-wise working and an example for each; draw the table; close with batch/online and instance/model-based. For supervised: define with labelled pairs, loss, examples, then contrast with unsupervised.
Asked: [7 marks] (Nov 2022) Briefly explain the types of Machine Learnings. Asked: [7 marks] (Nov 2023) Define supervised learning and give examples of tasks that can be solved using this approach.
Scope and limitations
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The scope of ML is the range of problems where learning from data replaces explicit rules; its limits come from data and interpretability.</mark>
Key points.
- Scope covers healthcare diagnosis, finance, recommendation, speech, vision and self-driving.
- ML needs large, good-quality data and much computation.
- Complex models are hard to interpret and can inherit bias from data.
- ML finds correlation, not causation, and fails outside the data it saw.
Challenges of Machine learning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Challenges are the data and model problems that reduce how well a model generalises.</mark>
Key points.
- Insufficient or unrepresentative training data gives poor and biased models.
- Poor-quality data (noise, errors, outliers, missing values) gives poor output.
- Irrelevant features hurt the model, so feature selection is needed.
- Overfitting means memorising training data and failing on new data; underfitting means the model is too simple to capture the pattern.
Data visualization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data visualization is presenting data graphically to reveal patterns, trends, outliers and relations before and after modelling.</mark>
Key points.
- A histogram shows a variable's distribution and a scatter plot shows the relation between two variables.
- Box plots show spread and outliers; bar and line charts compare categories and trends.
- A heat map of correlations shows which features relate.
- Matplotlib and Seaborn are the common Python libraries; good plots carry a title, labelled axes and units.
Hypothesis function and testing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A hypothesis function $h(x)$ is the model's proposed mapping from input $x$ to predicted output; hypothesis testing checks whether an observed result is statistically significant.</mark>
Key points.
- For linear regression, $h_\theta(x)=\theta_0+\theta_1 x$, and learning searches for the parameters $\theta$ giving the best fit.
- The set of all functions an algorithm can represent is its hypothesis space.
- Testing starts with a null hypothesis $H_0$ (no effect) against an alternative $H_1$.
- If the p-value is below the significance level (usually 0.05), $H_0$ is rejected.
Data pre-processing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data pre-processing converts raw data into a clean, consistent form suitable for a learning algorithm.</mark>
Key points.
- Cleaning fills missing values by mean, median or mode imputation, or drops such rows, and removes duplicates.
- Outliers are detected, for example by the IQR rule, and treated or removed.
- Categorical variables are encoded numerically by label or one-hot encoding.
- Data is scaled and split into training, validation and test sets.
Data augmentation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data augmentation artificially enlarges the training set by creating modified copies of existing data or synthetic samples.</mark>
Key points.
- For images, flipping, rotating, cropping, scaling and changing brightness create new samples.
- For text, synonym replacement and back-translation are used; for tabular data, SMOTE creates synthetic minority samples.
- It reduces overfitting and improves generalisation when data is scarce.
- Augmented samples must keep the original label.
Normalizing data sets
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Normalization rescales features to a common range or distribution so that no feature dominates because of its units.</mark>
Formula. Min-max: $x'=\dfrac{x-x_{min}}{x_{max}-x_{min}}$ gives values in $[0,1]$. Z-score: $z=\dfrac{x-\mu}{\sigma}$ gives mean 0 and standard deviation 1.
Key points.
- Scaling matters for distance-based and gradient-based methods such as k-NN, SVM and neural networks.
- For data 10, 20, 30, 50, min-max gives 0, 0.25, 0.5, 1.
- Compute the scaling parameters on training data only and reuse them on test data.
Bias-Variance tradeoff
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The bias-variance tradeoff is the balance between bias (error from over-simple assumptions) and variance (sensitivity to training data), which together decide generalisation error.</mark>
Key points.
- Expected error = bias$^2$ + variance + irreducible noise.
- High bias means underfitting: high error on training and test data.
- High variance means overfitting: low training error, high test error.
- Lowering one usually raises the other, so choose model complexity where total error is minimum, using regularisation or cross-validation.
Relation between AI (Artificial Intelligence), ML (Machine Learning), DL (Deep Learning) and DS (Data Science)
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Artificial Intelligence is the broad field of making machines act intelligently; Machine Learning is a subset of AI that learns from data; Deep Learning is a subset of ML that uses multi-layer neural networks; Data Science extracts knowledge and insight from data using statistics, programming and domain knowledge.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-02" viewBox="0 0 184 198" width="184" height="198" role="img" aria-label="AI contains ML, which contains DL; Data Science overlaps all three and uses them as tools"><style>#dsfig-u1-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-02 .t{fill:#16181D;font-weight:500}#dsfig-u1-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-02 .dot{fill:#16181D}#dsfig-u1-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-02 .ah{fill:#454C5A}#dsfig-u1-02 .ah.hi{fill:#2340B8}#dsfig-u1-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-02 .e{stroke:#B1B7C3}html.dark #dsfig-u1-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-02 .t{fill:#E6E8ED}html.dark #dsfig-u1-02 .t.inv{fill:#0F1115}html.dark #dsfig-u1-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-02 .dot{fill:#E6E8ED}html.dark #dsfig-u1-02 .ann{fill:#8FA3FF}html.dark #dsfig-u1-02 .lbl{fill:#858D9C}html.dark #dsfig-u1-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-02 .ah{fill:#B1B7C3}html.dark #dsfig-u1-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="124" y1="39" x2="80" y2="103"/><line class="e" x1="80" y1="103" x2="36" y2="167"/><circle class="n" cx="124" cy="39" r="17"/><text class="t" x="124" y="39" dy=".35em" text-anchor="middle">AI</text><circle class="n" cx="80" cy="103" r="17"/><text class="t" x="80" y="103" dy=".35em" text-anchor="middle">ML</text><circle class="n" cx="36" cy="167" r="17"/><text class="t" x="36" y="167" dy=".35em" text-anchor="middle">DL</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">AI contains ML, which contains DL; Data Science overlaps all three and uses them as tools</figcaption></figure>
Key points.
- AI covers any technique that imitates human intelligence, including rule-based expert systems, search and planning.
- ML is the part of AI that learns patterns from data, for example spam filters and price prediction.
- DL uses deep neural networks and learns features automatically, for example image recognition and speech assistants; it needs large data and GPUs.
- Data science is interdisciplinary: it collects, cleans, analyses and visualises data, and uses ML as one tool to support decisions.
- Data science applications: healthcare (disease prediction), finance (fraud detection, credit scoring), e-commerce (recommendation), social media (sentiment analysis), transport (route optimisation).
Answer frame. For AI, ML and DL: define each in order broad to narrow, draw the nested diagram, give one example each, and close that DL is a subset of ML which is a subset of AI. For data science: define it, list the applications with one line each on 3-4 of them, and close that it drives data-based decisions.
Asked: [7 marks] (Nov 2022) Explain in detail about of Artificial Intelligence, Machine Learning and Deep Learning. Asked: [7 marks] (Nov 2022) Define data science. List and explain different application areas of data science.
Last-minute revision
- ML learns from data without explicit programming (Mitchell: T, P, E).
- Life cycle: collect, preprocess, feature engineering, train, evaluate, deploy, monitor and retrain.
- Retrain on data drift, concept drift or performance drop.
- Supervised uses labels; unsupervised finds structure; reinforcement uses rewards.
- Batch learns offline; online learns incrementally.
- Instance-based memorises; model-based fits parameters.
- Overfitting is high variance; underfitting is high bias.
- Min-max: $(x-x_{min})/(x_{max}-x_{min})$; z-score: $(x-\mu)/\sigma$.
- AI contains ML contains DL; data science uses all.
- Reject $H_0$ when p-value is below 0.05.
Memory hooks
- "Can People Find Truth, Every Day, Maybe": Collect, Preprocess, Feature, Train, Evaluate, Deploy, Monitor.
- AI > ML > DL is like Russian dolls.
- Bias = too simple (underfit); variance = too jumpy (overfit).
- Supervised has a teacher, unsupervised has none, reinforcement has a reward.
Coverage checklist
- Introduction to machine learning: covered.
- Machine learning life cycle: Nov 2023 (2 questions).
- Types of Machine Learning System (supervised and unsupervised learning, Batch and online learning, Instance-Based and Model based Learning): Nov 2022, Nov 2023.
- scope and limitations: covered.
- Challenges of Machine learning: covered.
- data visualization: covered.
- hypothesis function and testing: covered.
- data pre-processing: covered.
- data augmentation: covered.
- normalizing data sets: covered.
- Bias-Variance tradeoff: covered.
- Relation between AI (Artificial Intelligence), ML (Machine Learning), DL (Deep Learning) and DS (Data Science): Nov 2022 (2 questions).