Skip to content
CS-601 · Machine Learning/Quick Revision Short Notes

Machine Learning (CS-601) - Unit 5 Short Notes

How unit 5 is examined

Covers SVM, Bayesian learning, and ML applications in vision, speech and language, plus the ImageNet case study; Bayesian learning (with the belief-network numerical) and computer vision carry the most marks.

Support Vector Machines

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>A Support Vector Machine (SVM) is a supervised learning algorithm that classifies data by finding the maximum-margin hyperplane, the boundary that separates the classes with the widest possible gap.</mark>

Key points.

  1. The hyperplane is $w\cdot x + b = 0$; a point is assigned to a class by the sign of $w\cdot x + b$.
  2. The margin is the distance between the two parallel planes $w\cdot x + b = \pm 1$ and equals $\frac{2}{\lVert w\rVert}$, so maximising the margin means minimising $\frac{1}{2}\lVert w\rVert^2$ subject to $y_i(w\cdot x_i + b)\ge 1$.
  3. Support vectors are the training points lying closest to the boundary, exactly on the margin planes; only they fix the position and orientation of the hyperplane.
  4. Points far from the margin can be moved or deleted without changing the boundary, which makes SVM efficient and robust to distant outliers.
  5. For noisy overlapping data a soft margin allows some violations, and the parameter $C$ trades a wide margin against training errors.
  6. For non-linear data the kernel trick maps points implicitly into a higher-dimensional space where a linear separator exists; common kernels are polynomial, RBF (Gaussian) and sigmoid.
  7. Role: SVM is used for classification (SVC), and the same idea gives regression (SVR) and outlier detection.
  8. Applications: text and spam classification, image and face recognition, handwriting recognition, and bioinformatics such as protein and cancer classification.

Answer frame. Open with the definition; draw two classes, the separating line, the two margin lines and circle the support vectors; develop points 1-4, then 5-6 (soft margin, kernel), then role and applications; close with one line that SVM gives the best generalisation through the widest margin.

Asked: [7 marks] (Dec 2020, May 2022) Explain the concept and role of support vector machine in details; describe its application areas. What is SVM? Discuss in detail. Asked: [7 marks] (May 2024) Describe support vectors and their role in defining the decision boundary in SVM.

Bayesian learning

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Bayesian learning is a probabilistic approach in which a prior belief about a hypothesis is updated with observed data, using Bayes' theorem, to give a posterior probability.</mark>

Formula. $$P(h\mid D)=\frac{P(D\mid h)\,P(h)}{P(D)}$$

Key points.

  1. The prior $P(h)$ is the belief in hypothesis $h$ before seeing data.
  2. The likelihood $P(D\mid h)$ is the probability of the observed data if $h$ is true.
  3. The evidence $P(D)$ is the total probability of the data and normalises the result so posteriors sum to 1.
  4. The posterior $P(h\mid D)$ is the updated belief, and it becomes the prior for the next batch of data.
  5. Classification picks the class with the maximum posterior (MAP); Naive Bayes assumes features are independent given the class.
  6. Impact on ML: it quantifies uncertainty, the prior acts as regularisation that reduces overfitting, and it works with small data and gives probabilistic predictions.
  7. A Bayesian belief network is a directed acyclic graph of variables in which each node stores a conditional probability table (CPT) given its parents.
  8. The joint probability is $P(x_1,\dots,x_n)=\prod_i P(x_i\mid \text{parents}(x_i))$.

Example (Bayes). Disease affects 1% of people; the test is positive for 90% of the sick and 5% of the healthy. $P(\text{pos})=0.9(0.01)+0.05(0.99)=0.0585$, so $P(\text{sick}\mid\text{pos})=\frac{0.009}{0.0585}=$ 0.154.

Example (belief network, May 2023). Network: Mileage → Engine → Car Value, and Air Conditioner → Car Value. Total records $N=40$ (20 High, 20 Low).

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-01" viewBox="0 0 389.6 252" width="389.6" height="252" role="img" aria-label="M = Mileage, E = Engine, A = Air Conditioner, C = Car Value"><style>#dsfig-u5-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-01 .t{fill:#16181D;font-weight:500}#dsfig-u5-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-01 .dot{fill:#16181D}#dsfig-u5-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-01 .ah{fill:#454C5A}#dsfig-u5-01 .ah.hi{fill:#2340B8}#dsfig-u5-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-01 .e{stroke:#B1B7C3}html.dark #dsfig-u5-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-01 .t{fill:#E6E8ED}html.dark #dsfig-u5-01 .t.inv{fill:#0F1115}html.dark #dsfig-u5-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-01 .dot{fill:#E6E8ED}html.dark #dsfig-u5-01 .ann{fill:#8FA3FF}html.dark #dsfig-u5-01 .lbl{fill:#858D9C}html.dark #dsfig-u5-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-01 .ah{fill:#B1B7C3}html.dark #dsfig-u5-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L173.8,40" marker-end="url(#ah10)"/><path class="e" d="M211.4,49.2 L331.2,115.8" marker-end="url(#ah10)"/><path class="e" d="M211.4,202.8 L331.2,136.2" marker-end="url(#ah10)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">M</text><circle class="n" cx="194.8" cy="40" r="18"/><text class="t" x="194.8" y="40" dy=".35em" text-anchor="middle">E</text><circle class="n" cx="194.8" cy="212" r="18"/><text class="t" x="194.8" y="212" dy=".35em" text-anchor="middle">A</text><circle class="n" cx="349.6" cy="126" r="18"/><text class="t" x="349.6" y="126" dy=".35em" text-anchor="middle">C</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">M = Mileage, E = Engine, A = Air Conditioner, C = Car Value</figcaption></figure>

Node CPT (from counts)
Mileage P(Hi)=20/40=0.5, P(Lo)=0.5
Air Conditioner P(Working)=25/40=0.625, P(Broken)=15/40=0.375
Engine given Mileage Hi: Good 0.5, Bad 0.5; Lo: Good 0.75, Bad 0.25
Car Value=High given Engine, AC Good,Working 12/16=0.75; Good,Broken 6/9=0.667; Bad,Working 2/9=0.222; Bad,Broken 0/6=0

Car Value=Low is 1 minus the High entry. For Mileage=Lo, Engine=Bad, AC=Broken: $P(Lo)P(Bad\mid Lo)P(Br)=0.5\times0.25\times0.375=0.0469$. Then $P(High)=0.0469\times 0=0$ and $P(Low)=0.0469\times1=0.0469$. Predicted Car Value = Low.

Answer frame. Theory: open with the definition, write the formula, define each term with the example, then points 5-6, close with the MAP idea. Numerical: draw the network, then all four CPTs, then the chain rule product for both classes and compare.

Pitfall: Forgetting that a node's CPT is conditioned only on its parents, not on all other variables. Asked: [7 marks] (May 2022, Dec 2024) Explain the concept of Bayesian theorem with an example; explain the principles of Bayesian learning, Bayes' theorem and posterior probability. Asked: [7 marks] (May 2022) Define Bayesian learning and how it impacts machine learning. Asked: [7 marks] (May 2023) For the given Bayesian belief network and table, draw the probability table for each node and predict Car Value for P(Mileage=Lo, Engine=Bad, Air Conditioner=Broken).

Application of machine learning in computer vision

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Computer vision is the field that enables machines to interpret images and video, and machine learning, especially CNNs, learns the visual features automatically from labelled data.</mark>

Key points.

  1. Image classification assigns one label to a whole image, for example cat or dog.
  2. Object detection finds and labels each object with a bounding box, for example pedestrians and cars in a road scene.
  3. Segmentation labels every pixel, giving the exact shape of each object or region.
  4. CNNs learn a hierarchy of features (edges, then textures, then object parts) so no hand-made features are needed.
  5. Face recognition uses learned embeddings to unlock phones and mark attendance.
  6. Autonomous driving uses detection and segmentation to see lanes, signs and obstacles.
  7. Medical imaging detects tumours and fractures in X-ray, CT and MRI scans, and industry uses it for defect inspection.
  8. Challenges are the need for large labelled datasets, lighting and viewpoint changes, and bias.

Reinforcement learning (May 2023, part ii). An agent interacts with an environment, takes actions, and receives rewards; it learns a policy that maximises cumulative reward by trial and error. Example: a game-playing agent (chess, Go) or a robot learning to walk earns +1 for progress and a penalty for falling. Main elements: agent, environment, state, action, reward.

Answer frame. Open with the definition; draw a pipeline Image → CNN → Task output (classify, detect, segment); develop points 1-4, then applications 5-7, then challenges; close that ML made vision practical. For the 14-mark question, give computer vision with a CNN classification example, then reinforcement learning with the agent-reward example.

Asked: [7 marks] (Dec 2024, Jun 2026) Explain the role or applications of machine learning in computer vision. Asked: [14 marks] (May 2023) Explain with an example: i) Computer Vision ii) Reinforcement learning.

Speech processing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Speech processing applies machine learning to analyse, recognise, enhance and generate human speech signals.</mark>

Key points.

  1. Speech recognition converts audio to text by combining an acoustic model (HMM or deep neural network) with a language model; modern end-to-end systems use seq2seq and attention.
  2. Feature extraction first turns the waveform into features such as MFCCs (mel-frequency cepstral coefficients) or spectrograms.
  3. Speaker identification recognises who is speaking from voice features, using SVM, GMM or deep networks.
  4. Speech enhancement removes noise using autoencoders or RNNs to improve audio quality.
  5. Emotion detection classifies the speaker's emotional state from prosody (pitch, energy, rhythm) and acoustic features.
  6. Speech synthesis (text-to-speech) generates natural voice using neural models such as WaveNet-style networks.
  7. Applications: voice assistants, transcription, translation and call-centre analytics; challenges are accents, noise and low-resource languages.

Answer frame. Open with the definition; draw the flow Audio → Features (MFCC) → Acoustic model → Language model → Text; develop points 1, 3, 4, 5 and 6 as the ML tasks, then applications; close with the challenges.

Asked: [9 marks] (May 2024, Jun 2025) Explain how machine learning algorithms are utilized in speech processing; discuss applications of ML in speech processing.

Natural language processing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Natural language processing (NLP) is the branch of AI that enables computers to understand, interpret and generate human language.</mark>

Key points.

  1. NLP has two components: natural language understanding (NLU), which extracts meaning, and natural language generation (NLG), which produces text.
  2. A typical pipeline is text cleaning, tokenization, stop-word removal, stemming or lemmatization, POS tagging, parsing, then a model.
  3. Tokenization splits text into units called tokens; it is the first step of the pipeline.
  4. Types of tokenization are word ("I love ML" → I, love, ML), sentence (split at full stops) and subword ("unhappy" → un, happy), which handles rare words.
  5. Text is then converted into numbers by bag-of-words, TF-IDF or word embeddings, and modelled with RNNs or transformers.
  6. Applications: machine translation, sentiment analysis, chatbots, summarisation, spam detection and question answering.
  7. Challenges are ambiguity, sarcasm, context and many languages.

Answer frame. Open with the definition; draw Text → Tokenize → Clean → Features → Model → Output; develop points 1-2, 5, then applications; close with challenges. For the 5-mark tokenization note, give definition, the three types with the example, and its role as step one.

Asked: [7 marks] (Dec 2020, Jun 2025) Write short notes on any two: i) Natural language processing ii) Application of ML in computer vision iii) Bayesian networks (answer from this topic, the computer vision topic and the Bayesian learning topic). Asked: [5 marks] (Jun 2025) Write a short note on Tokenization.

Case Study: ImageNet Competition

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>The ImageNet Large Scale Visual Recognition Challenge (ILSVRC) was an annual contest, run from 2010 to 2017, in which models classified images from ImageNet, a database of over 14 million labelled images in about 1,000 classes for the challenge.</mark>

Key points.

  1. In 2012 AlexNet, a deep CNN trained on GPUs with ReLU and dropout, cut the top-5 error from about 26% to about 15%, a breakthrough that started the deep learning boom.
  2. Later winners were VGG, GoogLeNet (Inception) and ResNet, which reached about 3.6% top-5 error in 2015, better than the typical human figure of about 5%.
  3. ImageNet showed that large data, GPUs and deep networks together beat hand-made features, and pretrained ImageNet models became the standard for transfer learning.
  4. Significance: it drove progress in vision and AI generally and set the benchmark culture for the field.

Asked: [7 marks] (Jun 2026) Explain the significance of the ImageNet Competition in the development of deep learning.

Last-minute revision

  • SVM finds the maximum-margin hyperplane; margin $=2/\lVert w\rVert$.
  • Support vectors are the points nearest the boundary; only they define it.
  • Kernel trick handles non-linear data (polynomial, RBF); $C$ controls the soft margin.
  • Bayes: $P(h\mid D)=P(D\mid h)P(h)/P(D)$ = likelihood times prior over evidence.
  • Belief network joint $=\prod P(x_i\mid \text{parents})$; May 2023 answer: $P=0.0469$ for Low, 0 for High, so Low.
  • CV tasks: classification, detection, segmentation; CNN learns features.
  • Reinforcement learning: agent, environment, action, reward, policy.
  • Speech: MFCC features, HMM/DNN acoustic model plus language model.
  • NLP = NLU + NLG; tokenization types: word, sentence, subword.
  • ImageNet 2012: AlexNet, top-5 error about 15%; ResNet 2015 about 3.6%.

Memory hooks

  • SVM: "widest street" between classes, and only the kerb-side points (support vectors) matter.
  • Bayes: Posterior = Prior x Likelihood / Evidence, "PLE" (posterior from likelihood and evidence).
  • CV tasks in order of detail: Classify, Detect, Segment = whole image, boxes, pixels.
  • Speech: "Audio to MFCC to Model to Text".
  • ImageNet: AlexNet 2012 lit the deep learning fire.

Coverage checklist

  • Support Vector Machines: Dec 2020/May 2022 concept and role; May 2024 support vectors.
  • Bayesian learning: May 2022/Dec 2024 Bayes theorem; May 2022 definition and impact; May 2023 belief network numerical; Bayesian networks short note (Dec 2020, Jun 2025).
  • application of machine learning in computer vision: Dec 2024/Jun 2026 role; May 2023 computer vision and reinforcement learning; short-note option.
  • speech processing: May 2024/Jun 2025 ML in speech processing.
  • natural language processing: Dec 2020/Jun 2025 short notes; Jun 2025 tokenization.
  • Case Study: ImageNet Competition: Jun 2026 significance of ImageNet.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in