Skip to content
CS-803 (A) · Image Processing and Computer Vision/Quick Revision Short Notes

Image Processing and Computer Vision (CS-803 (A)) - Unit 5 Short Notes

How unit 5 is examined

This unit covers how prior knowledge, matching and learning are used to recognise objects; Hough-based recognition and neural networks carry the most marks, then PCA.

Knowledge representation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Knowledge representation is the way facts, models and rules about objects and scenes are encoded in a form a vision system can store, search and reason with.</mark>

Key points.

  1. A knowledge-based vision system uses a stored knowledge base and an inference mechanism, not only pixels, to interpret an image.
  2. Semantic networks store objects as nodes and relations such as "is-a" or "part-of" as labelled links.
  3. Frames hold an object's slots (shape, colour, size, position) with default values that are filled in from the image.
  4. Production rules of the form IF condition THEN action encode expert knowledge, for example IF region is red and octagonal THEN it is a stop sign.
  5. Logic and ontologies give formal, unambiguous statements about scene objects and their relations.
  6. Prior knowledge constrains recognition: it reduces the search, rejects impossible interpretations and resolves ambiguity, so a road-scene system expects cars on roads and not in the sky.

Answer frame. Open with the definition; list the four schemes (semantic net, frames, rules, logic/ontology) with one line each; explain how knowledge-based vision uses the knowledge base plus inference to guide recognition; close with a road-sign or face example.

Asked: [7 marks] (May 2022) Define and explain knowledge representation and Information Integration. Asked: [7 marks] (Dec 2024) What is knowledge-based vision and how does it support object recognition? Asked: [14 marks] (May 2024) Write brief notes: a) Feature extraction b) Knowledge Representation c) Backtracking algorithm Asked: [14 marks] (Jun 2025) Write short notes on (any two): a) Machine learning and neural networks in image shape recognition b) Knowledge-Based Vision c) Curve Fitting (Least-square fitting)

Control-strategies

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Control strategy decides which knowledge or operation is applied next and in what order during interpretation.

Key points.

  1. Bottom-up (data-driven) control starts from pixels and builds edges, regions and then objects.
  2. Top-down (model-driven) control starts from an expected object model and looks for its evidence in the image.
  3. Hybrid control mixes both, so hypotheses from models are verified by image data.
  4. Backtracking is a search control: try a label, and if it violates a constraint, undo it and try the next choice, as in consistent labelling of line drawings.

Information Integration

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Information integration is combining evidence from several cues, sensors or processing levels into one consistent interpretation.

Key points.

  1. Sources include colour, texture, shape, motion, stereo and range data.
  2. Combination may be by rules, weighted voting or probabilistic evidence such as Bayes' rule.
  3. Redundant cues make the result robust when one cue is noisy or missing.
  4. Conflicts between cues are resolved using the knowledge base and constraints.

Object recognition: Hough transform and simple methods

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Object recognition is deciding which known object a detected region or image belongs to; the Hough transform detects parametric shapes such as lines and circles by letting each edge point vote in a parameter space.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-01" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="Simple recognition pipeline: Image, edges, Hough voting in accumulator, peak detection, classification"><style>#dsfig-u5-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-01 .t{fill:#16181D;font-weight:500}#dsfig-u5-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-01 .dot{fill:#16181D}#dsfig-u5-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-01 .ah{fill:#454C5A}#dsfig-u5-01 .ah.hi{fill:#2340B8}#dsfig-u5-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-01 .e{stroke:#B1B7C3}html.dark #dsfig-u5-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-01 .t{fill:#E6E8ED}html.dark #dsfig-u5-01 .t.inv{fill:#0F1115}html.dark #dsfig-u5-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-01 .dot{fill:#E6E8ED}html.dark #dsfig-u5-01 .ann{fill:#8FA3FF}html.dark #dsfig-u5-01 .lbl{fill:#858D9C}html.dark #dsfig-u5-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-01 .ah{fill:#B1B7C3}html.dark #dsfig-u5-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah9)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah9)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah9)"/><path class="e" d="M446,40 L535,40" marker-end="url(#ah9)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Img</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">Edg</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Vot</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">Pk</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">Cls</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Simple recognition pipeline: Image, edges, Hough voting in accumulator, peak detection, classification</figcaption></figure>

Key points.

  1. Object recognition steps are detection, description (features) and classification against stored models.
  2. Simple methods are template matching (slide a stored template and maximise correlation), feature-based matching (compare shape, colour, moments) and learned classifiers or deep learning.
  3. Hough principle: every edge point votes for all parameter values of shapes that could pass through it, and true shapes appear as peaks in the accumulator.
  4. For a line use $\rho = x\cos\theta + y\sin\theta$; each point $(x,y)$ traces a sinusoid in $(\rho,\theta)$ space and collinear points intersect at one cell.
  5. For a circle $(x-a)^2+(y-b)^2=r^2$ the accumulator is $(a,b,r)$; for known radius each edge point votes on a circle of possible centres, and the peak gives the centre.
  6. Accurate centre location: use the gradient direction so each point votes only along one line, smooth the accumulator, take the peak and refine by the weighted mean of neighbouring cells for sub-pixel accuracy.
  7. Advantage: the method tolerates noise, gaps and partial occlusion because only the vote count matters. Drawbacks are memory and time growing with the number of parameters.
  8. Road-sign identification in a vehicle: segment by colour (red, blue), find circles or polygons by Hough, then classify by template matching or a classifier; illumination change, occlusion and motion blur are the challenges.

Example. Points $(1,1),(2,2),(3,3)$ at $\theta=135^\circ$ all give $\rho = 0$, so the cell $(0,135^\circ)$ collects 3 votes and the line is $y=x$.

Answer frame. Open with the voting definition; draw the pipeline and the parameter space; develop points 3, 4, 5, 7 in that order; add the example; close with the advantage under noise. For the centre-location question stress points 5 and 6; for road signs use point 8; for the summary of methods use points 1 and 2 with a short comparison.

Asked: [7 marks] (Dec 2024, Jun 2025) Explain the Hough transform and discuss how it is used in detecting simple objects in images. Asked: [7 marks] (May 2022) Explain the identifying road signs in vehicle vision system. Asked: [7 marks] (May 2022) Discuss about accurate centre location by using Hough transform. Asked: [7 marks] (May 2023) Give a brief summary on object recognition methods.

Pitfall: Writing the line as $y=mx+c$ only; vertical lines need the $(\rho,\theta)$ form.

Shape correspondence and shape matching

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. Shape correspondence finds which parts of one shape match which parts of another, and shape matching measures how similar the two shapes are.

Key points.

  1. Descriptors are global (area, moments, Fourier descriptors) or local (curvature, keypoints, spin images).
  2. Evaluate a descriptor by retrieval: query with each shape and compute precision and recall of the returned matches.
  3. Also test robustness to pose change, noise, missing parts and sampling.
  4. Benchmark on a labelled database with the same queries for every descriptor, and compare precision-recall curves and speed.

Asked: [7 marks] (May 2022) How to evaluate extracted shape descriptors in 3D vision? Discuss.

Principal component analysis

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Principal Component Analysis is a linear transform that projects data onto orthogonal directions of maximum variance, the eigenvectors of the covariance matrix, to reduce dimensions while keeping most information.</mark>

Steps.

Step 1: Arrange each image or feature vector as a column and compute the mean vector.
Step 2: Subtract the mean from every vector.
Step 3: Compute the covariance matrix C.
Step 4: Find eigenvalues and eigenvectors of C and sort by decreasing eigenvalue.
Step 5: Keep the top k eigenvectors as principal components.
Step 6: Project the data, y = W^T (x - mean), to get the k-dimensional features.

Key points.

  1. The first component carries the largest variance and each next one is orthogonal to the earlier ones.
  2. The retained fraction of variance is the sum of the top $k$ eigenvalues divided by the sum of all.
  3. As feature extraction, the projections $y$ are a compact descriptor that replaces raw pixels.
  4. In face recognition the eigenvectors are called eigenfaces; a new face is projected and matched to the nearest stored projection.
  5. Benefits are compact representation, noise reduction and faster classification; the limit is that it is linear and not class-aware.

Answer frame. Open with the definition and the dimensionality-reduction goal; list the six steps; explain the role in feature extraction; close with the eigenface application.

Asked: [7 marks] (Dec 2024, Jun 2025) Define Principal Component Analysis (PCA) and its application in feature extraction / describe its role in feature extraction for object recognition.

Feature extraction

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Feature extraction converts an image or region into a compact set of measurements that describe it and separate one class from another.

Key points.

  1. Low-level features are edges, corners, colour and texture; shape features are area, perimeter, moments and Fourier descriptors.
  2. Global features describe the whole region, and local features describe small neighbourhoods around keypoints.
  3. Good features are discriminative, invariant to scale, rotation and illumination, and cheap to compute.
  4. The feature vector is the input to a classifier or neural network; PCA can reduce it further.

Neural network and machine learning for image shape recognition

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>A neural network is a layered network of weighted neurons that learns from labelled examples to map an input feature vector or image to a shape class.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-02" viewBox="0 0 424 252" width="424" height="252" role="img" aria-label="I = input layer (feature vector or pixels), H = hidden layer, O = output layer (one neuron per shape class)"><style>#dsfig-u5-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-02 .t{fill:#16181D;font-weight:500}#dsfig-u5-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-02 .dot{fill:#16181D}#dsfig-u5-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-02 .ah{fill:#454C5A}#dsfig-u5-02 .ah.hi{fill:#2340B8}#dsfig-u5-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-02 .e{stroke:#B1B7C3}html.dark #dsfig-u5-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-02 .t{fill:#E6E8ED}html.dark #dsfig-u5-02 .t.inv{fill:#0F1115}html.dark #dsfig-u5-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-02 .dot{fill:#E6E8ED}html.dark #dsfig-u5-02 .ann{fill:#8FA3FF}html.dark #dsfig-u5-02 .lbl{fill:#858D9C}html.dark #dsfig-u5-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-02 .ah{fill:#B1B7C3}html.dark #dsfig-u5-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L191,40" marker-end="url(#ah10)"/><path class="e" d="M59,126 L191,126" marker-end="url(#ah10)"/><path class="e" d="M59,212 L191,212" marker-end="url(#ah10)"/><path class="e" d="M57,48.5 L193.2,116.6" marker-end="url(#ah10)"/><path class="e" d="M57,203.5 L193.2,135.4" marker-end="url(#ah10)"/><path class="e" d="M230.4,44.6 L363.6,77.9" marker-end="url(#ah10)"/><path class="e" d="M230.4,121.4 L363.6,88.1" marker-end="url(#ah10)"/><path class="e" d="M230.4,130.6 L363.6,163.9" marker-end="url(#ah10)"/><path class="e" d="M230.4,207.4 L363.6,174.1" marker-end="url(#ah10)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">I1</text><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">I2</text><circle class="n" cx="40" cy="212" r="18"/><text class="t" x="40" y="212" dy=".35em" text-anchor="middle">I3</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">H1</text><circle class="n" cx="212" cy="126" r="18"/><text class="t" x="212" y="126" dy=".35em" text-anchor="middle">H2</text><circle class="n" cx="212" cy="212" r="18"/><text class="t" x="212" y="212" dy=".35em" text-anchor="middle">H3</text><circle class="n" cx="384" cy="83" r="18"/><text class="t" x="384" y="83" dy=".35em" text-anchor="middle">O1</text><circle class="n" cx="384" cy="169" r="18"/><text class="t" x="384" y="169" dy=".35em" text-anchor="middle">O2</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">I = input layer (feature vector or pixels), H = hidden layer, O = output layer (one neuron per shape class)</figcaption></figure>

Key points.

  1. The input layer takes the feature vector (moments, Fourier descriptors, PCA scores) or raw pixels.
  2. Each hidden neuron computes $a=f(\sum w_i x_i + b)$ with a nonlinear activation such as sigmoid or ReLU, which lets the network learn curved decision boundaries.
  3. The output layer has one neuron per class, for example circle, square, triangle, and the largest output gives the recognised shape.
  4. Training is supervised: forward pass, error against the target, then backpropagation of the error to update weights by gradient descent, repeated over many epochs.
  5. A Convolutional Neural Network (CNN) uses convolution layers to learn shape features, pooling to gain translation invariance, and fully connected layers to classify, so no hand-made features are needed.
  6. Other machine learning options are SVM, k-nearest neighbour and decision trees on extracted features.
  7. Advantages are learning from data, tolerance to noise and distortion, and high accuracy; limitations are the need for large labelled data, training cost and poor interpretability.

Example. Character recognition: binarise the letter image, extract a feature vector or feed the pixel grid, train on labelled letters with backpropagation, and at test time the output neuron with the highest activation names the letter.

Answer frame. Open with the definition; draw the three-layer network with the caption; develop points 1-4 for structure and training, then 5 for CNN; give the character example; close with advantages and limits. For the 14-mark short note add applications and one more topic.

Asked: [7 marks] (May 2022, May 2023, May 2024, Dec 2024) Explain the use of neural network structures for pattern recognition with an example; discuss how neural networks are used for image shape recognition. Asked: [14 marks] (Jun 2025) Write short notes on (any two): a) Machine learning and neural networks in image shape recognition b) Knowledge-Based Vision c) Curve Fitting (Least-square fitting)

Last-minute revision

  • Knowledge representation schemes: semantic nets, frames, production rules, logic/ontologies.
  • Control: bottom-up (data-driven), top-down (model-driven), hybrid; backtracking undoes failed labels.
  • Information integration combines cues and sensors into one consistent interpretation.
  • Hough line: $\rho = x\cos\theta + y\sin\theta$; circle accumulator is $(a,b,r)$; peaks mean shapes.
  • Hough tolerates noise and occlusion; cost grows with parameters.
  • Road signs: colour segmentation, Hough circles or polygons, then classification.
  • Shape descriptors are evaluated by precision, recall and robustness on a benchmark.
  • PCA: mean, covariance, eigenvectors, sort, top $k$, project.
  • Eigenfaces are PCA components of face images.
  • Neural network: input, hidden, output layers; trained by backpropagation; CNN = convolution, pooling, fully connected.

Memory hooks

  • SFRO for knowledge schemes: Semantic net, Frames, Rules, Ontology.
  • Hough = "every point votes, the peak wins".
  • PCA steps: Mean, Cov, Eigen, Keep, Project (MCEKP).
  • Neural net: forward to guess, backward to fix.
  • Top-down asks "where is the car?", bottom-up asks "what are these pixels?".

Coverage checklist

  • Knowledge representation: Q3 (May 2022), Q11 knowledge-based vision (Dec 2024), Q2b (May 2024), Q1b (Jun 2025).
  • Control-strategies: control types and backtracking (Q2c, May 2024).
  • Information Integration: Q3 (May 2022).
  • Object recognition-Hough transforms and other simple object recognition methods: Q5, Q6, Q7, Q8.
  • Shape correspondence and shape matching: Q10.
  • Principal component analysis: Q9.
  • feature extraction: Q2a (May 2024).
  • Neural network and Machine learning for image shape recognition: Q4, Q1a.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in