How unit 1 is examined
This unit covers what CVIP is, its models, filtering, digitisation and the whole of binary and gray-scale morphology; dilation and thinning carry the most marks, then the medium topics (basics, CV models, filtering, representations).
Basics of CVIP
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Computer vision is the science of making a machine understand a scene from images, while image processing takes an image in and gives an improved or transformed image out.</mark>
Key points.
- Image processing is a low-level activity: input and output are both images, as in noise removal, enhancement and restoration.
- Computer vision is a high-level activity: the input is an image and the output is a description or decision, such as "this is a face" or "the obstacle is 2 m away".
- Image processing is needed by vision as pre-processing, because a noisy, blurred or badly lit image gives wrong segmentation and recognition.
- The usual chain is acquisition, pre-processing, segmentation, feature extraction, recognition and interpretation.
- Applications include medical imaging, face and fingerprint recognition, remote sensing, robotics and autonomous vehicles.
Diagram. A digital image processing system has five elements.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-01" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="Elements of a digital image processing system: Acq acquisition (sensor and digitiser), Sto storage, Pro processing, Com communication, Dis display"><style>#dsfig-u1-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-01 .t{fill:#16181D;font-weight:500}#dsfig-u1-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-01 .dot{fill:#16181D}#dsfig-u1-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-01 .ah{fill:#454C5A}#dsfig-u1-01 .ah.hi{fill:#2340B8}#dsfig-u1-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-01 .e{stroke:#B1B7C3}html.dark #dsfig-u1-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-01 .t{fill:#E6E8ED}html.dark #dsfig-u1-01 .t.inv{fill:#0F1115}html.dark #dsfig-u1-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-01 .dot{fill:#E6E8ED}html.dark #dsfig-u1-01 .ann{fill:#8FA3FF}html.dark #dsfig-u1-01 .lbl{fill:#858D9C}html.dark #dsfig-u1-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-01 .ah{fill:#B1B7C3}html.dark #dsfig-u1-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah1)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah1)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah1)"/><path class="e" d="M446,40 L535,40" marker-end="url(#ah1)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Acq</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">Sto</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Pro</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">Com</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">Dis</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Elements of a digital image processing system: Acq acquisition (sensor and digitiser), Sto storage, Pro processing, Com communication, Dis display</figcaption></figure>
Elements.
- Acquisition uses a sensor (camera, scanner) and a digitiser to convert the scene into a digital image.
- Storage holds the image in short-term memory (frame buffer), online disks and archival media.
- Processing is the computer with software that enhances, restores, segments or compresses the image.
- Communication sends images over networks, so compression is used to save bandwidth.
- Display shows the result on monitors or printers.
Answer frame. Open with the definition of computer vision and how it differs from image processing; draw the five-block diagram; develop points 1-5 (low versus high level, then why processing is required); close with "image processing prepares the data and vision interprets it". For the system-elements question, draw the block diagram first and give one sentence per element.
Asked: [7 marks] (May 2022, May 2024) Describe the concept of computer vision which is mandatory in image processing. Asked: [7 marks] (May 2022) Describe various elements of digital image processing system. Asked: [7 marks] (May 2024) What are the components of digital image processing system? Explain each in detail.
History of CVIP
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. History of CVIP is the timeline of how image processing (improving images) and computer vision (understanding images) developed and merged.
Key points.
- In the 1920s-50s pictures were sent by cable and early scanners appeared; the 1960s space programme (Ranger and Apollo images) made computer-based enhancement important.
- In 1963 Roberts' thesis on recognising 3-D blocks from 2-D images started computer vision.
- In the 1970s Marr's theory of vision, edge detectors and medical CT scanning appeared.
- The 1980s brought active vision, stereo, motion analysis and industrial inspection, and the 1990s brought statistical methods, face recognition and cheap digital cameras.
- Since 2012 deep learning (CNNs such as AlexNet) has dominated both fields.
Asked: [7 marks] (Dec 2024) Discuss the history and evolution of Computer Vision and Image Processing (CVIP).
Evolution of CVIP
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. Evolution of CVIP is the shift from hand-crafted image operations to learned, intelligent scene understanding, driven by advances in hardware, sensors and algorithms.
Key points.
- The stages are early digital image processing, feature-based vision, statistical and machine learning methods, and deep learning.
- Hardware advances (faster CPUs, GPUs, large memory) made real-time processing of large images possible.
- Better sensors (CCD and CMOS, depth cameras, satellites, medical scanners) supplied high-resolution and 3-D data.
- Algorithms moved from filters and edges to SIFT and SVM, and then to CNNs that learn features automatically.
- The impact is that applications grew from enhancement to self-driving cars, medical diagnosis and face recognition, and the boundary between CV and IP has blurred.
Asked: [7 marks] (Jun 2025) Discuss the evolution and history of CVIP. How have advancements in technology influenced the development of CVIP?
CV Models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>A computer vision model is a mathematical description of how a scene forms an image (geometry, light, colour) and how that image can be interpreted.</mark>
Key points.
- Geometric models describe image formation: the pinhole camera maps a 3-D point $(X,Y,Z)$ to $x=fX/Z,\ y=fY/Z$ and gives perspective and reconstruction.
- Photometric models describe how illumination, surface reflectance and viewing angle decide pixel brightness, and are used for shape from shading.
- Shape (object) models represent objects as 2-D or 3-D structures such as wireframes, generalised cylinders and CAD models, used in recognition.
- Statistical and learning-based models (Bayes classifiers, SVM, CNN) learn the appearance of classes from data.
- Applications are recognition, 3-D reconstruction, robot navigation and inspection.
Colour vision models.
| Model | Basis | Use |
|---|---|---|
| RGB | Additive red, green, blue; unit cube | Cameras, monitors |
| CMY/CMYK | Subtractive cyan, magenta, yellow (+ black) | Printing |
| HSI/HSV | Hue, saturation, intensity separated | Segmentation by colour, enhancement |
| YIQ/YCbCr | Luminance plus two chrominance | TV, JPEG |
Colour models are needed because RGB mixes intensity with colour, whereas HSI separates them and matches human perception. Conversion is $I=(R+G+B)/3$.
Answer frame. Open with the definition; list geometric, photometric, shape and statistical models with one line each and an application; for the colour question start with the need for colour models, draw the table above and end by comparing RGB with HSI; close with the applications.
Asked: [7 marks] (May 2022, May 2023) Describe various computer vision models and their applications / different types of computer vision models. Asked: [7 marks] (May 2024) Discuss various color vision models.
Image Filtering
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Image filtering modifies pixel values using a neighbourhood or frequency operation to remove noise or enhance features.</mark>
Key points.
- Spatial filtering works directly on pixels: $g(x,y)=\sum_s\sum_t w(s,t)f(x+s,y+t)$ with a mask $w$.
- Smoothing filters (mean, Gaussian, median) reduce noise; the median filter removes salt-and-pepper noise best and preserves edges.
- Sharpening filters (Laplacian, unsharp masking, Sobel gradient) highlight edges and fine detail.
- Frequency filtering follows $G(u,v)=H(u,v)F(u,v)$: take the DFT, multiply by the filter, take the inverse DFT.
- Low-pass filters (ideal, Butterworth, Gaussian) smooth; ideal causes ringing, Gaussian has none, and Butterworth is in between with order $n$.
- High-pass filters sharpen, and homomorphic filtering treats $f=i\cdot r$ by taking the log, then filtering to reduce illumination and boost reflectance.
| Basis | Spatial domain | Frequency domain |
|---|---|---|
| Operates on | Pixels directly | Fourier coefficients |
| Method | Convolution with mask | DFT, multiply, inverse DFT |
| Complexity | Low for small masks | Higher, better for large masks |
| Control | Hard to design exact cutoff | Exact control of frequencies |
| Examples | Median, Laplacian | Butterworth, homomorphic |
Answer frame. Open with the definition and purpose; draw the DFT, filter, inverse DFT block chain for the frequency question; develop spatial then frequency techniques with examples; close with the table. For the "distinguish" question write the table first.
Asked: [7 marks] (May 2023) Discuss in detail the various types of image filtering techniques. Asked: [7 marks] (May 2024) Show the various techniques in frequency domain to enhance an image with necessary examples. Asked: [7 marks] (May 2024) Distinguish between spatial and frequency domain image enhancement.
Image Representations
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>A digital image is a two-dimensional function $f(x,y)$ whose spatial coordinates and amplitude are both discrete, stored as an $M\times N$ matrix.</mark>
Key points.
- A continuous image is converted by sampling followed by quantization.
- Sampling discretises the spatial coordinates $x,y$ into $M\times N$ pixels; more samples give higher spatial resolution.
- Quantization discretises the intensity into $L=2^k$ gray levels; more bits $k$ give more gray levels and less false contouring.
- Coarse sampling causes checkerboard patterns and coarse quantization causes false contours.
- Storage size is $b=M\times N\times k$ bits, so a $256\times256$ 8-bit image needs $524288$ bits, which is $64$ KB.
- An image is represented as a matrix, or as binary, gray-scale, colour or indexed types, and transforms (Fourier, DCT, wavelet) give another representation.
- Image transformation converts an image into another domain to give decorrelation, energy compaction, easier filtering and compression.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-02" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="Digitisation: scene, sensor (continuous image), sampling (spatial), quantization (gray levels), digital image matrix"><style>#dsfig-u1-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-02 .t{fill:#16181D;font-weight:500}#dsfig-u1-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-02 .dot{fill:#16181D}#dsfig-u1-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-02 .ah{fill:#454C5A}#dsfig-u1-02 .ah.hi{fill:#2340B8}#dsfig-u1-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-02 .e{stroke:#B1B7C3}html.dark #dsfig-u1-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-02 .t{fill:#E6E8ED}html.dark #dsfig-u1-02 .t.inv{fill:#0F1115}html.dark #dsfig-u1-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-02 .dot{fill:#E6E8ED}html.dark #dsfig-u1-02 .ann{fill:#8FA3FF}html.dark #dsfig-u1-02 .lbl{fill:#858D9C}html.dark #dsfig-u1-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-02 .ah{fill:#B1B7C3}html.dark #dsfig-u1-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah2)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah2)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah2)"/><path class="e" d="M446,40 L535,40" marker-end="url(#ah2)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Sc</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">Sen</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Smp</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">Qnt</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">Img</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Digitisation: scene, sensor (continuous image), sampling (spatial), quantization (gray levels), digital image matrix</figcaption></figure>
Answer frame. Open with the definition of a digital image; draw the digitisation chain and a small matrix; explain sampling then quantization with resolution effects; close with the storage formula. For the significance question, say representation decides what filtering can do and add filtering (noise removal, enhancement).
Asked: [7 marks] (May 2022, May 2024) Define digital image. What do you mean by image sampling and quantization? Illustrate how the image is digitized. Asked: [7 marks] (Dec 2024) Explain the significance of image representations and image filtering. Asked: [7 marks] (May 2023) What is meant by Image Transformation? Explain its needs in digital image processing.
Image Statistics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Image statistics are numerical measures of the gray-level distribution, taken from the histogram $p(r_k)=n_k/n$.
Key points.
- Mean gray level $\mu=\sum r_kp(r_k)$ shows overall brightness.
- Variance $\sigma^2=\sum (r_k-\mu)^2p(r_k)$ shows contrast.
- Skewness and entropy describe histogram shape and information content.
- They guide thresholding, equalisation and noise estimation.
Conditioning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. Conditioning is the pre-processing that cleans an image so that later stages work reliably.
Key points.
- Conditioning steps are enhancement (contrast stretching, equalisation), filtering (noise removal) and normalisation (size, brightness, orientation).
- Labeling then segments the image and assigns each region a label, typically with connected component labeling.
- Grouping combines related pixels, edges or regions into objects.
- The recognition pipeline is conditioning, labeling, grouping, extracting and matching.
Asked: [7 marks] (Jun 2025) Explain the steps involved in image conditioning, labeling, and grouping for image recognition.
Labeling
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Labeling assigns each pixel or region an identity, such as an object number.
Key points.
- Pixels are first classified by thresholding or edge detection.
- Connected pixels are given the same label using 4- or 8-connectivity.
- Labels let each region be counted and measured separately.
Grouping
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Grouping gathers labeled pixels, edge segments or regions that belong together into one object.
Key points.
- Similarity of intensity, colour, texture or proximity decides the grouping.
- Examples are joining edge fragments into lines and merging regions.
- The output is a set of objects ready for measurement.
Extracting
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Extracting measures properties (features) of each grouped object.
Key points.
- Typical features are area, perimeter, centroid, moments and shape descriptors.
- Good features are compact and invariant to shift, rotation and scale.
- They form the feature vector for matching.
Matching
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Matching compares the extracted features with stored model features to recognise the object.
Key points.
- Methods include template matching, minimum distance classification and structural matching.
- The best-scoring model gives the object's identity.
- A threshold rejects unknown objects.
Morphological Image Processing: Introduction
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Morphological image processing is a set-theory based technique that probes an image with a small shape called a structuring element (SE) to extract or modify shapes.</mark>
Key points.
- Binary images are treated as sets of foreground pixels, and gray-scale images use min and max.
- The SE (for example a $3\times3$ square or cross) has an origin and moves over the image like a mask.
- The primary operations are dilation, erosion, opening and closing; hit-or-miss, thinning and thickening are built from them.
- Applications are noise removal, boundary extraction, skeletons, hole filling and shape analysis.
Asked: [7 marks] (Dec 2024) Define morphological image processing and describe its primary operations.
Dilation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. ==Dilation of $A$ by structuring element $B$ is $A\oplus B=\{z\mid (\hat B)_z\cap A\neq\emptyset\}$: it grows objects.==
Key points.
- The SE is reflected and slid over the image, and the output is 1 wherever it overlaps at least one foreground pixel.
- It expands object boundaries, so objects get larger.
- It fills small holes and gaps and joins nearby broken parts.
- In gray-scale it is a local maximum: $(f\oplus b)(x,y)=\max\{f(x-s,y-t)+b(s,t)\}$, so bright regions grow and dark details shrink.
- It is commutative and increasing, and it is the dual of erosion of the complement.
- Bridging broken characters is a typical use.
Example. A $5\times5$ image has a $3\times3$ block of ones in the centre and the SE is a $3\times3$ cross.
| Operation | Result |
|---|---|
| Dilation | Block plus four pixels at $(0,2),(2,0),(2,4),(4,2)$, so 13 ones |
| Erosion | Only the centre $(2,2)$ survives, so 1 one |
Dilation versus erosion.
| Basis | Dilation | Erosion |
|---|---|---|
| Effect | Expands, grows objects | Shrinks objects |
| Binary set logic | Overlap of SE with A is non-empty (union) | SE fits completely inside A (subset) |
| Gray-scale | Local max | Local min |
| Holes and gaps | Fills them | Enlarges them |
| Small specks | Keeps or enlarges | Removes them |
| Relation | Dual of erosion | Dual of dilation |
Answer frame. Open with the formula and meaning; give the 5x5 example as a small grid; develop points 1-4; for the four-operation question define dilation, erosion, closing and thinning with expressions in one or two lines each. For the differentiate question write the table and add binary and gray-scale rows.
Asked: [14 marks] (May 2022) Explain the following operations: i) dilation ii) erosion iii) closing iv) thinning. Asked: [7 marks] (Dec 2024) Differentiate between dilation and erosion in binary and grayscale images.
Erosion
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Erosion of $A$ by $B$ is $A\ominus B=\{z\mid B_z\subseteq A\}$: it shrinks objects.
Key points.
- A pixel stays 1 only if the SE placed there fits entirely inside the object.
- It removes thin lines and small noise specks and separates touching objects.
- In gray-scale it is a local minimum, so bright details shrink.
- Boundary extraction is $\beta(A)=A-(A\ominus B)$.
Opening
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Opening is erosion followed by dilation: $A\circ B=(A\ominus B)\oplus B$.
Key points.
- It removes small objects, thin protrusions and noise while keeping the main shape size.
- It smooths contours and breaks narrow connections.
- It is idempotent: applying it twice equals applying it once.
Closing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Closing is dilation followed by erosion: $A\bullet B=(A\oplus B)\ominus B$.
Key points.
- It fills small holes and narrow gaps and joins near parts.
- It smooths the contour from outside without much change in size.
- It is idempotent and is the dual of opening.
Hit-or-Miss transformation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Hit-or-miss detects a shape: $A\circledast B=(A\ominus B_1)\cap(A^c\ominus B_2)$, where $B_1$ is the foreground pattern and $B_2$ the background pattern.
Key points.
- It outputs 1 only where the foreground pattern hits the object and the background pattern hits the background.
- It finds specific shapes, corners and isolated points.
- Thinning and thickening are built from it.
Morphological algorithm operations on binary images
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. Binary morphological algorithms apply set operations on foreground pixels using an SE.
Key points.
- Dilation adds pixels around objects wherever the SE touches them, and erosion keeps only pixels where the SE fits.
- Opening and closing combine them for noise removal and gap filling.
- Further algorithms are boundary extraction, region filling, connected component extraction, convex hull, thinning, thickening and skeletons.
- Dilation and erosion are duals: $(A\ominus B)^c=A^c\oplus\hat B$.
Asked: [7 marks] (Jun 2025) Define morphological image processing and describe the operations of dilation and erosion on binary images.
Morphological algorithm operations on gray-scale images
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. Gray-scale morphology uses a structuring function, with max replacing union and min replacing intersection.
Key points.
- Dilation takes the local maximum of $f+b$ and erosion takes the local minimum of $f-b$.
- Opening removes bright small details and closing removes dark small details.
- Smoothing is opening followed by closing, and the morphological gradient is dilation minus erosion.
- Top-hat is $f-(f\circ b)$ and shows bright objects on uneven background.
- Reconstruction rebuilds shapes from a marker image under a mask.
Asked: [7 marks] (May 2022) Discuss morphological algorithm operations performed on gray scale images.
Thinning
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. ==Thinning removes boundary pixels of a binary object, without breaking connectivity, until a one-pixel-wide skeleton remains: $A\otimes B=A-(A\circledast B)$.==
Key points.
- It uses the hit-or-miss transform to find boundary pixels that match a pattern, then subtracts them from the set.
- A sequence of structuring elements $\{B\}=\{B^1,B^2,\dots,B^8\}$ (rotations of the pattern) is applied one after another: $A\otimes\{B\}=((A\otimes B^1)\otimes B^2)\dots\otimes B^n$.
- The process repeats until no further change occurs.
- It preserves the topology and the end points, so the result is the medial line of the shape.
- Thickening is the dual: $A\odot B=A\cup(A\circledast B)$, which adds pixels to grow the object, and equals thinning of the background.
- Applications are fingerprint ridge analysis, OCR stroke thinning, skeletonisation and road-map extraction.
Step 1: Pick the first rotation of the pattern B.
Step 2: Find pixels where A hit-or-miss B is true.
Step 3: Delete them from A.
Step 4: Repeat with the next rotation.
Step 5: Cycle through all rotations until A stops changing.
Answer frame. Open with the thinning definition and expression; write the algo steps; explain points 1-4 in order; for thinning and thickening add point 5 and an application such as fingerprints; close with "thinning gives skeletons for shape analysis". For the two-of-four question write the definition, steps and applications.
Asked: [14 marks] (May 2023) Explain in brief any two of the following: a) Thinning b) Motion-based segmentation c) View class matching d) Knowledge representation. Asked: [7 marks] (Jun 2025) Describe thinning and thickening operations. Provide an example of an application where these operations are useful.
Thickening
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Thickening is $A\odot B=A\cup(A\circledast B)$, the dual of thinning.
Key points.
- It adds pixels to the object using hit-or-miss.
- It is done with rotating structuring elements until stable.
- It fills concavities and grows regions without merging them.
Region growing, region shrinking
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Region growing starts from seed pixels and adds neighbours that satisfy a similarity rule; region shrinking removes boundary pixels of a region.
Key points.
- Region growing needs seeds, a similarity criterion (intensity difference below a threshold) and a stopping rule.
- Growing is an iterative dilation restricted to similar pixels.
- Shrinking is repeated erosion and reduces regions to points or skeletons.
- Growing gives connected regions but depends on seed choice.
Last-minute revision
- Computer vision = understanding a scene; image processing = image in, image out.
- DIP system elements: acquisition, storage, processing, communication, display.
- Digital image: $f(x,y)$ sampled in space and quantized in amplitude; $L=2^k$; size $=M\times N\times k$.
- Frequency filtering: $G=HF$; ideal has ringing, Butterworth is the compromise, Gaussian has none.
- Dilation $A\oplus B$ grows (max); erosion $A\ominus B$ shrinks (min).
- Opening = erosion then dilation; closing = dilation then erosion.
- Hit-or-miss $=(A\ominus B_1)\cap(A^c\ominus B_2)$.
- Thinning $A-(A\circledast B)$; thickening $A\cup(A\circledast B)$.
- 3x3 block with cross SE: dilation gives 13 pixels, erosion gives 1.
- Gradient = dilation minus erosion; top-hat = $f-$ opening.
- Recognition chain: conditioning, labeling, grouping, extracting, matching.
Memory hooks
- "Dilate = Do more, Erode = Eat away."
- "Opening opens gaps (removes specks); closing closes holes."
- "Ideal rings, Butterworth bends, Gaussian glides."
- Pipeline mnemonic: "Clean, Label, Group, Extract, Match" (C-L-G-E-M).
- Sampling is space (x,y); quantization is amplitude (gray).
Coverage checklist
- Basics of CVIP: computer vision and image processing, DIP system elements (Q3, Q4, Q21)
- History of CVIP: timeline (Q10)
- Evolution of CVIP: stages and technology impact (Q9)
- CV Models: geometric, photometric, shape, colour models (Q5, Q6)
- Image Filtering: spatial, frequency, comparison (Q11, Q12, Q13)
- Image Representations: digital image, sampling, quantization, transformation (Q14, Q15, Q20)
- Image Statistics: mean, variance, histogram
- Conditioning: conditioning steps (Q7)
- Labeling: connected labeling (Q7)
- Grouping: grouping (Q7)
- Extracting: features
- Matching: recognition
- Morphological Image Processing: Introduction (Q16)
- Dilation (Q1, Q8)
- Erosion (Q1, Q8)
- Opening
- Closing (Q1)
- Hit-or-Miss transformation
- Morphological algorithm operations on binary images (Q17)
- Morphological algorithm operations on gray-scale images (Q18)
- Thinning (Q1, Q2, Q19)
- Thickening (Q19)
- Region growing, region shrinking