How unit 1 is examined
This unit covers what computer vision is, how images form (geometry, light, camera) and the basic operations on images; the marks sit in Computer Vision, geometric transformations, linear filtering and pyramids with wavelets.
Computer Vision
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Computer vision is the field of AI that enables machines to acquire, process and interpret visual data such as images and video, and to make decisions from what they see.</mark>
Key points.
- Computer vision is the inverse of image formation: a camera turns a 3D scene into a 2D image, and vision tries to recover the scene's objects, shape and meaning from that image.
- Core tasks are image classification (what is in the image), object detection (what and where), segmentation (which pixels belong to which object) and tracking (following objects across frames).
- In healthcare it detects tumours in X-ray, CT and MRI scans; in autonomous vehicles it finds lanes, pedestrians and signs; in surveillance it does face recognition and anomaly detection.
- Other uses are OCR, industrial defect inspection, agriculture, retail, AR/VR and robotics.
- Deep learning (CNNs, then vision transformers) replaced hand-made features and gave large accuracy gains, helped by big datasets and GPUs.
- Future trends are self-supervised learning (less labelled data), multimodal vision-language foundation models, edge AI on phones and cameras, and 3D vision such as NeRF and depth sensing.
- Their impact is faster diagnosis, safer autonomy, smarter factories and more capable robots.
- Challenges are privacy and ethics, bias in training data, high compute and energy cost, and lack of explainability.
Answer frame. Open with the definition; list the four tasks; give three examples (healthcare, vehicles, surveillance); close with the role of deep learning. For the trends question, open with the current state (CNNs, transformers, foundation models), then trends, then impact, then challenges, and close with a line on responsible AI.
Asked: [7 marks] (Dec 2024) Define Computer Vision and explain its significance in modern technology. Asked: [7 marks] (Dec 2024) Discuss future trends in AI for computer vision and their potential impact.
Geometric primitives
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Geometric primitives are the basic elements (points, lines, planes) used to describe shapes in 2D and 3D image geometry.
Key points.
- A 2D point is $(x,y)$; its homogeneous form $(x,y,1)$ lets translation be written as a matrix.
- A 2D line is $ax+by+c=0$, written as the vector $\mathbf{l}=(a,b,c)$; the point lies on it when $\mathbf{l}\cdot\tilde{\mathbf{x}}=0$.
- Two points give a line by the cross product $\mathbf{l}=\tilde{\mathbf{x}}_1\times\tilde{\mathbf{x}}_2$, and two lines meet at $\tilde{\mathbf{x}}=\mathbf{l}_1\times\mathbf{l}_2$.
- A 3D plane is $ax+by+cz+d=0$; 3D lines are given by two points or a point and a direction.
Geometric transformations
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>A geometric transformation maps each pixel position $(x,y)$ to a new position $(x',y')$, changing the geometry of the image by a fixed rule.</mark>
Formula. In homogeneous coordinates $\tilde{\mathbf{x}}'=H\tilde{\mathbf{x}}$.
| Type | Matrix (2D) | Preserves |
|---|---|---|
| Translation | $\begin{bmatrix}1&0&t_x\\0&1&t_y\end{bmatrix}$ | orientation, size, shape |
| Rotation | $\begin{bmatrix}\cos\theta&-\sin\theta\\\sin\theta&\cos\theta\end{bmatrix}$ | lengths, angles |
| Scaling | $\mathrm{diag}(s_x,s_y)$ | angles (if uniform) |
| Affine | $2\times3$, six parameters | parallel lines |
| Perspective | $3\times3$ homography | straight lines only |
Key points.
- Translation, rotation, scaling and affine keep parallel lines parallel; perspective does not, since it models a camera's view.
- Example: translating $(2,3)$ by $(4,1)$ gives $(6,4)$.
- Uses are image registration, panorama stitching, and data augmentation by flipping, rotating and scaling.
Asked: [7 marks] (Dec 2024) What are geometric transformations? List common types.
Photometric image formation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Photometric image formation explains how light from sources, reflected by surfaces and gathered by a lens, produces the brightness value of each pixel.
Key points.
- Image brightness (irradiance) depends on the light source, the surface reflectance and the camera's optics and sensor.
- Reflectance has a diffuse (Lambertian) part, brightness $\propto \cos\theta$ between normal and light, and a specular part that gives highlights.
- The lens focuses light so that irradiance falls with the square of the f-number and with $\cos^4$ off-axis (vignetting).
- The sensor then converts this irradiance to pixel values.
Digital camera
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A digital camera focuses light through a lens onto a CCD or CMOS sensor whose photosites convert photons to charge, which is then digitised into pixels.
Key points.
- CCD sensors shift charge out to a common amplifier, giving low noise; CMOS sensors amplify at each pixel, giving lower power, speed and cost.
- A Bayer colour filter array puts one red, green or blue filter per photosite, and demosaicing interpolates the missing colours.
- The pipeline is sampling, quantisation, white balance, gamma correction and compression (JPEG).
- Noise sources are shot noise, read noise and dark current.
Point operators
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A point operator computes each output pixel from the input pixel at the same location only: $g(x,y)=h(f(x,y))$.
Key points.
- Brightness and contrast change is $g=a f+b$, where $a$ is gain (contrast) and $b$ is bias (brightness).
- Gamma correction is $g=f^{1/\gamma}$ and fixes non-linear display or sensor response.
- Histogram equalisation spreads the intensity histogram to raise contrast.
- Thresholding, negatives and colour-space changes are also point operations.
Linear filtering
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Linear filtering replaces each pixel by a weighted sum of its neighbourhood, the weights being given by a kernel (mask) that slides over the image.</mark>
Formula. Correlation: $g(i,j)=\sum_{k,l} f(i+k,j+l)\,h(k,l)$. Convolution flips the kernel: $g(i,j)=\sum_{k,l} f(i-k,j-l)\,h(k,l)$.
Key points.
- The filter is linear because $h*(af_1+bf_2)=a\,h*f_1+b\,h*f_2$, and shift-invariant because the same kernel is applied everywhere.
- The mean (box) filter has equal weights and blurs; the Gaussian filter weights the centre more and blurs without ringing, and both remove noise.
- Sobel kernels give the horizontal and vertical derivatives and detect edges.
- Borders need padding (zero, replicate or reflect), and a separable kernel (Gaussian) is applied as two 1D passes to save computation.
Answer frame. Open with the definition; write the convolution formula; draw a $3\times3$ kernel; list mean, Gaussian, Sobel; close with linearity and shift-invariance.
Asked: [7 marks] (Dec 2024) Describe the concept of linear filtering in image processing.
More neighborhood operators
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Neighbourhood operators compute a pixel from its surrounding window; non-linear ones such as median and morphology do not use a weighted sum.
Key points.
- The median filter takes the middle value of the window and removes salt-and-pepper noise while keeping edges.
- Morphology uses a structuring element: erosion shrinks objects and dilation grows them.
- Opening (erosion then dilation) removes small specks; closing (dilation then erosion) fills small holes.
- Distance transforms and connected components also work on neighbourhoods.
Fourier transforms
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. The Fourier transform expresses an image as a sum of sinusoids of different spatial frequencies, moving it from the spatial to the frequency domain.
Formula. $F(u,v)=\sum_{x=0}^{M-1}\sum_{y=0}^{N-1} f(x,y)\,e^{-j2\pi(ux/M+vy/N)}$
Key points.
- Low frequencies carry smooth regions and high frequencies carry edges and noise.
- Convolution in space equals multiplication in frequency, so large filters are done faster with the FFT.
- A low-pass filter blurs, a high-pass filter sharpens.
Pyramids
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>An image pyramid is a set of copies of an image at successively lower resolution, built by repeated smoothing and subsampling.</mark>
Key points.
- A Gaussian pyramid blurs with a Gaussian and drops every second row and column, so each level is half the size.
- A Laplacian pyramid stores the difference between a level and the upsampled next level, $L_i=G_i-\mathrm{expand}(G_{i+1})$.
- Pyramids give scale-space representation for coarse-to-fine search, blending and detecting features at any scale.
- Wavelets (below) are the closely related multi-resolution tool, used for compression and denoising.
Answer frame. Define pyramid and wavelet; describe Gaussian and Laplacian levels; then state roles (compression, denoising, feature extraction) and the scale-space advantage.
Asked: [7 marks] (Dec 2024) Define pyramids and wavelets and their role in image processing.
Wavelets
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A wavelet transform splits an image into approximation and detail sub-bands at several scales using short localised waves, giving multi-resolution analysis.
Key points.
- The Haar wavelet takes pairwise averages and differences: $[4,6,10,12]$ gives averages $[5,11]$ and details $[-1,-1]$.
- Unlike Fourier, wavelets keep both frequency and position information.
- Small detail coefficients can be dropped or thresholded, which gives compression (JPEG 2000) and denoising.
- Applying the transform to rows and columns gives LL, LH, HL and HH sub-bands.
Global optimization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Global optimization solves vision problems by minimising an energy function over the whole image rather than deciding pixels locally.
Key points.
- The energy is $E=E_{data}+\lambda E_{smooth}$: the data term fits the observations and the smoothness term penalises differences between neighbours.
- Markov Random Fields (MRFs) model this neighbour structure.
- Solvers include graph cuts, belief propagation and simulated annealing.
- Uses are denoising, stereo, segmentation and inpainting.
Last-minute revision
- Computer vision: machines interpret visual data; tasks are classification, detection, segmentation, tracking.
- Trends: self-supervised, multimodal, edge AI, 3D vision; challenges are ethics, bias, compute.
- Homogeneous transform: $\tilde{\mathbf{x}}'=H\tilde{\mathbf{x}}$; types are translation, rotation, scaling, affine, perspective.
- Line: $\mathbf{l}=\mathbf{x}_1\times\mathbf{x}_2$; intersection: $\mathbf{x}=\mathbf{l}_1\times\mathbf{l}_2$.
- Point operator: $g=af+b$; $a$ is contrast, $b$ is brightness.
- Convolution: $g(i,j)=\sum f(i-k,j-l)h(k,l)$; linear and shift-invariant.
- Mean and Gaussian blur; Sobel finds edges; median removes salt-and-pepper noise.
- Convolution in space equals multiplication in frequency.
- Gaussian pyramid: blur and halve; Laplacian: $G_i-\mathrm{expand}(G_{i+1})$.
- Haar: average and difference; $[4,6,10,12]\to[5,11]$ and $[-1,-1]$.
- Energy: $E=E_{data}+\lambda E_{smooth}$.
Memory hooks
- Task ladder: Classify, Detect, Segment, Track (what, where, which pixels, over time).
- TRS-AP for transformations: Translation, Rotation, Scaling, Affine, Perspective.
- Gaussian pyramid goes down (blur and shrink); Laplacian stores what was lost.
- Point operator looks at one pixel; neighbourhood operator looks at a window.
- Wavelet keeps "what frequency and where"; Fourier keeps only "what frequency".
Coverage checklist
- Computer Vision: Q1 (define and significance), Q2 (future trends).
- Geometric primitives: not asked, covered.
- Geometric transformations: Q3 (definition and types).
- Photometric image formation: not asked, covered.
- Digital camera: not asked, covered.
- Point operators: not asked, covered.
- Linear filtering: Q4.
- More neighborhood operators: not asked, covered.
- Fourier transforms: not asked, covered.
- Pyramids: Q5 (pyramids and wavelets).
- Wavelets: not asked, covered (also in Q5).
- Global optimization: not asked, covered.