Skip to content
AD-701 · AI for Computer Vision/Quick Revision Short Notes

AI for Computer Vision (AD-701) - Unit 1 Short Notes

How unit 1 is examined

This unit covers what computer vision is, how images form (geometry, light, camera) and the basic operations on images; the marks sit in Computer Vision, geometric transformations, linear filtering and pyramids with wavelets.

Computer Vision

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Computer vision is the field of AI that enables machines to acquire, process and interpret visual data such as images and video, and to make decisions from what they see.</mark>

Key points.

  1. Computer vision is the inverse of image formation: a camera turns a 3D scene into a 2D image, and vision tries to recover the scene's objects, shape and meaning from that image.
  2. Core tasks are image classification (what is in the image), object detection (what and where), segmentation (which pixels belong to which object) and tracking (following objects across frames).
  3. In healthcare it detects tumours in X-ray, CT and MRI scans; in autonomous vehicles it finds lanes, pedestrians and signs; in surveillance it does face recognition and anomaly detection.
  4. Other uses are OCR, industrial defect inspection, agriculture, retail, AR/VR and robotics.
  5. Deep learning (CNNs, then vision transformers) replaced hand-made features and gave large accuracy gains, helped by big datasets and GPUs.
  6. Future trends are self-supervised learning (less labelled data), multimodal vision-language foundation models, edge AI on phones and cameras, and 3D vision such as NeRF and depth sensing.
  7. Their impact is faster diagnosis, safer autonomy, smarter factories and more capable robots.
  8. Challenges are privacy and ethics, bias in training data, high compute and energy cost, and lack of explainability.

Answer frame. Open with the definition; list the four tasks; give three examples (healthcare, vehicles, surveillance); close with the role of deep learning. For the trends question, open with the current state (CNNs, transformers, foundation models), then trends, then impact, then challenges, and close with a line on responsible AI.

Asked: [7 marks] (Dec 2024) Define Computer Vision and explain its significance in modern technology. Asked: [7 marks] (Dec 2024) Discuss future trends in AI for computer vision and their potential impact.

Geometric primitives

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Geometric primitives are the basic elements (points, lines, planes) used to describe shapes in 2D and 3D image geometry.

Key points.

  1. A 2D point is $(x,y)$; its homogeneous form $(x,y,1)$ lets translation be written as a matrix.
  2. A 2D line is $ax+by+c=0$, written as the vector $\mathbf{l}=(a,b,c)$; the point lies on it when $\mathbf{l}\cdot\tilde{\mathbf{x}}=0$.
  3. Two points give a line by the cross product $\mathbf{l}=\tilde{\mathbf{x}}_1\times\tilde{\mathbf{x}}_2$, and two lines meet at $\tilde{\mathbf{x}}=\mathbf{l}_1\times\mathbf{l}_2$.
  4. A 3D plane is $ax+by+cz+d=0$; 3D lines are given by two points or a point and a direction.

Geometric transformations

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>A geometric transformation maps each pixel position $(x,y)$ to a new position $(x',y')$, changing the geometry of the image by a fixed rule.</mark>

Formula. In homogeneous coordinates $\tilde{\mathbf{x}}'=H\tilde{\mathbf{x}}$.

Type Matrix (2D) Preserves
Translation $\begin{bmatrix}1&0&t_x\\0&1&t_y\end{bmatrix}$ orientation, size, shape
Rotation $\begin{bmatrix}\cos\theta&-\sin\theta\\\sin\theta&\cos\theta\end{bmatrix}$ lengths, angles
Scaling $\mathrm{diag}(s_x,s_y)$ angles (if uniform)
Affine $2\times3$, six parameters parallel lines
Perspective $3\times3$ homography straight lines only

Key points.

  1. Translation, rotation, scaling and affine keep parallel lines parallel; perspective does not, since it models a camera's view.
  2. Example: translating $(2,3)$ by $(4,1)$ gives $(6,4)$.
  3. Uses are image registration, panorama stitching, and data augmentation by flipping, rotating and scaling.

Asked: [7 marks] (Dec 2024) What are geometric transformations? List common types.

Photometric image formation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Photometric image formation explains how light from sources, reflected by surfaces and gathered by a lens, produces the brightness value of each pixel.

Key points.

  1. Image brightness (irradiance) depends on the light source, the surface reflectance and the camera's optics and sensor.
  2. Reflectance has a diffuse (Lambertian) part, brightness $\propto \cos\theta$ between normal and light, and a specular part that gives highlights.
  3. The lens focuses light so that irradiance falls with the square of the f-number and with $\cos^4$ off-axis (vignetting).
  4. The sensor then converts this irradiance to pixel values.

Digital camera

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. A digital camera focuses light through a lens onto a CCD or CMOS sensor whose photosites convert photons to charge, which is then digitised into pixels.

Key points.

  1. CCD sensors shift charge out to a common amplifier, giving low noise; CMOS sensors amplify at each pixel, giving lower power, speed and cost.
  2. A Bayer colour filter array puts one red, green or blue filter per photosite, and demosaicing interpolates the missing colours.
  3. The pipeline is sampling, quantisation, white balance, gamma correction and compression (JPEG).
  4. Noise sources are shot noise, read noise and dark current.

Point operators

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. A point operator computes each output pixel from the input pixel at the same location only: $g(x,y)=h(f(x,y))$.

Key points.

  1. Brightness and contrast change is $g=a f+b$, where $a$ is gain (contrast) and $b$ is bias (brightness).
  2. Gamma correction is $g=f^{1/\gamma}$ and fixes non-linear display or sensor response.
  3. Histogram equalisation spreads the intensity histogram to raise contrast.
  4. Thresholding, negatives and colour-space changes are also point operations.

Linear filtering

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Linear filtering replaces each pixel by a weighted sum of its neighbourhood, the weights being given by a kernel (mask) that slides over the image.</mark>

Formula. Correlation: $g(i,j)=\sum_{k,l} f(i+k,j+l)\,h(k,l)$. Convolution flips the kernel: $g(i,j)=\sum_{k,l} f(i-k,j-l)\,h(k,l)$.

Key points.

  1. The filter is linear because $h*(af_1+bf_2)=a\,h*f_1+b\,h*f_2$, and shift-invariant because the same kernel is applied everywhere.
  2. The mean (box) filter has equal weights and blurs; the Gaussian filter weights the centre more and blurs without ringing, and both remove noise.
  3. Sobel kernels give the horizontal and vertical derivatives and detect edges.
  4. Borders need padding (zero, replicate or reflect), and a separable kernel (Gaussian) is applied as two 1D passes to save computation.

Answer frame. Open with the definition; write the convolution formula; draw a $3\times3$ kernel; list mean, Gaussian, Sobel; close with linearity and shift-invariance.

Asked: [7 marks] (Dec 2024) Describe the concept of linear filtering in image processing.

More neighborhood operators

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Neighbourhood operators compute a pixel from its surrounding window; non-linear ones such as median and morphology do not use a weighted sum.

Key points.

  1. The median filter takes the middle value of the window and removes salt-and-pepper noise while keeping edges.
  2. Morphology uses a structuring element: erosion shrinks objects and dilation grows them.
  3. Opening (erosion then dilation) removes small specks; closing (dilation then erosion) fills small holes.
  4. Distance transforms and connected components also work on neighbourhoods.

Fourier transforms

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. The Fourier transform expresses an image as a sum of sinusoids of different spatial frequencies, moving it from the spatial to the frequency domain.

Formula. $F(u,v)=\sum_{x=0}^{M-1}\sum_{y=0}^{N-1} f(x,y)\,e^{-j2\pi(ux/M+vy/N)}$

Key points.

  1. Low frequencies carry smooth regions and high frequencies carry edges and noise.
  2. Convolution in space equals multiplication in frequency, so large filters are done faster with the FFT.
  3. A low-pass filter blurs, a high-pass filter sharpens.

Pyramids

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>An image pyramid is a set of copies of an image at successively lower resolution, built by repeated smoothing and subsampling.</mark>

Key points.

  1. A Gaussian pyramid blurs with a Gaussian and drops every second row and column, so each level is half the size.
  2. A Laplacian pyramid stores the difference between a level and the upsampled next level, $L_i=G_i-\mathrm{expand}(G_{i+1})$.
  3. Pyramids give scale-space representation for coarse-to-fine search, blending and detecting features at any scale.
  4. Wavelets (below) are the closely related multi-resolution tool, used for compression and denoising.

Answer frame. Define pyramid and wavelet; describe Gaussian and Laplacian levels; then state roles (compression, denoising, feature extraction) and the scale-space advantage.

Asked: [7 marks] (Dec 2024) Define pyramids and wavelets and their role in image processing.

Wavelets

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. A wavelet transform splits an image into approximation and detail sub-bands at several scales using short localised waves, giving multi-resolution analysis.

Key points.

  1. The Haar wavelet takes pairwise averages and differences: $[4,6,10,12]$ gives averages $[5,11]$ and details $[-1,-1]$.
  2. Unlike Fourier, wavelets keep both frequency and position information.
  3. Small detail coefficients can be dropped or thresholded, which gives compression (JPEG 2000) and denoising.
  4. Applying the transform to rows and columns gives LL, LH, HL and HH sub-bands.

Global optimization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Global optimization solves vision problems by minimising an energy function over the whole image rather than deciding pixels locally.

Key points.

  1. The energy is $E=E_{data}+\lambda E_{smooth}$: the data term fits the observations and the smoothness term penalises differences between neighbours.
  2. Markov Random Fields (MRFs) model this neighbour structure.
  3. Solvers include graph cuts, belief propagation and simulated annealing.
  4. Uses are denoising, stereo, segmentation and inpainting.

Last-minute revision

  • Computer vision: machines interpret visual data; tasks are classification, detection, segmentation, tracking.
  • Trends: self-supervised, multimodal, edge AI, 3D vision; challenges are ethics, bias, compute.
  • Homogeneous transform: $\tilde{\mathbf{x}}'=H\tilde{\mathbf{x}}$; types are translation, rotation, scaling, affine, perspective.
  • Line: $\mathbf{l}=\mathbf{x}_1\times\mathbf{x}_2$; intersection: $\mathbf{x}=\mathbf{l}_1\times\mathbf{l}_2$.
  • Point operator: $g=af+b$; $a$ is contrast, $b$ is brightness.
  • Convolution: $g(i,j)=\sum f(i-k,j-l)h(k,l)$; linear and shift-invariant.
  • Mean and Gaussian blur; Sobel finds edges; median removes salt-and-pepper noise.
  • Convolution in space equals multiplication in frequency.
  • Gaussian pyramid: blur and halve; Laplacian: $G_i-\mathrm{expand}(G_{i+1})$.
  • Haar: average and difference; $[4,6,10,12]\to[5,11]$ and $[-1,-1]$.
  • Energy: $E=E_{data}+\lambda E_{smooth}$.

Memory hooks

  • Task ladder: Classify, Detect, Segment, Track (what, where, which pixels, over time).
  • TRS-AP for transformations: Translation, Rotation, Scaling, Affine, Perspective.
  • Gaussian pyramid goes down (blur and shrink); Laplacian stores what was lost.
  • Point operator looks at one pixel; neighbourhood operator looks at a window.
  • Wavelet keeps "what frequency and where"; Fourier keeps only "what frequency".

Coverage checklist

  • Computer Vision: Q1 (define and significance), Q2 (future trends).
  • Geometric primitives: not asked, covered.
  • Geometric transformations: Q3 (definition and types).
  • Photometric image formation: not asked, covered.
  • Digital camera: not asked, covered.
  • Point operators: not asked, covered.
  • Linear filtering: Q4.
  • More neighborhood operators: not asked, covered.
  • Fourier transforms: not asked, covered.
  • Pyramids: Q5 (pyramids and wavelets).
  • Wavelets: not asked, covered (also in Q5).
  • Global optimization: not asked, covered.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in