UNIT 5: COMPUTER VISION - COMPREHENSIVE NOTES
I. FOUNDATIONS & CORE CONCEPTS
Definition & Scope of Computer Vision
-
Goal: Enable machines to extract, analyze, and understand meaningful information from digital images or videos.
-
Key Distinction from Image Processing:
-
Image Processing: Focuses on manipulating an image (e.g., enhancing, filtering, compressing). Input and output are both images.
-
Computer Vision: Focuses on interpreting and understanding the image content (e.g., recognizing objects, reconstructing 3D scenes). Output is symbolic information, decisions, or measurements.
-
-
Examples: Facial recognition, autonomous vehicle perception, medical image analysis, augmented reality.
Digital Image Fundamentals
-
Image Representation: A 2D discrete function $f(x, y)$ where $(x, y)$ are spatial coordinates and $f$ is the intensity (grayscale) or color value.
-
Grayscale vs. Color:
-
Grayscale: Single channel, intensity values (e.g., 0-255). $I(x,y) \in \mathbb{R}$.
-
Color (e.g., RGB): Multi-channel (typically 3). Each pixel is a vector $[R, G, B]$. Requires more storage and computation.
-
-
Data Types & Conversions:
-
Common types:
uint8(0-255),uint16,float32. -
Impact: Arithmetic operations can cause overflow/underflow (e.g.,
uint8200 + 100 = 44 due to modulo 256 wrap-around). -
Conversion: Use
cv2.convertScaleAbs(src, alpha, beta)for safe scaling: $$\displaystyle dst = \text{saturate}(| \alpha \cdot src + \beta |) $$.
-
[!TIP] Exam often asks for numerical examples of overflow/underflow. Always check data type before arithmetic.
II. IMAGE ENHANCEMENT & PRE-PROCESSING
Point Processing Operations
-
Brightness/Contrast Adjustment:
-
Brightness: $$\displaystyle I_{\text{new}}(x,y) = I(x,y) + \beta $$. $\beta$ is constant (can be negative).
-
Contrast Stretching: Linear mapping to a new range $[a, b]$.
-
$$I_{\text{new}} = a + \frac{(I - I_{\min}) \cdot (b - a)}{I_{\max} - I_{\min}}$$
* **Numerical Example (Dec 2025)**:
Given $$\displaystyle I_{\min}=0, I_{\max}=200 $$, map to $[0,255]$:
For pixel value 100: $$\displaystyle I_{\text{new}} = 0 + \frac{(100-0) \cdot 255}{200-0} = 127.5 \approx 128 $$.
Histogram & Histogram Equalization (HE)
-
Histogram: Counts frequency of each intensity value $$\displaystyle n_k $$.
-
Histogram Equalization Goal: Transform image to have a uniform histogram, maximizing contrast.
-
Steps (Numerical - Dec 2025):
-
Compute PDF: $$\displaystyle p(r_k) = n_k / (M \times N) $$.
-
Compute CDF: $$\displaystyle s_k = T(r_k) = (L-1) \sum_{j=0}^{k} p(r_j) $$.
-
Map each original intensity $$\displaystyle r_k $$ to $$\displaystyle s_k $$ (round to nearest integer).
-
-
Example (Intensities: [52,55,61,66,70,61,64,73], L=256):
-
Sorted unique: 52(1), 55(1), 61(2), 64(1), 66(1), 70(1), 73(1). Total pixels=8.
-
CDF for 61: $$\displaystyle s = 255 \times ( (1+1+2)/8 ) = 255 \times 0.5 = 127.5 \rightarrow 128 $$.
-
All pixels with original 61 become 128.
-
Adaptive Histogram Equalization & CLAHE
-
Global HE Limitation: Amplifies noise in flat/dark regions.
-
CLAHE:
-
Applies HE locally to small tiles (e.g., 8x8).
-
Contrast Limiting: Clip histogram at a
clip_limit(e.g., 40). Excess pixels redistributed uniformly. -
Tile Interpolation: To avoid border artifacts, use bilinear interpolation between tile transformation functions.
-
-
Parameters:
clip_limit(controls noise amplification),tile_grid_size(controls locality).
Spatial Filtering (Convolution)
-
Convolution Operation:
-
Kernel $K$ (size $m \times n$) slid over image.
-
Output pixel: $$\displaystyle g(x,y) = \sum_{i=-a}^{a} \sum_{j=-b}^{b} f(x+i, y+j) \cdot K(i,j) $$.
-
Often followed by offset addition: $g(x,y) + \text{bias}$.
-
Image Smoothing (Low-Pass Filters)
| Filter | Kernel | Properties | Best For | Drawbacks |
|---|---|---|---|---|
| Box Blur | Uniform $$\displaystyle \frac{1}{m n} $$ | Simple, fast | General smoothing | Poor noise reduction, creates "blocky" artifacts |
| Gaussian Blur | $$\displaystyle \frac{1}{2\pi\sigma^2}e^{-\frac{x^2+y^2}{2\sigma^2}} $$ | Weighted, separable | General smoothing, preserves edges better | Slightly slower |
| Median Blur | Non-linear (median of neighborhood) | Preserves edges | Salt-and-pepper noise | Computationally expensive, not linear |
Edge Detection & Gradient Filters (High-Pass)
-
First-Order Derivative (Gradient):
- Sobel Operator (3x3 kernels):
$$G_x = \begin{bmatrix} -1 & 0 & 1 \\ -2 & 0 & 2 \\ -1 & 0 & 1 \end{bmatrix}, \quad G_y = \begin{bmatrix} -1 & -2 & -1 \\ 0 & 0 & 0 \\ 1 & 2 & 1 \end{bmatrix}$$
* Gradient Magnitude (approximation): $$\displaystyle |G| = |G_x| + |G_y| $$ or $$\displaystyle \sqrt{G_x^2 + G_y^2} $$.
* **Numerical Application**: Apply $$\displaystyle G_x, G_y $$ to a 3x3 image patch, compute magnitude.
-
Second-Order Derivative:
-
Laplacian: $$\displaystyle \nabla^2 f = f(x+1,y) + f(x-1,y) + f(x,y+1) + f(x,y-1) - 4f(x,y) $$.
-
Kernel: $$\displaystyle \begin{bmatrix} 0 & 1 & 0 \\ 1 & -4 & 1 \\ 0 & 1 & 0 \end{bmatrix} $$ or $$\displaystyle \begin{bmatrix} 1 & 1 & 1 \\ 1 & -8 & 1 \\ 1 & 1 & 1 \end{bmatrix} $$.
-
Sensitive to noise, often applied after smoothing (Laplacian of Gaussian - LoG).
-
Canny Edge Detector (Nov 2023)
Multi-stage process:
-
Gaussian Smoothing: Reduce noise.
-
Gradient Magnitude & Direction: Compute $|G|$, $$\displaystyle \theta = \arctan(G_y/G_x) $$.
-
Non-Maximum Suppression: Thin edges. Keep pixel if its gradient magnitude is maximum along $\theta$ direction.
-
Hysteresis Thresholding:
-
High Threshold ($$\displaystyle T_H $$): Strong edges.
-
Low Threshold ($$\displaystyle T_L $$): Weak edges connected to strong edges are kept.
-
Parameters: $$\displaystyle T_H $$, $$\displaystyle T_L $$ (typically $$\displaystyle T_L = 0.4 T_H $$), kernel size ($\sigma$).
-
[!TIP] Canny is optimal for finding thin, well-localized edges. Tuning thresholds is critical.
III. BINARY IMAGE PROCESSING & MORPHOLOGY
Thresholding
-
Global Thresholding: $$\displaystyle I_{\text{bin}}(x,y) = \begin{cases} 1 & \text{if } I(x,y) > T \\ 0 & \text{otherwise} \end{cases} $$.
- Optimal T: Otsu's method (maximizes inter-class variance).
-
Adaptive Thresholding:
-
Local threshold $T(x,y)$ computed from neighborhood (mean or Gaussian-weighted).
-
Numerical Example (Dec 2025): For a 4x4 grid, compute mean of 3x3 neighborhood for each pixel, then threshold.
-
Morphological Operations (Structuring Element $B$)
-
Erosion ($A \ominus B$): "Shrinks" foreground. Pixel in output is 1 only if $B$ is fully contained in $A$.
- Removes small objects, separates connected objects.
-
Dilation ($A \oplus B$): "Expands" foreground. Pixel in output is 1 if any part of $B$ overlaps with $A$.
- Fills small holes/gaps, connects components.
-
Numerical Example (Dec 2025):
$$\displaystyle I = \begin{bmatrix} 0 & 1 & 0 \\ 1 & 1 & 0 \\ 0 & 1 & 1 \end{bmatrix}, B = \begin{bmatrix} 1 & 1 \\ 1 & 1 \end{bmatrix} $$ (origin at top-left).
-
Dilation: Place $B$'s origin at each pixel of $I$. If any 1 in $B$ overlaps a 1 in $I$, output pixel=1.
-
Result: $$\displaystyle \begin{bmatrix} 1 & 1 & 1 \\ 1 & 1 & 1 \\ 0 & 1 & 1 \end{bmatrix} $$.
-
-
Composite Operations:
-
Opening: Erosion followed by Dilation ($A \circ B$). Removes small objects, smoothens contours.
-
Closing: Dilation followed by Erosion ($A \bullet B$). Fills small holes, connects narrow breaks.
-
Connected Component Analysis (CCA)
-
Goal: Label and extract properties of distinct objects in a binary image.
-
Algorithm (4-connected):
-
First Pass: Scan image top-left to bottom-right. Assign provisional labels. Record label equivalences in a lookup table.
-
Second Pass: Resolve equivalences using lookup table. Assign final, consistent labels.
-
-
Output: Labeled image where each connected region has a unique integer label.
-
Properties per Label: Area (pixel count), Centroid ($$\displaystyle \bar{x} = m_{10}/m_{00}, \bar{y} = m_{01}/m_{00} $$), Bounding box, Perimeter.
-
Applications: Object counting, filtering by size/shape.
IV. FEATURE EXTRACTION & SHAPE ANALYSIS
Contour Analysis
-
Contour Detection:
cv2.findContours(binary_image, mode, method)finds boundaries of white blobs.-
Modes:
RETR_EXTERNAL(outer contours only),RETR_TREE(hierarchy). -
Methods:
CHAIN_APPROX_SIMPLE(compresses horizontal/vertical segments),CHAIN_APPROX_NONE(all points).
-
-
Contour Descriptors for Shape Recognition:
-
Moments: $$\displaystyle m_{pq} = \sum_x \sum_y x^p y^q I(x,y) $$. Centroid: $$\displaystyle (\bar{x}, \bar{y}) = (m_{10}/m_{00}, m_{01}/m_{00}) $$.
-
Area:
cv2.contourArea(contour). -
Perimeter:
cv2.arcLength(contour, closed). -
Bounding Shapes:
cv2.boundingRect(axis-aligned),cv2.minAreaRect(rotated). -
Aspect Ratio: $\text{width}/\text{height}$.
-
Extent: $\text{Area} / \text{(bounding box area)}$.
-
Solidity: $\text{Area} / \text{(convex hull area)}$.
-
Approximation:
cv2.approxPolyDP(contour, epsilon, closed)simplifies contour using Douglas-Peucker algorithm.
-
-
Shape Matching:
- Hu Moments: 7 rotation, translation, and scale-invariant moments derived from central moments. Used for shape comparison via
cv2.matchShapes.
- Hu Moments: 7 rotation, translation, and scale-invariant moments derived from central moments. Used for shape comparison via
V. COLOR SPACES & TRANSFORMATIONS
Common Color Models
| Color Space | Channels | Key Properties | Primary Use |
|---|---|---|---|
| RGB | R, G, B | Additive, device-dependent | Displays, standard image storage |
| HSV/HSB | Hue, Saturation, Value | Hue: color type (0-180 in OpenCV). Saturation: color purity. Value: brightness. | Intuitive color segmentation (e.g., "red objects"). |
| CIELAB (L*a*b*) | L* (lightness), a* (green-red), b* (blue-yellow) | Perceptually uniform, device-independent. $$\displaystyle L^* \in [0,100] $$, $$\displaystyle a^*,b^* \in [-128,127] $$. | Precise color correction, delta-E color difference. |
Color Transformation Equations
-
RGB to HSV (Normalized $R,G,B \in [0,1]$):
-
$$\displaystyle V' = \max(R,G,B) $$.
-
$$\displaystyle S' = \begin{cases} \frac{V' - \min(R,G,B)}{V'} & V' \neq 0 \\ 0 & \text{otherwise} \end{cases} $$.
-
$$\displaystyle H' = \begin{cases} \text{undefined} & S'=0 \\ 60^\circ \times \frac{G-B}{V'-\min} \mod 360 & V'=R \\ 120^\circ + 60^\circ \times \frac{B-R}{V'-\min} & V'=G \\ 240^\circ + 60^\circ \times \frac{R-G}{V'-\min} & V'=B \end{cases} $$.
- OpenCV: $H \in [0,179]$, $S,V \in [0,255]$.
-
-
RGB to LAB: Complex, involves:
-
Linear transformation from RGB to XYZ (using reference white D65).
-
Non-linear transformation: $$\displaystyle f(t) = \begin{cases} t^{1/3} & t > \delta^3 \\ \frac{t}{3\delta^2} + \frac{4}{29} & \text{otherwise} \end{cases} $$ where $$\displaystyle \delta = 6/29 $$.
-
$$\displaystyle L^* = 116 f(Y/Y_n) - 16 $$, $$\displaystyle a^* = 500 [f(X/X_n) - f(Y/Y_n)] $$, $$\displaystyle b^* = 200 [f(Y/Y_n) - f(Z/Z_n)] $$.
- Conceptual Goal: Separate lightness ($$\displaystyle L^* $$) from color ($$\displaystyle a^*,b^* $$) in a perceptually uniform way.
-
Applications
-
HSV: Skin detection ($H \in [0,25]$, $$\displaystyle S > 0.2 $$, $$\displaystyle V > 0.5 $$ approx), color-based object tracking.
-
LAB: Curves adjustment in Lightroom/Photoshop (adjust $$\displaystyle L^* $$ for brightness, $$\displaystyle a^*,b^* $$ for color casts).
VI. IMAGE SEGMENTATION
-
Definition: Partitioning an image into multiple homogeneous and meaningful regions/objects.
-
Relation to Recognition: Segmentation provides the regions of interest for subsequent feature extraction and classification. Poor segmentation leads to poor recognition.
-
Methods Overview:
-
Thresholding-based: Global (Otsu), Adaptive, Multi-level.
-
Edge-based: Detect edges (Canny), then link to form closed boundaries.
-
Region-based: Start from seeds (Region Growing), or recursively split/merge (Split-and-Merge).
-
Watershed: Treat gradient magnitude as topographic surface. Flood from markers (minima) to segment.
-
Clustering-based: K-Means in color/feature space. Pixels assigned to nearest cluster centroid.
-
Deep Learning-based: Semantic Segmentation (U-Net, DeepLab) - classifies each pixel. Instance Segmentation (Mask R-CNN) - distinguishes individual object instances.
-
VII. ADVANCED TOPICS & APPLICATIONS
Object Detection Techniques
-
Traditional:
-
Sliding Window + Classifier: Exhaustively scan image at multiple scales with a window, classify each window (e.g., using HOG + Linear SVM).
-
Drawbacks: Computationally expensive, many false positives.
-
-
Deep Learning-based:
-
Two-stage: R-CNN family (Faster R-CNN). First, generate region proposals (RPN), then classify/refine each proposal. Accurate but slower.
-
One-stage: YOLO, SSD. Predict bounding boxes and class probabilities directly over a grid in a single pass. Faster, slightly less accurate.
-
Key Output: Bounding box $(x,y,w,h)$ and class probability $P(\text{class}|\text{object})$.
-
Gesture Recognition
-
Traditional Pipeline:
-
Skin Color Segmentation: Convert to HSV or YCrCb. Threshold on Hue/Saturation or Cr/Cb.
-
Morphology: Clean mask (opening/closing).
-
Contour Detection: Find largest contour (hand).
-
Convex Hull & Defects:
cv2.convexHull,cv2.convexityDefectsto find fingers. -
Recognition: Count defects + 1 for fingers, or use aspect ratio for gestures (e.g., "Open Palm", "Fist").
-
-
Deep Learning Approach:
-
Train CNN (e.g., MobileNet) on hand gesture datasets (e.g., ASL).
-
Use MediaPipe Hands for real-time 21 keypoint landmark detection, then classify based on finger positions.
-
Motion Estimation & Object Tracking
-
Optical Flow:
-
Estimates dense motion field $(u,v)$ for every pixel between two frames.
-
Lucas-Kanade Assumptions:
-
Brightness Constancy: $$\displaystyle I(x,y,t) = I(x+u, y+v, t+1) $$.
-
Small Motion: $u,v$ small.
-
Spatial Coherence: Neighboring pixels have similar motion.
-
-
Solves for $(u,v)$ in a local window using least squares.
-
-
Tracking Algorithms:
-
Correlation-based (Template Matching): Track by finding best match of object template in next frame. Sensitive to scale/rotation.
-
Kalman Filter: Predicts object state (position, velocity) and corrects with noisy measurements. Assumes linear motion and Gaussian noise.
-
Deep Learning-based:
-
SiamFC/SiamRPN: Learn a similarity metric between template and search region.
-
DeepSORT: Combines detection (YOLO) with tracking (Kalman Filter + Re-ID network) for robust multi-object tracking.
-
-
Applications of Deep Learning in Computer Vision
-
Image Classification: ResNet, EfficientNet - assign single label to entire image.
-
Object Detection: YOLO (v3-v8), Faster R-CNN - localize and classify multiple objects.
-
Semantic Segmentation: U-Net (medical imaging), DeepLab - classify every pixel by class (no instance distinction).
-
Instance Segmentation: Mask R-CNN - extends Faster R-CNN with a mask branch for pixel-level instance separation.
-
Image Generation: GANs (StyleGAN), Diffusion Models (Stable Diffusion) - generate novel, realistic images.
[!TIP] For exams, know the key difference: Semantic Segmentation (per-pixel class) vs. Instance Segmentation (per-object instance). Also, know the trade-off: Two-stage (accurate) vs. One-stage (fast) detectors.