Skip to content
AL-701 · Computer Vision/Quick Revision Short Notes

Computer Vision (AL-701) - Unit 5 Short Notes

UNIT 5: COMPUTER VISION - COMPREHENSIVE NOTES

I. FOUNDATIONS & CORE CONCEPTS

Definition & Scope of Computer Vision

  • Goal: Enable machines to extract, analyze, and understand meaningful information from digital images or videos.

  • Key Distinction from Image Processing:

    • Image Processing: Focuses on manipulating an image (e.g., enhancing, filtering, compressing). Input and output are both images.

    • Computer Vision: Focuses on interpreting and understanding the image content (e.g., recognizing objects, reconstructing 3D scenes). Output is symbolic information, decisions, or measurements.

  • Examples: Facial recognition, autonomous vehicle perception, medical image analysis, augmented reality.

Digital Image Fundamentals

  • Image Representation: A 2D discrete function $f(x, y)$ where $(x, y)$ are spatial coordinates and $f$ is the intensity (grayscale) or color value.

  • Grayscale vs. Color:

    • Grayscale: Single channel, intensity values (e.g., 0-255). $I(x,y) \in \mathbb{R}$.

    • Color (e.g., RGB): Multi-channel (typically 3). Each pixel is a vector $[R, G, B]$. Requires more storage and computation.

  • Data Types & Conversions:

    • Common types: uint8 (0-255), uint16, float32.

    • Impact: Arithmetic operations can cause overflow/underflow (e.g., uint8 200 + 100 = 44 due to modulo 256 wrap-around).

    • Conversion: Use cv2.convertScaleAbs(src, alpha, beta) for safe scaling: $$\displaystyle dst = \text{saturate}(| \alpha \cdot src + \beta |) $$.

[!TIP] Exam often asks for numerical examples of overflow/underflow. Always check data type before arithmetic.


II. IMAGE ENHANCEMENT & PRE-PROCESSING

Point Processing Operations

  • Brightness/Contrast Adjustment:

    • Brightness: $$\displaystyle I_{\text{new}}(x,y) = I(x,y) + \beta $$. $\beta$ is constant (can be negative).

    • Contrast Stretching: Linear mapping to a new range $[a, b]$.

$$I_{\text{new}} = a + \frac{(I - I_{\min}) \cdot (b - a)}{I_{\max} - I_{\min}}$$

*   **Numerical Example (Dec 2025)**:

    Given $$\displaystyle I_{\min}=0, I_{\max}=200 $$, map to $[0,255]$:

    For pixel value 100: $$\displaystyle I_{\text{new}} = 0 + \frac{(100-0) \cdot 255}{200-0} = 127.5 \approx 128 $$.

Histogram & Histogram Equalization (HE)

  • Histogram: Counts frequency of each intensity value $$\displaystyle n_k $$.

  • Histogram Equalization Goal: Transform image to have a uniform histogram, maximizing contrast.

  • Steps (Numerical - Dec 2025):

    1. Compute PDF: $$\displaystyle p(r_k) = n_k / (M \times N) $$.

    2. Compute CDF: $$\displaystyle s_k = T(r_k) = (L-1) \sum_{j=0}^{k} p(r_j) $$.

    3. Map each original intensity $$\displaystyle r_k $$ to $$\displaystyle s_k $$ (round to nearest integer).

  • Example (Intensities: [52,55,61,66,70,61,64,73], L=256):

    • Sorted unique: 52(1), 55(1), 61(2), 64(1), 66(1), 70(1), 73(1). Total pixels=8.

    • CDF for 61: $$\displaystyle s = 255 \times ( (1+1+2)/8 ) = 255 \times 0.5 = 127.5 \rightarrow 128 $$.

    • All pixels with original 61 become 128.

Adaptive Histogram Equalization & CLAHE

  • Global HE Limitation: Amplifies noise in flat/dark regions.

  • CLAHE:

    • Applies HE locally to small tiles (e.g., 8x8).

    • Contrast Limiting: Clip histogram at a clip_limit (e.g., 40). Excess pixels redistributed uniformly.

    • Tile Interpolation: To avoid border artifacts, use bilinear interpolation between tile transformation functions.

  • Parameters: clip_limit (controls noise amplification), tile_grid_size (controls locality).

Spatial Filtering (Convolution)

  • Convolution Operation:

    • Kernel $K$ (size $m \times n$) slid over image.

    • Output pixel: $$\displaystyle g(x,y) = \sum_{i=-a}^{a} \sum_{j=-b}^{b} f(x+i, y+j) \cdot K(i,j) $$.

    • Often followed by offset addition: $g(x,y) + \text{bias}$.

Image Smoothing (Low-Pass Filters)
Filter Kernel Properties Best For Drawbacks
Box Blur Uniform $$\displaystyle \frac{1}{m n} $$ Simple, fast General smoothing Poor noise reduction, creates "blocky" artifacts
Gaussian Blur $$\displaystyle \frac{1}{2\pi\sigma^2}e^{-\frac{x^2+y^2}{2\sigma^2}} $$ Weighted, separable General smoothing, preserves edges better Slightly slower
Median Blur Non-linear (median of neighborhood) Preserves edges Salt-and-pepper noise Computationally expensive, not linear
Edge Detection & Gradient Filters (High-Pass)
  • First-Order Derivative (Gradient):

    • Sobel Operator (3x3 kernels):

$$G_x = \begin{bmatrix} -1 & 0 & 1 \\ -2 & 0 & 2 \\ -1 & 0 & 1 \end{bmatrix}, \quad G_y = \begin{bmatrix} -1 & -2 & -1 \\ 0 & 0 & 0 \\ 1 & 2 & 1 \end{bmatrix}$$

*   Gradient Magnitude (approximation): $$\displaystyle |G| = |G_x| + |G_y| $$ or $$\displaystyle \sqrt{G_x^2 + G_y^2} $$.

*   **Numerical Application**: Apply $$\displaystyle G_x, G_y $$ to a 3x3 image patch, compute magnitude.
  • Second-Order Derivative:

    • Laplacian: $$\displaystyle \nabla^2 f = f(x+1,y) + f(x-1,y) + f(x,y+1) + f(x,y-1) - 4f(x,y) $$.

    • Kernel: $$\displaystyle \begin{bmatrix} 0 & 1 & 0 \\ 1 & -4 & 1 \\ 0 & 1 & 0 \end{bmatrix} $$ or $$\displaystyle \begin{bmatrix} 1 & 1 & 1 \\ 1 & -8 & 1 \\ 1 & 1 & 1 \end{bmatrix} $$.

    • Sensitive to noise, often applied after smoothing (Laplacian of Gaussian - LoG).

Canny Edge Detector (Nov 2023)

Multi-stage process:

  1. Gaussian Smoothing: Reduce noise.

  2. Gradient Magnitude & Direction: Compute $|G|$, $$\displaystyle \theta = \arctan(G_y/G_x) $$.

  3. Non-Maximum Suppression: Thin edges. Keep pixel if its gradient magnitude is maximum along $\theta$ direction.

  4. Hysteresis Thresholding:

    • High Threshold ($$\displaystyle T_H $$): Strong edges.

    • Low Threshold ($$\displaystyle T_L $$): Weak edges connected to strong edges are kept.

    • Parameters: $$\displaystyle T_H $$, $$\displaystyle T_L $$ (typically $$\displaystyle T_L = 0.4 T_H $$), kernel size ($\sigma$).

[!TIP] Canny is optimal for finding thin, well-localized edges. Tuning thresholds is critical.


III. BINARY IMAGE PROCESSING & MORPHOLOGY

Thresholding

  • Global Thresholding: $$\displaystyle I_{\text{bin}}(x,y) = \begin{cases} 1 & \text{if } I(x,y) > T \\ 0 & \text{otherwise} \end{cases} $$.

    • Optimal T: Otsu's method (maximizes inter-class variance).
  • Adaptive Thresholding:

    • Local threshold $T(x,y)$ computed from neighborhood (mean or Gaussian-weighted).

    • Numerical Example (Dec 2025): For a 4x4 grid, compute mean of 3x3 neighborhood for each pixel, then threshold.

Morphological Operations (Structuring Element $B$)

  • Erosion ($A \ominus B$): "Shrinks" foreground. Pixel in output is 1 only if $B$ is fully contained in $A$.

    • Removes small objects, separates connected objects.
  • Dilation ($A \oplus B$): "Expands" foreground. Pixel in output is 1 if any part of $B$ overlaps with $A$.

    • Fills small holes/gaps, connects components.
  • Numerical Example (Dec 2025):

    $$\displaystyle I = \begin{bmatrix} 0 & 1 & 0 \\ 1 & 1 & 0 \\ 0 & 1 & 1 \end{bmatrix}, B = \begin{bmatrix} 1 & 1 \\ 1 & 1 \end{bmatrix} $$ (origin at top-left).

    • Dilation: Place $B$'s origin at each pixel of $I$. If any 1 in $B$ overlaps a 1 in $I$, output pixel=1.

    • Result: $$\displaystyle \begin{bmatrix} 1 & 1 & 1 \\ 1 & 1 & 1 \\ 0 & 1 & 1 \end{bmatrix} $$.

  • Composite Operations:

    • Opening: Erosion followed by Dilation ($A \circ B$). Removes small objects, smoothens contours.

    • Closing: Dilation followed by Erosion ($A \bullet B$). Fills small holes, connects narrow breaks.

Connected Component Analysis (CCA)

  • Goal: Label and extract properties of distinct objects in a binary image.

  • Algorithm (4-connected):

    1. First Pass: Scan image top-left to bottom-right. Assign provisional labels. Record label equivalences in a lookup table.

    2. Second Pass: Resolve equivalences using lookup table. Assign final, consistent labels.

  • Output: Labeled image where each connected region has a unique integer label.

  • Properties per Label: Area (pixel count), Centroid ($$\displaystyle \bar{x} = m_{10}/m_{00}, \bar{y} = m_{01}/m_{00} $$), Bounding box, Perimeter.

  • Applications: Object counting, filtering by size/shape.


IV. FEATURE EXTRACTION & SHAPE ANALYSIS

Contour Analysis

  • Contour Detection: cv2.findContours(binary_image, mode, method) finds boundaries of white blobs.

    • Modes: RETR_EXTERNAL (outer contours only), RETR_TREE (hierarchy).

    • Methods: CHAIN_APPROX_SIMPLE (compresses horizontal/vertical segments), CHAIN_APPROX_NONE (all points).

  • Contour Descriptors for Shape Recognition:

    • Moments: $$\displaystyle m_{pq} = \sum_x \sum_y x^p y^q I(x,y) $$. Centroid: $$\displaystyle (\bar{x}, \bar{y}) = (m_{10}/m_{00}, m_{01}/m_{00}) $$.

    • Area: cv2.contourArea(contour).

    • Perimeter: cv2.arcLength(contour, closed).

    • Bounding Shapes: cv2.boundingRect (axis-aligned), cv2.minAreaRect (rotated).

    • Aspect Ratio: $\text{width}/\text{height}$.

    • Extent: $\text{Area} / \text{(bounding box area)}$.

    • Solidity: $\text{Area} / \text{(convex hull area)}$.

    • Approximation: cv2.approxPolyDP(contour, epsilon, closed) simplifies contour using Douglas-Peucker algorithm.

  • Shape Matching:

    • Hu Moments: 7 rotation, translation, and scale-invariant moments derived from central moments. Used for shape comparison via cv2.matchShapes.

V. COLOR SPACES & TRANSFORMATIONS

Common Color Models

Color Space Channels Key Properties Primary Use
RGB R, G, B Additive, device-dependent Displays, standard image storage
HSV/HSB Hue, Saturation, Value Hue: color type (0-180 in OpenCV). Saturation: color purity. Value: brightness. Intuitive color segmentation (e.g., "red objects").
CIELAB (L*a*b*) L* (lightness), a* (green-red), b* (blue-yellow) Perceptually uniform, device-independent. $$\displaystyle L^* \in [0,100] $$, $$\displaystyle a^*,b^* \in [-128,127] $$. Precise color correction, delta-E color difference.

Color Transformation Equations

  • RGB to HSV (Normalized $R,G,B \in [0,1]$):

    1. $$\displaystyle V' = \max(R,G,B) $$.

    2. $$\displaystyle S' = \begin{cases} \frac{V' - \min(R,G,B)}{V'} & V' \neq 0 \\ 0 & \text{otherwise} \end{cases} $$.

    3. $$\displaystyle H' = \begin{cases} \text{undefined} & S'=0 \\ 60^\circ \times \frac{G-B}{V'-\min} \mod 360 & V'=R \\ 120^\circ + 60^\circ \times \frac{B-R}{V'-\min} & V'=G \\ 240^\circ + 60^\circ \times \frac{R-G}{V'-\min} & V'=B \end{cases} $$.

    • OpenCV: $H \in [0,179]$, $S,V \in [0,255]$.
  • RGB to LAB: Complex, involves:

    1. Linear transformation from RGB to XYZ (using reference white D65).

    2. Non-linear transformation: $$\displaystyle f(t) = \begin{cases} t^{1/3} & t > \delta^3 \\ \frac{t}{3\delta^2} + \frac{4}{29} & \text{otherwise} \end{cases} $$ where $$\displaystyle \delta = 6/29 $$.

    3. $$\displaystyle L^* = 116 f(Y/Y_n) - 16 $$, $$\displaystyle a^* = 500 [f(X/X_n) - f(Y/Y_n)] $$, $$\displaystyle b^* = 200 [f(Y/Y_n) - f(Z/Z_n)] $$.

    • Conceptual Goal: Separate lightness ($$\displaystyle L^* $$) from color ($$\displaystyle a^*,b^* $$) in a perceptually uniform way.

Applications

  • HSV: Skin detection ($H \in [0,25]$, $$\displaystyle S > 0.2 $$, $$\displaystyle V > 0.5 $$ approx), color-based object tracking.

  • LAB: Curves adjustment in Lightroom/Photoshop (adjust $$\displaystyle L^* $$ for brightness, $$\displaystyle a^*,b^* $$ for color casts).


VI. IMAGE SEGMENTATION

  • Definition: Partitioning an image into multiple homogeneous and meaningful regions/objects.

  • Relation to Recognition: Segmentation provides the regions of interest for subsequent feature extraction and classification. Poor segmentation leads to poor recognition.

  • Methods Overview:

    • Thresholding-based: Global (Otsu), Adaptive, Multi-level.

    • Edge-based: Detect edges (Canny), then link to form closed boundaries.

    • Region-based: Start from seeds (Region Growing), or recursively split/merge (Split-and-Merge).

    • Watershed: Treat gradient magnitude as topographic surface. Flood from markers (minima) to segment.

    • Clustering-based: K-Means in color/feature space. Pixels assigned to nearest cluster centroid.

    • Deep Learning-based: Semantic Segmentation (U-Net, DeepLab) - classifies each pixel. Instance Segmentation (Mask R-CNN) - distinguishes individual object instances.


VII. ADVANCED TOPICS & APPLICATIONS

Object Detection Techniques

  • Traditional:

    • Sliding Window + Classifier: Exhaustively scan image at multiple scales with a window, classify each window (e.g., using HOG + Linear SVM).

    • Drawbacks: Computationally expensive, many false positives.

  • Deep Learning-based:

    • Two-stage: R-CNN family (Faster R-CNN). First, generate region proposals (RPN), then classify/refine each proposal. Accurate but slower.

    • One-stage: YOLO, SSD. Predict bounding boxes and class probabilities directly over a grid in a single pass. Faster, slightly less accurate.

    • Key Output: Bounding box $(x,y,w,h)$ and class probability $P(\text{class}|\text{object})$.

Gesture Recognition

  • Traditional Pipeline:

    1. Skin Color Segmentation: Convert to HSV or YCrCb. Threshold on Hue/Saturation or Cr/Cb.

    2. Morphology: Clean mask (opening/closing).

    3. Contour Detection: Find largest contour (hand).

    4. Convex Hull & Defects: cv2.convexHull, cv2.convexityDefects to find fingers.

    5. Recognition: Count defects + 1 for fingers, or use aspect ratio for gestures (e.g., "Open Palm", "Fist").

  • Deep Learning Approach:

    • Train CNN (e.g., MobileNet) on hand gesture datasets (e.g., ASL).

    • Use MediaPipe Hands for real-time 21 keypoint landmark detection, then classify based on finger positions.

Motion Estimation & Object Tracking

  • Optical Flow:

    • Estimates dense motion field $(u,v)$ for every pixel between two frames.

    • Lucas-Kanade Assumptions:

      1. Brightness Constancy: $$\displaystyle I(x,y,t) = I(x+u, y+v, t+1) $$.

      2. Small Motion: $u,v$ small.

      3. Spatial Coherence: Neighboring pixels have similar motion.

    • Solves for $(u,v)$ in a local window using least squares.

  • Tracking Algorithms:

    • Correlation-based (Template Matching): Track by finding best match of object template in next frame. Sensitive to scale/rotation.

    • Kalman Filter: Predicts object state (position, velocity) and corrects with noisy measurements. Assumes linear motion and Gaussian noise.

    • Deep Learning-based:

      • SiamFC/SiamRPN: Learn a similarity metric between template and search region.

      • DeepSORT: Combines detection (YOLO) with tracking (Kalman Filter + Re-ID network) for robust multi-object tracking.

Applications of Deep Learning in Computer Vision

  • Image Classification: ResNet, EfficientNet - assign single label to entire image.

  • Object Detection: YOLO (v3-v8), Faster R-CNN - localize and classify multiple objects.

  • Semantic Segmentation: U-Net (medical imaging), DeepLab - classify every pixel by class (no instance distinction).

  • Instance Segmentation: Mask R-CNN - extends Faster R-CNN with a mask branch for pixel-level instance separation.

  • Image Generation: GANs (StyleGAN), Diffusion Models (Stable Diffusion) - generate novel, realistic images.

[!TIP] For exams, know the key difference: Semantic Segmentation (per-pixel class) vs. Instance Segmentation (per-object instance). Also, know the trade-off: Two-stage (accurate) vs. One-stage (fast) detectors.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in