Skip to content
AL-701 · Computer Vision/Quick Revision Short Notes

Computer Vision (AL-701) - Unit 1 Short Notes

UNIT 1: FOUNDATIONS OF IMAGE PROCESSING & ANALYSIS


I. INTRODUCTION & FUNDAMENTAL CONCEPTS

Definition & Scope

  • Image Processing: Input is an image, output is an image. Focuses on modifying/analyzing an existing image (e.g., enhancement, restoration, compression).

  • Computer Vision: Input is an image, output is a description/decision. Focuses on extracting high-level information, making measurements, and enabling decisions (e.g., object recognition, scene understanding).

Aspect Image Processing Computer Vision
Primary Goal Improve image quality/utility Understand & interpret image content
Input/Output Image → Image Image → Description/Decision
Examples Denoising, contrast adjustment, compression Face detection, autonomous driving, medical diagnosis

Digital Image Representation

  • Pixel (Picture Element): Smallest addressable element in a raster image. Stores intensity/color value.

  • Voxel (Volume Pixel): 3D equivalent, used in volumetric data (e.g., CT/MRI scans).

  • Resolution: Number of pixels in width × height (e.g., 1920×1080). Higher resolution = more detail, larger file size.

  • Color vs. Grayscale:

    • Grayscale: Single channel; intensity values (e.g., 0=black, 255=white for 8-bit).

    • Color (RGB): Three channels (Red, Green, Blue). Each pixel is a triplet (R, G, B).

Basic Mathematical Operations

  • Pixel-wise Arithmetic: I_out(x,y) = α * I_in1(x,y) + β * I_in2(x,y) + γ. Used for brightness/contrast adjustment.

  • Logic Operations (AND, OR, NOT): Primarily on binary images. Used for masking, combining segmentation results.

[!TIP] Exam Focus: Be prepared to differentiate IP vs. CV with clear examples. Understand pixel/voxel as fundamental units.


II. IMAGE DATA TYPES & CONVERSIONS

Common Data Types & Ranges

Type (OpenCV) Range Typical Use
uint8 0 to 255 Standard 8-bit images (most common)
uint16 0 to 65535 High dynamic range, medical imaging
float32 ~-3.4e38 to 3.4e38 Filtering, FFT, scientific computation
int16 -32768 to 32767 Some depth maps, signed data

Need for Conversion

  • Filtering: Many filters (Gaussian, Sobel) require floating-point for precision to avoid overflow/underflow and preserve negative values from derivatives.

  • Display: Most displays expect uint8 [0,255].

  • Calculations: Histogram equalization often works on uint8 but intermediate cumulative sums need larger types (int/float) to prevent overflow.

Impact on Quality & Storage

  • Downcasting (e.g., float32 → uint8): Clipping (values <0 → 0, >255 → 255) and quantization loss. Can cause banding or loss of detail.

  • Upcasting (e.g., uint8 → float32): No loss of information, increases storage/memory.

  • Precision Loss: Irreversible when reducing bit depth (e.g., 16-bit → 8-bit).

Practical Implementation (OpenCV)


// Convert to float for filtering, then back to uint8 for display

Mat img_uint8 = imread("image.jpg", IMREAD_GRAYSCALE);

Mat img_float;

img_uint8.convertTo(img_float, CV_32F, 1.0/255.0); // Scale to [0,1]

// ... perform filtering on img_float ...

Mat result_uint8;

img_float.convertTo(result_uint8, CV_8U, 255.0, 0); // Scale back to [0,255]

[!TIP] Common Pitfall: Forgetting to scale float images back to [0,255] before converting to uint8 results in a black image. Always manage scaling factors explicitly.


III. POINT PROCESSING (INTENSITY TRANSFORMATIONS)

Contrast Stretching (Normalization)

  • Goal: Expand the dynamic range of an image to utilize full intensity levels (e.g., 0-255).

  • Linear Stretching:

$$ I_{new} = \frac{(I_{old} - I_{min})}{(I_{max} - I_{min})} \times (L_{new} - 1) + L_{min} $$

where `I_min, I_max` are original min/max, `L_new` is new max level (e.g., 256).
  • Non-linear: Use functions like log, power-law (gamma) for perceptual correction.

Brightness/Intensity Adjustment

  • Addition/Subtraction: I_new = I_old ± c. Simple but causes clipping if values exceed data type limits.

    • Example (from DEC 2025): I = [[100,150,200],[50,100,150],[0,50,100]], c=+50.

    Result: [[150,200,250],[100,150,200],[50,100,150]] → Last element 250 is valid for uint8 (0-255). If c=100, 300 would clip to 255.

Histogram Processing

  • Histogram Definition: Discrete function h(r_k) = n_k, where r_k is intensity level k, n_k is number of pixels with that intensity.

  • Normalized Histogram: p(r_k) = n_k / M*N (probability distribution).

  • Histogram Equalization (Global):

    1. Compute histogram h(r_k).

    2. Compute cumulative distribution function (CDF):

$$ s_k = T(r_k) = (L-1) \sum_{j=0}^{k} p(r_j) $$

3.  Map each pixel intensity `r_k` to `s_k` (rounded to integer).

*   **Effect:** Flattens histogram, maximizes contrast.

*   **Numerical Example (DEC 2025):** Input `[52,55,61,66,70,61,64,73]` (8 pixels).

    *   Sort: `[52,55,61,61,64,66,70,73]`

    *   Compute CDF and map to `[0,255]` (L=256). Step-by-step mapping required in exam.

CLAHE (Contrast Limited Adaptive Histogram Equalization)

  • Concept: Apply histogram equalization locally to small tiles (e.g., 8×8) instead of globally. Prevents over-amplification of noise in low-contrast regions.

  • Key Parameters:

    • clip_limit: Threshold for histogram clipping. Bins above limit are redistributed. Controls noise amplification.

    • tile_grid_size: Number of tiles to divide image into (e.g., (8,8)).

  • Process: For each tile → compute histogram → clip → redistribute excess → perform equalization → interpolate across tile borders to avoid discontinuities (bilinear interpolation).

[!TIP] Exam Must-Know: Derive and compute Histogram Equalization step-by-step. Explain CLAHE's need (local vs. global) and define clip_limit & tile_grid_size.


IV. COLOR SPACES & TRANSFORMATIONS

Overview of Models

  • RGB: Additive, device-dependent (monitors, cameras). 3 channels.

  • Grayscale: Single channel, Y = 0.299R + 0.587G + 0.114B (luminance).

  • Indexed: Palette-based (e.g., GIF). Stores index to color table.

Transformations

1. RGB to HSV/HSI

  • Purpose: Separate chromaticity (Hue, Saturation) from intensity (Value/Brightness). Intuitive for human perception.

  • H (Hue): Dominant wavelength (0°-360°). Red=0°, Green=120°, Blue=240°.

  • S (Saturation): Purity of color (0=gray, 1=full color).

  • V/I (Value/Intensity): Brightness (0=black, 1=maximum brightness).

  • Use Cases: Color-based segmentation (e.g., detect red objects), image editing, color thresholding.

2. RGB to CIELAB (L*a*b*)

  • Purpose: Perceptually uniform space. Equal Euclidean distance ≈ equal perceived color difference.

  • L*: Lightness (0=black, 100=white).

  • a*: Green(-) to Red(+).

  • b*: Blue(-) to Yellow(+).

  • Use Cases: Color difference calculation (ΔE), color-consistent processing, segmentation where human vision matters.

Practical Need for Conversion

  • Segmentation: HSV/HSI easier for "red object" segmentation than RGB (which mixes color with intensity).

  • Display: Convert LAB to RGB for monitor display.

  • Processing: Work in LAB for color-constant operations.

[!TIP] High Frequency: Be ready to write transformation equations (conceptual form is fine) and state purpose/use cases for HSV and LAB.


V. IMAGE FILTERING & SMOOTHING

Convolution Fundamentals

  • Kernel/Filter (W): Small matrix (e.g., 3×3, 5×5) with weights.

  • Anchor Point: Typically center of kernel. Placed over pixel (x,y).

  • Operation: I_out(x,y) = ∑∑ I_in(x+i, y+j) * W(i,j).

  • Border Handling: zero-padding, replicate, reflect, wrap. Affects output size/artifacts.

Smoothing Filters (Noise Reduction)

Filter Kernel Effect Best For
Box/Mean All weights equal (e.g., 1/9 for 3x3) Averages neighborhood. Simple, blocky edges. Mild uniform noise
Gaussian G(x,y) = (1/(2πσ²)) * exp(-(x²+y²)/(2σ²)) Weighted average, more weight to center. Smooth, preserves edges better than box. Gaussian noise, general purpose
Median Non-linear. Sort neighborhood, take median Removes outliers. Preserves edges strongly. Salt-and-pepper noise

Edge Detection & Derivative Filters

First Order Derivatives (Gradient)

  • Concept: Magnitude of intensity change. G = √(G_x² + G_y²), θ = arctan(G_y/G_x).

  • Roberts (2×2): Simple, sensitive to noise.

  • Prewitt (3×3): G_x = [[-1,0,1],[-1,0,1],[-1,0,1]], G_y = transpose.

  • Sobel (3×3) - DERIVATION REQUIRED:

    • G_x Kernel: [[-1,0,1],[-2,0,2],[-1,0,1]]

    • G_y Kernel: [[-1,-2,-1],[0,0,0],[1,2,1]]

    • Derivation: Combines differentiation (e.g., [-1,0,1]) with smoothing (e.g., [1,2,1] column) to reduce noise sensitivity.

    • Numerical Application (DEC 2025): Apply to sample 3×3 image matrix. Compute G_x, G_y, magnitude.

Second Order Derivatives

  • Laplacian: ∇²I = ∂²I/∂x² + ∂²I/∂y². Kernel: [[0,1,0],[1,-4,1],[0,1,0]] or [[1,1,1],[1,-8,1],[1,1,1]].

    • Zero-crossings indicate edges.

    • Very sensitive to noise (doubles noise effect).

  • LoG (Laplacian of Gaussian): Smooth with Gaussian first, then apply Laplacian. Reduces noise sensitivity. Kernel size depends on σ.

Comparison: First vs. Second Order

Property First Order (Gradient) Second Order (Laplacian)
Response Strong at start/end of edge step Strong at zero-crossing (mid-edge)
Edge Thickness Thick (double edge) Thin (single pixel)
Noise Sensitivity Moderate High (amplifies noise)
Directionality Yes (G_x, G_y) No (isotropic)
Typical Use Edge magnitude & orientation Fine edge localization, zero-crossing

Canny Edge Detector

  1. Gaussian Smoothing: Reduce noise (σ parameter).

  2. Gradient Magnitude & Direction: Compute G_x, G_y (often Sobel), then magnitude G and direction θ (quantized to 4 angles: 0°, 45°, 90°, 135°).

  3. Non-Maximum Suppression (NMS): For each pixel, compare G with neighbors along θ. Keep only if local maximum. Thins edges to single pixel.

  4. Hysteresis Thresholding:

    • High Threshold (T_high): Pixels > T_high are strong edges (sure).

    • Low Threshold (T_low): Pixels between T_low and T_high are weak edges.

    • Edge Tracking: Weak pixels connected to strong pixels are kept; others suppressed.

    • Parameters: T_high (typically 2-3× T_low). Controls edge continuity vs. noise.

[!TIP] VERY HIGH: Derive Sobel (explain smoothing+diff). Apply numerically to sample matrix. Compare 1st vs 2nd order with properties. List Canny steps and role of thresholds.


VI. IMAGE SEGMENTATION (THRESHOLDING & MORPHOLOGY)

Thresholding

  • Global Thresholding: I_bin(x,y) = 1 if I(x,y) > T else 0.

    • Fixed T: Manual choice. Simple but fails with varying illumination.

    • Otsu's Method: Automatically determines optimal T by maximizing inter-class variance (or minimizing intra-class variance). Assumes bimodal histogram.

  • Adaptive Thresholding: T varies across image (local).

    • Concept: Compute T(x,y) from neighborhood (mean, Gaussian-weighted mean).

    • Numerical Implementation (DEC 2025): Given 4×4 matrix and window (e.g., 3×3), for each pixel:

      1. Extract neighborhood.

      2. Compute local mean T = mean(neighborhood).

      3. Compare center pixel to T.

    • Use Case: Images with non-uniform illumination.

Morphological Operations (Binary Images)

  • Structuring Element (B): Shape (rectangle, disk, cross), size, origin (reference point, usually center).

  • Erosion:

    • Definition: A ⊖ B = { z | (B_z) ⊆ A }. Output pixel is 1 only if all pixels under B are 1 in input.

    • Effect: Shrinks foreground, removes small objects, separates connected objects.

    • Numerical Example (DEC 2025):

      
      I = [[0,1,0],
      
           [1,1,0],
      
           [0,1,1]]
      
      B = [[1,1],
      
           [1,1]]  (origin at top-left? clarify! Usually center for odd sizes. For 2x2, origin often at (0,0) or (1,1). Be explicit in exam.)
      
      

      Step-by-step: Slide B over I. For each position, check if all 1s in B overlap with 1s in I. Output 1 only if condition met.

  • Dilation:

    • Definition: A ⊕ B = { z | (B_z) ∩ A ≠ ∅ }. Output pixel is 1 if any pixel under B overlaps with 1 in input.

    • Effect: Expands foreground, fills small holes, connects components.

    • Relationship: Dilation(Erosion(A)) = Opening (removes small objects). Erosion(Dilation(A)) = Closing (fills small holes).

Connected Component Analysis (CCA)

  • Goal: Label distinct foreground objects in binary image.

  • Algorithm (2-pass):

    1. First Pass (Scan): Raster scan. For each foreground pixel:

      • If neighbors (typically 4-connected) have labels → assign min label.

      • If multiple different labels → assign min label, record equivalence (e.g., label 2 ≡ 3).

      • If no labeled neighbors → assign new label.

    2. Equivalence Resolution: Build equivalence table. Replace all equivalent labels with a single representative label.

    3. Second Pass (Relabel): Rescan image, replace each label with its final representative label.

  • Output: Labeled image where each connected component has unique integer label.

  • Properties Computed: Area (pixel count), centroid ((x̄, ȳ)), bounding box, etc.

Contour Analysis

  • Contour Detection: Find boundaries of labeled components (e.g., using border following algorithm like Suzuki85 in OpenCV).

  • Contour Approximation: Simplify contour using Douglas-Peucker algorithm (approximate polygon). Parameter epsilon (max distance from original contour).

  • Hierarchies: Contours can be nested (e.g., hole inside object). Represented as [Next, Previous, First_Child, Parent].

  • Shape Recognition Features:

    • Aspect Ratio: Width/Height of bounding rect.

    • Extent: Area / BoundingBoxArea.

    • Solidity: Area / ConvexHullArea.

    • Moments: Central moments (μ_pq) for centroid, orientation, etc.

    • Matching: Hu Moments (7 invariants to rotation, scale, translation), shape context, contour matching (matchShapes in OpenCV).

[!TIP] VERY HIGH: Morphology (Erosion/Dilation) – perform numerical on given matrix. Adaptive Thresholding – implement numerically. CCA steps (2-pass algorithm). Contour features for shape recognition.


VII. ADVANCED TOPICS & APPLICATIONS (From Recent Papers)

Object Detection Techniques

  • Classical:

    • Sliding Window + Classifier: Exhaustive search with SVM/HOG features. Slow.

    • HOG (Histogram of Oriented Gradients): Compute gradient histograms in local cells, normalize in blocks. Used with SVM (e.g., Dalal & Triggs for pedestrian detection).

  • Deep Learning-based:

    • Two-stage: R-CNN family (R-CNN, Fast R-CNN, Faster R-CNN). Propose regions → classify. Accurate but slower.

    • One-stage: YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector). Predict bounding boxes & classes directly from full image in single pass. Fast, real-time.

Motion Estimation & Object Tracking

  • Optical Flow (Lucas-Kanade): Assumes brightness constancy and spatial coherence.

    • Solves for flow vectors (u,v) in small local window using gradient constraints.

    • Equation: I_x u + I_y v + I_t = 0 (underdetermined). Add smoothness constraint.

  • Tracking Paradigms:

    • Kalman Filter: Predicts object state (position, velocity) and corrects with measurements. For linear Gaussian systems.

    • Correlation/Template Tracking: Track by maximizing correlation between template and search region (e.g., CAMShift).

    • Deep Learning Trackers: Siamese networks (e.g., SiamFC), track by learning appearance matching.

Gesture Recognition Pipeline

  1. Detection: Locate hand region (skin color segmentation, depth sensor, object detector).

  2. Tracking: Track hand across frames (optical flow, Kalman filter).

  3. Feature Extraction: Extract hand shape (contours, convex hull defects), keypoints ( fingertips via convexity defects), or use CNN features from ROI.

  4. Classification: Recognize gesture using:

    • Traditional: SVM/Random Forest on hand-crafted features (aspect ratio, number of fingers).

    • Deep Learning: CNN (static gesture), CNN+LSTM (dynamic gestures, sequences).

Applications of Deep Learning in CV

  • Image Classification: Assign label to entire image (e.g., ResNet, EfficientNet).

  • Object Detection: Locate & classify multiple objects (YOLO, Faster R-CNN).

  • Semantic Segmentation: Classify every pixel into a category (e.g., U-Net, DeepLab). No instance distinction.

  • Instance Segmentation: Detect + separate each object instance (e.g., Mask R-CNN). Combines detection + segmentation.

  • Style Transfer: Apply artistic style to image using CNN (e.g., Gatys et al., neural style transfer).

[!TIP] Emerging Focus: Know brief overviews of classical vs. DL object detection. Understand Lucas-Kanade assumptions. Outline Gesture Recognition pipeline steps. List DL applications with one example model each.


Final Exam Strategy:

  1. Definitions First: Always start with clear definitions (e.g., "Morphological dilation is...").

  2. Numerical Steps: For calculations (Histogram Eq, Adaptive Thresh, Morphology), write each intermediate step clearly on a separate line.

  3. Compare & Contrast: Use tables in your mind for (Filter types, Derivative orders, Thresholding methods).

  4. Equations: Memorize key formulas: Histogram Eq CDF, Sobel kernels, Gaussian function, RGB↔HSV (at least conceptually).

  5. Diagrams: Sketch small 3×3 examples for morphology/convolution if allowed. Label structuring element origin.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in