UNIT 3: COMPUTER VISION - SHORT NOTES
I. FOUNDATIONS & FUNDAMENTAL CONCEPTS
Core Definition & Scope
-
Computer Vision (CV): Field enabling machines to extract, interpret, and make decisions from visual data (images/videos). Goal is high-level understanding.
-
Image Processing vs. Computer Vision:
| Aspect | Image Processing | Computer Vision | | :--- | :--- | :--- | | Input | Image | Image/Video | | Output | Modified Image | Information/Description (e.g., labels, coordinates, decisions) | | Goal | Enhancement, restoration, compression | Interpretation, recognition, analysis | | Example | Noise removal, contrast adjustment | Face detection, object recognition, scene reconstruction |
[!TIP] Exam Focus: This distinction is a very frequent 7-mark question. Always illustrate with clear examples (e.g., smoothing an image vs. identifying the smoothed object).
Image Representation & Mathematical Operations
-
Grayscale: Single channel, intensity matrix $I(x,y) \in [0, L-1]$ (e.g., $$\displaystyle L=256 $$ for 8-bit).
-
Color: Multi-channel (e.g., RGB: 3 channels). Pixel is a vector $$\displaystyle \vec{I} = (R, G, B) $$.
-
Mathematical Operations:
-
Point Operations: $$\displaystyle J(x,y) = f(I(x,y)) $$. E.g., brightness adjustment: $$\displaystyle J = I + c $$.
-
Neighborhood Operations (Filtering): $$\displaystyle J(x,y) = \sum_{i=-a}^{a} \sum_{j=-b}^{b} I(x+i, y+j) \cdot K(i,j) $$. This is convolution/correlation.
-
-
Data Type Conversion:
-
Need: Algorithms require specific types (e.g.,
float32for filtering to avoid overflow,uint8for display/storage). -
Impact: Clipping (values outside range are truncated), rounding errors, loss of dynamic range.
-
OpenCV Example:
cv2.convertScaleAbs(src, alpha=1.0, beta=0)performs $$\displaystyle dst = |src \cdot alpha + beta| $$ and converts touint8.
-
II. IMAGE ENHANCEMENT & HISTOGRAM MANIPULATION
Point Processing (Intensity Transformations)
-
Brightness Adjustment: $$\displaystyle J = I + c $$
-
$$\displaystyle c>0 $$: brighter, $$\displaystyle c<0 $$: darker.
-
Numerical Example (Dec 2025): For $$\displaystyle I = \begin{bmatrix} 100 & 150 & 200 \\ 50 & 100 & 150 \\ 0 & 50 & 100 \end{bmatrix} $$, $$\displaystyle J = I + 50 $$.
-
Result: $$\displaystyle \begin{bmatrix} 150 & 200 & 250 \\ 100 & 150 & 200 \\ 50 & 100 & 150 \end{bmatrix} $$. Clipping occurs for values >255 (250 is fine, but if +60, 260→255).
-
-
Contrast Stretching (Linear): $$\displaystyle J = \frac{(I - I_{min}) \cdot (L_{new\_max} - L_{new\_min})}{(I_{max} - I_{min})} + L_{new\_min}} $$
-
Goal: Expand intensity range to full dynamic range.
-
Numerical Example (Dec 2025): For same $I$, $$\displaystyle I_{min}=0, I_{max}=200 $$. New range (0,255).
-
$$\displaystyle J = I \cdot \frac{255}{200} = I \cdot 1.275 $$.
-
Result: $$\displaystyle \begin{bmatrix} 127.5 \to 128 & 191.25 \to 191 & 255 \\ 63.75 \to 64 & 127.5 \to 128 & 191.25 \to 191 \\ 0 & 63.75 \to 64 & 127.5 \to 128 \end{bmatrix} $$ (after rounding).
-
Histogram-Based Techniques
-
Histogram Equalization (HE):
-
Goal: Transform image to have uniform histogram (flat PDF). Enhances global contrast.
-
Process:
-
Compute PDF: $$\displaystyle p_r(r_k) = \frac{n_k}{N} $$.
-
Compute CDF: $$\displaystyle cdf(r_k) = \sum_{j=0}^{k} p_r(r_j) $$.
-
Mapping: $$\displaystyle s_k = (L-1) \cdot cdf(r_k) $$.
-
-
Numerical Example (Dec 2025): Intensities: $[52, 55, 61, 66, 70, 61, 64, 73]$ (N=8).
-
Sorted unique: $$\displaystyle r_0=52, r_1=55, r_2=61, r_3=64, r_4=66, r_5=70, r_6=73 $$.
-
Frequencies: $$\displaystyle n_0=1, n_1=1, n_2=2, n_3=1, n_4=1, n_5=1, n_6=1 $$.
-
PDF: $$\displaystyle p_r = [1/8, 1/8, 2/8, 1/8, 1/8, 1/8, 1/8] $$.
-
CDF: $[0.125, 0.25, 0.5, 0.625, 0.75, 0.875, 1.0]$.
-
Mapping (L=256): $$\displaystyle s_k = 255 \cdot CDF $$. E.g., $$\displaystyle s_0 = 255 \times 0.125 = 31.875 \approx 32 $$.
-
-
-
Adaptive Histogram Equalization (AHE) & CLAHE:
-
Problem with Global HE: Over-amplifies noise in low-contrast regions.
-
AHE: Computes histogram/CDF locally in small tiles around each pixel.
-
CLAHE (Contrast Limited AHE):
-
Divide image into tiles.
-
Compute histogram for each tile, clip histogram at a predefined limit (e.g., 40).
-
Redistribute clipped pixels uniformly.
-
Apply HE using tile's CDF.
-
Interpolate tile borders to avoid artifacts.
-
-
Image Smoothing/Blurring (Noise Reduction)
| Filter | Kernel Type | Effect | Best For | Example |
|---|---|---|---|---|
| Box Blur (Mean) | Uniform weights $$\displaystyle \frac{1}{k^2} $$ | Strong uniform smoothing, edge blurring | General noise reduction | $$\displaystyle K_{3x3} = \frac{1}{9}\begin{bmatrix}1&1&1\\1&1&1\\1&1&1\end{bmatrix} $$ |
| Gaussian Blur | Gaussian weights $$\displaystyle \propto e^{-(x^2+y^2)/(2\sigma^2)} $$ | Smooths while preserving edges better | Pre-processing for edge detection | Separable: $$\displaystyle K_x = \frac{1}{16}\begin{bmatrix}1&2&1\end{bmatrix} $$, $$\displaystyle K_y^T $$ |
| Median Blur | Non-linear (order-statistic) | Excellent for salt-and-pepper noise | Impulse noise | Replaces pixel with median of neighborhood. |
[!TIP] Key Difference: Box/Gaussian are linear (convolution). Median is non-linear (sorting neighborhood values).
III. FILTERING, EDGE DETECTION & GRADIENTS
Convolution & Correlation
-
Convolution: Kernel is flipped (180°) before sliding. Standard for linear filtering.
-
Correlation: Kernel not flipped. Used in template matching.
-
Process: Slide kernel over image, compute weighted sum at each position.
-
Border Handling: Zero-padding, replicate, wrap-around.
Edge Detection: Derivative-Based Filters
-
First Order (Gradient): Detect edges via rapid intensity change.
-
Sobel Operator (Dec 2025 Derivation):
-
Separable Kernels:
$$\displaystyle G_x = K_y^T * K_x = \begin{bmatrix} -1 & 0 & 1 \\ -2 & 0 & 2 \\ -1 & 0 & 1 \end{bmatrix} $$ (horizontal edges)
$$\displaystyle G_y = K_x^T * K_y = \begin{bmatrix} -1 & -2 & -1 \\ 0 & 0 & 0 \\ 1 & 2 & 1 \end{bmatrix} $$ (vertical edges)
where $$\displaystyle K_x = \begin{bmatrix} 1 \\ 2 \\ 1 \end{bmatrix} $$, $$\displaystyle K_y = \begin{bmatrix} 1 & 0 & -1 \end{bmatrix} $$.
-
Gradient Magnitude: $$\displaystyle |\nabla I| = \sqrt{G_x^2 + G_y^2} \approx |G_x| + |G_y| $$.
-
Gradient Direction: $$\displaystyle \theta = \arctan(G_y / G_x) $$.
-
-
Numerical Effect (Sample 3x3): For a step edge, Sobel gives a peak across the edge.
-
-
Second Order (Laplacian): Detect edges via zero-crossings of second derivative.
-
Laplacian Kernel: $$\displaystyle L = \begin{bmatrix} 0 & 1 & 0 \\ 1 & -4 & 1 \\ 0 & 1 & 0 \end{bmatrix} $$ or $$\displaystyle \begin{bmatrix} 1 & 1 & 1 \\ 1 & -8 & 1 \\ 1 & 1 & 1 \end{bmatrix} $$.
-
Laplacian of Gaussian (LoG): Smooth (Gaussian) then Laplacian to reduce noise sensitivity.
-
Difference of Gaussians (DoG): Approximation of LoG: $$\displaystyle DoG = G_{\sigma_1} - G_{\sigma_2} $$.
-
Canny Edge Detector (Nov 2023)
-
Smoothing: Gaussian blur ($\sigma$ controls scale).
-
Gradient: Compute $$\displaystyle G_x, G_y $$, magnitude $|\nabla I|$, direction $\theta$.
-
Non-Maximum Suppression (NMS): Thin edges. For each pixel, compare magnitude with neighbors along gradient direction. Keep if local max.
-
Hysteresis Thresholding:
-
High Threshold (T_high): Pixels > T_high are strong edges.
-
Low Threshold (T_low): Pixels between T_low and T_high are weak edges.
-
Edge Tracking: Weak edge pixel is kept only if connected to a strong edge pixel.
-
Parameters: $\sigma$ (smoothing), T_high, T_low (typically T_low : T_high ≈ 1:2 or 1:3).
-
IV. MORPHOLOGICAL IMAGE PROCESSING (Binary & Grayscale)
Fundamental Operations
-
Structuring Element (B): Small shape/pattern (e.g., 3x3 square) with defined origin. Defines neighborhood.
-
Dilation ($A \oplus B$): Expands foreground (1s).
-
Definition: $$\displaystyle A \oplus B = \{ z | (B_z) \cap A \neq \emptyset \} $$.
-
Effect: Fills holes, connects disjoint objects.
-
Numerical Example (Dec 2025):
$$\displaystyle A = \begin{bmatrix} 0 & 1 & 0 \\ 1 & 1 & 0 \\ 0 & 1 & 1 \end{bmatrix} $$, $$\displaystyle B = \begin{bmatrix} 1 & 1 \\ 1 & 1 \end{bmatrix} $$ (origin at top-left).
-
Place $B$ at each pixel of $A$. If any 1 in $B$ overlaps a 1 in $A$, output pixel = 1.
-
Result: $$\displaystyle \begin{bmatrix} 1 & 1 & 1 \\ 1 & 1 & 1 \\ 1 & 1 & 1 \end{bmatrix} $$.
-
-
-
Erosion ($A \ominus B$): Shrinks foreground (1s).
-
Definition: $$\displaystyle A \ominus B = \{ z | (B_z) \subseteq A \} $$.
-
Effect: Removes small objects, separates joined objects, erodes boundaries.
-
Use for Noise Removal (Nov 2023): Erosion removes small foreground noise (salt). Dilation after erosion (Opening) removes small background noise (pepper).
-
Composite Operations
-
Opening ($A \circ B$): Erosion followed by Dilation: $$\displaystyle A \circ B = (A \ominus B) \oplus B $$.
- Effect: Removes small objects, smoothens contours, breaks narrow connections.
-
Closing ($A \bullet B$): Dilation followed by Erosion: $$\displaystyle A \bullet B = (A \oplus B) \ominus B $$.
- Effect: Fills small holes, closes narrow gaps, connects nearby objects.
-
Boundary Extraction: $$\displaystyle \partial A = A - (A \ominus B) $$.
[!TIP] Mnemonic: Opening = Opening a bag (removes small grains inside). Closing = Closing a bag (fills small holes).
V. IMAGE SEGMENTATION & ANALYSIS
Thresholding
-
Binary Thresholding: $$\displaystyle J(x,y) = \begin{cases} 255 & \text{if } I(x,y) > T \\ 0 & \text{otherwise} \end{cases} $$.
-
Global Threshold (T): Single value for whole image.
-
Otsu's Method: Automatically computes optimal $T$ by maximizing inter-class variance (or minimizing intra-class variance). Assumes bimodal histogram.
-
-
Adaptive Thresholding: Computes $T$ locally for each pixel's neighborhood.
-
Methods: Mean ($$\displaystyle T = \text{mean}(N) $$), Gaussian-weighted mean.
-
Numerical Example (Dec 2025): For a 4x4 image with window size 3x3, compute local mean for each pixel (ignoring borders or padding), then threshold.
-
-
Choosing Optimal T: Analyze histogram (peak separation), Otsu's, trial-and-error.
Connected Component Analysis (CCA) / Labeling (Nov 2023)
-
Goal: Assign unique label to each connected foreground region.
-
Connectivity: 4-connected (N, S, E, W), 8-connected (includes diagonals).
-
Algorithm (Two-Pass):
-
First Pass: Scan image (row-wise). For each foreground pixel:
-
Check already-labeled neighbors (4/8).
-
If none: Assign new provisional label.
-
If one or more: Assign minimum label among neighbors. Record label equivalences (if multiple different labels found).
-
-
Resolve Equivalences: Use union-find/disjoint-set to merge equivalent labels into a single root label.
-
Second Pass: Relabel each pixel using the resolved root labels.
-
-
Output: Labeled image where each connected component has a unique integer value.
-
Applications: Object counting, size filtering (by pixel count), region properties.
Contour Analysis (Dec 2025)
-
Contour: Boundary points of a connected component (open/closed curve).
-
Detection: Find contours from binary image (e.g.,
cv2.findContours). -
Shape Recognition Features (from Contours):
-
Area: $$\displaystyle A = \text{cv2.contourArea}(C) $$.
-
Perimeter: $$\displaystyle P = \text{cv2.arcLength}(C, \text{closed}) $$.
-
Aspect Ratio: $$\displaystyle = \frac{\text{width}}{\text{height}} $$ (bounding rectangle).
-
Extent: $$\displaystyle = \frac{A}{\text{area of bounding rectangle}} $$.
-
Solidity: $$\displaystyle = \frac{A}{\text{area of convex hull}} $$.
-
Moments & Hu Moments: Translation, scale, rotation invariant shape descriptors.
-
VI. COLOR SPACES & TRANSFORMATIONS
Need for Different Color Models
-
RGB: Device-dependent (monitors, cameras). Not perceptually uniform.
-
HSV/HSL/YCrCb/LAB: Separate intensity (luminance) from color (chrominance). More intuitive, perceptually relevant, better for segmentation.
Key Color Spaces & Transformations
| Space | Components | Key Equations (from RGB) | Use Cases |
|---|---|---|---|
| RGB | R, G, B | - | Display, acquisition |
| HSV | H (0-180°), S (0-255), V (0-255) | $$\displaystyle V = \max(R,G,B) $$<br>$$\displaystyle S = \begin{cases} \frac{V - \min}{V} & V\neq0 \\ 0 & \text{else} \end{cases} $$<br>$$\displaystyle H = \begin{cases} \theta & G \geq B \\ 360-\theta & \text{else} \end{cases} $$, $$\displaystyle \theta = \cos^{-1}\left(\frac{(R-G)+(R-B)}{2\sqrt{(R-G)^2+(R-B)(G-B)}}\right) $$ | Color-based segmentation (skin, objects), tracking |
| CIELAB (L*a*b*) | L* (0-100), a* (-128 to 127), b* (-128 to 127) | Complex non-linear (requires RGB→XYZ→LAB).<br>$$\displaystyle \Delta E = \sqrt{(\Delta L^*)^2 + (\Delta a^*)^2 + (\Delta b^*)^2} $$ (color difference) | Perceptual uniformity, color correction, precise difference measurement |
[!TIP] Remember: In OpenCV, HSV H-channel range is [0,179] (not 360), S,V are [0,255]. LAB L* is [0,100], a*,b* roughly [-128,127].
Color Adjustment & Curves (Nov 2023)
-
Curves: Non-linear mapping of input intensity/color channel to output. Allows fine control over shadows, midtones, highlights.
-
Scenarios: White balance correction, color grading in photography/film, medical image enhancement where linear adjustments are insufficient.
VII. ADVANCED TOPICS & APPLICATIONS
Image Segmentation (Comprehensive - Dec 2025)
-
Methods:
-
Thresholding: Global/adaptive.
-
Edge-Based: Canny + contour closing.
-
Region-Based: CCA, region growing, split-and-merge.
-
Clustering: K-means on pixel features (color, texture).
-
-
Relation to Recognition: Segmentation isolates objects of interest from background, providing clean regions for subsequent feature extraction and classification.
Object Detection Techniques (Dec 2025)
-
Traditional:
-
Sliding Window + Classifier: Exhaustive search with hand-crafted features (Haar cascades, HOG) + SVM.
-
Example: Viola-Jones (Haar) for face detection.
-
-
Deep Learning:
-
Single-Stage (YOLO, SSD): Predict bounding boxes & classes directly in one pass. Fast, real-time.
-
Two-Stage (Faster R-CNN): 1) Region Proposal Network (RPN) generates candidate regions. 2) Classify & refine each region. More accurate, slower.
-
Motion Estimation & Object Tracking (Dec 2025)
-
Optical Flow: Estimates dense pixel motion between frames.
-
Lucas-Kanade: Assumes small motion, constant flow in neighborhood. Sparse.
-
Farneback: Dense, uses polynomial expansion.
-
-
Tracking Algorithms:
-
Tracking-by-Detection: Run detector every frame (simple, drift if detection fails).
-
Correlation Trackers (KCF, MOSSE): Learn a discriminative filter of the object, track via correlation.
-
Deep Trackers (SiamFC, SiamRPN): Siamese networks learn similarity metric between template and search region.
-
-
Example: Video surveillance (track person), autonomous vehicles (track pedestrians/other cars).
Gesture Recognition (High Priority - Both Papers)
-
Process (OpenCV + DL):
-
Hand Detection: Skin color segmentation (HSV/YCrCb), depth sensor (Kinect), or object detector (YOLO).
-
Preprocessing: Thresholding, morphological cleaning, contour extraction.
-
Feature Extraction:
-
Traditional: HOG on hand region, convex hull defects (fingers).
-
Deep Learning: CNN features (MediaPipe Hands, custom CNN).
-
-
Classification/Recognition:
-
ML: SVM, Random Forest on extracted features.
-
DL: Pre-trained CNNs (MobileNet), or custom models trained on gesture datasets.
-
-
-
HCI Enablement: Control devices (volume, cursor), sign language interpretation, gaming, VR/AR interaction.
Applications of Deep Learning in CV (Dec 2025)
-
Image Classification: ResNet, EfficientNet (identify what is in image).
-
Object Detection: YOLO, Faster R-CNN (localize & classify multiple objects).
-
Semantic Segmentation: U-Net, DeepLab (classify every pixel by category).
-
Instance Segmentation: Mask R-CNN (pixel-level mask for each object instance).
-
Image Synthesis: GANs (generate images), Diffusion Models (high-quality synthesis).
-
Pose Estimation: Detect human joints (OpenPose, HRNet).
-
Super-resolution: Enhance resolution (SRCNN, ESRGAN).