UNIT 2: COMPUTER VISION - CORE CONCEPTS & TECHNIQUES
I. FOUNDATIONS & IMAGE REPRESENTATION
Definition and Scope
-
Computer Vision (CV): Field enabling machines to interpret and understand visual data from the world (e.g., recognize objects, track motion, reconstruct scenes). Output is a description or decision.
-
Image Processing (IP): Field focused on manipulating an image to enhance it or extract useful information. Input and output are both images.
Example:
• IP: Adjusting brightness/contrast of a dark photo.
• CV: Detecting faces in that photo and labeling them.
Image Fundamentals
-
Grayscale Image: Single channel, intensity values per pixel (e.g., 0-255 for
uint8). Represented as a 2D matrixI[m,n]. -
Color Image: Multiple channels (typically 3 for RGB). Represented as a 3D matrix
I[m,n,c]wherecis channel index. -
Digital Image: A 2D discrete array of pixels. Pixel coordinate
(x,y)refers to column (width) and row (height).
Data Type Conversion (High Frequency)
-
Concept: Converting pixel data between types (
uint8,float32,bool) to suit operations. -
Impact:
-
uint8(0-255): Standard for storage/display. Arithmetic can cause overflow/underflow (wrap-around). -
float32(0.0-1.0 or 0-255): Allows precise math, no wrap-around. Essential for filtering, normalization. -
bool: Binary images (0/1 or True/False) for morphological operations.
-
-
OpenCV Example:
img_uint8 = cv2.imread('image.jpg') # Default uint8 img_float = img_uint8.astype('float32') / 255.0 # Normalize to [0,1] img_uint8_back = cv2.convertScaleAbs(img_float * 255) # Convert back safely
!TIP: Always be mindful of range.
uint8addition200 + 100 = 44(wrap-around). Usefloatfor intermediate calculations.
II. POINT PROCESSING & IMAGE ENHANCEMENT
Basic Intensity Transformations
- Brightness Adjustment: Add/subtract a constant
c.
$$g(x,y) = f(x,y) + c$$
* `c > 0`: Brighter. `c < 0`: Darker.
* **Must clamp** result to valid range (e.g., 0-255).
-
Contrast Stretching (Linear Contrast Adjustment): (High Frequency)
-
Goal: Map input intensity range
[a,b]to a new full range[c,d](usually[0,255]). -
Formula:
-
$$g(x,y) = \frac{(d-c)}{(b-a)} \left( f(x,y) - a \right) + c$$
* **Numerical Example (Dec 2025):** For image `I` with min=0, max=100, stretch to `[0,255]`.
$$g = \frac{255}{100} \times (f - 0) + 0 = 2.55 \times f$$
* Pixel `100` → `255`, `50` → `127.5 ≈ 128`.
Histogram-Based Processing
-
Histogram Equalization (HE): (High Frequency)
-
Goal: Transform image to have a uniform histogram, enhancing global contrast.
-
Process:
-
Compute Probability Density Function (PDF) of intensities.
-
Compute Cumulative Distribution Function (CDF).
-
Map each original intensity
r_kto new intensitys_k:
-
-
$$s_k = \text{round}\left( (L-1) \times \sum_{j=0}^{k} p_r(r_j) \right)$$
where `L` = number of possible intensity levels (e.g., 256).
* **Numerical Example (Dec 2025):** For intensities `[52,55,61,66,70,61,64,73]` (8 pixels).
1. Sort & bin: `[52,55,61,61,64,66,70,73]`. PDF: `52:1/8, 55:1/8, 61:2/8, 64:1/8, 66:1/8, 70:1/8, 73:1/8`.
2. CDF: `52:0.125, 55:0.25, 61:0.5, 64:0.625, 66:0.75, 70:0.875, 73:1.0`.
3. New `s_k = 255 * CDF`. E.g., `61 → 255*0.5 = 127.5 ≈ 128`.
* **Limitation:** Can over-amplify noise in homogeneous regions.
-
CLAHE (Contrast Limited Adaptive HE): (High Frequency)
-
Motivation: Overcome HE's global nature problem.
-
Process:
-
Divide image into small tiles (e.g., 8x8).
-
Compute histogram for each tile.
-
Clip histogram at a predefined limit (e.g., 40) to control noise amplification.
-
Perform HE on clipped histogram.
-
Interpolate across tile borders to eliminate artifacts.
-
-
Key Parameter:
clip_limit(controls amplification).
-
III. FILTERING, EDGES & GRADIENTS
Convolution Fundamentals
- Operation: Slide a small matrix (kernel/filter) over an image. At each position, compute the sum of element-wise products.
$$g(x,y) = \sum_{i=-a}^{a} \sum_{j=-b}^{b} f(x-i, y-j) \cdot h(i,j)$$
where `h` is the kernel.
- Border Handling: Zero-padding, replicate, wrap-around. Affects output size/artifacts.
Derivative Filters & Edge Detection
-
First-Order Derivative: (High Frequency)
-
Concept: Edge occurs at intensity gradient maximum. Sensitive to noise.
-
Examples: Roberts (
[[1,0],[0,-1]],[[0,1],[-1,0]]), Prewitt (3x3).
-
-
Second-Order Derivative: (High Frequency)
-
Concept: Edge occurs at zero-crossing of Laplacian. More sensitive to noise, but finer edges.
-
Laplacian Kernel:
[[0,1,0],[1,-4,1],[0,1,0]]or[[1,1,1],[1,-8,1],[1,1,1]].
-
-
Differentiation:
| Property | First-Order (Gradient) | Second-Order (Laplacian) | | :--- | :--- | :--- | | Response | Strong at edges | Strong at edges & isolated points | | Noise | Less sensitive | Very sensitive | | Edge Localization | Good (thick edges) | Better (thin edges) | | Polarity | Has direction (sign) | No direction (magnitude only) |
The Sobel Operator (High Frequency)
-
Derivation: Approximates gradient
∇f = [Gx, Gy]^T.-
Gx (horizontal edges):
[[-1,0,1],[-2,0,2],[-1,0,1]] -
Gy (vertical edges):
[[-1,-2,-1],[0,0,0],[1,2,1]]
-
-
Gradient Magnitude & Direction:
$$G = \sqrt{G_x^2 + G_y^2} \approx |G_x| + |G_y|$$
$$\theta = \tan^{-1}(G_y / G_x)$$
-
Numerical Demonstration (Dec 2025): Apply to sample 3x3 region.
!TIP: Sobel is a smoothed derivative (Gaussian + gradient), so it's more noise-robust than simple Prewitt.
Canny Edge Detector (High Frequency - Nov 2023)
-
Multi-stage Process:
-
Gaussian Smoothing: Reduce noise (sigma parameter).
-
Gradient Magnitude & Direction: Compute
Gandθ(Sobel). -
Non-Maximum Suppression (NMS): Thin edges. Keep pixel if its gradient magnitude is local max along
θdirection. -
Hysteresis Thresholding:
-
High Threshold (
T_high): Strong edges (sure edges). -
Low Threshold (
T_low): Weak edges. Only kept if connected to a strong edge. -
Result: Connected strong edges + relevant weak edges.
-
-
-
Parameters:
sigma(smoothing),T_low,T_high(typicallyT_high ≈ 2-3 * T_low).
IV. MORPHOLOGICAL IMAGE PROCESSING (Binary Images)
Fundamental Operations (Very High Frequency)
-
Structuring Element (SE): Small shape (kernel) defining neighborhood. Has origin (anchor point). Shape/size/origin drastically affect result.
-
Erosion:
-
Concept:
A ⊖ B. Output pixel is1only if all pixels under SEBare1in imageA. Shrinks foreground, removes small objects/connections. -
Numerical Example (Dec 2025):
-
$$I = \begin{bmatrix} 0 & 1 & 0 \\ 1 & 1 & 0 \\ 0 & 1 & 1 \end{bmatrix}, B = \begin{bmatrix} 1 & 1 \\ 1 & 1 \end{bmatrix}$$
* Place SE origin on each `1` in `I`. If all SE positions overlap `1`s in `I`, output `1`; else `0`.
* Result: `[[0,0,0],[0,1,0],[0,0,0]]` (only center pixel survives).
-
Dilation:
-
Concept:
A ⊕ B. Output pixel is1if any pixel under SEBoverlaps a1inA. Expands foreground, fills small holes, connects components. -
Numerical Example (same
I,B):-
Place SE. If any
1inIunder SE, output1. -
Result:
[[1,1,1],[1,1,1],[1,1,1]](all become 1).
-
-
-
Opening & Closing:
-
Opening:
A ∘ B = (A ⊖ B) ⊕ B. Erosion followed by Dilation.- Use: Remove small objects/noise, smooth contours, separate objects.
-
Closing:
A • B = (A ⊕ B) ⊖ B. Dilation followed by Erosion.- Use: Fill small holes, close narrow gaps, smooth contours.
-
!TIP: Erosion removes boundary pixels; Dilation adds boundary pixels. Opening removes protrusions; Closing fills indentations.
V. THRESHOLDING & BINARY IMAGE ANALYSIS
Global Thresholding
-
Concept: Simple comparison:
g(x,y) = 1 if f(x,y) > T else 0. -
Choosing
T:-
Visual inspection of histogram.
-
Otsu's Method: Automatically finds
Tthat maximizes inter-class variance (assumes bimodal histogram).
-
Adaptive/Local Thresholding (High Frequency)
-
Concept: Threshold
T(x,y)varies across image. For uneven illumination. -
Method (Mean of Neighborhood):
$$T(x,y) = \text{mean of pixels in } N_{x,y}$$
where `N` is a small window (e.g., 11x11).
* `g(x,y) = 1 if f(x,y) > T(x,y) else 0`
-
Numerical Example (Nov 2023): For 4x4 image, window=2x2 (with padding).
Compute
Tfor each pixel as mean of its 2x2 neighborhood (including itself), then threshold.
Connected Component Analysis (CCA/Labeling) (High Frequency)
-
Goal: Assign unique labels to all connected foreground pixels (8-connectivity common).
-
Algorithm (Two-Pass):
-
First Pass: Scan image top-left to bottom-right.
-
For each foreground pixel
p, examine its already-scanned neighbors (e.g., top, left, top-left, top-right for 8-conn). -
If no labeled neighbors → assign new label.
-
If one or more labeled neighbors → assign minimum label among them.
-
Record label equivalences (if two different neighbor labels found).
-
-
Second Pass: Resolve equivalences (using Disjoint Set/Union-Find). Relabel pixels with final root label.
-
-
Output: Labeled image. Can compute properties per label: area, centroid, bounding box, mean intensity.
-
Application: Object counting, filtering by size (remove small labels).
VI. COLOR SPACES & TRANSFORMATIONS
Common Color Models (High Frequency)
| Model | Type | Channels | Key Property | Use Case |
|---|---|---|---|---|
| RGB | Additive | R, G, B | Device-dependent (camera, screen) | Storage, display |
| HSV/HSI | Cylindrical | H (Hue), S (Saturation), V/I (Value/Intensity) | Separates color (H,S) from brightness (V/I) | Color selection, segmentation by hue |
| CIE LAB | Perceptual | L* (Lightness), a* (Green-Red), b* (Blue-Yellow) | Perceptually uniform, device-independent | Color difference (ΔE), color invariant processing |
Color Space Transformations (High Frequency)
- RGB → HSV:
$$V = \max(R,G,B)$$
$$S = \begin{cases} \frac{V - \min(R,G,B)}{V} & \text{if } V \neq 0 \\ 0 & \text{otherwise} \end{cases}$$
$$H = \begin{cases} 60^\circ \times \frac{G-B}{V-\min} & \text{if } V=R \\ 120^\circ + 60^\circ \times \frac{B-R}{V-\min} & \text{if } V=G \\ 240^\circ + 60^\circ \times \frac{R-G}{V-\min} & \text{if } V=B \end{cases}$$
(H in [0°,360°], often normalized to [0,1]).
-
RGB → CIE LAB: (Non-linear, involves XYZ intermediate space). Not typically computed by hand in exams, but know the pipeline:
RGB → sRGB → XYZ → LAB. Know thatL*is lightness,a*/b*are color opponents. -
Purpose: Segmentation (e.g., track red object in HSV by hue range), illumination invariance (use
a*,b*only), perceptual analysis.
Precise Color Adjustment (Curves)
-
Concept: Non-linear, point-wise mapping of input intensity to output intensity for each channel. Defined by a curve (e.g., diagonal = no change).
-
Use Cases: Photo editing (tone mapping), medical imaging (enhance specific tissue contrast), cinematic color grading.
-
Advantage over Levels: Allows localized adjustments (e.g., brighten only shadows without affecting highlights).
VII. CONTOURS, SEGMENTATION & OBJECT ANALYSIS
Contour Analysis (High Frequency)
-
Contours: Boundaries of objects in a binary image. Stored as a list of
(x,y)points. -
Retrieval & Approximation:
-
cv2.findContours(): Retrieves contours from binary image (modes:RETR_EXTERNAL,RETR_TREE). -
cv2.approxPolyDP(): Approximates contour to fewer vertices based on precision parameter (Douglas-Peucker algorithm).
-
-
Shape Recognition Features:
-
Aspect Ratio:
Width/Heightof bounding rect. (Circle ≈ 1, line → 0 or ∞). -
Extent:
Area / BoundingBoxArea. (Solidity of shape within its box). -
Solidity:
ContourArea / ConvexHullArea. (Holes make this < 1). -
Hu Moments: 7 rotation/scale/translation invariant moments. Used for shape matching.
-
Image Segmentation (High Frequency)
-
Goal: Partition image into meaningful, homogeneous regions.
-
Relation to Recognition: Segmentation is often a pre-processing step for recognition (isolate objects first).
-
Typical Steps:
-
Pre-processing: Denoise (filter), enhance contrast.
-
Segmentation Algorithm: Apply chosen method.
-
Post-processing: Morphological ops, CCA to clean/merge regions.
-
-
Types:
-
Thresholding-based: Global/Adaptive.
-
Edge-based: Canny + contour closing.
-
Region-based: Region Growing, Split-and-Merge.
-
Clustering-based: K-means on pixel features (color, texture).
-
VIII. ADVANCED TOPICS & APPLICATIONS
Object Detection Techniques (High Frequency)
| Traditional | Deep Learning-based |
|---|---|
| Sliding Window + Classifier: Exhaustive search. Slow. | R-CNN Family: Region proposals (Selective Search) + CNN. Accurate but slow (Faster R-CNN improves speed). |
| HOG + SVM: Extract Histogram of Oriented Gradients features, train linear SVM. | YOLO (You Only Look Once): Single-stage, predicts bounding boxes & classes in one pass. Fast, real-time. |
| Viola-Jones: Haar-like features + AdaBoost + cascade classifier. Fast for faces. | SSD (Single Shot MultiBox Detector): Single-stage, uses multi-scale feature maps. Balance speed/accuracy. |
| Comparison: Traditional: hand-crafted features, less accurate. DL: learned features, state-of-the-art accuracy, data-hungry. |
Motion Estimation & Object Tracking (High Frequency)
-
Motion Estimation:
-
Optical Flow (Lucas-Kanade): Assumes brightness constancy & spatial coherence. Solves for flow vectors
(u,v)in small neighborhoods. Equation:I_x u + I_y v + I_t = 0. -
Block Matching: Divides frame into blocks, searches for best match in next frame (e.g., for video compression).
-
-
Object Tracking Algorithms:
-
Classical:
-
Kalman Filter: Predicts object state (position, velocity) and corrects with measurement. For linear Gaussian systems.
-
Mean Shift/CAMShift: Iteratively moves window to densest region (based on color histogram). CAMShift adapts window size.
-
-
Deep Learning: DeepSORT (combines Kalman filter with CNN re-identification), SiamFC (fully-convolutional Siamese network for matching).
-
Gesture Recognition (High Frequency)
-
Process (OpenCV + DL):
-
Hand Detection: Skin color segmentation (YCbCr, HSV), background subtraction, or CNN (e.g., YOLO for hand bounding box).
-
Keypoint/Contour Extraction: Find hand contour, convex hull, defects (for fingertips), or use hand keypoint detector (MediaPipe Hands).
-
Feature Extraction: Handcrafted (aspect ratio, number of fingers) or CNN features from cropped hand image.
-
Classification: Train ML classifier (SVM, Random Forest) on features or use end-to-end DL model (CNN, LSTM for sequences) to classify gesture class (e.g., "thumbs up", "open palm").
-
-
HCI Role: Enables touchless control (presentations, gaming), sign language translation, VR/AR interaction.
Applications of Deep Learning in CV (High Frequency)
-
Image Classification:
ResNet,EfficientNet(identify main object in image). -
Object Detection:
YOLO,Faster R-CNN(localize & classify multiple objects). -
Semantic Segmentation:
U-Net,DeepLab(classify every pixel by category, e.g., road, car, person). -
Instance Segmentation:
Mask R-CNN(extends detection by predicting pixel-level mask for each object instance). -
Other: Face Recognition (
FaceNet), Pose Estimation (OpenPose), Image Captioning (CNN + RNN/Transformer).
BOXED KEY FORMULAS & RESULTS
- Contrast Stretching (Linear):
$$\boxed{g(x,y) = \frac{(d-c)}{(b-a)} \left( f(x,y) - a \right) + c}$$
- Histogram Equalization Mapping:
$$\boxed{s_k = \text{round}\left( (L-1) \times \sum_{j=0}^{k} p_r(r_j) \right)}$$
- Sobel Gradient Magnitude (approx):
$$\boxed{G \approx |G_x| + |G_y|}$$
- RGB to HSV (Value):
$$\boxed{V = \max(R, G, B)}$$
- RGB to HSV (Saturation):
$$\boxed{S = \frac{V - \min(R,G,B)}{V} \quad (\text{if } V \neq 0)}$$
- Morphological Dilation (Set Notation):
$$\boxed{A \oplus B = \{ z | (\hat{B})_z \cap A \neq \emptyset \}}$$
where $\hat{B}$ is reflection of `B`.