UNIT 4: Computer Vision
I. Fundamentals and Core Concepts
Computer Vision (CV) enables machines to interpret and understand visual data from the world, aiming for high-level semantic understanding (e.g., recognizing objects, scenes, activities).
Key Distinction: Image Processing vs. Computer Vision
| Aspect | Image Processing | Computer Vision |
|---------------------|------------------------------------------|-----------------------------------------|
| Input → Output | Image → Modified Image | Image → Description/Decision |
| Abstraction Level | Pixel-level manipulation | Semantic-level interpretation |
| Goal | Enhance, restore, compress | Analyze, recognize, reason |
| Example | Contrast stretching, noise removal | Face detection, autonomous driving |
Real-World Applications: Autonomous vehicles (object detection), medical imaging (tumor segmentation), surveillance (activity recognition), augmented reality.
II. Image Representation and Color Spaces
Grayscale vs. Color Images:
-
Grayscale: Single channel (intensity values 0–255). Storage: $H \times W \times 1$.
-
Color: Multiple channels (typically 3: R, G, B). Storage: $H \times W \times 3$.
Color Spaces & Transformations:
| Color Space | Properties | Use Cases |
|---|---|---|
| RGB | Additive, device-dependent | Displays, cameras |
| HSV/HSB | Hue (color), Saturation, Value (intensity) | Color selection, segmentation |
| LAB | Perceptually uniform, device-independent | Color correction, comparison |
Transformation Equations:
- RGB → HSV (for $R,G,B \in [0,1]$):
$$ V = \max(R,G,B), \quad S = \frac{V - \min(R,G,B)}{V}, \quad H = \begin{cases} 60^\circ \times \frac{G-B}{V-\min} & \text{if } V=R \\ 120^\circ + 60^\circ \times \frac{B-R}{V-\min} & \text{if } V=G \\ 240^\circ + 60^\circ \times \frac{R-G}{V-\min} & \text{if } V=B \end{cases} $$
- RGB → LAB: Non-linear; involves conversion to XYZ first (CIE standard). LAB is designed so that Euclidean distance approximates perceptual difference.
III. Mathematical Operations on Digital Images
Arithmetic/Logical Operations (pixel-wise):
-
Addition/Subtraction: Brightness adjustment ($$\displaystyle I_{\text{new}} = I + c $$, clip to range).
-
Multiplication/Division: Contrast scaling.
-
Logical (AND, OR, XOR): Masking, region extraction.
Convolution & Correlation:
- Convolution: Kernel $K$ flipped spatially before sliding over image $I$:
$$ (I * K)(x,y) = \sum_{i}\sum_{j} I(x-i, y-j) K(i,j) $$
-
Correlation: No kernel flip. In practice, many libraries (e.g., OpenCV
filter2D) use correlation. -
Border Handling: Zero-padding, replicate, reflect, wrap.
Datatype Conversion:
-
Common types:
uint8(0–255),float32(0.0–1.0 or normalized),int16(signed). -
Impact: Precision loss if converting from float→uint8 without scaling; storage increase for higher bit-depth.
-
OpenCV Example:
img_float = img_uint8.astype('float32') / 255.0 img_uint8 = cv2.convertScaleAbs(img_float, alpha=255.0)
IV. Image Enhancement Techniques
Contrast Stretching (Linear):
Maps input range $$\displaystyle [I_{\min}, I_{\max}] $$ to output $[0, 255]$:
$$ I_{\text{out}} = \frac{(I - I_{\min}) \times 255}{(I_{\max} - I_{\min})} $$
Numerical Example (DEC 2025):
$$ > I = \begin{bmatrix} > 100 & 150 & 200 \\ > 50 & 100 & 150 \\ > 0 & 50 & 100 > \end{bmatrix}, \quad I_{\min}=0, I_{\max}=200 > $$
$$ > I_{\text{out}} = I \times \frac{255}{200} = I \times 1.275 \Rightarrow \text{clip to 255} > $$
Result:
$$ > \begin{bmatrix} > 127 & 191 & 255 \\ > 63 & 127 & 191 \\ > 0 & 63 & 127 > \end{bmatrix} > $$
Brightness Adjustment: $$\displaystyle I_{\text{new}} = I + c $$ (add constant). Must clip to valid range.
Numerical Example (DEC 2025): $$\displaystyle c=+50 $$ on above $I$:
$$ > \begin{bmatrix} > 150 & 200 & 250 \\ > 100 & 150 & 200 \\ > 50 & 100 & 150 > \end{bmatrix} \text{ (clip 250→255 if uint8)} > $$
Histogram Equalization (HE):
-
Goal: Spread histogram to full dynamic range.
-
Steps:
-
Compute histogram $$\displaystyle h(r_k) $$, $$\displaystyle k=0,...,L-1 $$.
-
Normalized PDF: $$\displaystyle p(r_k) = h(r_k)/(M \times N) $$.
-
CDF: $$\displaystyle s_k = T(r_k) = (L-1) \sum_{j=0}^{k} p(r_j) $$.
-
Map each pixel: $$\displaystyle r_k \rightarrow s_k $$.
-
Numerical Example (DEC 2025):
Intensities: $[52,55,61,66,70,61,64,73]$, $$\displaystyle L=256 $$.
- Histogram: count each intensity.
- PDF: divide by 8.
- CDF: cumulative sum × 255.
- Round to nearest integer for mapping.
CLAHE:
-
Applies HE to small tiles (e.g., 8×8), then interpolates.
-
Clip limit: Caps histogram bin to prevent over-amplification in noise.
-
Advantage: Local contrast enhancement without noise explosion.
Curves & Levels Adjustment (NOV 2023):
-
Curves: Non-linear tone mapping via piecewise function (e.g., S-curve for contrast).
-
Levels: Set black point, white point, gamma (midtones).
V. Image Filtering and Edge Detection
Image Smoothing (Noise Reduction):
| Filter | Kernel | Effect | Best For |
|---|---|---|---|
| Box Blur | Uniform weights ($$\displaystyle 1/n^2 $$) | Averages neighborhood | Simple blurring |
| Gaussian | Gaussian weights ($$\displaystyle e^{-(x^2+y^2)/(2\sigma^2)} $$) | Smooth, separable (1D→2D) | Gaussian noise, preprocessing |
| Median | Rank-based (median of neighborhood) | Non-linear, preserves edges | Salt-and-pepper noise |
Edge Detection:
-
First-Order Derivatives (Gradient):
-
Sobel Operator: Approximates $$\displaystyle \frac{\partial I}{\partial x}, \frac{\partial I}{\partial y} $$.
Kernels:
-
$$ G_x = \begin{bmatrix} -1 & 0 & 1 \\ -2 & 0 & 2 \\ -1 & 0 & 1 \end{bmatrix}, \quad G_y = \begin{bmatrix} -1 & -2 & -1 \\ 0 & 0 & 0 \\ 1 & 2 & 1 \end{bmatrix} $$
Magnitude: $$\displaystyle G = \sqrt{G_x^2 + G_y^2} $$, direction: $$\displaystyle \theta = \arctan(G_y/G_x) $$.
Numerical Effect: Apply to 3×3 image patch; highlights vertical/horizontal edges.
-
Second-Order Derivatives (Laplacian):
-
$$\displaystyle \nabla^2 I = \frac{\partial^2 I}{\partial x^2} + \frac{\partial^2 I}{\partial y^2} $$.
-
Kernel (4-connected): $$\displaystyle \begin{bmatrix} 0 & 1 & 0 \\ 1 & -4 & 1 \\ 0 & 1 & 0 \end{bmatrix} $$.
-
Zero-crossing: Edge where Laplacian changes sign.
-
Canny Edge Detector (NOV 2023):
-
Gaussian smoothing (reduce noise).
-
Gradient magnitude & direction (Sobel).
-
Non-maximum suppression: Thin edges by keeping local maxima along gradient direction.
-
Hysteresis thresholding:
-
High threshold: Strong edges.
-
Low threshold: Weak edges connected to strong ones kept.
-
Parameters: $\sigma$ (smoothing), $$\displaystyle T_{high} $$, $$\displaystyle T_{low} $$ (typically $$\displaystyle T_{low} = 0.4 T_{high} $$).
-
VI. Thresholding and Binarization
Global Thresholding:
-
$$\displaystyle I_{\text{bin}}(x,y) = \begin{cases} 255 & I(x,y) > T \\ 0 & \text{otherwise} \end{cases} $$
-
Optimal Threshold (Otsu’s Method):
Maximizes between-class variance $$\displaystyle \sigma_b^2 = \omega_0 \omega_1 (\mu_1 - \mu_0)^2 $$, where $\omega$ = class probability, $\mu$ = class mean.
-
Assumes bimodal histogram.
-
Compute for all $T$, pick $T$ maximizing $$\displaystyle \sigma_b^2 $$.
-
Adaptive Thresholding:
-
Threshold varies locally: $$\displaystyle T(x,y) = \text{mean/Gaussian of neighborhood} - C $$.
-
Mean Adaptive: $$\displaystyle T = \text{mean}(N_{xy}) - C $$.
-
Gaussian Adaptive: $$\displaystyle T = \text{Gaussian weighted mean}(N_{xy}) - C $$.
-
Block size: Must be odd (e.g., 11×11), $C$ constant (typically 5–10).
Numerical Implementation (DEC 2025):
For 4×4 image, compute mean in 3×3 neighborhood (with padding), subtract $C$, threshold.
VII. Morphological Operations
Erosion & Dilation (on binary images):
-
Structuring Element (SE): Small shape (e.g., 3×3 square) with origin (reference point).
-
Erosion: $$\displaystyle A \ominus B = \{ z \mid B_z \subseteq A \} $$ → shrinks foreground.
-
Dilation: $$\displaystyle A \oplus B = \{ z \mid (B_z) \cap A \neq \emptyset \} $$ → expands foreground.
Numerical Example (DEC 2025):
$$ > I = \begin{bmatrix} > 0 & 1 & 0 \\ > 1 & 1 & 0 \\ > 0 & 1 & 1 > \end{bmatrix}, \quad > B = \begin{bmatrix} > 1 & 1 \\ > 1 & 1 > \end{bmatrix} \text{ (origin at top-left)} > $$
Dilation: Place SE on each pixel; if any 1 in SE overlaps 1 in $I$, output 1.
Result:
$$ > \begin{bmatrix} > 1 & 1 & 1 \\ > 1 & 1 & 1 \\ > 0 & 1 & 1 > \end{bmatrix} > $$
Compound Operations:
-
Opening: Erosion → Dilation. Removes small objects, smoothens contours.
-
Closing: Dilation → Erosion. Fills small holes, closes gaps.
VIII. Image Segmentation
Definition: Partitioning image into regions with similar attributes (intensity, texture, color).
Relationship to Recognition: Segmentation isolates objects for subsequent feature extraction/classification.
Methods:
-
Thresholding-based: Simple, intensity-driven.
-
Edge-based: Detect edges → link → closed boundaries.
-
Region-based: Growing (merge similar neighbors), splitting/merging.
Connected Component Analysis (CCA):
-
Two-Pass Algorithm:
-
First pass: Scan left→right, top→bottom.
-
Assign new label to foreground pixel if neighbors unlabeled.
-
If multiple labeled neighbors, assign one label and record equivalence.
-
-
Second pass: Resolve equivalences (union-find), relabel.
-
-
Properties: Area, centroid, bounding box, label.
Numerical Example: Apply to small binary image (e.g., 4×4 with two touching objects; show labeling and equivalence resolution).
IX. Contour Analysis and Shape Recognition
Contour Detection:
-
OpenCV:
contours, hierarchy = cv2.findContours(binary, mode, method). -
Hierarchy: [Next, Previous, First_Child, Parent] for nested contours.
Contour Properties:
-
Area:
cv2.contourArea(C). -
Perimeter:
cv2.arcLength(C, closed). -
Centroid: $$\displaystyle C_x = \frac{M_{10}}{M_{00}}, C_y = \frac{M_{01}}{M_{00}} $$ (spatial moments).
-
Aspect Ratio: $$\displaystyle \frac{\text{width}}{\text{height}} $$ of bounding rect.
-
Extent: $$\displaystyle \frac{\text{area}}{\text{bounding rect area}} $$.
Contour Approximation:
- Douglas-Peucker: Simplify contour by removing points with distance $$\displaystyle < \epsilon $$ from chord.
Shape Recognition:
-
Hu Moments: 7 rotation/scale/translation invariant moments.
-
Convex Hull & Defects:
cv2.convexHull(), defects indicate concavities (e.g., finger gaps). -
Template Matching:
cv2.matchShapes()(Hu moments comparison).
X. Object Detection and Recognition Techniques
Traditional Methods:
-
Haar Cascades (Viola-Jones):
-
Features: Haar-like rectangles (edge, line, center-surround).
-
Integral Image: Fast sum computation in constant time.
-
AdaBoost: Selects best features, trains strong classifier from weak learners.
-
Cascade: Stages of classifiers to reject non-objects quickly.
-
-
HOG (Histogram of Oriented Gradients):
-
Compute gradient orientation histograms in dense grids.
-
Normalize across blocks (contrast normalization).
-
Classifier: SVM trained on HOG descriptors.
-
Deep Learning-Based:
| Type | Examples | Mechanism | Speed vs Accuracy |
|---|---|---|---|
| Two-stage | Faster R-CNN | Region proposals (RPN) → classify/regress | High accuracy, slower |
| One-stage | YOLO, SSD | Direct prediction on grid (no proposals) | Fast, slightly less accurate |
XI. Motion Estimation and Object Tracking
Motion Estimation:
-
Optical Flow: Assumes brightness constancy: $$\displaystyle I(x,y,t) = I(x+dx, y+dy, t+dt) $$.
-
Lucas-Kanade (sparse): Solves for flow in small window assuming similar motion.
-
Farneback (dense): Polynomial expansion for dense flow field.
-
Object Tracking:
-
Tracking-by-Detection: Detect each frame → associate (SORT: Kalman + Hungarian; DeepSORT: adds appearance features).
-
Tracking-by-Matching:
-
Mean Shift: Iteratively shift window to densest region (color histogram similarity).
-
Correlation Filters: Learn discriminative filter (e.g., MOSSE, CSRT).
-
-
Kalman Filter:
-
Predict: $$\displaystyle \hat{x}_k^- = F \hat{x}_{k-1} $$, $$\displaystyle P_k^- = F P_{k-1} F^T + Q $$.
-
Update: $$\displaystyle K_k = P_k^- H^T (H P_k^- H^T + R)^{-1} $$, $$\displaystyle \hat{x}_k = \hat{x}_k^- + K_k (z_k - H \hat{x}_k^-) $$.
-
XII. Gesture Recognition
Definition: Interpreting human hand/body movements for HCI (e.g., touchless control, sign language).
Techniques:
-
Traditional:
-
Skin color segmentation (HSV/YCrCb thresholding).
-
Contour extraction → convex hull defects → finger counting.
-
-
Deep Learning:
-
CNN: Classify static hand gestures from cropped images.
-
Keypoint Detection: MediaPipe Hands (21 3D landmarks).
-
Temporal Models: LSTM/GRU on landmark sequences for dynamic gestures.
-
OpenCV Implementation:
-
Hand detection (skin mask or object detector).
-
Segmentation (morphology, contour).
-
Feature extraction (landmarks, Hu moments).
-
Classification (pre-trained CNN or SVM).
XIII. Deep Learning in Computer Vision
Key Architectures:
-
CNNs: Hierarchical features via conv → ReLU → pool. Examples: ResNet (skip connections), EfficientNet (compound scaling).
-
Transformers (ViT): Split image into patches, self-attention for global context.
-
RNNs/LSTMs: Sequential data (video frames, gesture sequences).
Major Applications:
-
Classification: ResNet, EfficientNet.
-
Detection: YOLO (single-stage), Faster R-CNN (two-stage), DETR (Transformer-based).
-
Segmentation: U-Net (encoder-decoder, skip connections), DeepLab (atrous conv).
-
Generation: GANs (generator/discriminator), Diffusion models (forward/backward process).
Transfer Learning: Freeze early layers of pre-trained model (e.g., ImageNet), fine-tune last layers on target dataset.
XIV. OpenCV Integration and Practical Implementation
OpenCV Core Modules:
-
cv2: Main library (image I/O, filtering, color conversion). -
numpy: Array operations (images asndarray). -
cv2.dnn: Deep learning inference (load Caffe/TensorFlow/PyTorch models).
Typical Pipeline:
# 1. Load
img = cv2.imread('image.jpg')
# 2. Preprocess (resize, normalize, blobFromImage)
blob = cv2.dnn.blobFromImage(img, scalefactor=1/255.0, size=(224,224))
# 3. Load model
net = cv2.dnn.readNet('model.pb')
# 4. Inference
net.setInput(blob)
output = net.forward()
# 5. Post-process (NMS, thresholding)
# 6. Visualize (cv2.rectangle, cv2.putText)
Performance Optimization:
-
Use GPU:
net.setPreferableBackend(cv2.dnn.DNN_BACKEND_CUDA). -
Batch processing for multiple images.
-
Lower precision (FP16) if supported.
\boxed{\text{Exam Focus Summary}}
-
Morphological Operations: Dilation/erosion with SE, opening/closing applications.
-
Thresholding: Otsu’s (maximize between-class variance), adaptive (local mean/Gaussian).
-
Color Spaces: RGB↔HSV/LAB equations, use cases.
-
Histogram Equalization: CDF-based mapping; CLAHE (clip limit, tiles).
-
Filtering & Edge Detection: Sobel derivation ($$\displaystyle G_x, G_y $$ kernels), Canny steps (NMS, hysteresis).
-
Connected Component Analysis: Two-pass algorithm, equivalence resolution.
-
Deep Learning: Compare two-stage (R-CNN) vs one-stage (YOLO); CNN basics.
-
IP vs CV: Analysis vs manipulation, semantic vs pixel-level.
[!TIP] Common Pitfalls
- Otsu’s: Assumes bimodal histogram; fails on uniform images.
- Adaptive Thresholding: Block size must be odd; $C$ too large → over-thresholding.
- Morphology: SE origin affects result; test with
cv2.MORPH_CROSSvs square.
- Sobel: Kernels already include derivative approximation; no need to flip.
- CLAHE: Clip limit too high → no effect; too low → noise amplification.