Skip to content
AL-701 · Computer Vision/Quick Revision Short Notes

Computer Vision (AL-701) - Unit 4 Short Notes

UNIT 4: Computer Vision


I. Fundamentals and Core Concepts

Computer Vision (CV) enables machines to interpret and understand visual data from the world, aiming for high-level semantic understanding (e.g., recognizing objects, scenes, activities).

Key Distinction: Image Processing vs. Computer Vision

| Aspect | Image Processing | Computer Vision |

|---------------------|------------------------------------------|-----------------------------------------|

| Input → Output | Image → Modified Image | Image → Description/Decision |

| Abstraction Level | Pixel-level manipulation | Semantic-level interpretation |

| Goal | Enhance, restore, compress | Analyze, recognize, reason |

| Example | Contrast stretching, noise removal | Face detection, autonomous driving |

Real-World Applications: Autonomous vehicles (object detection), medical imaging (tumor segmentation), surveillance (activity recognition), augmented reality.


II. Image Representation and Color Spaces

Grayscale vs. Color Images:

  • Grayscale: Single channel (intensity values 0–255). Storage: $H \times W \times 1$.

  • Color: Multiple channels (typically 3: R, G, B). Storage: $H \times W \times 3$.

Color Spaces & Transformations:

Color Space Properties Use Cases
RGB Additive, device-dependent Displays, cameras
HSV/HSB Hue (color), Saturation, Value (intensity) Color selection, segmentation
LAB Perceptually uniform, device-independent Color correction, comparison

Transformation Equations:

  1. RGB → HSV (for $R,G,B \in [0,1]$):

$$ V = \max(R,G,B), \quad S = \frac{V - \min(R,G,B)}{V}, \quad H = \begin{cases} 60^\circ \times \frac{G-B}{V-\min} & \text{if } V=R \\ 120^\circ + 60^\circ \times \frac{B-R}{V-\min} & \text{if } V=G \\ 240^\circ + 60^\circ \times \frac{R-G}{V-\min} & \text{if } V=B \end{cases} $$

  1. RGB → LAB: Non-linear; involves conversion to XYZ first (CIE standard). LAB is designed so that Euclidean distance approximates perceptual difference.

III. Mathematical Operations on Digital Images

Arithmetic/Logical Operations (pixel-wise):

  • Addition/Subtraction: Brightness adjustment ($$\displaystyle I_{\text{new}} = I + c $$, clip to range).

  • Multiplication/Division: Contrast scaling.

  • Logical (AND, OR, XOR): Masking, region extraction.

Convolution & Correlation:

  • Convolution: Kernel $K$ flipped spatially before sliding over image $I$:

$$ (I * K)(x,y) = \sum_{i}\sum_{j} I(x-i, y-j) K(i,j) $$

  • Correlation: No kernel flip. In practice, many libraries (e.g., OpenCV filter2D) use correlation.

  • Border Handling: Zero-padding, replicate, reflect, wrap.

Datatype Conversion:

  • Common types: uint8 (0–255), float32 (0.0–1.0 or normalized), int16 (signed).

  • Impact: Precision loss if converting from float→uint8 without scaling; storage increase for higher bit-depth.

  • OpenCV Example:

    
    img_float = img_uint8.astype('float32') / 255.0
    
    img_uint8 = cv2.convertScaleAbs(img_float, alpha=255.0)
    
    

IV. Image Enhancement Techniques

Contrast Stretching (Linear):

Maps input range $$\displaystyle [I_{\min}, I_{\max}] $$ to output $[0, 255]$:

$$ I_{\text{out}} = \frac{(I - I_{\min}) \times 255}{(I_{\max} - I_{\min})} $$

Numerical Example (DEC 2025):

$$ > I = \begin{bmatrix} > 100 & 150 & 200 \\ > 50 & 100 & 150 \\ > 0 & 50 & 100 > \end{bmatrix}, \quad I_{\min}=0, I_{\max}=200 > $$

$$ > I_{\text{out}} = I \times \frac{255}{200} = I \times 1.275 \Rightarrow \text{clip to 255} > $$

Result:

$$ > \begin{bmatrix} > 127 & 191 & 255 \\ > 63 & 127 & 191 \\ > 0 & 63 & 127 > \end{bmatrix} > $$

Brightness Adjustment: $$\displaystyle I_{\text{new}} = I + c $$ (add constant). Must clip to valid range.

Numerical Example (DEC 2025): $$\displaystyle c=+50 $$ on above $I$:

$$ > \begin{bmatrix} > 150 & 200 & 250 \\ > 100 & 150 & 200 \\ > 50 & 100 & 150 > \end{bmatrix} \text{ (clip 250→255 if uint8)} > $$

Histogram Equalization (HE):

  • Goal: Spread histogram to full dynamic range.

  • Steps:

    1. Compute histogram $$\displaystyle h(r_k) $$, $$\displaystyle k=0,...,L-1 $$.

    2. Normalized PDF: $$\displaystyle p(r_k) = h(r_k)/(M \times N) $$.

    3. CDF: $$\displaystyle s_k = T(r_k) = (L-1) \sum_{j=0}^{k} p(r_j) $$.

    4. Map each pixel: $$\displaystyle r_k \rightarrow s_k $$.

Numerical Example (DEC 2025):

Intensities: $[52,55,61,66,70,61,64,73]$, $$\displaystyle L=256 $$.

  • Histogram: count each intensity.
  • PDF: divide by 8.
  • CDF: cumulative sum × 255.
  • Round to nearest integer for mapping.

CLAHE:

  • Applies HE to small tiles (e.g., 8×8), then interpolates.

  • Clip limit: Caps histogram bin to prevent over-amplification in noise.

  • Advantage: Local contrast enhancement without noise explosion.

Curves & Levels Adjustment (NOV 2023):

  • Curves: Non-linear tone mapping via piecewise function (e.g., S-curve for contrast).

  • Levels: Set black point, white point, gamma (midtones).


V. Image Filtering and Edge Detection

Image Smoothing (Noise Reduction):

Filter Kernel Effect Best For
Box Blur Uniform weights ($$\displaystyle 1/n^2 $$) Averages neighborhood Simple blurring
Gaussian Gaussian weights ($$\displaystyle e^{-(x^2+y^2)/(2\sigma^2)} $$) Smooth, separable (1D→2D) Gaussian noise, preprocessing
Median Rank-based (median of neighborhood) Non-linear, preserves edges Salt-and-pepper noise

Edge Detection:

  1. First-Order Derivatives (Gradient):

    • Sobel Operator: Approximates $$\displaystyle \frac{\partial I}{\partial x}, \frac{\partial I}{\partial y} $$.

      Kernels:

$$ G_x = \begin{bmatrix} -1 & 0 & 1 \\ -2 & 0 & 2 \\ -1 & 0 & 1 \end{bmatrix}, \quad G_y = \begin{bmatrix} -1 & -2 & -1 \\ 0 & 0 & 0 \\ 1 & 2 & 1 \end{bmatrix} $$

 Magnitude: $$\displaystyle G = \sqrt{G_x^2 + G_y^2} $$, direction: $$\displaystyle \theta = \arctan(G_y/G_x) $$.

Numerical Effect: Apply to 3×3 image patch; highlights vertical/horizontal edges.

  1. Second-Order Derivatives (Laplacian):

    • $$\displaystyle \nabla^2 I = \frac{\partial^2 I}{\partial x^2} + \frac{\partial^2 I}{\partial y^2} $$.

    • Kernel (4-connected): $$\displaystyle \begin{bmatrix} 0 & 1 & 0 \\ 1 & -4 & 1 \\ 0 & 1 & 0 \end{bmatrix} $$.

    • Zero-crossing: Edge where Laplacian changes sign.

Canny Edge Detector (NOV 2023):

  1. Gaussian smoothing (reduce noise).

  2. Gradient magnitude & direction (Sobel).

  3. Non-maximum suppression: Thin edges by keeping local maxima along gradient direction.

  4. Hysteresis thresholding:

    • High threshold: Strong edges.

    • Low threshold: Weak edges connected to strong ones kept.

    • Parameters: $\sigma$ (smoothing), $$\displaystyle T_{high} $$, $$\displaystyle T_{low} $$ (typically $$\displaystyle T_{low} = 0.4 T_{high} $$).


VI. Thresholding and Binarization

Global Thresholding:

  • $$\displaystyle I_{\text{bin}}(x,y) = \begin{cases} 255 & I(x,y) > T \\ 0 & \text{otherwise} \end{cases} $$

  • Optimal Threshold (Otsu’s Method):

    Maximizes between-class variance $$\displaystyle \sigma_b^2 = \omega_0 \omega_1 (\mu_1 - \mu_0)^2 $$, where $\omega$ = class probability, $\mu$ = class mean.

    • Assumes bimodal histogram.

    • Compute for all $T$, pick $T$ maximizing $$\displaystyle \sigma_b^2 $$.

Adaptive Thresholding:

  • Threshold varies locally: $$\displaystyle T(x,y) = \text{mean/Gaussian of neighborhood} - C $$.

  • Mean Adaptive: $$\displaystyle T = \text{mean}(N_{xy}) - C $$.

  • Gaussian Adaptive: $$\displaystyle T = \text{Gaussian weighted mean}(N_{xy}) - C $$.

  • Block size: Must be odd (e.g., 11×11), $C$ constant (typically 5–10).

Numerical Implementation (DEC 2025):

For 4×4 image, compute mean in 3×3 neighborhood (with padding), subtract $C$, threshold.


VII. Morphological Operations

Erosion & Dilation (on binary images):

  • Structuring Element (SE): Small shape (e.g., 3×3 square) with origin (reference point).

  • Erosion: $$\displaystyle A \ominus B = \{ z \mid B_z \subseteq A \} $$ → shrinks foreground.

  • Dilation: $$\displaystyle A \oplus B = \{ z \mid (B_z) \cap A \neq \emptyset \} $$ → expands foreground.

Numerical Example (DEC 2025):

$$ > I = \begin{bmatrix} > 0 & 1 & 0 \\ > 1 & 1 & 0 \\ > 0 & 1 & 1 > \end{bmatrix}, \quad > B = \begin{bmatrix} > 1 & 1 \\ > 1 & 1 > \end{bmatrix} \text{ (origin at top-left)} > $$

Dilation: Place SE on each pixel; if any 1 in SE overlaps 1 in $I$, output 1.

Result:

$$ > \begin{bmatrix} > 1 & 1 & 1 \\ > 1 & 1 & 1 \\ > 0 & 1 & 1 > \end{bmatrix} > $$

Compound Operations:

  • Opening: Erosion → Dilation. Removes small objects, smoothens contours.

  • Closing: Dilation → Erosion. Fills small holes, closes gaps.


VIII. Image Segmentation

Definition: Partitioning image into regions with similar attributes (intensity, texture, color).

Relationship to Recognition: Segmentation isolates objects for subsequent feature extraction/classification.

Methods:

  • Thresholding-based: Simple, intensity-driven.

  • Edge-based: Detect edges → link → closed boundaries.

  • Region-based: Growing (merge similar neighbors), splitting/merging.

Connected Component Analysis (CCA):

  • Two-Pass Algorithm:

    1. First pass: Scan left→right, top→bottom.

      • Assign new label to foreground pixel if neighbors unlabeled.

      • If multiple labeled neighbors, assign one label and record equivalence.

    2. Second pass: Resolve equivalences (union-find), relabel.

  • Properties: Area, centroid, bounding box, label.

Numerical Example: Apply to small binary image (e.g., 4×4 with two touching objects; show labeling and equivalence resolution).


IX. Contour Analysis and Shape Recognition

Contour Detection:

  • OpenCV: contours, hierarchy = cv2.findContours(binary, mode, method).

  • Hierarchy: [Next, Previous, First_Child, Parent] for nested contours.

Contour Properties:

  • Area: cv2.contourArea(C).

  • Perimeter: cv2.arcLength(C, closed).

  • Centroid: $$\displaystyle C_x = \frac{M_{10}}{M_{00}}, C_y = \frac{M_{01}}{M_{00}} $$ (spatial moments).

  • Aspect Ratio: $$\displaystyle \frac{\text{width}}{\text{height}} $$ of bounding rect.

  • Extent: $$\displaystyle \frac{\text{area}}{\text{bounding rect area}} $$.

Contour Approximation:

  • Douglas-Peucker: Simplify contour by removing points with distance $$\displaystyle < \epsilon $$ from chord.

Shape Recognition:

  • Hu Moments: 7 rotation/scale/translation invariant moments.

  • Convex Hull & Defects: cv2.convexHull(), defects indicate concavities (e.g., finger gaps).

  • Template Matching: cv2.matchShapes() (Hu moments comparison).


X. Object Detection and Recognition Techniques

Traditional Methods:

  • Haar Cascades (Viola-Jones):

    • Features: Haar-like rectangles (edge, line, center-surround).

    • Integral Image: Fast sum computation in constant time.

    • AdaBoost: Selects best features, trains strong classifier from weak learners.

    • Cascade: Stages of classifiers to reject non-objects quickly.

  • HOG (Histogram of Oriented Gradients):

    • Compute gradient orientation histograms in dense grids.

    • Normalize across blocks (contrast normalization).

    • Classifier: SVM trained on HOG descriptors.

Deep Learning-Based:

Type Examples Mechanism Speed vs Accuracy
Two-stage Faster R-CNN Region proposals (RPN) → classify/regress High accuracy, slower
One-stage YOLO, SSD Direct prediction on grid (no proposals) Fast, slightly less accurate

XI. Motion Estimation and Object Tracking

Motion Estimation:

  • Optical Flow: Assumes brightness constancy: $$\displaystyle I(x,y,t) = I(x+dx, y+dy, t+dt) $$.

    • Lucas-Kanade (sparse): Solves for flow in small window assuming similar motion.

    • Farneback (dense): Polynomial expansion for dense flow field.

Object Tracking:

  • Tracking-by-Detection: Detect each frame → associate (SORT: Kalman + Hungarian; DeepSORT: adds appearance features).

  • Tracking-by-Matching:

    • Mean Shift: Iteratively shift window to densest region (color histogram similarity).

    • Correlation Filters: Learn discriminative filter (e.g., MOSSE, CSRT).

  • Kalman Filter:

    • Predict: $$\displaystyle \hat{x}_k^- = F \hat{x}_{k-1} $$, $$\displaystyle P_k^- = F P_{k-1} F^T + Q $$.

    • Update: $$\displaystyle K_k = P_k^- H^T (H P_k^- H^T + R)^{-1} $$, $$\displaystyle \hat{x}_k = \hat{x}_k^- + K_k (z_k - H \hat{x}_k^-) $$.


XII. Gesture Recognition

Definition: Interpreting human hand/body movements for HCI (e.g., touchless control, sign language).

Techniques:

  • Traditional:

    • Skin color segmentation (HSV/YCrCb thresholding).

    • Contour extraction → convex hull defects → finger counting.

  • Deep Learning:

    • CNN: Classify static hand gestures from cropped images.

    • Keypoint Detection: MediaPipe Hands (21 3D landmarks).

    • Temporal Models: LSTM/GRU on landmark sequences for dynamic gestures.

OpenCV Implementation:

  1. Hand detection (skin mask or object detector).

  2. Segmentation (morphology, contour).

  3. Feature extraction (landmarks, Hu moments).

  4. Classification (pre-trained CNN or SVM).


XIII. Deep Learning in Computer Vision

Key Architectures:

  • CNNs: Hierarchical features via conv → ReLU → pool. Examples: ResNet (skip connections), EfficientNet (compound scaling).

  • Transformers (ViT): Split image into patches, self-attention for global context.

  • RNNs/LSTMs: Sequential data (video frames, gesture sequences).

Major Applications:

  • Classification: ResNet, EfficientNet.

  • Detection: YOLO (single-stage), Faster R-CNN (two-stage), DETR (Transformer-based).

  • Segmentation: U-Net (encoder-decoder, skip connections), DeepLab (atrous conv).

  • Generation: GANs (generator/discriminator), Diffusion models (forward/backward process).

Transfer Learning: Freeze early layers of pre-trained model (e.g., ImageNet), fine-tune last layers on target dataset.


XIV. OpenCV Integration and Practical Implementation

OpenCV Core Modules:

  • cv2: Main library (image I/O, filtering, color conversion).

  • numpy: Array operations (images as ndarray).

  • cv2.dnn: Deep learning inference (load Caffe/TensorFlow/PyTorch models).

Typical Pipeline:


# 1. Load

img = cv2.imread('image.jpg')

# 2. Preprocess (resize, normalize, blobFromImage)

blob = cv2.dnn.blobFromImage(img, scalefactor=1/255.0, size=(224,224))

# 3. Load model

net = cv2.dnn.readNet('model.pb')

# 4. Inference

net.setInput(blob)

output = net.forward()

# 5. Post-process (NMS, thresholding)

# 6. Visualize (cv2.rectangle, cv2.putText)

Performance Optimization:

  • Use GPU: net.setPreferableBackend(cv2.dnn.DNN_BACKEND_CUDA).

  • Batch processing for multiple images.

  • Lower precision (FP16) if supported.


\boxed{\text{Exam Focus Summary}}

  1. Morphological Operations: Dilation/erosion with SE, opening/closing applications.

  2. Thresholding: Otsu’s (maximize between-class variance), adaptive (local mean/Gaussian).

  3. Color Spaces: RGB↔HSV/LAB equations, use cases.

  4. Histogram Equalization: CDF-based mapping; CLAHE (clip limit, tiles).

  5. Filtering & Edge Detection: Sobel derivation ($$\displaystyle G_x, G_y $$ kernels), Canny steps (NMS, hysteresis).

  6. Connected Component Analysis: Two-pass algorithm, equivalence resolution.

  7. Deep Learning: Compare two-stage (R-CNN) vs one-stage (YOLO); CNN basics.

  8. IP vs CV: Analysis vs manipulation, semantic vs pixel-level.

[!TIP] Common Pitfalls

  • Otsu’s: Assumes bimodal histogram; fails on uniform images.
  • Adaptive Thresholding: Block size must be odd; $C$ too large → over-thresholding.
  • Morphology: SE origin affects result; test with cv2.MORPH_CROSS vs square.
  • Sobel: Kernels already include derivative approximation; no need to flip.
  • CLAHE: Clip limit too high → no effect; too low → noise amplification.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in