Skip to content
IT-703 (D) · Digital Image Processing/Quick Revision Short Notes

Digital Image Processing (IT-703 (D)) - Unit 2 Short Notes

UNIT 2: Digital Image Processing


1. Image Fundamentals

Image Formation in the Human Eye

  • Structure: Light enters cornea → pupil (controlled by iris) → lens (focuses) → retina (photoreceptors: rods & cones).

  • Brightness Adaptation: Eye adjusts sensitivity to wide luminance ranges (~10^10:1) via pupil size & photoreceptor response.

    [!TIP] Often confused with discrimination; adaptation is dynamic range adjustment, discrimination is just noticeable difference.

  • Brightness Discrimination: Ability to detect intensity differences. Follows Weber’s Law: ΔI/I = constant (≈0.02 for vision).

Image Sampling and Quantization

  • Sampling: Digitizing continuous spatial coordinates (x,y).

    Sampling rate determines spatial resolution. Undersampling → aliasing.

  • Quantization: Digitizing amplitude (intensity).

    Number of bits (b) → levels L = 2^b.

    Mean Square Error (MSE) for uniform quantization:

$$MSE = \frac{\Delta^2}{12}, \quad \Delta = \frac{\text{max} - \text{min}}{L}$$

  • Key trade-off: Higher sampling/quantization → better quality but larger file size.

Fourier Transform Properties

2-D Continuous Fourier Transform (CFT)

Definition:

$$F(u,v) = \int_{-\infty}^{\infty} \int_{-\infty}^{\infty} f(x,y) e^{-j2\pi(ux + vy)} \, dx \, dy$$

Inverse:

$$f(x,y) = \int_{-\infty}^{\infty} \int_{-\infty}^{\infty} F(u,v) e^{j2\pi(ux + vy)} \, du \, dv$$

2-D Discrete Fourier Transform (DFT)

Definition for M×N image:

$$F(u,v) = \sum_{x=0}^{M-1} \sum_{y=0}^{N-1} f(x,y) e^{-j2\pi(ux/M + vy/N)}$$

Inverse:

$$f(x,y) = \frac{1}{MN} \sum_{u=0}^{M-1} \sum_{v=0}^{N-1} F(u,v) e^{j2\pi(ux/M + vy/N)}$$

Linearity Property (Proof Required in Exams)

For any images f₁, f₂ and scalars a, b:

$$\mathcal{F}\{a f_1(x,y) + b f_2(x,y)\} = a F_1(u,v) + b F_2(u,v)$$

Proof sketch: Substitute linear combination into transform definition; use integral/sum linearity.

\boxed{\mathcal{F}{a f_1 + b f_2} = a F_1 + b F_2}


2. Image Enhancement

A. Spatial Domain Techniques

Histogram Processing

  • Grayscale:

    • Histogram equalization: Redistribute intensities to uniform histogram.

      Cumulative distribution:

$$s_k = T(r_k) = (L-1) \sum_{j=0}^{k} p_r(r_j)$$

where \(p_r\) = normalized histogram.
  • Histogram specification: Match to given histogram via equalization + inverse mapping.

  • Color Images:

    • Method 1: Process each RGB channel independently → may cause color artifacts.

    • Method 2: Convert to HSV/HSL → equalize V (value) or S (saturation) only → preserve hue.

    • Method 3: Use luminance (Y in YCbCr) → equalize Y, keep Cb, Cr unchanged.

Image Point Operations (Intensity Transformations)

  • Definition: \(s = T(r)\), where r = input intensity, s = output.

  • Common forms:

    • Negative: \(s = L - 1 - r\)

    • Log: \(s = c \log(1 + r)\) → compress high intensities.

    • Power-law (Gamma): \(s = c r^\gamma\)

      \(\gamma < 1\) → expand dark values; \(\gamma > 1\) → expand bright values.

  • Contrast stretching: Piecewise linear to increase dynamic range in ROI.

B. Frequency Domain Techniques

Highpass Filtering (Image Sharpening)

  • Ideal Highpass Filter (IHPF):

$$H_{ihp}(u,v) = \begin{cases} 0 & \text{if } D(u,v) \le D_0 \\ 1 & \text{otherwise} \end{cases}$$

→ Ringing artifacts (Gibbs phenomenon).

  • Butterworth Highpass Filter (BHPF) of order n:

$$H_{bhp}(u,v) = \frac{1}{1 + \left( \frac{D_0}{D(u,v)} \right)^{2n}}$$

→ Smooth transition, no ringing.

  • Gaussian Highpass Filter (GHPF):

$$H_{ghp}(u,v) = 1 - e^{-\frac{D^2(u,v)}{2D_0^2}}$$

→ No ringing, smoothest.

Homomorphic Filtering

  • Concept: Separate illumination (i(x,y)) and reflectance (r(x,y)) components:

$$f(x,y) = i(x,y) \cdot r(x,y)$$

  • Steps:

    1. Take natural log: \(\ln f = \ln i + \ln r\)

    2. Fourier transform: \(F(u,v) = I(u,v) + R(u,v)\)

    3. Apply filter \(H(u,v)\): \(G(u,v) = H(u,v) F(u,v)\)

    4. Inverse Fourier: \(g(x,y) = e^{\mathcal{F}^{-1}\{G(u,v)\}} = \hat{i}(x,y) \cdot \hat{r}(x,y)\)

  • Filter choice:

    \(H(u,v) = (H_{lp}(u,v) - H_{hp}(u,v)) + 1\) or

    \(H(u,v) = (\gamma_L - \gamma_H) [1 - e^{-c(D^2/D_0^2)}] + \gamma_H\)

    where \(\gamma_L < 1\) (attenuate low freq/illumination), \(\gamma_H > 1\) (boost high freq/reflectance).

  • Goal: Compress illumination range, enhance reflectance details.


3. Image Restoration

Noise Parameter Estimation Methods

  1. Local Statistics:

    Estimate noise variance \(\sigma_\eta^2\) from flat regions (low variance) using:

$$\sigma_\eta^2 = \frac{1}{MN} \sum_{x,y} [f(x,y) - \bar{f}]^2 \quad \text{(if region is noise-only)}$$

  1. Training Areas: Manually select noise-only regions → compute sample variance.

  2. Using Image Differences: For additive noise, \(f(x,y) - f(x+1,y)\) or \(f(x,y) - f(x,y+1)\) approximates noise → compute variance.

  3. Robust Estimators: Median absolute deviation (MAD) from median:

    \(\sigma_\eta \approx 1.4826 \cdot \text{MAD}\).

Filtering Techniques

Minimum Mean Square Error (MMSE) Filtering

  • Objective: Minimize \(E\{(\hat{f} - f)^2\}\).

  • Assumes: Image is random field with known covariance.

  • Solution: Wiener filter is a special case when signal and noise are jointly wide-sense stationary.

Wiener Filtering

  • Model: \(f = h * s + \eta\) (convolution + additive noise).

  • Wiener filter in frequency domain:

$$H_w(u,v) = \frac{H^*(u,v)}{|H(u,v)|^2 + K}$$

where \(K = \frac{\sigma_\eta^2}{\sigma_s^2}\) (noise-to-signal power ratio), \(H^*\) = complex conjugate of degradation PSF.

  • Interpretation: Balances inverse filtering (1/H) and noise smoothing.

  • Requires: Estimation of power spectra \(\sigma_s^2\), \(\sigma_\eta^2\) and PSF H.


4. Morphological Image Processing

A. Fundamental Operations

Let A = image (set of foreground pixels), B = structuring element.

  • Erosion: \(A \ominus B = \{ z \mid (B)_z \subseteq A \}\)

    → Shrinks foreground, removes small objects, widens gaps.

  • Dilation: \(A \oplus B = \{ z \mid (B)_z \cap A \neq \emptyset \}\)

    → Expands foreground, fills small holes, connects broken parts.

[!TIP] Erosion = Minkowski subtraction; Dilation = Minkowski addition. Use origin of B as reference.

B. Morphological Algorithms

Boundary Extraction

  • Formula: \(\text{Boundary}(A) = A - (A \ominus B)\)

    where B is small (e.g., 3×3 square or cross).

  • Process: Erode A by one pixel → subtract from original → get outer boundary.

Hole Filling

  • Concept: Hole = background region completely enclosed by foreground.

  • Algorithm:

    1. Find a background point inside the hole (seed).

    2. \(X_0 = \text{seed}\)

    3. \(X_{k+1} = (X_k \oplus B) \cap A^c\)

      (dilate seed, intersect with complement of A)

    4. Iterate until \(X_{k+1} = X_k\) → filled hole = \(X_k\).

  • Result: \(A \cup X_k\) = image with hole filled.


5. Image Segmentation

Region-Based Segmentation

  • Approach: Group pixels into homogeneous regions based on intensity, color, texture.

  • Methods:

    • Thresholding: Global (single T), local (adaptive T), multi-level.

    • Region growing: Start with seed, add neighboring pixels with similar properties (intensity, color).

    • Region splitting & merging: Quad-tree decomposition → merge similar adjacent regions.

    • Watershed: Treat gradient magnitude as topographic surface → catch basins = segments.

Texture Analysis

  • Definition: Visual pattern of spatial variation in intensity/color.

  • Approaches:

    • Statistical: Gray-level co-occurrence matrix (GLCM) → compute contrast, correlation, homogeneity, energy.

    • Structural: Identify primitives (e.g., tiles) and placement rules.

    • Model-based: Use fractals, Markov random fields.

  • Key: Texture descriptor should be invariant to rotation, scale if needed.

Object Recognition

  • Steps:

    1. Segmentation → isolate object.

    2. Feature extraction: Shape (Fourier descriptors, moments), color histograms, texture descriptors, SIFT/ORB keypoints.

    3. Classification/ matching:

      • Template matching (correlation).

      • Machine learning (SVM, CNN).

      • Distance metrics (Euclidean, Mahalanobis).

  • Challenges: Occlusion, viewpoint changes, illumination.

Motion-Based Segmentation

  • Principle: Segment based on temporal changes in image sequence.

  • Methods:

    • Frame differencing: \(|I_t - I_{t-1}| > T\) → moving pixels.

    • Optical flow: Estimate velocity field → cluster by motion vectors.

    • Background subtraction: Model static background → foreground = moving objects.

  • Differentiation from Region-Based:

Feature Region-Based Segmentation Motion-Based Segmentation
Input Single image Image sequence (≥2 frames)
Basis Spatial homogeneity (intensity, color, texture) Temporal change (motion)
Output Static regions/objects Moving objects (foreground)
Sensitivity To spatial features only To motion; static objects ignored
Applications Object detection in still images Video surveillance, activity analysis

[!TIP] Motion segmentation cannot detect stationary objects; region-based cannot distinguish moving from static without temporal info.


6. Image Compression

Need for Compression

  • Reduce storage requirements (memory, disk).

  • Reduce transmission bandwidth (faster transfer, lower cost).

  • Enable real-time processing (video streaming, conferencing).

  • Trade-off: Compression ratio vs. distortion (PSNR, SSIM).

  • Types:

    Lossless (reconstruction identical: PNG, GIF) → lower ratio.

    Lossy (approximate: JPEG, MPEG) → higher ratio, perceptual quality.

Vector Quantization (VQ)

  • Concept: Group pixels (or blocks) into vectors; represent each by a codeword from a codebook.

  • Steps:

    1. Training: Collect representative vectors from image set.

    2. Codebook design (e.g., LBG algorithm):

      • Initialize codebook (random/K-means).

      • Assign vectors to nearest codeword (using Euclidean distance).

      • Update codeword = centroid of assigned vectors.

      • Repeat until convergence.

    3. Encoding: For each input vector, find nearest codeword → output its index.

    4. Decoding: Use index to retrieve codeword from codebook.

  • Advantages: Simple decoding, fixed bit rate.

    Disadvantages: Codebook design complex; blocking artifacts at low bit rates.

Predictive Coding (Lossy Model)

  • Principle: Predict current pixel from past (causal) neighbors → encode prediction error (residual).

  • Encoder Block Diagram:

    
    Input f(x,y) → Predictor → ŷ(x,y) → [-][+] → Quantizer → Encoder → Output
    
                       ↑                          ↑
    
                       └────── Feedback ──────────┘
    
    

    Working:

    1. Predictor uses causal neighbors (e.g., left, top, diagonal) to compute \(\hat{f}(x,y)\).

    2. Compute error: \(e(x,y) = f(x,y) - \hat{f}(x,y)\).

    3. Quantize \(e(x,y)\) → \(\hat{e}(x,y)\) (lossy step).

    4. Encode \(\hat{e}(x,y)\) (entropy coding like Huffman).

    5. Feedback: Decoded error \(\hat{e}\) added to prediction to reconstruct \(\hat{f}\) for next pixels (prevent error propagation).

  • Decoder Block Diagram:

    
    Input (encoded bits) → Decoder → Quantized error ê(x,y) → [+]
    
                                                        ↓
    
                                                    Predictor → Output f̂(x,y)
    
    

    Working:

    1. Decode bits → \(\hat{e}(x,y)\).

    2. Same predictor as encoder → \(\hat{f}(x,y)\).

    3. Add: \(f̂(x,y) = \hat{f}(x,y) + \hat{e}(x,y)\).

  • Key: Predictor must be identical in encoder/decoder; quantization introduces loss.

[!TIP] Common exam question: Draw and explain block diagrams. Remember: encoder has quantizer; decoder has inverse quantizer (but often just "quantizer" in diagram). Feedback loop is critical.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in