UNIT 2: Digital Image Processing
1. Image Fundamentals
Image Formation in the Human Eye
-
Structure: Light enters cornea → pupil (controlled by iris) → lens (focuses) → retina (photoreceptors: rods & cones).
-
Brightness Adaptation: Eye adjusts sensitivity to wide luminance ranges (~10^10:1) via pupil size & photoreceptor response.
[!TIP] Often confused with discrimination; adaptation is dynamic range adjustment, discrimination is just noticeable difference.
-
Brightness Discrimination: Ability to detect intensity differences. Follows Weber’s Law: ΔI/I = constant (≈0.02 for vision).
Image Sampling and Quantization
-
Sampling: Digitizing continuous spatial coordinates (x,y).
Sampling rate determines spatial resolution. Undersampling → aliasing.
-
Quantization: Digitizing amplitude (intensity).
Number of bits (b) → levels L = 2^b.
Mean Square Error (MSE) for uniform quantization:
$$MSE = \frac{\Delta^2}{12}, \quad \Delta = \frac{\text{max} - \text{min}}{L}$$
- Key trade-off: Higher sampling/quantization → better quality but larger file size.
Fourier Transform Properties
2-D Continuous Fourier Transform (CFT)
Definition:
$$F(u,v) = \int_{-\infty}^{\infty} \int_{-\infty}^{\infty} f(x,y) e^{-j2\pi(ux + vy)} \, dx \, dy$$
Inverse:
$$f(x,y) = \int_{-\infty}^{\infty} \int_{-\infty}^{\infty} F(u,v) e^{j2\pi(ux + vy)} \, du \, dv$$
2-D Discrete Fourier Transform (DFT)
Definition for M×N image:
$$F(u,v) = \sum_{x=0}^{M-1} \sum_{y=0}^{N-1} f(x,y) e^{-j2\pi(ux/M + vy/N)}$$
Inverse:
$$f(x,y) = \frac{1}{MN} \sum_{u=0}^{M-1} \sum_{v=0}^{N-1} F(u,v) e^{j2\pi(ux/M + vy/N)}$$
Linearity Property (Proof Required in Exams)
For any images f₁, f₂ and scalars a, b:
$$\mathcal{F}\{a f_1(x,y) + b f_2(x,y)\} = a F_1(u,v) + b F_2(u,v)$$
Proof sketch: Substitute linear combination into transform definition; use integral/sum linearity.
\boxed{\mathcal{F}{a f_1 + b f_2} = a F_1 + b F_2}
2. Image Enhancement
A. Spatial Domain Techniques
Histogram Processing
-
Grayscale:
-
Histogram equalization: Redistribute intensities to uniform histogram.
Cumulative distribution:
-
$$s_k = T(r_k) = (L-1) \sum_{j=0}^{k} p_r(r_j)$$
where \(p_r\) = normalized histogram.
-
Histogram specification: Match to given histogram via equalization + inverse mapping.
-
Color Images:
-
Method 1: Process each RGB channel independently → may cause color artifacts.
-
Method 2: Convert to HSV/HSL → equalize V (value) or S (saturation) only → preserve hue.
-
Method 3: Use luminance (Y in YCbCr) → equalize Y, keep Cb, Cr unchanged.
-
Image Point Operations (Intensity Transformations)
-
Definition: \(s = T(r)\), where r = input intensity, s = output.
-
Common forms:
-
Negative: \(s = L - 1 - r\)
-
Log: \(s = c \log(1 + r)\) → compress high intensities.
-
Power-law (Gamma): \(s = c r^\gamma\)
\(\gamma < 1\) → expand dark values; \(\gamma > 1\) → expand bright values.
-
-
Contrast stretching: Piecewise linear to increase dynamic range in ROI.
B. Frequency Domain Techniques
Highpass Filtering (Image Sharpening)
- Ideal Highpass Filter (IHPF):
$$H_{ihp}(u,v) = \begin{cases} 0 & \text{if } D(u,v) \le D_0 \\ 1 & \text{otherwise} \end{cases}$$
→ Ringing artifacts (Gibbs phenomenon).
- Butterworth Highpass Filter (BHPF) of order n:
$$H_{bhp}(u,v) = \frac{1}{1 + \left( \frac{D_0}{D(u,v)} \right)^{2n}}$$
→ Smooth transition, no ringing.
- Gaussian Highpass Filter (GHPF):
$$H_{ghp}(u,v) = 1 - e^{-\frac{D^2(u,v)}{2D_0^2}}$$
→ No ringing, smoothest.
Homomorphic Filtering
- Concept: Separate illumination (i(x,y)) and reflectance (r(x,y)) components:
$$f(x,y) = i(x,y) \cdot r(x,y)$$
-
Steps:
-
Take natural log: \(\ln f = \ln i + \ln r\)
-
Fourier transform: \(F(u,v) = I(u,v) + R(u,v)\)
-
Apply filter \(H(u,v)\): \(G(u,v) = H(u,v) F(u,v)\)
-
Inverse Fourier: \(g(x,y) = e^{\mathcal{F}^{-1}\{G(u,v)\}} = \hat{i}(x,y) \cdot \hat{r}(x,y)\)
-
-
Filter choice:
\(H(u,v) = (H_{lp}(u,v) - H_{hp}(u,v)) + 1\) or
\(H(u,v) = (\gamma_L - \gamma_H) [1 - e^{-c(D^2/D_0^2)}] + \gamma_H\)
where \(\gamma_L < 1\) (attenuate low freq/illumination), \(\gamma_H > 1\) (boost high freq/reflectance).
-
Goal: Compress illumination range, enhance reflectance details.
3. Image Restoration
Noise Parameter Estimation Methods
-
Local Statistics:
Estimate noise variance \(\sigma_\eta^2\) from flat regions (low variance) using:
$$\sigma_\eta^2 = \frac{1}{MN} \sum_{x,y} [f(x,y) - \bar{f}]^2 \quad \text{(if region is noise-only)}$$
-
Training Areas: Manually select noise-only regions → compute sample variance.
-
Using Image Differences: For additive noise, \(f(x,y) - f(x+1,y)\) or \(f(x,y) - f(x,y+1)\) approximates noise → compute variance.
-
Robust Estimators: Median absolute deviation (MAD) from median:
\(\sigma_\eta \approx 1.4826 \cdot \text{MAD}\).
Filtering Techniques
Minimum Mean Square Error (MMSE) Filtering
-
Objective: Minimize \(E\{(\hat{f} - f)^2\}\).
-
Assumes: Image is random field with known covariance.
-
Solution: Wiener filter is a special case when signal and noise are jointly wide-sense stationary.
Wiener Filtering
-
Model: \(f = h * s + \eta\) (convolution + additive noise).
-
Wiener filter in frequency domain:
$$H_w(u,v) = \frac{H^*(u,v)}{|H(u,v)|^2 + K}$$
where \(K = \frac{\sigma_\eta^2}{\sigma_s^2}\) (noise-to-signal power ratio), \(H^*\) = complex conjugate of degradation PSF.
-
Interpretation: Balances inverse filtering (1/H) and noise smoothing.
-
Requires: Estimation of power spectra \(\sigma_s^2\), \(\sigma_\eta^2\) and PSF H.
4. Morphological Image Processing
A. Fundamental Operations
Let A = image (set of foreground pixels), B = structuring element.
-
Erosion: \(A \ominus B = \{ z \mid (B)_z \subseteq A \}\)
→ Shrinks foreground, removes small objects, widens gaps.
-
Dilation: \(A \oplus B = \{ z \mid (B)_z \cap A \neq \emptyset \}\)
→ Expands foreground, fills small holes, connects broken parts.
[!TIP] Erosion = Minkowski subtraction; Dilation = Minkowski addition. Use origin of B as reference.
B. Morphological Algorithms
Boundary Extraction
-
Formula: \(\text{Boundary}(A) = A - (A \ominus B)\)
where B is small (e.g., 3×3 square or cross).
-
Process: Erode A by one pixel → subtract from original → get outer boundary.
Hole Filling
-
Concept: Hole = background region completely enclosed by foreground.
-
Algorithm:
-
Find a background point inside the hole (seed).
-
\(X_0 = \text{seed}\)
-
\(X_{k+1} = (X_k \oplus B) \cap A^c\)
(dilate seed, intersect with complement of A)
-
Iterate until \(X_{k+1} = X_k\) → filled hole = \(X_k\).
-
-
Result: \(A \cup X_k\) = image with hole filled.
5. Image Segmentation
Region-Based Segmentation
-
Approach: Group pixels into homogeneous regions based on intensity, color, texture.
-
Methods:
-
Thresholding: Global (single T), local (adaptive T), multi-level.
-
Region growing: Start with seed, add neighboring pixels with similar properties (intensity, color).
-
Region splitting & merging: Quad-tree decomposition → merge similar adjacent regions.
-
Watershed: Treat gradient magnitude as topographic surface → catch basins = segments.
-
Texture Analysis
-
Definition: Visual pattern of spatial variation in intensity/color.
-
Approaches:
-
Statistical: Gray-level co-occurrence matrix (GLCM) → compute contrast, correlation, homogeneity, energy.
-
Structural: Identify primitives (e.g., tiles) and placement rules.
-
Model-based: Use fractals, Markov random fields.
-
-
Key: Texture descriptor should be invariant to rotation, scale if needed.
Object Recognition
-
Steps:
-
Segmentation → isolate object.
-
Feature extraction: Shape (Fourier descriptors, moments), color histograms, texture descriptors, SIFT/ORB keypoints.
-
Classification/ matching:
-
Template matching (correlation).
-
Machine learning (SVM, CNN).
-
Distance metrics (Euclidean, Mahalanobis).
-
-
-
Challenges: Occlusion, viewpoint changes, illumination.
Motion-Based Segmentation
-
Principle: Segment based on temporal changes in image sequence.
-
Methods:
-
Frame differencing: \(|I_t - I_{t-1}| > T\) → moving pixels.
-
Optical flow: Estimate velocity field → cluster by motion vectors.
-
Background subtraction: Model static background → foreground = moving objects.
-
-
Differentiation from Region-Based:
| Feature | Region-Based Segmentation | Motion-Based Segmentation |
|---|---|---|
| Input | Single image | Image sequence (≥2 frames) |
| Basis | Spatial homogeneity (intensity, color, texture) | Temporal change (motion) |
| Output | Static regions/objects | Moving objects (foreground) |
| Sensitivity | To spatial features only | To motion; static objects ignored |
| Applications | Object detection in still images | Video surveillance, activity analysis |
[!TIP] Motion segmentation cannot detect stationary objects; region-based cannot distinguish moving from static without temporal info.
6. Image Compression
Need for Compression
-
Reduce storage requirements (memory, disk).
-
Reduce transmission bandwidth (faster transfer, lower cost).
-
Enable real-time processing (video streaming, conferencing).
-
Trade-off: Compression ratio vs. distortion (PSNR, SSIM).
-
Types:
Lossless (reconstruction identical: PNG, GIF) → lower ratio.
Lossy (approximate: JPEG, MPEG) → higher ratio, perceptual quality.
Vector Quantization (VQ)
-
Concept: Group pixels (or blocks) into vectors; represent each by a codeword from a codebook.
-
Steps:
-
Training: Collect representative vectors from image set.
-
Codebook design (e.g., LBG algorithm):
-
Initialize codebook (random/K-means).
-
Assign vectors to nearest codeword (using Euclidean distance).
-
Update codeword = centroid of assigned vectors.
-
Repeat until convergence.
-
-
Encoding: For each input vector, find nearest codeword → output its index.
-
Decoding: Use index to retrieve codeword from codebook.
-
-
Advantages: Simple decoding, fixed bit rate.
Disadvantages: Codebook design complex; blocking artifacts at low bit rates.
Predictive Coding (Lossy Model)
-
Principle: Predict current pixel from past (causal) neighbors → encode prediction error (residual).
-
Encoder Block Diagram:
Input f(x,y) → Predictor → ŷ(x,y) → [-][+] → Quantizer → Encoder → Output ↑ ↑ └────── Feedback ──────────┘Working:
-
Predictor uses causal neighbors (e.g., left, top, diagonal) to compute \(\hat{f}(x,y)\).
-
Compute error: \(e(x,y) = f(x,y) - \hat{f}(x,y)\).
-
Quantize \(e(x,y)\) → \(\hat{e}(x,y)\) (lossy step).
-
Encode \(\hat{e}(x,y)\) (entropy coding like Huffman).
-
Feedback: Decoded error \(\hat{e}\) added to prediction to reconstruct \(\hat{f}\) for next pixels (prevent error propagation).
-
-
Decoder Block Diagram:
Input (encoded bits) → Decoder → Quantized error ê(x,y) → [+] ↓ Predictor → Output f̂(x,y)Working:
-
Decode bits → \(\hat{e}(x,y)\).
-
Same predictor as encoder → \(\hat{f}(x,y)\).
-
Add: \(f̂(x,y) = \hat{f}(x,y) + \hat{e}(x,y)\).
-
-
Key: Predictor must be identical in encoder/decoder; quantization introduces loss.
[!TIP] Common exam question: Draw and explain block diagrams. Remember: encoder has quantizer; decoder has inverse quantizer (but often just "quantizer" in diagram). Feedback loop is critical.