UNIT 5: Advanced Image Processing Techniques
1. Image Acquisition and Perceptual Fundamentals
Image Formation in Human Eye
-
Cornea & Lens: Focus light onto retina.
-
Retina: Contains photoreceptors (rods for low-light, cones for color).
-
Brightness Adaptation: Eye adjusts sensitivity to light over wide range (~10¹⁰:1). Mechanism: pupil size change + photoreceptor sensitivity adjustment.
-
Brightness Discrimination: Ability to distinguish intensity differences. Weber's Law: ΔI/I = constant (just noticeable difference proportional to background intensity).
Image Sampling & Quantization
-
Sampling: Digitizing spatial coordinates (x,y). Sampling Theorem: To avoid aliasing, sampling rate ≥ 2× highest spatial frequency (Nyquist rate).
- Spatial Resolution: Minimum distance between samples (pixel size).
-
Quantization: Digitizing amplitude (intensity). Levels = 2ᵏ (k bits/pixel).
-
Quantization Error: Max error = ±Δ/2, where Δ = (max-min)/(L-1), L = levels.
-
Trade-off: More levels → smoother image, larger file size.
-
[!TIP] Exam often asks to relate sampling/quantization to image quality and storage.
2. Noise Modeling and Estimation
Types of Noise
| Noise Type | Probability Density Function (PDF) | Common Source |
|---|---|---|
| Gaussian | $$\displaystyle p(z) = \frac{1}{\sqrt{2\pi\sigma^2}} e^{-\frac{(z-\mu)^2}{2\sigma^2}} $$ | Electronic circuitry, sensor noise |
| Impulse (Salt-and-Pepper) | Random values: 0 (pepper) or 255 (salt) | Defective pixels, transmission errors |
| Rayleigh | $$\displaystyle p(z) = \frac{2(z-a)}{b^2} e^{-\frac{(z-a)^2}{b^2}} $$ (z≥a) | Radar, medical imaging |
| Erlang (Gamma) | $$\displaystyle p(z) = \frac{a^b z^{b-1} e^{-az}}{(b-1)!} $$ | Laser imaging |
Noise Parameter Estimation
-
Global Statistics: Estimate mean (μ) and variance (σ²) from smooth regions or using robust estimators (median, trimmed mean).
-
Local Statistics: Compute variance in small sliding windows (e.g., 3×3). High variance → likely edge/texture, not noise.
-
Method of Moments: Match sample moments to theoretical PDF.
-
Maximum Likelihood Estimation (MLE): Optimize parameters to maximize likelihood of observed data.
[!TIP] For impulse noise, use median of local window to estimate noise-free value.
3. Frequency Domain Processing
2-D Fourier Transform (FT) Properties & Linearity Proof
-
Continuous FT: $$\displaystyle F(u,v) = \int_{-\infty}^{\infty}\int_{-\infty}^{\infty} f(x,y) e^{-j2\pi(ux+vy)} dxdy $$
-
Discrete FT (DFT): $$\displaystyle F(u,v) = \sum_{x=0}^{M-1}\sum_{y=0}^{N-1} f(x,y) e^{-j2\pi(ux/M + vy/N)} $$
-
Linearity Proof:
Let $$\displaystyle g(x,y) = af_1(x,y) + bf_2(x,y) $$. Then:
$$G(u,v) = \mathcal{F}\{g(x,y)\} = \mathcal{F}\{af_1 + bf_2\} = a\mathcal{F}\{f_1\} + b\mathcal{F}\{f_2\} = aF_1(u,v) + bF_2(u,v)$$
Holds for both continuous and discrete cases due to linearity of integral/summation.
Homomorphic Filtering
- Concept: Separate illumination (slowly varying) and reflectance (details) via log transform.
$$f(x,y) = i(x,y) \cdot r(x,y) \quad \text{(multiplicative model)}$$
$$\ln f(x,y) = \ln i(x,y) + \ln r(x,y)$$
-
Steps:
-
Take log: $$\displaystyle g(x,y) = \ln f(x,y) $$
-
Apply filter $H(u,v)$ in frequency domain: $$\displaystyle G'(u,v) = H(u,v)G(u,v) $$
-
Exponentiate: $$\displaystyle f'(x,y) = \exp[g'(x,y)] $$
-
-
Filter Design: $H(u,v)$ suppresses low frequencies (illumination) and amplifies high frequencies (reflectance). Common: Butterworth highpass.
Image Sharpening in Frequency Domain
-
Highpass Filters:
- Butterworth Highpass (order n, cutoff D₀):
$$H(u,v) = \frac{1}{1 + \left(\frac{D_0}{D(u,v)}\right)^{2n}}$$
Smooth transition, no ringing.
- Gaussian Highpass:
$$H(u,v) = 1 - e^{-\frac{D^2(u,v)}{2D_0^2}}$$
No ringing, smooth everywhere.
- $$\displaystyle D(u,v) = \sqrt{(u-M/2)^2 + (v-N/2)^2} $$ (centered at origin).
[!TIP] Butterworth: sharper cutoff with higher n, but more ringing. Gaussian: always smooth.
4. Color Image Processing
Histogram Processing for Color Images
-
Individual Channel Processing: Apply grayscale histogram techniques (equalization, stretching) to R, G, B channels independently. Risk: Color distortion (hue shift).
-
Joint Histogram: 2-D or 3-D histogram considering multiple channels (e.g., RG, GB, RGB). Used for color segmentation, contrast enhancement preserving hue.
-
HSV/HSI Space: Process only Intensity (V or I) channel to preserve color relationships. Example: histogram equalize V channel in HSV, then convert back to RGB.
5. Image Restoration Techniques
Minimum Mean Square Error (MMSE) Filtering
-
Objective: Estimate original image $\hat{f}$ from degraded $g$ to minimize $$\displaystyle E[(\hat{f} - f)^2] $$.
-
Linear MMSE Estimator (for additive noise, linear degradation):
$$\hat{F}(u,v) = \frac{H^*(u,v) S_f(u,v)}{|H(u,v)|^2 S_f(u,v) + S_n(u,v)} G(u,v)$$
where:
-
$H(u,v)$: degradation function
-
$$\displaystyle S_f $$, $$\displaystyle S_n $$: power spectra of image and noise
-
$$\displaystyle H^* $$: complex conjugate
-
Assumes: Stationary signal/noise, known statistics.
Wiener Filtering
-
Special case of MMSE when degradation is linear and shift-invariant, noise is additive, stationary, and uncorrelated with image.
-
Wiener Filter:
$$W(u,v) = \frac{H^*(u,v)}{|H(u,v)|^2 + \frac{S_n(u,v)}{S_f(u,v)}}$$
-
Interpretation: Balances inverse filtering (amplifies noise) and smoothing.
-
Implementation: Often use $$\displaystyle S_f $$ estimated from $g$ (e.g., via periodogram) or assume $$\displaystyle S_f \propto 1/|u|^\alpha $$ (typical image spectrum).
[!TIP] Wiener filter optimal in MMSE sense; requires noise-to-signal power ratio.
6. Image Segmentation
Region-Based Segmentation
-
Thresholding: Partition based on intensity. Global (single T) vs. local (adaptive T). Otsu's method maximizes inter-class variance.
-
Region Growing: Start with seeds, merge neighboring pixels with similar properties (intensity, texture). Needs similarity criterion and stopping rule.
-
Splitting and Merging:
-
Split: Recursively divide image into quadrants until homogeneity.
-
Merge: Combine adjacent homogeneous regions.
-
Often combined: Split-merge algorithm (quadtree representation).
-
Motion-Based Segmentation
-
Uses motion information (optical flow, motion vectors) to separate moving objects from static background.
-
Optical Flow: Apparent motion field $(u,v)$ satisfying brightness constraint: $$\displaystyle I_x u + I_y v + I_t = 0 $$.
-
Segmentation: Cluster motion vectors (e.g., via k-means) or threshold magnitude of flow.
Texture Analysis for Segmentation
-
Co-occurrence Matrix $P(i,j,d,\theta)$: Counts pairs of pixels with intensity $i,j$, separated by distance $d$ at angle $\theta$.
-
Features (from GLCM):
-
Contrast: $$\displaystyle \sum_{i,j} (i-j)^2 P(i,j) $$
-
Correlation: $$\displaystyle \frac{\sum_{i,j} (i-\mu_i)(j-\mu_j)P(i,j)}{\sigma_i\sigma_j} $$
-
Homogeneity: $$\displaystyle \sum_{i,j} \frac{P(i,j)}{1+|i-j|} $$
-
Energy: $$\displaystyle \sum_{i,j} P(i,j)^2 $$
-
-
Use these features for classification (e.g., SVM) or thresholding.
7. Mathematical Morphology
Fundamental Operations (with Structuring Element $B$)
-
Erosion: $$\displaystyle A \ominus B = \{ z | (B)_z \subseteq A \} $$
-
Shrinks foreground, removes small objects, widens gaps.
-
Effect: $$\displaystyle \text{size}(A \ominus B) \approx \text{size}(A) - \text{size}(B) $$.
-
-
Dilation: $$\displaystyle A \oplus B = \{ z | (B)_z \cap A \neq \emptyset \} $$
-
Expands foreground, fills small holes/breaks.
-
Effect: $$\displaystyle \text{size}(A \oplus B) \approx \text{size}(A) + \text{size}(B) $$.
-
Morphological Algorithms
-
Boundary Extraction: $$\displaystyle \partial A = A - (A \ominus B) $$
-
$B$: 3×3 square or cross.
-
Application: object轮廓检测.
-
-
Hole Filling:
-
Complement image: $$\displaystyle A^c = E - A $$ (E: entire image)
-
Find a point $p$ inside hole (in $$\displaystyle A^c $$).
-
Compute $$\displaystyle X_0 = \{p\} $$
-
Iterate: $$\displaystyle X_{k+1} = (X_k \oplus B) \cap A^c $$ until convergence.
-
Hole = final $$\displaystyle X_k $$. Filled image: $$\displaystyle A \cup X_k $$.
-
[!TIP] Erosion/dilation are duals: $$\displaystyle (A \ominus B)^c = A^c \oplus B $$.
8. Image Compression
Need for Compression
-
Redundancy Types:
-
Spatial: Correlation between neighboring pixels.
-
Spectral: Correlation between color bands.
-
Psycho-visual: Human eye insensitive to high frequencies.
-
-
Goals: Reduce bits for storage/transmission, with acceptable loss (lossy) or no loss (lossless).
Vector Quantization (VQ)
-
Concept: Group pixels (or blocks) into vectors; map each to nearest codeword in codebook.
-
LBG Algorithm (Lloyd-Max for vectors):
-
Initialize codebook (e.g., split centroid method).
-
Partition: Assign each training vector to nearest codeword (Voronoi partition).
-
Update: Recompute each codeword as centroid of its partition.
-
Repeat 2-3 until convergence (distortion change < threshold).
-
-
Distortion Measure: Usually mean squared error (MSE).
Lossy Predictive Coding
-
Encoder:
Input f(x,y) → Predictor → Predicted ŷ → Subtract → Error e → Quantizer → Encoded bits → Output ↑ | └──────────────────────────────────────┘ (feedback from decoder) -
Decoder:
Encoded bits → Decoder → Quantized error eq → Add (with predictor) → Reconstructed f'(x,y) -
Working:
-
Predictor uses past/neighboring pixels (e.g., linear: $$\displaystyle \hat{f}(x,y) = \alpha f(x-1,y) + \beta f(x,y-1) $$).
-
Transmit only prediction error $$\displaystyle e = f - \hat{f} $$ (quantized).
-
Decoder reconstructs using same predictor and quantized error.
-
-
Advantage: Exploits spatial correlation; error often small → fewer bits.
9. Object Recognition
Basic Concepts & Pipeline
-
Preprocessing: Noise removal, normalization.
-
Segmentation: Isolate object of interest.
-
Feature Extraction: Describe object (e.g., shape moments, texture features, SIFT).
-
Matching/Classification: Compare features to templates or classify using learned models (e.g., SVM, neural networks).
-
Decision: Assign label.
Common Approaches
-
Template Matching: Directly compare input image (or region) with stored template(s).
-
Methods:
-
Cross-correlation: $$\displaystyle c(i,j) = \sum_{x,y} t(x,y) f(x+i, y+j) $$
-
Sum of Squared Differences (SSD): $$\displaystyle \sum (t - f)^2 $$
-
-
Issues: Sensitive to scale, rotation, illumination.
-
-
Feature-Based Matching: Extract invariant features (e.g., SIFT, SURF) and match keypoints.
-
Model-Based: Use geometric models (e.g., shape context) or machine learning classifiers.
[!TIP] Template matching is simple but not robust; use normalized cross-correlation for illumination invariance.
Key Formulas Summary
-
Nyquist Rate: $$\displaystyle f_s \geq 2f_{\max} $$
-
Quantization Error: $$\displaystyle \pm \frac{\Delta}{2}, \Delta = \frac{L_{\max}-L_{\min}}{2^k-1} $$
-
Wiener Filter: $$\displaystyle W(u,v) = \frac{H^*(u,v)}{|H(u,v)|^2 + K} $$, $$\displaystyle K = \frac{S_n}{S_f} $$
-
Butterworth HPF: $$\displaystyle H(u,v) = \frac{1}{1 + (D_0/D(u,v))^{2n}} $$
-
GLCM Contrast: $$\displaystyle \sum_{i,j} (i-j)^2 P(i,j) $$
-
MSE for VQ: $$\displaystyle \frac{1}{MN}\sum_{i,j} \|x_{ij} - C_{k(i,j)}\|^2 $$
Common Pitfalls
-
Forgetting to center DFT for filtering (shift zero-frequency to center).
-
Using histogram equalization on each RGB channel independently → color distortion.
-
Confusing erosion (shrinks) vs. dilation (expands) effects on object size.
-
Assuming Wiener filter works without noise estimation; need $$\displaystyle S_n/S_f $$ ratio.
Exam Focus Areas (from past papers):
-
Prove Fourier linearity (both continuous & discrete).
-
Derive/explain homomorphic filtering equations.
-
Compare Butterworth vs. Gaussian highpass filters.
-
Explain MMSE vs. Wiener filtering.
-
Draw and explain lossy predictive coding block diagrams.
-
Describe morphological boundary extraction and hole filling algorithms.
-
Explain LBG algorithm steps for vector quantization.
-
Define co-occurrence matrix and list texture features.