UNIT 3: Digital Image Processing
1. Image Fundamentals and Noise
Image Formation in Human Eye
-
Cornea & Lens: Focus light onto retina.
-
Retina: Contains photoreceptors—rods (low light, no color) and cones (color, bright light).
-
Brightness Adaptation: Eye's ability to adjust sensitivity over a wide range of illumination (≈ $$\displaystyle 10^{10} $$). Governed by pupil size and photopigment regeneration.
-
Brightness Discrimination: Ability to distinguish intensity differences. Follows Weber's Law: $$\displaystyle \frac{\Delta I}{I} = \text{constant} $$ (just noticeable difference proportional to background intensity $I$). At low intensities, Ricco's Law applies: $\Delta I \propto I$.
[!TIP] Exam often asks to differentiate adaptation (dynamic range adjustment) vs. discrimination (intensity resolution).
Image Sampling & Quantization
-
Sampling: Digitizing spatial coordinates $(x,y)$. Sampling rate must satisfy Nyquist criterion ($\geq 2 \times$ max spatial frequency) to avoid aliasing.
-
Quantization: Digitizing amplitude (intensity). Number of bits $k$ gives $$\displaystyle L = 2^k $$ gray levels. Quantization error = $$\displaystyle \pm \frac{1}{2} $$ gray level.
-
Key Formula: Bits per pixel (bpp) = $$\displaystyle \log_2 L $$.
Noise Parameter Estimation Approaches
| Approach | Principle | Typical Use |
|---|---|---|
| Training Samples | Estimate from known noise-only regions of image. | Stationary noise. |
| Local Statistics | Compute mean/variance in small neighborhoods; assume noise zero-mean. | Spatially varying noise. |
| Robust Estimators | Use median, trimmed mean to resist outliers. | Impulsive noise (salt-and-pepper). |
| Method of Moments | Fit parametric noise model (e.g., Gaussian $$\displaystyle \mathcal{N}(\mu,\sigma^2) $$) via sample moments. | Additive white noise. |
2. Frequency Domain Processing
2-D Fourier Transform (DFT)
-
Continuous: $$\displaystyle F(u,v) = \int_{-\infty}^{\infty} \int_{-\infty}^{\infty} f(x,y) e^{-j2\pi(ux+vy)} \,dx\,dy $$
-
Discrete (for $M \times N$ image): $$\displaystyle F(u,v) = \sum_{x=0}^{M-1} \sum_{y=0}^{N-1} f(x,y) e^{-j2\pi(ux/M + vy/N)} $$
-
Linearity Proof: If $$\displaystyle f_1(x,y) \leftrightarrow F_1(u,v) $$ and $$\displaystyle f_2(x,y) \leftrightarrow F_2(u,v) $$, then $$\displaystyle af_1 + bf_2 \leftrightarrow aF_1 + bF_2 $$ (follows from linearity of summation and exponential).
Homomorphic Filtering
-
Concept: Separate illumination ($i(x,y)$) and reflectance ($r(x,y)$) components: $$\displaystyle f(x,y) = i(x,y) \cdot r(x,y) $$. Apply log to convert to additive: $$\displaystyle \ln f = \ln i + \ln r $$. Filter in Fourier domain, then exponentiate.
-
Steps:
-
Compute $$\displaystyle F(u,v) = \mathcal{F}\{\ln f(x,y)\} $$
-
Apply filter $H(u,v)$: $$\displaystyle G(u,v) = H(u,v) \cdot F(u,v) $$
-
Inverse DFT: $$\displaystyle g(x,y) = \mathcal{F}^{-1}\{G(u,v)\} $$
-
Exponential: $$\displaystyle \hat{f}(x,y) = e^{g(x,y)} $$
-
-
Filter Design: $$\displaystyle H(u,v) = H_{\text{hp}}(u,v) $$ to suppress low-frequency illumination, enhance high-frequency reflectance.
Highpass Filters for Sharpening
- Butterworth Highpass (order $n$, cutoff $$\displaystyle D_0 $$):
$$H_{\text{BHP}}(u,v) = \frac{1}{1 + \left(\frac{D_0}{D(u,v)}\right)^{2n}}$$
where $$\displaystyle D(u,v) = \sqrt{(u-M/2)^2 + (v-N/2)^2} $$.
- Gaussian Highpass:
$$H_{\text{GHP}}(u,v) = 1 - e^{-\frac{D^2(u,v)}{2D_0^2}}$$
-
Comparison:
| Filter | Roll-off | Ringing Artifacts | |--------|----------|-------------------| | Butterworth | Sharp (high $n$) | Visible | | Gaussian | Smooth | Minimal |
[!TIP] Sharpening in frequency domain = highpass filtering; inverse transform yields spatially enhanced image.
3. Intensity Transformations and Histogram Analysis
Histogram Processing for Color Images
-
Separate Channel Processing: Apply grayscale histogram techniques (equalization, stretching) independently to R, G, B channels. Risk: color distortion.
-
Joint Histogram: Consider 3-D histogram (R,G,B) or transform to HSV/HSI space, process Intensity (V/I) channel only, preserve hue/saturation. Preferred for natural color preservation.
-
Uniform HSI Histogram Equalization: $$\displaystyle H_i = \text{round}\left( (L-1) \cdot \frac{\text{cumulative count}_i}{\text{total pixels}} \right) $$ for intensity channel.
Image Point Operations
-
Definition: $$\displaystyle s = T(r) $$ where output pixel $s$ depends only on input pixel $r$.
-
Types:
-
Negative: $$\displaystyle s = L-1 - r $$
-
Log: $$\displaystyle s = c \log(1+r) $$ (compresses dynamic range)
-
Power-Law (Gamma): $$\displaystyle s = c r^\gamma $$ ($$\displaystyle \gamma < 1 $$ brightens, $$\displaystyle \gamma > 1 $$ darkens)
-
Contrast Stretching: Piecewise linear (e.g., thresholding).
-
4. Image Restoration Techniques
Minimum Mean Square Error (MMSE) Filtering
-
Criterion: Minimize $$\displaystyle E\{ (f - \hat{f})^2 \} $$ where $\hat{f}$ is estimate of original image $f$ from degraded $g$.
-
Assumption: Image and noise are jointly wide-sense stationary.
-
Wiener Filter (special MMSE solution for linear degradation):
$$H_{\text{W}}(u,v) = \frac{P_f(u,v)}{P_f(u,v) + P_n(u,v)} \cdot H^*(u,v) / |H(u,v)|^2$$
where $$\displaystyle P_f $$, $$\displaystyle P_n $$ are power spectra, $H$ is degradation function, $*$ denotes complex conjugate.
- Simplified Form (if $H$ known, noise zero-mean):
$$H_{\text{W}}(u,v) = \frac{H^*(u,v)}{|H(u,v)|^2 + \frac{P_n(u,v)}{P_f(u,v)}}$$
[!TIP] Wiener filter is optimal in MSE sense; requires knowledge of signal and noise power spectra.
5. Image Segmentation
Region-Based Segmentation
-
Approach: Group pixels into homogeneous regions based on intensity/texture.
-
Methods:
-
Thresholding: $$\displaystyle g(x,y) = \begin{cases} 1 & \text{if } f(x,y) > T \\ 0 & \text{otherwise} \end{cases} $$
-
Region Growing: Start with seed, add neighboring pixels with similar intensity (e.g., $$\displaystyle |f_{\text{new}} - f_{\text{region}}| < \delta $$).
-
Split & Merge: Recursively split homogeneous quadrants, then merge adjacent similar regions.
-
-
Key Challenge: Selection of homogeneity criterion and seeds.
Motion-Based Segmentation
-
Principle: Exploit temporal changes between frames.
-
Techniques:
-
Frame Differencing: $$\displaystyle |I_t(x,y) - I_{t-1}(x,y)| > \text{threshold} $$ indicates motion.
-
Optical Flow: Estimate velocity field $$\displaystyle \mathbf{v}(x,y) = (u,v) $$ from brightness constancy: $$\displaystyle I(x,y,t) = I(x+u,y+v,t+1) $$.
-
-
Application: Video object segmentation, tracking.
Morphological Operations (on binary image $A$, structuring element $B$)
-
Erosion: $$\displaystyle A \ominus B = \{ z \mid (B)_z \subseteq A \} $$
Shrinks foreground; removes small objects; boundary moves inward.
-
Dilation: $$\displaystyle A \oplus B = \{ z \mid (B)_z \cap A \neq \emptyset \} $$
Expands foreground; fills small holes; boundary moves outward.
Morphological Algorithms
-
Boundary Extraction: $$\displaystyle \beta(A) = A - (A \ominus B) $$
(Original minus eroded image).
-
Hole Filling:
-
Let $A$ = image with holes, $B$ = structuring element.
-
Find complement $\overline{A}$ (background with holes as objects).
-
Seed $$\displaystyle X_0 $$ = point inside a hole (known background point).
-
Iterate: $$\displaystyle X_{k+1} = (X_k \oplus B) \cap \overline{A} $$ until $$\displaystyle X_{k+1} = X_k $$.
-
Filled holes = $$\displaystyle A \cup X_k $$.
-
[!TIP] Erosion/dilation are duals: $$\displaystyle A \oplus B = (A^c \ominus B)^c $$. Boundary extraction uses erosion difference.
6. Image Compression
Need for Compression
-
Reduce storage space and transmission bandwidth.
-
Lossless: Exact reconstruction (e.g., PNG). Compression ratio limited.
-
Lossy: Higher compression, acceptable distortion (e.g., JPEG). Exploits perceptual redundancy.
Vector Quantization (VQ)
-
Concept: Group pixel blocks (vectors) into codebook of representative vectors.
-
Steps:
-
Training: Collect sample vectors $$\displaystyle \{\mathbf{x}_i\} $$ from images.
-
Codebook Generation: Use Lloyd-Max algorithm ( clustering ) to obtain $K$ codewords $$\displaystyle \{\mathbf{c}_j\} $$.
-
Encoding: For input vector $\mathbf{x}$, find nearest codeword (min Euclidean distance): $$\displaystyle j = \arg\min_k \|\mathbf{x} - \mathbf{c}_k\| $$. Transmit index $j$.
-
Decoding: Replace index with corresponding codeword $$\displaystyle \mathbf{c}_j $$.
-
-
Complexity: Encoding $O(K \cdot \text{dim})$, decoding $O(1)$.
Lossy Predictive Coding
-
Encoder Block Diagram:
Input f(x,y) → Predictor → Predicted ŷ → Subtract → Error e → Quantizer → Entropy Encoder → Bitstream ↑ | └────── Feedback (reconstructed) ←─────┘ -
Decoder:
Bitstream → Entropy Decoder → Quantized error → Add (from predictor) → Reconstructed image -
Working Principle:
-
Prediction: $$\displaystyle \hat{f}(x,y) = P\{ f(\text{neighbors}) \} $$ (e.g., linear: $$\displaystyle \hat{f} = a f_{N} + b f_{W} + c $$).
-
Error Computation: $$\displaystyle e = f - \hat{f} $$.
-
Quantization: $$\displaystyle e_q = Q(e) $$ (lossy step).
-
Reconstruction: $$\displaystyle \hat{f}_{\text{rec}} = \hat{f} + e_q $$.
-
Entropy Coding: Compress $$\displaystyle e_q $$ (e.g., Huffman).
-
-
Advantage: Exploits spatial redundancy; lower bit rate than PCM.
7. Image Analysis and Recognition
Texture Analysis
-
Definition: Visual patterns of repeated arrangements of pixels/gray levels.
-
Approaches:
-
Statistical: Second-order moments (contrast, correlation), Gray-Level Co-occurrence Matrix (GLCM) features (energy, entropy, homogeneity).
-
Structural: Primitive elements (e.g., bricks) and placement rules.
-
Spectral: Frequency domain energy distribution (e.g., Fourier power spectrum).
-
-
Application: Classification, segmentation.
Object Recognition
-
Steps:
-
Segmentation: Isolate objects.
-
Feature Extraction: Shape (Fourier descriptors), moments (Hu's invariants), texture (GLCM), color histograms.
-
Classification: Match features to stored models using:
-
Template matching (correlation, SSD).
-
Statistical classifiers (Bayes, SVM).
-
Neural networks (CNN).
-
-
-
Challenges: Scale, rotation, occlusion, illumination variation.
[!TIP] Recognition = detection + classification. Features must be invariant to transformations.