Skip to content
AD-803 (A) · AI for Remote Sensing/Quick Revision Short Notes

AI for Remote Sensing (AD-803 (A)) - Unit 3 Short Notes

How unit 3 is examined

This unit covers how supervised, unsupervised and deep learning are applied to satellite imagery, and the image classification techniques built on them; no question appeared in the supplied papers, so every topic is taught in full for a possible fresh question.

Supervised Learning for Remote Sensing Analysis

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Supervised learning trains a classifier on pixels whose land cover class is already known, called training samples, and then uses the learned rule to label every other pixel of the image.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-01" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="Supervised workflow. T = training samples, F = spectral features, M = trained model, C = classified map, A = accuracy check"><style>#dsfig-u3-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-01 .t{fill:#16181D;font-weight:500}#dsfig-u3-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-01 .dot{fill:#16181D}#dsfig-u3-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-01 .ah{fill:#454C5A}#dsfig-u3-01 .ah.hi{fill:#2340B8}#dsfig-u3-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-01 .e{stroke:#B1B7C3}html.dark #dsfig-u3-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-01 .t{fill:#E6E8ED}html.dark #dsfig-u3-01 .t.inv{fill:#0F1115}html.dark #dsfig-u3-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-01 .dot{fill:#E6E8ED}html.dark #dsfig-u3-01 .ann{fill:#8FA3FF}html.dark #dsfig-u3-01 .lbl{fill:#858D9C}html.dark #dsfig-u3-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-01 .ah{fill:#B1B7C3}html.dark #dsfig-u3-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah6)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah6)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah6)"/><path class="e" d="M446,40 L535,40" marker-end="url(#ah6)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">T</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">F</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">M</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">A</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Supervised workflow. T = training samples, F = spectral features, M = trained model, C = classified map, A = accuracy check</figcaption></figure>

Key points.

  1. The analyst first selects training areas of known cover, such as water, forest, crop and built-up land, using field visits or high-resolution reference images, and records their spectral values in every band.
  2. The classifier learns the rule that links a spectral signature to a class; common classifiers are minimum distance, maximum likelihood, support vector machine (SVM), decision tree and random forest.
  3. Minimum distance is the simplest: each pixel goes to the class whose mean vector is nearest in spectral space, but it ignores the spread of the class.
  4. Maximum likelihood assumes each class is normally distributed and assigns the pixel to the class with the highest probability, so it is more accurate but needs enough samples per class.
  5. Machine learning classifiers such as SVM and random forest make no normality assumption and handle many bands and mixed data well.
  6. The model is tested on separate, unseen test samples, because accuracy measured on the training pixels is over-optimistic.
  7. Regression, the numeric form of supervised learning, predicts a continuous value such as crop yield, biomass or soil moisture from the bands.
  8. The main weakness is the cost and effort of collecting good labelled samples, and a classifier trained in one region may fail in another.

Formula. Minimum distance rule: pixel $x$ belongs to class $k$ if $d_k=\lVert x-\mu_k\rVert$ is the smallest over all classes, where $\mu_k$ is the class mean vector.

Example. A pixel $x=(60,40)$ in two bands, with water mean $(50,30)$ and soil mean $(80,70)$, has $d_{water}=\sqrt{10^2+10^2}=14.14$ and $d_{soil}=\sqrt{20^2+30^2}=36.06$, so it is classified as water.

Answer frame. Open with the definition; draw the workflow box diagram; develop training samples, classifier choice, minimum distance, maximum likelihood, testing and regression in that order; close with the labelling cost as the main limitation.

Unsupervised Learning for Remote Sensing Analysis

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Unsupervised learning groups image pixels into natural spectral clusters automatically, without any training labels, and the analyst assigns a land cover name to each cluster afterwards.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-02" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="Unsupervised workflow. I = input image, K = choose number of clusters, C = cluster pixels, L = label clusters, M = final map"><style>#dsfig-u3-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-02 .t{fill:#16181D;font-weight:500}#dsfig-u3-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-02 .dot{fill:#16181D}#dsfig-u3-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-02 .ah{fill:#454C5A}#dsfig-u3-02 .ah.hi{fill:#2340B8}#dsfig-u3-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-02 .e{stroke:#B1B7C3}html.dark #dsfig-u3-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-02 .t{fill:#E6E8ED}html.dark #dsfig-u3-02 .t.inv{fill:#0F1115}html.dark #dsfig-u3-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-02 .dot{fill:#E6E8ED}html.dark #dsfig-u3-02 .ann{fill:#8FA3FF}html.dark #dsfig-u3-02 .lbl{fill:#858D9C}html.dark #dsfig-u3-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-02 .ah{fill:#B1B7C3}html.dark #dsfig-u3-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah7)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah7)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah7)"/><path class="e" d="M446,40 L535,40" marker-end="url(#ah7)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">I</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">K</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">L</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">M</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Unsupervised workflow. I = input image, K = choose number of clusters, C = cluster pixels, L = label clusters, M = final map</figcaption></figure>

Key points.

  1. Pixels with similar spectral values fall into the same cluster, so the method needs no prior ground knowledge of the area.
  2. K-means is the basic algorithm: it picks $K$ starting means, assigns each pixel to the nearest mean, recomputes the means, and repeats until the clusters stop changing.
  3. ISODATA extends K-means by splitting large or spread-out clusters and merging close clusters, so the number of clusters can change during the run.
  4. The analyst decides the number of clusters and the stopping rule, and afterwards labels each cluster by comparing it with reference maps or field data.
  5. It is useful for unfamiliar or inaccessible areas and as a first look before choosing training areas for supervised work.
  6. Its weakness is that one cluster may mix two real classes, or one class may be split across several clusters, so results need careful interpretation.

Formula. K-means minimises the within-cluster sum of squares $J=\sum_{k=1}^{K}\sum_{x\in C_k}\lVert x-\mu_k\rVert^2$.

Example. One-band pixels $2,4,10,12$ with starting means $2$ and $10$ give clusters $\{2,4\}$ and $\{10,12\}$; the new means are $3$ and $11$, and no pixel changes cluster, so the algorithm stops with means 3 and 11.

Answer frame. Open with the definition; draw the unsupervised workflow; develop the K-means steps, then ISODATA, then labelling of clusters; close with a one-line comparison against supervised learning.

Point Supervised Unsupervised
Training samples Required Not required
Analyst effort Before classification After classification
Classes Fixed in advance Found from the data
Typical methods Maximum likelihood, SVM K-means, ISODATA
Main risk Poor or biased samples Clusters do not match real classes

Deep Learning for Remote Sensing Analysis

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Deep learning uses multi-layer neural networks, mainly convolutional neural networks (CNN), that learn useful features directly from the image instead of relying on hand-designed features.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-03" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="CNN. I = image patch, C1 and C2 = convolution with ReLU, P1 = pooling, FC = fully connected layer, O = class output"><style>#dsfig-u3-03 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-03 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-03 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-03 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-03 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-03 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-03 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-03 .t{fill:#16181D;font-weight:500}#dsfig-u3-03 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-03 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-03 .dot{fill:#16181D}#dsfig-u3-03 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-03 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-03 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-03 .ah{fill:#454C5A}#dsfig-u3-03 .ah.hi{fill:#2340B8}#dsfig-u3-03 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-03 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-03 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-03 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-03 .e{stroke:#B1B7C3}html.dark #dsfig-u3-03 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-03 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-03 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-03 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-03 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-03 .t{fill:#E6E8ED}html.dark #dsfig-u3-03 .t.inv{fill:#0F1115}html.dark #dsfig-u3-03 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-03 .dot{fill:#E6E8ED}html.dark #dsfig-u3-03 .ann{fill:#8FA3FF}html.dark #dsfig-u3-03 .lbl{fill:#858D9C}html.dark #dsfig-u3-03 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-03 .ah{fill:#B1B7C3}html.dark #dsfig-u3-03 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-03 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-03 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-03 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L122.2,40" marker-end="url(#ah8)"/><path class="e" d="M162.2,40 L225.4,40" marker-end="url(#ah8)"/><path class="e" d="M265.4,40 L328.6,40" marker-end="url(#ah8)"/><path class="e" d="M368.6,40 L431.8,40" marker-end="url(#ah8)"/><path class="e" d="M471.8,40 L535,40" marker-end="url(#ah8)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">I</text><circle class="n" cx="143.2" cy="40" r="18"/><text class="t" x="143.2" y="40" dy=".35em" text-anchor="middle">C1</text><circle class="n" cx="246.4" cy="40" r="18"/><text class="t" x="246.4" y="40" dy=".35em" text-anchor="middle">P1</text><circle class="n" cx="349.6" cy="40" r="18"/><text class="t" x="349.6" y="40" dy=".35em" text-anchor="middle">C2</text><circle class="n" cx="452.8" cy="40" r="18"/><text class="t" x="452.8" y="40" dy=".35em" text-anchor="middle">FC</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">O</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">CNN. I = image patch, C1 and C2 = convolution with ReLU, P1 = pooling, FC = fully connected layer, O = class output</figcaption></figure>

Key points.

  1. A CNN is built from convolution layers, a non-linear activation such as ReLU, pooling layers and a final fully connected layer that outputs class probabilities.
  2. A convolution layer slides small filters over the image and produces feature maps, so the same filter is reused across the whole scene with few parameters.
  3. Early layers learn edges and textures, and deeper layers learn shapes, roofs, roads and fields, so both spatial and spectral context are used.
  4. Pooling reduces the size of the feature maps and gives small shift invariance, which keeps the computation manageable on large scenes.
  5. Popular uses are scene classification, object detection of ships and buildings, and pixel-wise semantic segmentation with U-Net or similar encoder-decoder networks.
  6. Recurrent networks and transformers handle time series of images, for example crop growth or change over seasons.
  7. Deep learning needs a large labelled dataset and GPUs, and is often trained by transfer learning from a pre-trained model to save data.
  8. It usually gives higher accuracy than classical classifiers, but the model behaves as a black box and can overfit small datasets.

Formula. Convolution: $y(i,j)=\sum_{m}\sum_{n}w(m,n)\,x(i+m,\,j+n)+b$, followed by $\text{ReLU}(y)=\max(0,y)$.

Example. A $5\times5$ image convolved with a $3\times3$ filter, stride 1 and no padding, gives an output of size $(5-3+1)\times(5-3+1)=3\times3$; output is 3 x 3.

Answer frame. Open with the definition of a CNN; draw the layer diagram; develop convolution, ReLU, pooling, the fully connected layer, then remote sensing uses; close with the data requirement and accuracy advantage.

Pitfall: Do not say deep learning needs no labels; standard CNN classification is still supervised and needs labelled images.

Image Classification Techniques in Remote Sensing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Image classification is the process of assigning every pixel, or group of pixels, of a satellite image to a land cover class to produce a thematic map.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-04" viewBox="0 0 1074 198" width="1074" height="198" role="img" aria-label="Classification techniques"><style>#dsfig-u3-04 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-04 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-04 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-04 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-04 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-04 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-04 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-04 .t{fill:#16181D;font-weight:500}#dsfig-u3-04 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-04 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-04 .dot{fill:#16181D}#dsfig-u3-04 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-04 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-04 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-04 .ah{fill:#454C5A}#dsfig-u3-04 .ah.hi{fill:#2340B8}#dsfig-u3-04 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-04 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-04 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-04 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-04 .e{stroke:#B1B7C3}html.dark #dsfig-u3-04 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-04 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-04 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-04 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-04 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-04 .t{fill:#E6E8ED}html.dark #dsfig-u3-04 .t.inv{fill:#0F1115}html.dark #dsfig-u3-04 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-04 .dot{fill:#E6E8ED}html.dark #dsfig-u3-04 .ann{fill:#8FA3FF}html.dark #dsfig-u3-04 .lbl{fill:#858D9C}html.dark #dsfig-u3-04 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-04 .ah{fill:#B1B7C3}html.dark #dsfig-u3-04 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-04 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-04 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-04 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="598.3" y1="39" x2="227.8" y2="103"/><line class="e" x1="598.3" y1="39" x2="485" y2="103"/><line class="e" x1="598.3" y1="39" x2="723.5" y2="103"/><line class="e" x1="598.3" y1="39" x2="968.8" y2="103"/><line class="e" x1="227.8" y1="103" x2="86.5" y2="167"/><line class="e" x1="227.8" y1="103" x2="255.5" y2="167"/><line class="e" x1="227.8" y1="103" x2="369" y2="167"/><line class="e" x1="485" y1="103" x2="439.5" y2="167"/><line class="e" x1="485" y1="103" x2="530.5" y2="167"/><line class="e" x1="723.5" y1="103" x2="641" y2="167"/><line class="e" x1="723.5" y1="103" x2="806" y2="167"/><line class="e" x1="968.8" y1="103" x2="937.5" y2="167"/><line class="e" x1="968.8" y1="103" x2="1000" y2="167"/><rect class="n" x="510.3" y="24" width="176" height="30" rx="8"/><text class="t" x="598.3" y="39" dy=".35em" text-anchor="middle">Image classification</text><rect class="n" x="178.8" y="88" width="98" height="30" rx="8"/><text class="t" x="227.8" y="103" dy=".35em" text-anchor="middle">Supervised</text><rect class="n" x="14" y="152" width="145" height="30" rx="8"/><text class="t" x="86.5" y="167" dy=".35em" text-anchor="middle">Minimum distance</text><rect class="n" x="175" y="152" width="161" height="30" rx="8"/><text class="t" x="255.5" y="167" dy=".35em" text-anchor="middle">Maximum likelihood</text><circle class="n" cx="369" cy="167" r="17"/><text class="t" x="369" y="167" dy=".35em" text-anchor="middle">SVM</text><rect class="n" x="428" y="88" width="114" height="30" rx="8"/><text class="t" x="485" y="103" dy=".35em" text-anchor="middle">Unsupervised</text><rect class="n" x="402" y="152" width="75" height="30" rx="8"/><text class="t" x="439.5" y="167" dy=".35em" text-anchor="middle">K-means</text><rect class="n" x="493" y="152" width="75" height="30" rx="8"/><text class="t" x="530.5" y="167" dy=".35em" text-anchor="middle">ISODATA</text><rect class="n" x="666.5" y="88" width="114" height="30" rx="8"/><text class="t" x="723.5" y="103" dy=".35em" text-anchor="middle">Object-based</text><rect class="n" x="584" y="152" width="114" height="30" rx="8"/><text class="t" x="641" y="167" dy=".35em" text-anchor="middle">Segmentation</text><rect class="n" x="714" y="152" width="184" height="30" rx="8"/><text class="t" x="806" y="167" dy=".35em" text-anchor="middle">Rule or ML on objects</text><rect class="n" x="907.8" y="88" width="122" height="30" rx="8"/><text class="t" x="968.8" y="103" dy=".35em" text-anchor="middle">Deep learning</text><circle class="n" cx="937.5" cy="167" r="17"/><text class="t" x="937.5" y="167" dy=".35em" text-anchor="middle">CNN</text><rect class="n" x="970.5" y="152" width="59" height="30" rx="8"/><text class="t" x="1000" y="167" dy=".35em" text-anchor="middle">U-Net</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Classification techniques</figcaption></figure>

Steps.

Step 1: Select the image and bands, then pre-process (radiometric and geometric correction).
Step 2: Decide the classes, for example water, forest, crop, urban, bare soil.
Step 3: Train the classifier (supervised) or cluster the pixels (unsupervised).
Step 4: Classify all pixels into classes.
Step 5: Post-process by smoothing and merging tiny patches.
Step 6: Assess accuracy with a confusion matrix.

Key points.

  1. Pixel-based classification uses only the spectrum of each pixel, which is fast but gives a noisy salt-and-pepper look in high-resolution images.
  2. Object-based classification first segments the image into homogeneous objects and then classifies them using colour, shape, size and texture, which suits high-resolution data.
  3. Hard classification gives each pixel exactly one class, while soft or fuzzy classification gives class fractions to handle mixed pixels.
  4. Hybrid classification combines clustering with supervised labelling to get the benefits of both.
  5. Accuracy is assessed on independent reference samples using a confusion matrix, whose diagonal holds the correct pixels.
  6. Overall accuracy is correct pixels divided by total pixels, and the kappa coefficient removes the agreement expected by chance.
  7. Producer's accuracy shows how well a real class was captured, and user's accuracy shows how reliable a mapped class is.

Formula. $\text{OA}=\dfrac{\sum \text{diagonal}}{N}\times100$ and $\kappa=\dfrac{p_o-p_e}{1-p_e}$, where $p_o$ is observed agreement and $p_e$ is chance agreement.

Example. For a 3-class matrix with rows $[50,3,2]$, $[4,45,1]$, $[2,5,38]$ and $N=150$: diagonal $=133$, so $p_o=0.887$; row totals $55,50,45$ and column totals $56,53,41$ give $p_e=7575/22500=0.337$; $\kappa=(0.887-0.337)/(1-0.337)=0.83$. OA = 88.7 percent, kappa = 0.83.

Answer frame. Open with the definition; draw the classification tree; give the six steps; compare pixel-based with object-based; close with accuracy assessment by confusion matrix, overall accuracy and kappa.

Point Pixel-based Object-based
Unit classified Single pixel Segment of similar pixels
Information used Spectrum only Spectrum, shape, texture, context
Best for Medium and coarse resolution High resolution
Output look Can be noisy Smooth, map-like
Effort Low Higher (segmentation step)

Last-minute revision

  • Supervised learning uses labelled training pixels; unsupervised learning uses none and the analyst names clusters afterwards.
  • Minimum distance assigns a pixel to the nearest class mean: $d_k=\lVert x-\mu_k\rVert$.
  • Maximum likelihood assumes each class is normally distributed and picks the class with the highest probability.
  • K-means minimises $J=\sum\sum\lVert x-\mu_k\rVert^2$; ISODATA also splits and merges clusters.
  • A CNN has convolution, ReLU, pooling and a fully connected layer; U-Net does pixel-wise segmentation.
  • Convolution output size without padding is $(n-f+1)$ for stride 1.
  • Deep learning needs large labelled data and GPUs; transfer learning reduces the data need.
  • Pixel-based classification uses the spectrum; object-based classification uses segments, shape and texture.
  • Hard classification gives one class per pixel; soft classification gives fractions for mixed pixels.
  • Overall accuracy $=$ diagonal $/\,N$; kappa $=(p_o-p_e)/(1-p_e)$.
  • Producer's accuracy relates to omission errors; user's accuracy relates to commission errors.

Memory hooks

  • Supervised = "teacher shows the answers"; unsupervised = "sort into piles, then name the piles".
  • CNN order: Convolve, Activate, Pool, Fully connect (CAPF).
  • Classification steps: Bands, Classes, Train, Classify, Clean, Check.
  • Accuracy trio: matrix, overall accuracy, kappa.
  • Producer asks "did you find my class", user asks "can I trust your map".

Coverage checklist

  • Supervised Learning for Remote Sensing Analysis: definition, workflow, classifiers, minimum distance example; no past questions.
  • Unsupervised Learning for Remote Sensing Analysis: definition, K-means, ISODATA, comparison table; no past questions.
  • Deep Learning for Remote Sensing Analysis: definition, CNN diagram, uses, convolution formula; no past questions.
  • Image Image Classification Techniques in Remote Sensing: classification tree, steps, pixel versus object table, accuracy and kappa; no past questions.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in