How unit 3 is examined
This unit covers how supervised, unsupervised and deep learning are applied to satellite imagery, and the image classification techniques built on them; no question appeared in the supplied papers, so every topic is taught in full for a possible fresh question.
Supervised Learning for Remote Sensing Analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Supervised learning trains a classifier on pixels whose land cover class is already known, called training samples, and then uses the learned rule to label every other pixel of the image.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-01" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="Supervised workflow. T = training samples, F = spectral features, M = trained model, C = classified map, A = accuracy check"><style>#dsfig-u3-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-01 .t{fill:#16181D;font-weight:500}#dsfig-u3-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-01 .dot{fill:#16181D}#dsfig-u3-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-01 .ah{fill:#454C5A}#dsfig-u3-01 .ah.hi{fill:#2340B8}#dsfig-u3-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-01 .e{stroke:#B1B7C3}html.dark #dsfig-u3-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-01 .t{fill:#E6E8ED}html.dark #dsfig-u3-01 .t.inv{fill:#0F1115}html.dark #dsfig-u3-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-01 .dot{fill:#E6E8ED}html.dark #dsfig-u3-01 .ann{fill:#8FA3FF}html.dark #dsfig-u3-01 .lbl{fill:#858D9C}html.dark #dsfig-u3-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-01 .ah{fill:#B1B7C3}html.dark #dsfig-u3-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh6" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah6)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah6)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah6)"/><path class="e" d="M446,40 L535,40" marker-end="url(#ah6)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">T</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">F</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">M</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">A</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Supervised workflow. T = training samples, F = spectral features, M = trained model, C = classified map, A = accuracy check</figcaption></figure>
Key points.
- The analyst first selects training areas of known cover, such as water, forest, crop and built-up land, using field visits or high-resolution reference images, and records their spectral values in every band.
- The classifier learns the rule that links a spectral signature to a class; common classifiers are minimum distance, maximum likelihood, support vector machine (SVM), decision tree and random forest.
- Minimum distance is the simplest: each pixel goes to the class whose mean vector is nearest in spectral space, but it ignores the spread of the class.
- Maximum likelihood assumes each class is normally distributed and assigns the pixel to the class with the highest probability, so it is more accurate but needs enough samples per class.
- Machine learning classifiers such as SVM and random forest make no normality assumption and handle many bands and mixed data well.
- The model is tested on separate, unseen test samples, because accuracy measured on the training pixels is over-optimistic.
- Regression, the numeric form of supervised learning, predicts a continuous value such as crop yield, biomass or soil moisture from the bands.
- The main weakness is the cost and effort of collecting good labelled samples, and a classifier trained in one region may fail in another.
Formula. Minimum distance rule: pixel $x$ belongs to class $k$ if $d_k=\lVert x-\mu_k\rVert$ is the smallest over all classes, where $\mu_k$ is the class mean vector.
Example. A pixel $x=(60,40)$ in two bands, with water mean $(50,30)$ and soil mean $(80,70)$, has $d_{water}=\sqrt{10^2+10^2}=14.14$ and $d_{soil}=\sqrt{20^2+30^2}=36.06$, so it is classified as water.
Answer frame. Open with the definition; draw the workflow box diagram; develop training samples, classifier choice, minimum distance, maximum likelihood, testing and regression in that order; close with the labelling cost as the main limitation.
Unsupervised Learning for Remote Sensing Analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Unsupervised learning groups image pixels into natural spectral clusters automatically, without any training labels, and the analyst assigns a land cover name to each cluster afterwards.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-02" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="Unsupervised workflow. I = input image, K = choose number of clusters, C = cluster pixels, L = label clusters, M = final map"><style>#dsfig-u3-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-02 .t{fill:#16181D;font-weight:500}#dsfig-u3-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-02 .dot{fill:#16181D}#dsfig-u3-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-02 .ah{fill:#454C5A}#dsfig-u3-02 .ah.hi{fill:#2340B8}#dsfig-u3-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-02 .e{stroke:#B1B7C3}html.dark #dsfig-u3-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-02 .t{fill:#E6E8ED}html.dark #dsfig-u3-02 .t.inv{fill:#0F1115}html.dark #dsfig-u3-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-02 .dot{fill:#E6E8ED}html.dark #dsfig-u3-02 .ann{fill:#8FA3FF}html.dark #dsfig-u3-02 .lbl{fill:#858D9C}html.dark #dsfig-u3-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-02 .ah{fill:#B1B7C3}html.dark #dsfig-u3-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L148,40" marker-end="url(#ah7)"/><path class="e" d="M188,40 L277,40" marker-end="url(#ah7)"/><path class="e" d="M317,40 L406,40" marker-end="url(#ah7)"/><path class="e" d="M446,40 L535,40" marker-end="url(#ah7)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">I</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">K</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">L</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">M</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Unsupervised workflow. I = input image, K = choose number of clusters, C = cluster pixels, L = label clusters, M = final map</figcaption></figure>
Key points.
- Pixels with similar spectral values fall into the same cluster, so the method needs no prior ground knowledge of the area.
- K-means is the basic algorithm: it picks $K$ starting means, assigns each pixel to the nearest mean, recomputes the means, and repeats until the clusters stop changing.
- ISODATA extends K-means by splitting large or spread-out clusters and merging close clusters, so the number of clusters can change during the run.
- The analyst decides the number of clusters and the stopping rule, and afterwards labels each cluster by comparing it with reference maps or field data.
- It is useful for unfamiliar or inaccessible areas and as a first look before choosing training areas for supervised work.
- Its weakness is that one cluster may mix two real classes, or one class may be split across several clusters, so results need careful interpretation.
Formula. K-means minimises the within-cluster sum of squares $J=\sum_{k=1}^{K}\sum_{x\in C_k}\lVert x-\mu_k\rVert^2$.
Example. One-band pixels $2,4,10,12$ with starting means $2$ and $10$ give clusters $\{2,4\}$ and $\{10,12\}$; the new means are $3$ and $11$, and no pixel changes cluster, so the algorithm stops with means 3 and 11.
Answer frame. Open with the definition; draw the unsupervised workflow; develop the K-means steps, then ISODATA, then labelling of clusters; close with a one-line comparison against supervised learning.
| Point | Supervised | Unsupervised |
|---|---|---|
| Training samples | Required | Not required |
| Analyst effort | Before classification | After classification |
| Classes | Fixed in advance | Found from the data |
| Typical methods | Maximum likelihood, SVM | K-means, ISODATA |
| Main risk | Poor or biased samples | Clusters do not match real classes |
Deep Learning for Remote Sensing Analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Deep learning uses multi-layer neural networks, mainly convolutional neural networks (CNN), that learn useful features directly from the image instead of relying on hand-designed features.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-03" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="CNN. I = image patch, C1 and C2 = convolution with ReLU, P1 = pooling, FC = fully connected layer, O = class output"><style>#dsfig-u3-03 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-03 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-03 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-03 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-03 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-03 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-03 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-03 .t{fill:#16181D;font-weight:500}#dsfig-u3-03 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-03 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-03 .dot{fill:#16181D}#dsfig-u3-03 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-03 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-03 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-03 .ah{fill:#454C5A}#dsfig-u3-03 .ah.hi{fill:#2340B8}#dsfig-u3-03 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-03 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-03 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-03 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-03 .e{stroke:#B1B7C3}html.dark #dsfig-u3-03 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-03 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-03 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-03 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-03 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-03 .t{fill:#E6E8ED}html.dark #dsfig-u3-03 .t.inv{fill:#0F1115}html.dark #dsfig-u3-03 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-03 .dot{fill:#E6E8ED}html.dark #dsfig-u3-03 .ann{fill:#8FA3FF}html.dark #dsfig-u3-03 .lbl{fill:#858D9C}html.dark #dsfig-u3-03 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-03 .ah{fill:#B1B7C3}html.dark #dsfig-u3-03 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-03 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-03 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-03 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L122.2,40" marker-end="url(#ah8)"/><path class="e" d="M162.2,40 L225.4,40" marker-end="url(#ah8)"/><path class="e" d="M265.4,40 L328.6,40" marker-end="url(#ah8)"/><path class="e" d="M368.6,40 L431.8,40" marker-end="url(#ah8)"/><path class="e" d="M471.8,40 L535,40" marker-end="url(#ah8)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">I</text><circle class="n" cx="143.2" cy="40" r="18"/><text class="t" x="143.2" y="40" dy=".35em" text-anchor="middle">C1</text><circle class="n" cx="246.4" cy="40" r="18"/><text class="t" x="246.4" y="40" dy=".35em" text-anchor="middle">P1</text><circle class="n" cx="349.6" cy="40" r="18"/><text class="t" x="349.6" y="40" dy=".35em" text-anchor="middle">C2</text><circle class="n" cx="452.8" cy="40" r="18"/><text class="t" x="452.8" y="40" dy=".35em" text-anchor="middle">FC</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">O</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">CNN. I = image patch, C1 and C2 = convolution with ReLU, P1 = pooling, FC = fully connected layer, O = class output</figcaption></figure>
Key points.
- A CNN is built from convolution layers, a non-linear activation such as ReLU, pooling layers and a final fully connected layer that outputs class probabilities.
- A convolution layer slides small filters over the image and produces feature maps, so the same filter is reused across the whole scene with few parameters.
- Early layers learn edges and textures, and deeper layers learn shapes, roofs, roads and fields, so both spatial and spectral context are used.
- Pooling reduces the size of the feature maps and gives small shift invariance, which keeps the computation manageable on large scenes.
- Popular uses are scene classification, object detection of ships and buildings, and pixel-wise semantic segmentation with U-Net or similar encoder-decoder networks.
- Recurrent networks and transformers handle time series of images, for example crop growth or change over seasons.
- Deep learning needs a large labelled dataset and GPUs, and is often trained by transfer learning from a pre-trained model to save data.
- It usually gives higher accuracy than classical classifiers, but the model behaves as a black box and can overfit small datasets.
Formula. Convolution: $y(i,j)=\sum_{m}\sum_{n}w(m,n)\,x(i+m,\,j+n)+b$, followed by $\text{ReLU}(y)=\max(0,y)$.
Example. A $5\times5$ image convolved with a $3\times3$ filter, stride 1 and no padding, gives an output of size $(5-3+1)\times(5-3+1)=3\times3$; output is 3 x 3.
Answer frame. Open with the definition of a CNN; draw the layer diagram; develop convolution, ReLU, pooling, the fully connected layer, then remote sensing uses; close with the data requirement and accuracy advantage.
Pitfall: Do not say deep learning needs no labels; standard CNN classification is still supervised and needs labelled images.
Image Classification Techniques in Remote Sensing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Image classification is the process of assigning every pixel, or group of pixels, of a satellite image to a land cover class to produce a thematic map.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-04" viewBox="0 0 1074 198" width="1074" height="198" role="img" aria-label="Classification techniques"><style>#dsfig-u3-04 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-04 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-04 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-04 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-04 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-04 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-04 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-04 .t{fill:#16181D;font-weight:500}#dsfig-u3-04 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-04 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-04 .dot{fill:#16181D}#dsfig-u3-04 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-04 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-04 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-04 .ah{fill:#454C5A}#dsfig-u3-04 .ah.hi{fill:#2340B8}#dsfig-u3-04 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-04 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-04 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-04 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-04 .e{stroke:#B1B7C3}html.dark #dsfig-u3-04 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-04 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-04 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-04 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-04 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-04 .t{fill:#E6E8ED}html.dark #dsfig-u3-04 .t.inv{fill:#0F1115}html.dark #dsfig-u3-04 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-04 .dot{fill:#E6E8ED}html.dark #dsfig-u3-04 .ann{fill:#8FA3FF}html.dark #dsfig-u3-04 .lbl{fill:#858D9C}html.dark #dsfig-u3-04 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-04 .ah{fill:#B1B7C3}html.dark #dsfig-u3-04 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-04 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-04 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-04 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="598.3" y1="39" x2="227.8" y2="103"/><line class="e" x1="598.3" y1="39" x2="485" y2="103"/><line class="e" x1="598.3" y1="39" x2="723.5" y2="103"/><line class="e" x1="598.3" y1="39" x2="968.8" y2="103"/><line class="e" x1="227.8" y1="103" x2="86.5" y2="167"/><line class="e" x1="227.8" y1="103" x2="255.5" y2="167"/><line class="e" x1="227.8" y1="103" x2="369" y2="167"/><line class="e" x1="485" y1="103" x2="439.5" y2="167"/><line class="e" x1="485" y1="103" x2="530.5" y2="167"/><line class="e" x1="723.5" y1="103" x2="641" y2="167"/><line class="e" x1="723.5" y1="103" x2="806" y2="167"/><line class="e" x1="968.8" y1="103" x2="937.5" y2="167"/><line class="e" x1="968.8" y1="103" x2="1000" y2="167"/><rect class="n" x="510.3" y="24" width="176" height="30" rx="8"/><text class="t" x="598.3" y="39" dy=".35em" text-anchor="middle">Image classification</text><rect class="n" x="178.8" y="88" width="98" height="30" rx="8"/><text class="t" x="227.8" y="103" dy=".35em" text-anchor="middle">Supervised</text><rect class="n" x="14" y="152" width="145" height="30" rx="8"/><text class="t" x="86.5" y="167" dy=".35em" text-anchor="middle">Minimum distance</text><rect class="n" x="175" y="152" width="161" height="30" rx="8"/><text class="t" x="255.5" y="167" dy=".35em" text-anchor="middle">Maximum likelihood</text><circle class="n" cx="369" cy="167" r="17"/><text class="t" x="369" y="167" dy=".35em" text-anchor="middle">SVM</text><rect class="n" x="428" y="88" width="114" height="30" rx="8"/><text class="t" x="485" y="103" dy=".35em" text-anchor="middle">Unsupervised</text><rect class="n" x="402" y="152" width="75" height="30" rx="8"/><text class="t" x="439.5" y="167" dy=".35em" text-anchor="middle">K-means</text><rect class="n" x="493" y="152" width="75" height="30" rx="8"/><text class="t" x="530.5" y="167" dy=".35em" text-anchor="middle">ISODATA</text><rect class="n" x="666.5" y="88" width="114" height="30" rx="8"/><text class="t" x="723.5" y="103" dy=".35em" text-anchor="middle">Object-based</text><rect class="n" x="584" y="152" width="114" height="30" rx="8"/><text class="t" x="641" y="167" dy=".35em" text-anchor="middle">Segmentation</text><rect class="n" x="714" y="152" width="184" height="30" rx="8"/><text class="t" x="806" y="167" dy=".35em" text-anchor="middle">Rule or ML on objects</text><rect class="n" x="907.8" y="88" width="122" height="30" rx="8"/><text class="t" x="968.8" y="103" dy=".35em" text-anchor="middle">Deep learning</text><circle class="n" cx="937.5" cy="167" r="17"/><text class="t" x="937.5" y="167" dy=".35em" text-anchor="middle">CNN</text><rect class="n" x="970.5" y="152" width="59" height="30" rx="8"/><text class="t" x="1000" y="167" dy=".35em" text-anchor="middle">U-Net</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Classification techniques</figcaption></figure>
Steps.
Step 1: Select the image and bands, then pre-process (radiometric and geometric correction).
Step 2: Decide the classes, for example water, forest, crop, urban, bare soil.
Step 3: Train the classifier (supervised) or cluster the pixels (unsupervised).
Step 4: Classify all pixels into classes.
Step 5: Post-process by smoothing and merging tiny patches.
Step 6: Assess accuracy with a confusion matrix.
Key points.
- Pixel-based classification uses only the spectrum of each pixel, which is fast but gives a noisy salt-and-pepper look in high-resolution images.
- Object-based classification first segments the image into homogeneous objects and then classifies them using colour, shape, size and texture, which suits high-resolution data.
- Hard classification gives each pixel exactly one class, while soft or fuzzy classification gives class fractions to handle mixed pixels.
- Hybrid classification combines clustering with supervised labelling to get the benefits of both.
- Accuracy is assessed on independent reference samples using a confusion matrix, whose diagonal holds the correct pixels.
- Overall accuracy is correct pixels divided by total pixels, and the kappa coefficient removes the agreement expected by chance.
- Producer's accuracy shows how well a real class was captured, and user's accuracy shows how reliable a mapped class is.
Formula. $\text{OA}=\dfrac{\sum \text{diagonal}}{N}\times100$ and $\kappa=\dfrac{p_o-p_e}{1-p_e}$, where $p_o$ is observed agreement and $p_e$ is chance agreement.
Example. For a 3-class matrix with rows $[50,3,2]$, $[4,45,1]$, $[2,5,38]$ and $N=150$: diagonal $=133$, so $p_o=0.887$; row totals $55,50,45$ and column totals $56,53,41$ give $p_e=7575/22500=0.337$; $\kappa=(0.887-0.337)/(1-0.337)=0.83$. OA = 88.7 percent, kappa = 0.83.
Answer frame. Open with the definition; draw the classification tree; give the six steps; compare pixel-based with object-based; close with accuracy assessment by confusion matrix, overall accuracy and kappa.
| Point | Pixel-based | Object-based |
|---|---|---|
| Unit classified | Single pixel | Segment of similar pixels |
| Information used | Spectrum only | Spectrum, shape, texture, context |
| Best for | Medium and coarse resolution | High resolution |
| Output look | Can be noisy | Smooth, map-like |
| Effort | Low | Higher (segmentation step) |
Last-minute revision
- Supervised learning uses labelled training pixels; unsupervised learning uses none and the analyst names clusters afterwards.
- Minimum distance assigns a pixel to the nearest class mean: $d_k=\lVert x-\mu_k\rVert$.
- Maximum likelihood assumes each class is normally distributed and picks the class with the highest probability.
- K-means minimises $J=\sum\sum\lVert x-\mu_k\rVert^2$; ISODATA also splits and merges clusters.
- A CNN has convolution, ReLU, pooling and a fully connected layer; U-Net does pixel-wise segmentation.
- Convolution output size without padding is $(n-f+1)$ for stride 1.
- Deep learning needs large labelled data and GPUs; transfer learning reduces the data need.
- Pixel-based classification uses the spectrum; object-based classification uses segments, shape and texture.
- Hard classification gives one class per pixel; soft classification gives fractions for mixed pixels.
- Overall accuracy $=$ diagonal $/\,N$; kappa $=(p_o-p_e)/(1-p_e)$.
- Producer's accuracy relates to omission errors; user's accuracy relates to commission errors.
Memory hooks
- Supervised = "teacher shows the answers"; unsupervised = "sort into piles, then name the piles".
- CNN order: Convolve, Activate, Pool, Fully connect (CAPF).
- Classification steps: Bands, Classes, Train, Classify, Clean, Check.
- Accuracy trio: matrix, overall accuracy, kappa.
- Producer asks "did you find my class", user asks "can I trust your map".
Coverage checklist
- Supervised Learning for Remote Sensing Analysis: definition, workflow, classifiers, minimum distance example; no past questions.
- Unsupervised Learning for Remote Sensing Analysis: definition, K-means, ISODATA, comparison table; no past questions.
- Deep Learning for Remote Sensing Analysis: definition, CNN diagram, uses, convolution formula; no past questions.
- Image Image Classification Techniques in Remote Sensing: classification tree, steps, pixel versus object table, accuracy and kappa; no past questions.