How unit 4 is examined
This unit covers four topics: detecting and tracking objects, detecting change over time, hyperspectral and LiDAR analysis, and multi-modal data fusion. None was asked in the supplied papers, so each is taught with a definition, developed points, formulas and a frame for a 7-mark answer in case it appears this year.
Object detection and tracking in remote sensing imagery
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Object detection locates and classifies objects such as ships, aircraft, vehicles and buildings in an image by drawing a bounding box and a class label around each one; tracking links the same object across successive images to follow its motion.</mark>
Diagram.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-01" viewBox="0 0 467 252" width="467" height="252" role="img" aria-label="Detection pipeline. Img = satellite image, CNN = feature extractor, Box = bounding-box regression, Cls = class label, Trk = tracker linking detections across frames"><style>#dsfig-u4-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-01 .t{fill:#16181D;font-weight:500}#dsfig-u4-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-01 .dot{fill:#16181D}#dsfig-u4-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-01 .ah{fill:#454C5A}#dsfig-u4-01 .ah.hi{fill:#2340B8}#dsfig-u4-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-01 .e{stroke:#B1B7C3}html.dark #dsfig-u4-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-01 .t{fill:#E6E8ED}html.dark #dsfig-u4-01 .t.inv{fill:#0F1115}html.dark #dsfig-u4-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-01 .dot{fill:#E6E8ED}html.dark #dsfig-u4-01 .ann{fill:#8FA3FF}html.dark #dsfig-u4-01 .lbl{fill:#858D9C}html.dark #dsfig-u4-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-01 .ah{fill:#B1B7C3}html.dark #dsfig-u4-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh10" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,126 L148,126" marker-end="url(#ah10)"/><path class="e" d="M184.8,115.5 L280.5,51.6" marker-end="url(#ah10)"/><path class="e" d="M184.8,136.5 L280.5,200.4" marker-end="url(#ah10)"/><path class="e" d="M313.8,50.5 L409.5,114.4" marker-end="url(#ah10)"/><path class="e" d="M313.8,201.5 L409.5,137.6" marker-end="url(#ah10)"/><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">Img</text><circle class="n" cx="169" cy="126" r="18"/><text class="t" x="169" y="126" dy=".35em" text-anchor="middle">CNN</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Box</text><circle class="n" cx="298" cy="212" r="18"/><text class="t" x="298" y="212" dy=".35em" text-anchor="middle">Cls</text><circle class="n" cx="427" cy="126" r="18"/><text class="t" x="427" y="126" dy=".35em" text-anchor="middle">Trk</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Detection pipeline. Img = satellite image, CNN = feature extractor, Box = bounding-box regression, Cls = class label, Trk = tracker linking detections across frames</figcaption></figure>
Key points.
- Object detection answers two questions at once: what the object is (classification) and where it is (localisation by a bounding box).
- Modern detectors are deep CNN models, either two-stage (Faster R-CNN proposes regions, then classifies them) or one-stage (YOLO, SSD predict boxes and classes in a single pass, so they are faster).
- Remote sensing objects are small, densely packed (parked cars, ships in a port), arbitrarily rotated and seen from above, so oriented bounding boxes and high-resolution inputs are often used.
- Training needs many labelled images; benchmark datasets such as DOTA and xView provide labelled aircraft, ships and vehicles, and transfer learning from ImageNet reduces the labelling cost.
- Accuracy is judged by intersection over union, $IoU = \dfrac{\text{area of overlap}}{\text{area of union}}$; a detection counts as correct when IoU exceeds a threshold, usually 0.5.
- Precision is $TP/(TP+FP)$, recall is $TP/(TP+FN)$, and mean average precision (mAP) averages the area under the precision-recall curve over all classes.
- Tracking follows an object over time in satellite video or drone footage by matching detections between frames using a Kalman filter for predicted motion and appearance features for identity.
- Applications are ship monitoring, vehicle counting, aircraft detection at airfields, building extraction and wildlife counting.
Example. A predicted box overlaps the true box by 60 units of area, and their union is 100 units, so $IoU = 60/100 = 0.6$, which is above 0.5 and is counted as a true positive.
Answer frame. Open with the definition; draw the pipeline diagram; then develop points 1-2 (task and detector types), 3-4 (why remote sensing is hard), 5-6 (IoU and metrics), 7 (tracking); close with one application line such as ship monitoring.
Pitfall: Do not confuse detection with classification; classification gives one label per image, detection gives a label and a location for every object.
Change detection and time-series analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Change detection compares co-registered images of the same area taken at different dates to identify what has changed on the ground; time-series analysis studies a sequence of many dates to find trends, seasonal cycles and abrupt events.</mark>
Diagram.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-02" viewBox="0 0 467 252" width="467" height="252" role="img" aria-label="Change detection workflow. T1, T2 = images at two dates, Reg = co-registration and radiometric correction, Cmp = comparison method, Map = change map"><style>#dsfig-u4-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-02 .t{fill:#16181D;font-weight:500}#dsfig-u4-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-02 .dot{fill:#16181D}#dsfig-u4-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-02 .ah{fill:#454C5A}#dsfig-u4-02 .ah.hi{fill:#2340B8}#dsfig-u4-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-02 .e{stroke:#B1B7C3}html.dark #dsfig-u4-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-02 .t{fill:#E6E8ED}html.dark #dsfig-u4-02 .t.inv{fill:#0F1115}html.dark #dsfig-u4-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-02 .dot{fill:#E6E8ED}html.dark #dsfig-u4-02 .ann{fill:#8FA3FF}html.dark #dsfig-u4-02 .lbl{fill:#858D9C}html.dark #dsfig-u4-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-02 .ah{fill:#B1B7C3}html.dark #dsfig-u4-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah11" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh11" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.8,50.5 L151.5,114.4" marker-end="url(#ah11)"/><path class="e" d="M55.8,201.5 L151.5,137.6" marker-end="url(#ah11)"/><path class="e" d="M188,126 L277,126" marker-end="url(#ah11)"/><path class="e" d="M317,126 L406,126" marker-end="url(#ah11)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">T1</text><circle class="n" cx="40" cy="212" r="18"/><text class="t" x="40" y="212" dy=".35em" text-anchor="middle">T2</text><circle class="n" cx="169" cy="126" r="18"/><text class="t" x="169" y="126" dy=".35em" text-anchor="middle">Reg</text><circle class="n" cx="298" cy="126" r="18"/><text class="t" x="298" y="126" dy=".35em" text-anchor="middle">Cmp</text><circle class="n" cx="427" cy="126" r="18"/><text class="t" x="427" y="126" dy=".35em" text-anchor="middle">Map</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Change detection workflow. T1, T2 = images at two dates, Reg = co-registration and radiometric correction, Cmp = comparison method, Map = change map</figcaption></figure>
Key points.
- Change detection needs images of the same place, sensor type and season, because different sun angle or season creates apparent change that is not real.
- Images must first be co-registered (aligned to sub-pixel accuracy) and radiometrically corrected; misregistration is the main source of false change.
- Image differencing computes $D = X_2 - X_1$ for each pixel and marks pixels whose difference exceeds a threshold as changed.
- Image ratioing uses $R = X_2 / X_1$, where a value near 1 means no change and it reduces the effect of illumination differences.
- Post-classification comparison classifies each date separately and compares the two maps, which gives a from-to change matrix such as forest to urban.
- Deep learning methods use Siamese networks: both dates pass through the same shared-weight CNN and the difference of their features is decoded into a change map.
- Time-series analysis uses many dates of an index such as $NDVI = \dfrac{NIR - Red}{NIR + Red}$ to track crop growth, deforestation, urban growth and drought.
- Recurrent networks (LSTM) and algorithms such as BFAST and LandTrendr model the series to separate seasonal cycles from real disturbance.
Example. A pixel has NDVI 0.72 in 2020 and 0.15 in 2024; $D = 0.15 - 0.72 = -0.57$, a large fall beyond any sensible threshold, so it is marked as vegetation loss.
Answer frame. Open with the definition; draw the workflow diagram; then develop points 1-2 (preconditions), 3-5 (classical methods), 6 (deep learning), 7-8 (time series); close with an application such as deforestation or urban growth mapping.
Pitfall: Skipping co-registration; small misalignment alone produces false change along every edge.
Hyperspectral and LiDAR data analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Hyperspectral imaging records hundreds of narrow, contiguous spectral bands for every pixel, giving a continuous spectrum for material identification; LiDAR (Light Detection and Ranging) is an active laser system that measures distance to give a 3D point cloud of the surface.</mark>
Comparison.
| Feature | Multispectral | Hyperspectral | LiDAR |
|---|---|---|---|
| Source | Passive sunlight | Passive sunlight | Active laser pulse |
| Data | 3-10 broad bands | 100-250+ narrow bands (about 10 nm) | 3D points with x, y, z |
| Main output | Land cover map | Material spectrum | Elevation, height, structure |
| Main problem | Low spectral detail | High dimensionality | Irregular, unstructured points |
Key points.
- Hyperspectral data forms a 3D cube (rows, columns, bands); each pixel holds a full spectrum, so materials such as minerals, crop types and water quality are told apart by their spectral signature.
- The large number of correlated bands causes the Hughes effect (curse of dimensionality): with few training samples, accuracy falls as bands increase, so dimensionality is reduced by PCA, MNF or band selection.
- Spectral unmixing splits a mixed pixel into pure endmembers and their abundance fractions.
- Spectral Angle Mapper matches a pixel spectrum $t$ to a reference $r$ using $\theta = \cos^{-1}\dfrac{t \cdot r}{\lVert t \rVert \lVert r \rVert}$; a small angle means a match, and it is insensitive to brightness.
- Deep models used are 1D CNN on spectra, 3D CNN on the cube, and hybrid spectral-spatial networks or transformers.
- LiDAR range is $R = \dfrac{c\,t}{2}$, where $t$ is the two-way travel time of the pulse and $c = 3 \times 10^8$ m/s; division by 2 is because the pulse goes to the target and back.
- LiDAR gives a DSM (surface including trees and buildings) and a DTM (bare ground), so canopy or building height is $CHM = DSM - DTM$.
- Point clouds are classified into ground, vegetation and buildings using deep networks such as PointNet or PointNet++, which work directly on unordered points.
Example. A LiDAR pulse returns after $t = 6.67 \times 10^{-6}$ s, so $R = \dfrac{3 \times 10^8 \times 6.67 \times 10^{-6}}{2} \approx 1000$ m.
Answer frame. Open with both definitions; draw or write the comparison table; then develop hyperspectral points 1-5 (cube, dimensionality, unmixing, SAM, deep models) and LiDAR points 6-8 (range formula, DSM and DTM, point-cloud networks); close with applications such as mineral mapping and forest height.
Pitfall: Forgetting the factor of 2 in the LiDAR range formula.
Fusion of multi-modal remote sensing data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Multi-modal data fusion combines data from different sensors, such as optical, SAR, hyperspectral and LiDAR, so that the result is more accurate and informative than any single source.</mark>
Diagram.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-03" viewBox="0 0 442 130" width="442" height="130" role="img" aria-label="tree diagram"><style>#dsfig-u4-03 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-03 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-03 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-03 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-03 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-03 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-03 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-03 .t{fill:#16181D;font-weight:500}#dsfig-u4-03 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-03 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-03 .dot{fill:#16181D}#dsfig-u4-03 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-03 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-03 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-03 .ah{fill:#454C5A}#dsfig-u4-03 .ah.hi{fill:#2340B8}#dsfig-u4-03 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-03 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-03 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-03 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-03 .e{stroke:#B1B7C3}html.dark #dsfig-u4-03 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-03 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-03 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-03 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-03 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-03 .t{fill:#E6E8ED}html.dark #dsfig-u4-03 .t.inv{fill:#0F1115}html.dark #dsfig-u4-03 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-03 .dot{fill:#E6E8ED}html.dark #dsfig-u4-03 .ann{fill:#8FA3FF}html.dark #dsfig-u4-03 .lbl{fill:#858D9C}html.dark #dsfig-u4-03 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-03 .ah{fill:#B1B7C3}html.dark #dsfig-u4-03 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-03 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-03 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-03 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah12" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh12" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="203" y1="37" x2="67" y2="101"/><line class="e" x1="203" y1="37" x2="197" y2="101"/><line class="e" x1="203" y1="37" x2="339" y2="101"/><rect class="n" x="142" y="22" width="122" height="30" rx="8"/><text class="t" x="203" y="37" dy=".35em" text-anchor="middle">Fusion levels</text><rect class="n" x="14" y="86" width="106" height="30" rx="8"/><text class="t" x="67" y="101" dy=".35em" text-anchor="middle">Pixel level</text><rect class="n" x="136" y="86" width="122" height="30" rx="8"/><text class="t" x="197" y="101" dy=".35em" text-anchor="middle">Feature level</text><rect class="n" x="274" y="86" width="130" height="30" rx="8"/><text class="t" x="339" y="101" dy=".35em" text-anchor="middle">Decision level</text></svg></figure>
Key points.
- Sensors complement each other: optical imagery gives spectral detail but fails under cloud, SAR sees through cloud and works day and night, hyperspectral gives material identity and LiDAR adds height.
- Pixel-level fusion merges raw data, for example pan-sharpening, where a high-resolution panchromatic band sharpens a lower-resolution multispectral image using IHS, Brovey or PCA methods.
- Feature-level fusion extracts features (texture, spectral indices, height) from each source and concatenates them into one vector before a single classifier.
- Decision-level fusion classifies each source separately and combines the outputs by majority voting or weighted averaging.
- Deep learning uses two-branch networks, one encoder per modality, joined by concatenation or attention, and trained end to end.
- All data must be co-registered to the same grid and resampled to a common resolution before fusion, or the fused product is misleading.
- Typical uses are land cover mapping with optical plus SAR, flood mapping under cloud, and urban or forest mapping with hyperspectral plus LiDAR.
- Challenges are different resolutions, different noise (speckle in SAR), missing modalities and the cost of labelled data.
Answer frame. Open with the definition and why one sensor is not enough; draw the three-level diagram; then develop points 2-4 (the levels, with an example each), 5 (deep fusion), 6 (co-registration); close with one application, for example flood mapping using SAR plus optical.
Pitfall: Confusing pixel-level (raw data merged) with decision-level (final labels merged) fusion.
Last-minute revision
- Object detection = bounding box plus class label; tracking = linking the same object across frames.
- $IoU$ = overlap area divided by union area; a detection is correct when IoU exceeds 0.5.
- Detectors: two-stage Faster R-CNN (accurate), one-stage YOLO and SSD (fast).
- Change detection needs co-registered images of the same area at two or more dates.
- Differencing $D = X_2 - X_1$ then threshold; ratioing $X_2/X_1$; post-classification gives a from-to matrix.
- $NDVI = (NIR - Red)/(NIR + Red)$ is the usual time-series index.
- Hyperspectral = hundreds of narrow contiguous bands, a 3D data cube; the Hughes effect is handled by PCA or MNF.
- Spectral Angle Mapper: a small angle means a close match.
- LiDAR range $R = ct/2$; $CHM = DSM - DTM$.
- Fusion levels: pixel, feature, decision; pan-sharpening is pixel-level.
Memory hooks
- Detection: "box, label, follow" (localise, classify, track).
- Change: "two dates, align, subtract, threshold".
- LiDAR: "light echo, half the trip" (divide by 2).
- Fusion levels: "Pixel, Feature, Decision" = P-F-D, raw data to final labels.
- Hughes effect: "too many bands, too few samples".
Coverage checklist
- Object detection and tracking in remote sensing imagery: definition, pipeline diagram, detectors, IoU, tracking; no past questions.
- Change detection and time-series analysis: definition, workflow diagram, differencing, post-classification, Siamese networks, NDVI series; no past questions.
- Hyperspectral and LiDAR data analysis: definition, comparison table, Hughes effect, SAM, range formula, DSM and DTM; no past questions.
- Fusion of multi-modal remote sensing data: definition, levels diagram, pan-sharpening, deep fusion, co-registration; no past questions.