How unit 5 is examined
This unit covers image-based rendering (view interpolation, layered depth images, light fields, environment mattes, video-based rendering) and recognition (detection, faces, instances, categories, scenes, datasets); the marks sit in view interpolation vs layered depth images, video-based rendering, category recognition, recognition databases and IBR for VR.
View interpolation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Image-based rendering (IBR) synthesises new views of a scene directly from captured photographs instead of from an explicit 3D model, and view interpolation is the IBR method that creates an in-between view by warping and blending nearby reference images using their depth or correspondences.</mark>
Key points.
- Each reference image carries per-pixel depth or optical-flow correspondences, and every pixel is moved (forward-warped) to its position in the new view.
- The warped images are blended with weights that favour the nearer camera, so the view morphs smoothly from one reference to the other.
- Holes appear where the new camera sees regions that were hidden in the reference, and they are filled from the other image or by inpainting.
- For VR, the steps are capture (many photos or a 360-degree rig), reconstruction (depth or correspondences), and rendering (warp and blend to the head pose), which gives photorealism without modelling and lets the user look around with realism and immersion.
| Basis | View interpolation | Layered depth image (LDI) |
|---|---|---|
| Stores | Several separate reference images with depth | One image whose pixels hold many depth layers |
| Rendering quality | Good between views, blending can blur | Good, one consistent warp |
| Storage | Grows with number of images | Compact, one view-centred structure |
| Occlusion handling | Holes filled from another image | Hidden surfaces already stored in layers |
| Use | Walkthroughs, morphing between views | Fast novel views for IBR and VR |
Answer frame. Open with the definitions of both; draw the comparison table above; close by saying LDI trades storage and hole-free rendering against the flexibility of many images. For the VR question, open with the IBR definition, then capture, reconstruction, rendering, and close with realism and immersion.
Asked: [7 marks] (Dec 2024) Compare and contrast view interpolation and layered depth images. Asked: [7 marks] (Dec 2024) How can image-based rendering be applied in virtual reality?
Layered depth images
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A layered depth image (LDI) is a single view-centred image in which each pixel stores a list of depth samples, one for every surface the ray through that pixel hits.</mark>
Key points.
- Each sample holds colour and depth, and the samples are sorted front to back along the ray.
- Hidden surfaces behind the first one are kept, so occlusion holes are reduced when the viewpoint moves.
- It is rendered by warping the samples to the new view in a fixed occlusion-compatible order, so no depth sort is needed.
Light fields and Lumigraphs
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A light field (Lumigraph) is a 4D function $L(u,v,s,t)$ of the rays crossing two parallel planes, from which any new view is made by sampling the rays, without any geometry.</mark>
Key points.
- A ray is fixed by its point $(u,v)$ on the camera plane and $(s,t)$ on the image plane, giving four parameters.
- A novel view is rendered by looking up and interpolating the stored rays, so rendering cost does not depend on scene complexity.
- The Lumigraph adds a rough geometric model to correct the ray lookup, which needs fewer images.
- Many images are needed for a dense capture, so storage is large.
Environment mattes
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>An environment matte records how a transparent or reflective object bends and reflects the background, so it can be composited over a new background with correct refraction and reflection.</mark>
Key points.
- The object is photographed against known structured backdrops, and each pixel is modelled as a foreground colour $F$ plus a weighted region of the background.
- The compositing model is $C = F + (1-\alpha)B$, extended so $B$ is a sampled area of the environment rather than one point.
- It is used for glass, water and shiny objects that ordinary alpha mattes cannot handle.
Video-based rendering
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Video-based rendering (VBR) creates new, controllable views or animations from captured video streams, so dynamic scenes can be replayed from viewpoints that were never filmed.</mark>
Key points.
- The pipeline is: capture with synchronised multiple cameras, estimate depth or flow per frame, then synthesise the new view by warping and blending, frame by frame.
- Variants include video textures (looping a clip), 3D video from a camera array, and free-viewpoint video.
- Its significance is for dynamic scenes, telepresence, sports replay and VR, where the scene moves and a static model would fail.
- The main challenge is temporal consistency: depth errors that differ from frame to frame cause flicker, and large data volume needs compression.
Answer frame. Open with the definition; list capture, depth estimation, view synthesis; then significance; close with temporal consistency.
Asked: [7 marks] (Dec 2024) Explain the process of video-based rendering and its significance.
Object detection
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Object detection finds every object of interest in an image and outputs, for each, a bounding box and a class label.</mark>
Key points.
- Two-stage detectors (R-CNN family) first propose regions and then classify them, and they are accurate but slower.
- One-stage detectors (YOLO, SSD) predict boxes and classes in one pass, and they are fast and real-time.
- Overlap is measured by $\text{IoU} = \dfrac{\text{area of overlap}}{\text{area of union}}$, and non-maximum suppression removes duplicate boxes.
Face recognition
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Face recognition identifies or verifies a person from a face image, after face detection has located the face.</mark>
Key points.
- The steps are detection, alignment, feature extraction and matching.
- Eigenfaces (PCA) and Fisherfaces (LDA) are classic holistic methods.
- Modern systems use deep CNN embeddings compared by distance, and pose, lighting and expression are the main difficulties.
Instance recognition
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Instance recognition identifies a specific known object (this building, this book) in a new image, despite changes in viewpoint, scale and lighting.</mark>
Key points.
- Local features such as SIFT are detected and matched against a database of stored objects.
- Speed comes from indexing features with a vocabulary tree or inverted index.
- Geometric verification (RANSAC with a homography) rejects wrong matches.
Category recognition
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Category recognition assigns an image or object to a general class (car, dog) rather than a specific instance, so it must generalise over large within-class variation.</mark>
Key points.
- Template matching compares the image with a stored pattern; it is simple but fragile to pose and appearance change.
- Bag-of-words builds a histogram of quantised local features (visual words) and classifies it with an SVM, ignoring spatial layout.
- CNNs learn features end to end, and give the best accuracy given enough data and compute.
- Transformers (ViT) use self-attention over patches, scale well and generalise strongly, but need very large data.
- Applications are object detection, image retrieval and robotics; evaluation uses accuracy, precision, recall and mean average precision on datasets such as ImageNet.
Answer frame. Open with the definition; discuss the four techniques in order of history, comparing accuracy, scalability and generalisation; close with applications and metrics.
Asked: [7 marks] (Dec 2024) Analyze the techniques for category recognition and their applications.
Context and scene understanding
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Scene understanding interprets the whole image, using context (surroundings, co-occurring objects and layout) to decide what each part is.</mark>
Key points.
- Context helps: a keyboard is more likely near a monitor than in a forest.
- Scene parsing labels every pixel (semantic segmentation).
- Scene classification and object relationships give the global scene type.
Recognition databases and test sets
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Recognition databases are large, human-annotated image collections used to train, benchmark and compare vision algorithms.</mark>
Key points.
- ImageNet (millions of labelled images) made deep learning possible, and COCO adds boxes, masks and captions for detection.
- Training on large data lets models learn general features, and pretraining allows transfer learning to small tasks.
- Fixed test sets and metrics give fair benchmarking of different methods, so progress can be measured.
- Diverse data improves generalisation, and auditing datasets helps reduce bias, though annotation bias can persist.
Answer frame. Open with the definition; name ImageNet, COCO; develop training, benchmarking, transfer learning; close with evaluation and bias.
Asked: [7 marks] (Dec 2024) How do recognition databases aid in improving computer vision algorithms?
Last-minute revision
- IBR makes new views from photos without a full 3D model.
- View interpolation warps and blends reference images; holes come from occlusion.
- LDI stores many depth samples per pixel, so hidden surfaces are kept.
- Light field is a 4D ray function $L(u,v,s,t)$.
- Environment matte models refraction and reflection of a background.
- VBR pipeline: capture, depth estimation, view synthesis; challenge is temporal consistency.
- IoU is overlap area divided by union area.
- Instance recognition uses SIFT, indexing and RANSAC.
- Category recognition: template, bag-of-words, CNN, transformer.
- ImageNet and COCO drive training, benchmarking and transfer learning.
Memory hooks
- LDI: "Layers Of Depth per pixel".
- Light field: two planes, four numbers.
- VBR pipeline: "Capture, Depth, Synthesize".
- Category history: Template, Bag, CNN, Transformer.
Coverage checklist
- View interpolation: Dec 2024 compare with LDI, IBR in VR.
- Layered depth images: definition and key points.
- Light fields and Lumigraphs: definition and key points.
- Environment mattes: definition and key points.
- Video-based rendering: Dec 2024 process and significance.
- Object detection: definition and key points.
- Face recognition: definition and key points.
- Instance recognition: definition and key points.
- Category recognition: Dec 2024 techniques and applications.
- Context and scene understanding: definition and key points.
- Recognition databases and test sets: Dec 2024 role of databases.