Skip to content
AD-701 · AI for Computer Vision/Quick Revision Short Notes

AI for Computer Vision (AD-701) - Unit 5 Short Notes

How unit 5 is examined

This unit covers image-based rendering (view interpolation, layered depth images, light fields, environment mattes, video-based rendering) and recognition (detection, faces, instances, categories, scenes, datasets); the marks sit in view interpolation vs layered depth images, video-based rendering, category recognition, recognition databases and IBR for VR.

View interpolation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Image-based rendering (IBR) synthesises new views of a scene directly from captured photographs instead of from an explicit 3D model, and view interpolation is the IBR method that creates an in-between view by warping and blending nearby reference images using their depth or correspondences.</mark>

Key points.

  1. Each reference image carries per-pixel depth or optical-flow correspondences, and every pixel is moved (forward-warped) to its position in the new view.
  2. The warped images are blended with weights that favour the nearer camera, so the view morphs smoothly from one reference to the other.
  3. Holes appear where the new camera sees regions that were hidden in the reference, and they are filled from the other image or by inpainting.
  4. For VR, the steps are capture (many photos or a 360-degree rig), reconstruction (depth or correspondences), and rendering (warp and blend to the head pose), which gives photorealism without modelling and lets the user look around with realism and immersion.
Basis View interpolation Layered depth image (LDI)
Stores Several separate reference images with depth One image whose pixels hold many depth layers
Rendering quality Good between views, blending can blur Good, one consistent warp
Storage Grows with number of images Compact, one view-centred structure
Occlusion handling Holes filled from another image Hidden surfaces already stored in layers
Use Walkthroughs, morphing between views Fast novel views for IBR and VR

Answer frame. Open with the definitions of both; draw the comparison table above; close by saying LDI trades storage and hole-free rendering against the flexibility of many images. For the VR question, open with the IBR definition, then capture, reconstruction, rendering, and close with realism and immersion.

Asked: [7 marks] (Dec 2024) Compare and contrast view interpolation and layered depth images. Asked: [7 marks] (Dec 2024) How can image-based rendering be applied in virtual reality?

Layered depth images

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A layered depth image (LDI) is a single view-centred image in which each pixel stores a list of depth samples, one for every surface the ray through that pixel hits.</mark>

Key points.

  1. Each sample holds colour and depth, and the samples are sorted front to back along the ray.
  2. Hidden surfaces behind the first one are kept, so occlusion holes are reduced when the viewpoint moves.
  3. It is rendered by warping the samples to the new view in a fixed occlusion-compatible order, so no depth sort is needed.

Light fields and Lumigraphs

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A light field (Lumigraph) is a 4D function $L(u,v,s,t)$ of the rays crossing two parallel planes, from which any new view is made by sampling the rays, without any geometry.</mark>

Key points.

  1. A ray is fixed by its point $(u,v)$ on the camera plane and $(s,t)$ on the image plane, giving four parameters.
  2. A novel view is rendered by looking up and interpolating the stored rays, so rendering cost does not depend on scene complexity.
  3. The Lumigraph adds a rough geometric model to correct the ray lookup, which needs fewer images.
  4. Many images are needed for a dense capture, so storage is large.

Environment mattes

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>An environment matte records how a transparent or reflective object bends and reflects the background, so it can be composited over a new background with correct refraction and reflection.</mark>

Key points.

  1. The object is photographed against known structured backdrops, and each pixel is modelled as a foreground colour $F$ plus a weighted region of the background.
  2. The compositing model is $C = F + (1-\alpha)B$, extended so $B$ is a sampled area of the environment rather than one point.
  3. It is used for glass, water and shiny objects that ordinary alpha mattes cannot handle.

Video-based rendering

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Video-based rendering (VBR) creates new, controllable views or animations from captured video streams, so dynamic scenes can be replayed from viewpoints that were never filmed.</mark>

Key points.

  1. The pipeline is: capture with synchronised multiple cameras, estimate depth or flow per frame, then synthesise the new view by warping and blending, frame by frame.
  2. Variants include video textures (looping a clip), 3D video from a camera array, and free-viewpoint video.
  3. Its significance is for dynamic scenes, telepresence, sports replay and VR, where the scene moves and a static model would fail.
  4. The main challenge is temporal consistency: depth errors that differ from frame to frame cause flicker, and large data volume needs compression.

Answer frame. Open with the definition; list capture, depth estimation, view synthesis; then significance; close with temporal consistency.

Asked: [7 marks] (Dec 2024) Explain the process of video-based rendering and its significance.

Object detection

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Object detection finds every object of interest in an image and outputs, for each, a bounding box and a class label.</mark>

Key points.

  1. Two-stage detectors (R-CNN family) first propose regions and then classify them, and they are accurate but slower.
  2. One-stage detectors (YOLO, SSD) predict boxes and classes in one pass, and they are fast and real-time.
  3. Overlap is measured by $\text{IoU} = \dfrac{\text{area of overlap}}{\text{area of union}}$, and non-maximum suppression removes duplicate boxes.

Face recognition

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Face recognition identifies or verifies a person from a face image, after face detection has located the face.</mark>

Key points.

  1. The steps are detection, alignment, feature extraction and matching.
  2. Eigenfaces (PCA) and Fisherfaces (LDA) are classic holistic methods.
  3. Modern systems use deep CNN embeddings compared by distance, and pose, lighting and expression are the main difficulties.

Instance recognition

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Instance recognition identifies a specific known object (this building, this book) in a new image, despite changes in viewpoint, scale and lighting.</mark>

Key points.

  1. Local features such as SIFT are detected and matched against a database of stored objects.
  2. Speed comes from indexing features with a vocabulary tree or inverted index.
  3. Geometric verification (RANSAC with a homography) rejects wrong matches.

Category recognition

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Category recognition assigns an image or object to a general class (car, dog) rather than a specific instance, so it must generalise over large within-class variation.</mark>

Key points.

  1. Template matching compares the image with a stored pattern; it is simple but fragile to pose and appearance change.
  2. Bag-of-words builds a histogram of quantised local features (visual words) and classifies it with an SVM, ignoring spatial layout.
  3. CNNs learn features end to end, and give the best accuracy given enough data and compute.
  4. Transformers (ViT) use self-attention over patches, scale well and generalise strongly, but need very large data.
  5. Applications are object detection, image retrieval and robotics; evaluation uses accuracy, precision, recall and mean average precision on datasets such as ImageNet.

Answer frame. Open with the definition; discuss the four techniques in order of history, comparing accuracy, scalability and generalisation; close with applications and metrics.

Asked: [7 marks] (Dec 2024) Analyze the techniques for category recognition and their applications.

Context and scene understanding

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Scene understanding interprets the whole image, using context (surroundings, co-occurring objects and layout) to decide what each part is.</mark>

Key points.

  1. Context helps: a keyboard is more likely near a monitor than in a forest.
  2. Scene parsing labels every pixel (semantic segmentation).
  3. Scene classification and object relationships give the global scene type.

Recognition databases and test sets

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Recognition databases are large, human-annotated image collections used to train, benchmark and compare vision algorithms.</mark>

Key points.

  1. ImageNet (millions of labelled images) made deep learning possible, and COCO adds boxes, masks and captions for detection.
  2. Training on large data lets models learn general features, and pretraining allows transfer learning to small tasks.
  3. Fixed test sets and metrics give fair benchmarking of different methods, so progress can be measured.
  4. Diverse data improves generalisation, and auditing datasets helps reduce bias, though annotation bias can persist.

Answer frame. Open with the definition; name ImageNet, COCO; develop training, benchmarking, transfer learning; close with evaluation and bias.

Asked: [7 marks] (Dec 2024) How do recognition databases aid in improving computer vision algorithms?

Last-minute revision

  • IBR makes new views from photos without a full 3D model.
  • View interpolation warps and blends reference images; holes come from occlusion.
  • LDI stores many depth samples per pixel, so hidden surfaces are kept.
  • Light field is a 4D ray function $L(u,v,s,t)$.
  • Environment matte models refraction and reflection of a background.
  • VBR pipeline: capture, depth estimation, view synthesis; challenge is temporal consistency.
  • IoU is overlap area divided by union area.
  • Instance recognition uses SIFT, indexing and RANSAC.
  • Category recognition: template, bag-of-words, CNN, transformer.
  • ImageNet and COCO drive training, benchmarking and transfer learning.

Memory hooks

  • LDI: "Layers Of Depth per pixel".
  • Light field: two planes, four numbers.
  • VBR pipeline: "Capture, Depth, Synthesize".
  • Category history: Template, Bag, CNN, Transformer.

Coverage checklist

  • View interpolation: Dec 2024 compare with LDI, IBR in VR.
  • Layered depth images: definition and key points.
  • Light fields and Lumigraphs: definition and key points.
  • Environment mattes: definition and key points.
  • Video-based rendering: Dec 2024 process and significance.
  • Object detection: definition and key points.
  • Face recognition: definition and key points.
  • Instance recognition: definition and key points.
  • Category recognition: Dec 2024 techniques and applications.
  • Context and scene understanding: definition and key points.
  • Recognition databases and test sets: Dec 2024 role of databases.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in