How unit 4 is examined
This unit covers how depth and shape are recovered from images or range sensors and then stored as meshes, point clouds, voxels, fitted models and textured surfaces. No topic has been asked recently, so learn the definitions, the formulas and the key points of each.
Shape from X
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Shape from X is the family of methods that recover 3D shape from a single cue in the image, such as shading, stereo, texture, focus or silhouettes.</mark>
Key points.
- Shape from shading uses the brightness variation of a surface under known lighting, and for a diffuse surface the intensity is $I = \rho\,(\mathbf{n}\cdot\mathbf{l})$, so the normals, and then the depth, are estimated from $I$.
- Shape from stereo uses the disparity $d$ between two views of the same point, giving depth $Z = fB/d$, where $f$ is the focal length and $B$ the baseline between the cameras.
- Shape from texture uses the way a repeated pattern is foreshortened and stretched on a slanted surface, and shape from focus picks, for each pixel, the focus setting at which it is sharpest.
- Shape from silhouettes intersects the viewing cones formed by the object outline in many views, giving the visual hull of the object.
- Each cue is ambiguous alone, so practical systems combine several cues or add smoothness constraints.
Active range finding
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Active range finding measures depth directly by sending out its own signal, such as a laser beam or a light pattern, and analysing what comes back.</mark>
Key points.
- Time of flight sends a pulse and measures the round-trip time $t$, so the distance is $d = ct/2$, where $c$ is the speed of light.
- LiDAR sweeps a laser over the scene and returns a dense set of accurate distance points, and it is used in mapping and self-driving cars.
- Structured light projects a known stripe or grid pattern, and depth is found by triangulation from how the pattern is deformed by the surface.
- Phase-shift and amplitude-modulated devices measure the phase change of a modulated beam instead of timing a pulse.
- Active methods work on textureless surfaces where stereo fails, but they struggle in bright sunlight, on shiny or transparent objects, and at long range.
Surface representations
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A surface representation describes a 3D object by its boundary surface only, most commonly as a polygon mesh of vertices, edges and faces.</mark>
Key points.
- A triangle mesh stores a list of vertices $(x, y, z)$ and a list of faces, each naming three vertex indices, so shared vertices are stored once.
- Parametric surfaces such as splines, Bezier patches and NURBS describe smooth surfaces with a few control points.
- Implicit surfaces define the shape as the zero set of a function, $f(x, y, z) = 0$.
- A mesh can be built from range data by connecting neighbouring points into triangles, and then simplified to reduce the number of faces.
- Meshes are compact and are the standard input for graphics rendering, but they are hard to change in topology.
Point-based representations
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A point-based representation models a surface as an unordered set of 3D points, called a point cloud, with no connectivity between them.</mark>
Key points.
- Each point stores its position $(x, y, z)$ and optionally colour and a surface normal.
- Point clouds come directly from LiDAR, depth cameras, stereo matching and structure from motion.
- A surfel, or surface element, is a point with a normal and a radius, so it acts as a small oriented disc and a dense set of surfels looks like a continuous surface.
- Points are simple to store, merge and render, and they need no topology, but they need an extra step such as surface fitting to become a mesh.
- Raw clouds are noisy and uneven in density, so filtering and normal estimation are usually applied first.
Volumetric representations
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A volumetric representation divides space into a regular 3D grid of voxels and stores in each voxel whether it is occupied or its signed distance to the surface.</mark>
Key points.
- A voxel is the 3D counterpart of a pixel, a small cube of space.
- An occupancy grid marks every voxel as empty, occupied or unknown, and it is widely used in robot mapping.
- A signed distance function stores the distance to the nearest surface, negative inside and positive outside, so the surface is where the value crosses zero.
- Memory grows as $n^3$ for an $n \times n \times n$ grid, and octrees cut this by subdividing space finely only near the surface.
- The surface is extracted from the volume with marching cubes, which places triangles in each voxel the surface crosses.
Example. A $256^3$ grid has about 16.8 million voxels, which shows why octrees are used.
Model-based reconstruction
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Model-based reconstruction recovers 3D shape by fitting a known parametric model, such as a face or a body model, to the image or range data.</mark>
Key points.
- A prior model supplies knowledge of the object class, which resolves ambiguity when the data is noisy, partial or from only one image.
- Fitting adjusts the model parameters, such as shape, pose and scale, to minimise the error between the model projection and the image features.
- A morphable model describes shape as a mean shape plus a linear combination of basis shapes, $S = \bar{S} + \sum_i \alpha_i S_i$.
- Articulated models add joints, so a human body or hand can be fitted with pose parameters.
- The result is complete and smooth even from few views, but it cannot represent shapes outside the range of the model.
Recovering texture maps
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A texture map is an image draped over a 3D surface through UV coordinates, and recovering it means computing that image from the input photographs.</mark>
Key points.
- Each mesh vertex is given a 2D coordinate $(u, v)$ that points into the texture image, which is called UV mapping.
- The colour for each surface point is obtained by projecting it into the calibrated input photographs.
- When several photographs see the same point, their colours are blended, with weights that favour views facing the surface directly.
- Differences in exposure and small misalignments cause visible seams, which are reduced by blending and by correcting brightness.
- A texture map adds realistic detail without adding geometry, so a coarse mesh can still look rich.
Albedos
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Albedo is the fraction of incident light that a surface reflects, an intrinsic property that does not depend on the lighting.</mark>
Key points.
- For a diffuse Lambertian surface, $I = \rho\,(\mathbf{n}\cdot\mathbf{l})$, where $\rho$ is the albedo, $\mathbf{n}$ the surface normal and $\mathbf{l}$ the light direction.
- Albedo lies between 0 for a surface that absorbs all light and 1 for one that reflects all light.
- The BRDF gives the reflected light as a function of the incoming and outgoing directions, and albedo is its simplest, direction-independent form.
- Recovering albedo separates the true surface colour from shading and shadows, so the model can be relit under new lighting.
- It is an ill-posed problem from one image, so methods use several lighting conditions or priors such as smooth shading.
Last-minute revision
- Shape from X recovers 3D from shading, stereo, texture, focus or silhouettes.
- Stereo depth is $Z = fB/d$, so larger disparity means a nearer point.
- Time of flight distance is $d = ct/2$.
- Structured light projects a known pattern and triangulates depth from its deformation.
- LiDAR sweeps a laser to give a dense point cloud.
- A mesh stores vertices, edges and faces; a point cloud stores only points.
- A surfel is a point with a normal and radius, acting as an oriented disc.
- A voxel is a 3D pixel; grid memory grows as $n^3$ and octrees save space.
- Marching cubes extracts a mesh from a volume.
- Lambertian shading is $I = \rho\,(\mathbf{n}\cdot\mathbf{l})$ and albedo lies from 0 to 1.
- A texture map uses $(u, v)$ coordinates to place a photo on the surface.
Memory hooks
- Passive methods use the scene as it is, active methods send their own signal.
- Mesh means faces, cloud means dots, voxel means cubes.
- Model-based means the prior fills what the data misses.
- Albedo is the true colour with the light taken away.
- UV is the address of a surface point in the texture image.
Coverage checklist
- Shape from X: covered, no past questions.
- Active range finding: covered, no past questions.
- Surface representations: covered, no past questions.
- Point-based representations: covered, no past questions.
- Volumetric representations: covered, no past questions.
- Model-based reconstruction: covered, no past questions.
- Recovering texture maps: covered, no past questions.
- Albedos: covered, no past questions.