UNIT 4: Computer Graphics & Multimedia
I. Display Systems & Output Devices
Raster Scan Display
-
Architecture: Displays images as a grid of pixels (raster). An electron beam sweeps across the screen row by row (left to right, top to bottom), illuminating phosphors to create the image.
-
Block Diagram:
-
Central Processing Unit (CPU): Controls overall system.
-
Display Controller / Raster-Scan Processor: Fetches pixel data from Frame Buffer (a dedicated memory area storing pixel intensities/colors) and converts it to analog signals for the CRT.
-
Display Buffer Memory (Frame Buffer): Stores the entire image as a 2D array of pixel values. Size =
resolution_x * resolution_y * bits_per_pixel. -
Video Amplifier: Amplifies the analog signal.
-
CRT: The output device.
-
-
Key Feature: Stores the complete image definition in memory. Refreshing is automatic from the frame buffer.
Random Scan Display (Vector Display)
-
Architecture: Electron beam directly draws lines (vectors) by deflecting to specific coordinate points. Only the lines of the image are drawn, not the entire screen.
-
Display List/File: The image is stored as a set of line-drawing commands (coordinates, intensity) in a display file. The display processor reads this list and generates deflection signals to draw each vector.
-
Comparison with Raster Scan:
| Feature | Raster Scan | Random Scan |
|---|---|---|
| Image Storage | Full frame buffer (pixel-based) | Display list (vector-based) |
| Beam Movement | Fixed, systematic scan (raster) | Direct, arbitrary deflection |
| Refresh | Constant (60-80 Hz) from memory | Only redraws vectors (can flicker if complex) |
| Realism | Better for complex, shaded scenes | Better for line drawings, CAD |
| Cost/Memory | High memory cost | Lower memory, higher cost for deflection circuitry |
| Artifacts | Aliasing (jaggies) | No aliasing on lines, but limited resolution |
Direct View Storage Tube (DVST)
-
Working: Uses a special CRT with a storage mesh (or grid) behind the phosphor coating.
-
Write Cycle: A high-energy "writing" electron beam charges the storage mesh at specific points, creating a latent electrostatic image.
-
Viewing Cycle: A low-energy "viewing" beam scans the entire screen. It is modulated by the charged areas on the storage mesh, causing the phosphor to glow only where the mesh is charged. The stored image persists without refreshing.
-
-
Advantages:
-
No need for a large frame buffer (image stored in the tube).
-
No flicker (image is continuously visible).
-
High resolution, no aliasing.
-
-
Disadvantages:
-
Cannot modify part of the image without redrawing the entire screen (erases old image).
-
Cannot display different intensity levels (only on/off).
-
Cannot display color easily.
-
Cannot display dynamic/animating scenes effectively.
-
Cathode Ray Tube (CRT) - Basic Principle & Structure
-
Principle: An electron gun emits a focused beam of electrons. This beam is accelerated and deflected by magnetic or electric fields to strike specific phosphor-coated points on the screen, causing them to emit light.
-
Structure:
-
Electron Gun: Cathode (electron emitter), control grid (intensity control), focusing system, accelerating anode.
-
Deflection System: Magnetic (yoke coils) or Electrostatic (plates) to steer the beam horizontally and vertically.
-
Phosphor Coated Screen: Emits light (photons) when struck by electrons. Different phosphors have different colors and persistence (glow duration).
-
Aquadag Coating: Conductive coating on the inside of the screen to collect secondary electrons and prevent charge buildup.
-
II. Interactive Input Devices & Techniques
Interactive Input Devices (Types)
-
Keyboard: Text and command entry.
-
Mouse: 2D positional input (relative movement). Buttons for selection.
-
Trackball: Stationary mouse; ball is rolled with fingers/hand.
-
Joystick: 2D positional input with force feedback; often used for games.
-
Digitizer/Graphics Tablet: Absolute positional input. Stylus puck moves over a flat surface to input coordinates. Used for drawing, CAD.
-
Touch Panel: Detects touch on screen (resistive, capacitive, infrared). Provides direct screen interaction.
-
Light Pen: Photoelectric cell in a pen-shaped device. Detects electron beam's position when pointed at screen. Requires high-intensity CRT.
-
Data Glove: Wearable glove with sensors (flex sensors, trackers) to capture hand/finger position and orientation for VR/3D manipulation.
-
Motion Capture (Mocap): Systems (optical, magnetic, inertial) to record movement of objects or actors for animation.
Rubber Band Techniques
-
Concept: A visual feedback technique where a temporary line (the "rubber band") is dynamically drawn between the current cursor position and a fixed anchor point (e.g., where a drag operation started) as the user moves the mouse.
-
Application: Used for interactive drawing of lines, rectangles, circles, or for selecting a region. The user sees the shape's size/orientation change in real-time before releasing the mouse button to finalize it.
-
Example: To draw a rectangle, user clicks at one corner (anchor), drags mouse. A rectangle outline stretches from anchor to current cursor position, like a stretched rubber band.
Position Input Devices
-
Definition: Devices that provide absolute (x, y) coordinates.
-
Examples: Digitizer/Graphics Tablet, Touch Screen, Light Pen.
-
Key: Directly specify a point in a coordinate system.
Motion Input Devices
-
Definition: Devices that capture movement, velocity, or orientation over time.
-
Examples: Data Glove (hand/finger motion), Motion Capture Suit (full body), Joystick (force/direction), Trackball/Mouse (relative motion).
-
Key: Provide continuous stream of positional/directional data for animation or control.
III. Geometric Primitive Generation Algorithms
DDA (Digital Differential Analyzer) Line Algorithm
-
Concept: Uses the line equation
y = mx + corx = (y - c)/m. Calculates subsequent points by incrementing either x or y and using the differential equationΔy = mΔxorΔx = Δy/m. -
Steps:
-
Calculate
dx = x2 - x1,dy = y2 - y1. -
Determine steps:
steps = max(|dx|, |dy|). -
Calculate increments:
x_inc = dx/steps,y_inc = dy/steps. -
Initialize
x = x1,y = y1. -
For
i = 0tosteps: Plot(round(x), round(y));x += x_inc;y += y_inc.
-
-
Example: Line from (1,1) to (5,5).
-
dx=4, dy=4, steps=4.
-
x_inc=1, y_inc=1.
-
Points: (1,1), (2,2), (3,3), (4,4), (5,5).
-
-
Disadvantage: Uses floating-point operations and rounding; slower than Bresenham's.
Bresenham's Line Drawing Algorithm
-
Concept: Uses only integer arithmetic (addition, subtraction, bit shifts). Based on the decision parameter
pto choose the next pixel closest to the true line. -
For slope
0 ≤ m ≤ 1:-
dx = x2 - x1,dy = y2 - y1,p0 = 2*dy - dx. -
x = x1,y = y1. -
For
i = 0todx:-
Plot
(x, y). -
If
p_k < 0:p_{k+1} = p_k + 2*dy;x++. -
Else:
p_{k+1} = p_k + 2*(dy - dx);x++;y++.
-
-
-
Example: Line (1,1) to (5,5). dx=4, dy=4.
-
p0 = 2*4 - 4 = 4.
-
k=0: p0=4≥0 -> (2,2), p1=4+2*(4-4)=4.
-
k=1: p1=4≥0 -> (3,3), p2=4+2*(0)=4.
-
... Points: (1,1),(2,2),(3,3),(4,4),(5,5).
-
-
Advantages: Fast (integer ops only), no rounding error accumulation, widely used.
Midpoint Circle Drawing Algorithm
-
Concept: Uses the circle equation
x² + y² = r². At each step, decides between two possible pixels (East or South-East) based on the value of the decision parameterpat the midpoint between them. -
Steps (for radius
r, center (0,0)):-
x = 0,y = r,p = 1 - r. -
Plot 8 symmetric points
(±x, ±y),(±y, ±x). -
While
x < y:-
If
p < 0:p += 2*x + 3;x++. -
Else:
p += 2*(x - y) + 5;x++;y--. -
Plot 8 symmetric points.
-
-
-
Example: Radius 5:
-
(0,5), p=1-5=-4.
-
p<0 -> (1,5), p=-4+2*0+3=-1.
-
p<0 -> (2,5), p=-1+2*1+3=4.
-
p≥0 -> (3,4), p=4+2*(2-5)+5=4-6+5=3.
-
p≥0 -> (4,3), p=3+2*(3-4)+5=3-2+5=6.
-
p≥0 -> (5,3) but x<y? 5<3 false. Stop.
-
Points: (0,5),(1,5),(2,5),(3,4),(4,3),(5,0) and symmetries.
-
Bezier Curves
-
Definition: A parametric curve defined by a set of control points. The curve is contained within the convex hull of these points.
-
Control Points:
P0, P1, ..., Pn.P0andPnare endpoints; others "pull" the curve. -
Equation (Cubic, n=3):
$$B(t) = (1-t)^3 P_0 + 3(1-t)^2 t P_1 + 3(1-t) t^2 P_2 + t^3 P_3, \quad t \in [0,1]$$
-
Midpoint Calculation: For
t=0.5, the midpoint is the average of the control points after one level of linear interpolation (De Casteljau's algorithm). -
Drawing: Use De Casteljau's algorithm or evaluate the Bernstein polynomial at incremental
tvalues.
B-Spline Curves
-
Definition: A piecewise polynomial curve defined by control points and a knot vector. Offers local control (moving one control point affects only a curve segment).
-
Properties:
-
Degree
k: Curve isC^{k-1}continuous (smooth). -
Convex Hull: Curve lies within convex hull of control points.
-
Local Control: Changing a control point affects only
ksegments. -
Affine Invariance: Transformations apply to control points.
-
-
Difference from Bezier:
-
Bezier: Single polynomial segment; global control (all points affect entire curve); endpoints interpolated.
-
B-Spline: Multiple polynomial segments; local control; endpoints generally not interpolated (unless knot vector is clamped).
-
IV. 2D & 3D Transformations
2D Transformations (Matrix Form)
-
Point Representation: Column vector
P = [x, y, 1]^T(using homogeneous coordinates). -
Translation
T(tx, ty):
$$T = \begin{bmatrix} 1 & 0 & t_x \\ 0 & 1 & t_y \\ 0 & 0 & 1 \end{bmatrix}$$
`P' = T * P`
- Rotation
R(θ)about origin:
$$R = \begin{bmatrix} \cos\theta & -\sin\theta & 0 \\ \sin\theta & \cos\theta & 0 \\ 0 & 0 & 1 \end{bmatrix}$$
- Scaling
S(sx, sy)about origin:
$$S = \begin{bmatrix} s_x & 0 & 0 \\ 0 & s_y & 0 \\ 0 & 0 & 1 \end{bmatrix}$$
2D Rotation about an Arbitrary Point P(a,b)
-
Translate point
Pto origin:T(-a, -b). -
Rotate by
θ:R(θ). -
Translate back:
T(a, b).
-
Composite Matrix:
M = T(a,b) * R(θ) * T(-a,-b). -
Example: Rotate triangle A(0,0), B(2,2), C(4,2) about origin and P(-2,-2) by 45°.
-
About origin: Apply
R(45°)to each vertex. -
About P(-2,-2): For each vertex
V, computeV' = M * VwhereM = T(-2,-2)*R(45°)*T(2,2).
-
Viewing Transformation (Window-to-Viewport Mapping)
-
Concept: Transforms a window (world-coordinate area of interest) to a viewport (device-coordinate area on screen).
-
Steps:
-
Translate window so its minimum x,y is at (0,0):
T(-w_xmin, -w_ymin). -
Scale to viewport size:
S( v_xwidth / w_xwidth , v_yheight / w_yheight ). -
Translate to viewport minimum:
T(v_xmin, v_ymin).
-
-
Composite Matrix:
M = T(v_xmin, v_ymin) * S(...) * T(-w_xmin, -w_ymin).
Homogeneous Coordinates
-
Need: To represent translation as a matrix multiplication (affine transformations). Standard 2D/3D coordinates cannot combine translation with rotation/scaling in a single matrix.
-
Use: Represent 2D point
(x,y)as(x, y, 1). 3D point(x,y,z)as(x,y,z,1). Allows all transformations (translation, rotation, scaling, perspective) to be represented as 3x3 or 4x4 matrices. Composition is simple matrix multiplication.
V. Clipping Algorithms
Line Clipping Approaches (General Comparison)
| Approach | Principle | Example Algorithm | Pros | Cons |
|---|---|---|---|---|
| Inside-Outside Test | Trivial accept/reject using region codes. | Cohen-Sutherland | Simple, fast for many lines. | Inefficient for lines crossing window (many intersections). |
| Parametric | Represent line parametrically P(t) = P1 + t(P2-P1). Solve for t where line enters/exits clip region. |
Cyrus-Beck, Liang-Barsky | Efficient, handles all cases uniformly. | More complex math. |
| Area/Subdivision | Recursively subdivide space until segments are entirely inside/outside. | Not common for lines. | Generalizable to polygons. | Overhead for simple lines. |
Cohen-Sutherland Line Clipping Algorithm
-
Region Codes: 4-bit code for each endpoint. Bits:
LEFT(1), RIGHT(2), BOTTOM(4), TOP(8). 9 regions (0=inside). -
Steps:
-
Assign region codes to
P1,P2. -
Trivial Accept: Both codes = 0 -> line entirely inside.
-
Trivial Reject:
code1 & code2 != 0-> line entirely outside. -
Clip: Else, choose an endpoint with non-zero code. Determine which window edge it is outside. Find intersection
Iof line with that edge. Replace the endpoint withI, recalculate its code. Repeat from step 2.
-
-
Example: Window P(0,0), Q(340,0), R(340,340), S(0,340). Line AB[(-170,595),(170,255)].
-
Code(A):
x=-170<0(LEFT=1),y=595>340(TOP=8) -> 1001 (9). -
Code(B):
x=170(in),y=255(in) -> 0000 (0). -
9 & 0 = 0-> Not trivial reject. Clip A with TOP edge (y=340). -
Parametric:
y = 595 + t*(255-595) = 340->t = (340-595)/(-340) = 0.75.x = -170 + 0.75*(170+170) = -170 + 255 = 85. -
New point A'(85,340) code=0. Line A'B is inside? Check B code=0 -> Accept. Visible segment from (85,340) to (170,255).
-
Cyrus-Beck Algorithm (Parametric Line Clipping)
-
Principle: For a convex polygon clip window. Represent line parametrically:
P(t) = P1 + t(P2-P1),t ∈ [0,1]. Findt_enter(max of allt_nfor entering edges) andt_exit(min of allt_efor exiting edges). Visible ift_enter < t_exit. -
Dot Product Use: For each edge
iwith outward normaln_iand a vertexV_ion the edge:-
n_i · (P(t) - V_i) ≤ 0for points inside. -
Substitute
P(t):n_i · (P1 - V_i) + t * n_i · (P2-P1) ≤ 0. -
Let
w = P1 - V_i,d = P2-P1. -
If
n_i · d == 0: line parallel to edge. Ifn_i · w < 0-> outside (reject). -
Else:
t_i = - (n_i · w) / (n_i · d).-
If
n_i · d < 0:t_iis entering (t_enter = max(t_enter, t_i)). -
If
n_i · d > 0:t_iis exiting (t_exit = min(t_exit, t_i)).
-
-
Sutherland-Hodgman Polygon Clipping
-
Concept: Clips a polygon against one edge of the convex clip window at a time (left, right, bottom, top). Output of one stage is input to next.
-
For each clip edge:
-
Process polygon vertices sequentially (S, E = Start, End of current edge).
-
Cases:
-
S in, E in: Output E.
-
S in, E out: Output intersection I.
-
S out, E in: Output I, then E.
-
S out, E out: Output nothing.
-
-
-
Result: Clipped polygon (may be degenerate or empty). Only for convex clip polygons.
Painter's Algorithm
-
Concept: Depth sorting algorithm for visible surface determination. Polygons are sorted by their maximum depth (z-coordinate) and drawn back to front (farthest first). Closer polygons paint over farther ones.
-
Steps:
-
Sort all polygons by their farthest z-value (or centroid depth).
-
Draw polygons in sorted order (from farthest to nearest).
-
-
Limitations:
-
Not general: Fails for cyclic overlaps (e.g., three polygons each behind one other).
-
Computationally expensive: Sorting O(n log n) per frame.
-
Requires splitting: If polygon A overlaps B and B overlaps C but C overlaps A, must split polygons to resolve.
-
Only for convex polygons? Can handle concave but sorting is harder.
-
VI. Projection Techniques
Parallel Projection
-
Principle: Projection lines are parallel to each other and to the projection direction. No vanishing point. Preserves relative lengths and angles (affine transformation).
-
Types:
-
Orthographic: Projection direction is perpendicular to the view plane.
-
Multi-view: Front, top, side (principal projections).
-
Axonometric: Single view with rotated object (isometric, dimetric, trimetric).
-
-
Oblique: Projection direction is not perpendicular to the view plane. One face is shown true shape.
-
Cavalier: Projection lines at 45°, foreshortening factor = 1.
-
Cabinet: Projection lines at 45°, foreshortening factor = 0.5 (more realistic).
-
-
-
Characteristics: No perspective foreshortening; parallel lines remain parallel; good for measurements, engineering drawings.
Perspective Projection
-
Principle: Projection lines converge at a center of projection (COP) or eye point. Simulates human vision. Objects appear smaller with distance.
-
Types (by number of principal axes/vanishing points):
-
1-point: One set of parallel lines is perpendicular to view plane (no vanishing point). Other two sets converge to one vanishing point on horizon.
-
2-point: Two sets of parallel lines converge to two vanishing points. Most common for general objects.
-
3-point: All three sets of parallel lines converge to three vanishing points. Used for dramatic views (looking up/down).
-
-
Characteristics: Foreshortening (objects shrink with distance), vanishing points, non-parallel lines may converge, realistic depth perception.
Differentiate: Parallel vs Perspective Projection
| Feature | Parallel Projection | Perspective Projection |
|---|---|---|
| Projection Lines | Parallel to each other | Converge at COP |
| Vanishing Points | None | 1, 2, or 3 |
| Foreshortening | No (constant scale) | Yes (scale decreases with distance) |
| Parallelism | Preserved (parallel lines stay parallel) | Not preserved (converge) |
| Realism | Low (synthetic look) | High (simulates eye) |
| Use Case | Engineering/Architectural drawings (measurements) | Art, visualization, simulations (realism) |
| Transformation | Affine (linear + translation) | Non-linear (division by z) |
VII. Visible Surface Detection & Rendering
Z-Buffer (Depth Buffer) Technique
-
Algorithm:
-
Initialize Depth Buffer (
Z_buffer[x][y]) withz_min(far plane) for every pixel. -
Initialize Frame Buffer (
Frame_buffer[x][y]) with background color. -
For each polygon in scene:
-
For each pixel
(x,y)inside polygon's projection:-
Compute polygon's depth
zat(x,y). -
If
z < Z_buffer[x][y](closer to viewer):-
Z_buffer[x][y] = z -
Frame_buffer[x][y] = polygon's color at (x,y)
-
-
-
-
-
Use of Buffers:
-
Depth Buffer (Z-Buffer): Stores depth value (z-coordinate) of the closest object found so far for each pixel.
-
Frame Buffer: Stores the final color (RGB) for each pixel to be displayed.
-
-
Advantages: Simple, easy to implement, works for any polygon order.
-
Disadvantages: Requires extra memory for depth buffer, cannot handle transparent objects correctly, no spatial coherence exploitation.
Back-Face Detection Algorithm
-
Concept: For a convex polyhedron, any polygon whose normal vector points away from the viewer (i.e., has a positive dot product with the view direction) is a back face and cannot be visible.
-
**Steps for a polygon with vertices
V0, V1, ...in counter-clockwise order when viewed from outside:-
Compute normal vector
Nusing cross product of two edges:N = (V1 - V0) × (V2 - V0). -
Compute view vector
Vfrom polygon to viewer (eye). For simple cases, if viewer is at(0,0,0)looking down -Z,V = (0,0,-1)or use(V0 - Eye_position). -
Compute dot product
N · V. -
Decision:
-
If
N · V < 0(angle > 90°): Front face (potentially visible). -
If
N · V > 0(angle < 90°): Back face (invisible, discard). -
If
N · V = 0: Edge-on (may be visible depending on fill rules).
-
-
-
Note: Requires consistent vertex ordering (usually CCW for front faces).
Shading Models (Comparison)
| Model | Computation | Smoothness | Speed | Use |
|---|---|---|---|---|
| Flat Shading | Compute one normal per polygon (usually geometric normal). Apply lighting model per polygon. | Faceted appearance (sharp edges between polygons). | Fastest (1 lighting calc per polygon). | Low-poly models, speed-critical. |
| Gouraud Shading | Compute vertex normals (average of adjacent polygon normals). Interpolate intensity (I) across polygon from vertex values. | Smooth intensity variation, but highlights/specular may be distorted. | Moderate (lighting at vertices, linear interpolation). | General smooth surfaces. |
| Phong Shading | Compute vertex normals. Interpolate normals (N) across polygon. Apply lighting model per pixel using interpolated N. | Smooth highlights and specular, most realistic. | Slowest (lighting per pixel). | High-quality rendering, specular highlights. |
Reflection Models
-
Diffuse Reflection (Lambertian):
-
Light is scattered equally in all directions.
-
Intensity:
I_diffuse = k_d * I_light * (N · L) -
k_d: diffuse reflection coefficient (material). -
I_light: intensity of light source. -
N: surface normal. -
L: direction to light source. -
Depends only on angle between N and L.
-
-
Specular Reflection (Phong Model):
-
Mirror-like reflection. Bright highlights.
-
Intensity:
I_specular = k_s * I_light * (R · V)^n -
k_s: specular reflection coefficient. -
R: reflection direction of light vector (R = 2(N·L)N - L). -
V: direction to viewer. -
n: Phong exponent (shininess). Highern= smaller, sharper highlight.
-
Color Models (Brief Overview)
-
RGB (Red, Green, Blue): Additive model. Primary colors of light. Used in monitors, cameras.
(0,0,0)=Black,(1,1,1)=White. -
CMY/CMYK (Cyan, Magenta, Yellow, Key/Black): Subtractive model. Primary colors of pigments/printing.
(0,0,0)=White(paper),(1,1,1)=Black(ideally, but uses K for true black). -
HSV/HSB (Hue, Saturation, Value/Brightness): Cylindrical model. Intuitive for humans.
-
Hue: Color type (0-360° on color wheel).
-
Saturation: Purity (0=gray, 1=full color).
-
Value/Brightness: Intensity (0=black, 1=brightest).
-
VIII. Multimedia System Fundamentals
Multimedia: Definition, Components, Characteristics
-
Definition: Integration of multiple media types (text, graphics, audio, video, animation) in a computer-controlled interactive environment.
-
Components:
-
Text: Alphanumeric characters.
-
Graphics: 2D/3D vector/raster images.
-
Audio: Speech, music, sound effects.
-
Video: Moving images (sequence of frames).
-
Animation: Synthetic moving images (2D/3D).
-
-
Characteristics:
-
Discreteness: Media elements are distinct, independent objects.
-
Synchronization: Temporal coordination between media streams (e.g., audio with video, animation with text).
-
Integrity: Combined presentation creates a unified experience.
-
Interactivity: User control over presentation (navigation, manipulation).
-
High Bandwidth/Storage Requirements: Especially for audio/video.
-
Applications of Multimedia
-
Education: E-learning, simulations, virtual labs, tutorials.
-
Entertainment: Video games, movies, streaming services, VR experiences.
-
Medicine: Medical imaging (MRI, CT), surgical simulation, telemedicine.
-
Business: Presentations, digital signage, web conferencing, kiosks, advertising.
-
Engineering/Design: CAD/CAE visualization, architectural walkthroughs.
-
Scientific Visualization: Visualizing complex data (weather, molecules).
Multimedia System Architecture (Detailed Layers)
-
Capture/Production Layer: Input devices and software to create media (camera, microphone, scanner, authoring tools).
-
Storage Layer: Databases and file systems to store large multimedia objects (RAID, optical media, cloud storage). Includes Multimedia Databases with content-based retrieval.
-
Processing Layer: Tools for editing, compressing, indexing, and synchronizing media (codecs, editors, synchronization engines).
-
Presentation Layer: Output devices (display, speakers) and rendering software (players, browsers). Handles synchronization during playback.
-
Communication/Network Layer: Transmits media over networks (streaming protocols, QoS management). Critical for distributed multimedia.
IX. Multimedia Data & Storage
Multimedia Databases
-
Need: To efficiently store, manage, and retrieve large volumes of heterogeneous multimedia data (images, video, audio) which is content-rich and unstructured compared to traditional relational data.
-
Challenges:
-
Huge Volume: High storage and bandwidth needs.
-
Content-Based Retrieval: Query by example (QBE) or by features (color, shape, texture, motion), not just text keywords.
-
Temporal Aspects: Synchronization requirements for video/audio.
-
Diversity: Different formats, resolutions, compression schemes.
-
Indexing: Efficient indexing of high-dimensional feature vectors.
-
-
Structure: Often uses a relational DBMS for metadata (text descriptions, timestamps) and a file server for actual media objects. May use object-oriented or object-relational models.
-
Content-Based Retrieval (CBR): Extract features (color histogram, texture co-occurrence, shape descriptors, audio spectral features) during ingestion. Store features in index. Query: extract features from query object, search index for similar feature vectors.
Multimedia Data File Formats Standards
-
Still Image:
-
JPEG (Joint Photographic Experts Group): Lossy compression, DCT-based.
.jpg,.jpeg. Widely used for photos. -
PNG (Portable Network Graphics): Lossless compression, supports transparency (alpha channel).
.png. Web graphics, lossless needs.
-
-
Audio:
-
WAV (Waveform Audio File Format): Uncompressed (PCM), large size.
.wav. High-quality master copies. -
MP3 (MPEG-1 Audio Layer 3): Lossy compression, perceptual coding.
.mp3. Music distribution.
-
-
Video:
-
AVI (Audio Video Interleave): Container format (can contain various codecs).
.avi. Legacy. -
MPEG (Moving Picture Experts Group): Standards family.
-
MPEG-1: VCD, early web video.
-
MPEG-2: DVD, digital TV.
-
MPEG-4: Modern (
.mp4), efficient compression, interactive features.
-
-
-
Animation:
-
GIF (Graphics Interchange Format): Lossless, supports simple 2D animation and transparency.
.gif. Web. -
FLV (Flash Video): Legacy container for Flash video.
.flv.
-
X. Compression Techniques
Lossless Compression
-
Goal: Reconstruct original data exactly. No information loss.
-
Techniques:
-
Run-Length Encoding (RLE): Replace consecutive identical symbols (runs) with
(value, count). Effective for simple graphics, bitmaps with large uniform areas. -
Huffman Coding: Variable-length prefix code. Assign shorter codes to more frequent symbols. Builds optimal prefix tree from symbol frequencies.
-
LZW (Lempel-Ziv-Welch): Dictionary-based. Builds a string table dynamically during encoding. Used in GIF, TIFF, UNIX compress.
-
-
Compression Ratio: Typically 2:1 to 5:1 for text/data. Limited for complex images/audio.
Lossy Compression
-
Goal: Achieve much higher compression by discarding perceptually less important information. Irreversible.
-
Techniques:
-
Transform Coding (DCT - Discrete Cosine Transform): Core of JPEG and MPEG. Converts spatial data to frequency domain. High-frequency coefficients (less perceptible) are quantized coarsely or zeroed out.
-
Vector Quantization (VQ): Groups similar data vectors (e.g., image blocks) and represents each by a single codeword from a pre-trained codebook. Encodes index of codeword.
-
Predictive Coding: Encode difference between sample and its prediction (e.g., DPCM). Residual often has lower entropy.
-
-
Compression Ratio: 10:1 to 100:1+ for images/audio/video with acceptable quality loss.
Compression Standards
-
JPEG (Still Images): Lossy (baseline), lossless, progressive modes. Uses DCT, Huffman coding, quantization.
-
MPEG (Video & Audio):
-
MPEG-1/2: Inter-frame prediction (I, P, B frames), motion compensation, DCT.
-
MPEG-4 (Part 2 & H.264/AVC, H.265/HEVC): Advanced prediction, transform, entropy coding. Used in streaming (YouTube, Netflix), Blu-ray.
-
-
AAC (Advanced Audio Coding): Successor to MP3. Better efficiency at low bitrates. Standard for Apple, broadcasting.
XI. Animation & Authoring
Animation: Definition and Principles
-
Definition: The rapid display of a sequence of slightly different images (frames) to create the illusion of motion.
-
12 Principles of Animation (Disney, Ollie Johnston & Frank Thomas):
-
Squash and Stretch: Deform objects to show weight, flexibility.
-
Anticipation: Prepare the audience for an action (wind-up).
-
Staging: Direct audience attention to what's important.
-
Straight Ahead Action & Pose to Pose: Two approaches to drawing frames.
-
Follow Through & Overlapping Action: Parts continue moving after main action stops.
-
Slow In & Slow Out: More frames at start/end of action for natural acceleration/deceleration.
-
Arc: Natural motion follows curved paths.
-
Secondary Action: Supporting actions that add life.
-
Timing: Spacing of frames to convey weight, emotion.
-
Exaggeration: Enhance extremes for clarity/impact.
-
Solid Drawing: 3D form, weight, volume.
-
Appeal: Charismatic, interesting characters.
-
Types of Animation
-
2D Animation: Traditional cel, digital vector (Flash/Animate), stop-motion (clay, cut-out). Flat, drawn.
-
3D Animation: Computer-generated imagery (CGI). Models rigged with skeletons, keyframed. Used in films, games, visualization.
-
Motion Capture (Mocap): Record real actor/object movement and apply to digital character. High realism for human motion.
-
Stop Motion: Physically manipulate real objects (clay, puppets, models) and photograph frame-by-frame.
Authoring Tools
-
Definition: Software that enables non-programmers to create multimedia applications by assembling media elements, defining logic, and setting timing without writing code.
-
Types & Examples:
-
Card-and-Stack (Page-based): Slides/pages linked by buttons. Example: Microsoft PowerPoint (basic), Adobe Persuasion.
-
Icon-based/Flowchart: Visual programming with icons representing actions/events. Example: Macromedia Authorware, ToolBook.
-
Time-based: Arrange media on a timeline (like video editing). Example: Adobe Flash (Animate), Macromedia Director (Shockwave), Adobe After Effects.
-
Object-based: Focus on object properties and behaviors. Example: Visual Basic (with multimedia controls), Java-based tools.
-
-
Key Features: Media import, timeline, interactivity (buttons, hotspots), scripting support (for advanced logic).
XII. Evolving Technologies & Visualization
Evolving Technologies for Multimedia
-
Virtual Reality (VR): Immersive, computer-generated simulation of a 3D environment. User is isolated from real world (head-mounted display, headphones). Applications: training, gaming, therapy.
-
Augmented Reality (AR): Overlay digital information (graphics, text) onto the real world in real-time. Uses cameras/sensors. Example: Pokémon GO, Microsoft HoloLens, industrial maintenance.
-
360° Video: Spherical video captured by omnidirectional camera. User can look around freely (on VR headsets or by dragging on screen). Immersive but not interactive 3D.
-
Holography: Recording and reconstructing light fields to create true 3D images that can be viewed without glasses. Uses interference patterns. Still emerging for consumer use.
-
Interactive Media: Broad term encompassing any media where user input affects the content (games, interactive TV, choose-your-own-adventure videos, responsive installations).
Data Visualization
-
High-Dimensional Data Visualization:
-
Problem: Data with >3 dimensions (attributes) cannot be visualized directly.
-
Techniques:
-
Parallel Coordinates: Each dimension is a vertical axis (parallel). A data item is a polyline crossing each axis at its value. Patterns reveal clusters/correlations.
-
Scatterplot Matrices: Grid of 2D scatterplots showing all pairwise combinations of dimensions. Reveals pairwise relationships.
-
Glyphs/Icons: Encode multiple variables in shape, size, color of a single icon (e.g., star plot, Chernoff faces).
-
Dimensionality Reduction: Project high-D data to 2D/3D while preserving structure (PCA, t-SNE, UMAP).
-
-
-
Applications of Visualization:
-
Scientific Visualization: Visualizing physical phenomena (fluid flow, medical scans, molecular structures).
-
Information Visualization: Abstract data (networks, hierarchies, text, time-series). Example: Stock charts, social network graphs, infographics.
-
Business Intelligence (BI): Dashboards, KPIs, sales performance maps. Supports decision-making.
-
Visual Analytics: Combines automated analysis (data mining) with interactive visual exploration.
-
[!TIP]
Exam Focus: Past papers heavily test Cohen-Sutherland (with calculation), Cyrus-Beck (principle), Bresenham/DDA/Midpoint Circle (with examples), Bezier/B-Spline (equations, differences), 2D Rotation about arbitrary point, Parallel vs Perspective, Z-Buffer/Back-Face, Shading Models, Multimedia Architecture/Compression, Animation Principles, and Evolving Technologies. Always be ready to differentiate (e.g., raster vs random, Bezier vs B-Spline, shading models, projections) and explain algorithms with a small example. For clipping, practice region codes and parametric
tcalculation.