Skip to content
AL-603 (A) · Image and Video Processing/Quick Revision Short Notes

Image and Video Processing (AL-603 (A)) - Unit 5 Short Notes

UNIT 5: Data and Visual Analytics (for Image and Video Processing)


I. Fundamentals of Data

Variables and Data Categorization

  • Variable: A characteristic or attribute that can be measured or counted (e.g., pixel intensity, frame rate, video duration).

  • Data Categorization:

    • Categorical (Qualitative): Represents categories or groups.

      • Nominal: No inherent order (e.g., video format: MP4, AVI, MOV).

      • Ordinal: Has a meaningful order but no fixed interval (e.g., video quality rating: Low, Medium, High).

    • Numerical (Quantitative): Represents measurable quantities.

      • Discrete: Countable values (e.g., number of frames, object count in a video).

      • Continuous: Any value within a range (e.g., pixel brightness, video bitrate).

Levels of Measurement

Level Description Example (Image/Video Context) Mathematical Operations
Nominal Labels/categories only. No order. Video codec type (H.264, HEVC). Equality/inequality checks.
Ordinal Ordered categories. Unknown intervals. Compression level (Low, Medium, High). Greater/less than comparisons.
Interval Ordered with equal intervals. No true zero. Timestamp in a video (seconds). Addition/subtraction.
Ratio All interval properties + a true zero. Image width/height (pixels), file size (MB). All arithmetic operations.

[!TIP] Exam Focus: Be prepared to classify given scenarios into the correct level of measurement and state which operations are valid.


II. Descriptive Statistics

Measures of Central Tendency

  • Mean (Arithmetic Average): Sum of all values divided by count.

$$ \bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_i $$

  • Median: Middle value in an ordered dataset. Robust to outliers.

  • Mode: Most frequently occurring value. Can be used for categorical data.

Measures of Dispersion (Variability)

Measure Formula (Sample) Description
Range $\text{Max} - \text{Min}$ Simplest measure, sensitive to outliers.
Variance $$\displaystyle s^2 = \frac{1}{n-1} \sum_{i=1}^{n} (x_i - \bar{x})^2 $$ Average squared deviation from mean.
Standard Deviation $$\displaystyle s = \sqrt{s^2} $$ Square root of variance. In same units as data.
IQR $$\displaystyle Q_3 - Q_1 $$ Spread of middle 50% of data. Robust to outliers.

[!TIP] Common Pitfall: Remember to use $n-1$ (Bessel's correction) for sample variance/standard deviation when inferring about a population.


III. Inferential Statistics

A. Hypothesis Testing

Core Steps:

  1. State Null ($$\displaystyle H_0 $$) and Alternative ($$\displaystyle H_1 $$) hypotheses.

  2. Choose significance level ($\alpha$, typically 0.05).

  3. Select appropriate test statistic and calculate its value.

  4. Determine p-value or critical region.

  5. Make a decision: Reject $$\displaystyle H_0 $$ if p-value < $\alpha$.

t-test

Used to compare means when population standard deviation is unknown and/or sample size is small.

Test Type Purpose Test Statistic
One-sample Compare sample mean ($\bar{x}$) to a known population mean ($$\displaystyle \mu_0 $$). $$\displaystyle t = \frac{\bar{x} - \mu_0}{s/\sqrt{n}} $$
Two-sample (Independent) Compare means of two independent groups. $$\displaystyle t = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} $$ (Welch's t)
Paired (Dependent) Compare means of two related measurements (before/after). $$\displaystyle t = \frac{\bar{d}}{s_d/\sqrt{n}} $$ where $d$ is difference.

Example (From May 2024 Paper): Test if potato yield from 12 farms is significantly better than standard $$\displaystyle \mu=20 $$.

Data: $$\displaystyle X = [21.5, 24.5, 18.5, 17.2, 14.5, 23.2, 22.1, 20.5, 19.4, 18.1, 24.1, 18.5] $$

Solution:

  1. $$\displaystyle H_0: \mu = 20 $$ (No improvement), $$\displaystyle H_1: \mu > 20 $$ (Better yield - one-tailed).
  1. $$\displaystyle \alpha = 0.05 $$, df = 11.
  1. $$\displaystyle \bar{x} = 20.15 $$, $s \approx 3.15$.
  1. $$\displaystyle t = \frac{20.15 - 20}{3.15/\sqrt{12}} \approx 0.164 $$.
  1. Critical $$\displaystyle t_{0.05,11} \approx 1.796 $$. Since $$\displaystyle 0.164 < 1.796 $$, Fail to reject $$\displaystyle H_0 $$. No significant evidence yield is better.

Chi-Square ($$\displaystyle \chi^2 $$) Test

Used for categorical data.

  1. Goodness of Fit: Tests if observed frequencies fit an expected distribution.

$$ \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$

df = (number of categories - 1).
  1. Test of Independence: Tests if two categorical variables are associated (contingency table).

    df = (rows - 1) * (columns - 1).

    Condition: All expected frequencies $$\displaystyle E_i \ge 5 $$.

[!TIP] Key Difference: t-test compares means of numerical data. Chi-square compares frequencies of categorical data.

B. Regression Analysis

Models relationship between a dependent variable (Y) and one or more independent variables (X).

Type Equation Use Case
Simple Linear $$\displaystyle Y = \beta_0 + \beta_1 X + \epsilon $$ One predictor (e.g., predict image quality score from compression ratio).
Multiple $$\displaystyle Y = \beta_0 + \beta_1 X_1 + ... + \beta_k X_k + \epsilon $$ Multiple predictors (e.g., predict video bitrate from resolution, frame rate, codec).
  • Dependent Variable (Y): Outcome being predicted.

  • Independent Variable (X): Predictor(s).

  • Dummy Variables: Categorical X converted to binary (0/1) for regression (e.g., Is_MP4, Is_AVI).

C. Maximum Likelihood Estimation (MLE)

Method to estimate model parameters by finding the values that maximize the likelihood of observing the given data.

Steps:

  1. Assume a probability distribution for the data (e.g., Normal, Bernoulli).

  2. Write the Likelihood Function $L(\theta)$: probability of data given parameters $\theta$.

  3. Take the Log-Likelihood $$\displaystyle \ell(\theta) = \log L(\theta) $$ (simplifies multiplication to addition).

  4. Differentiate $\ell(\theta)$ w.r.t. $\theta$ and set to zero.

  5. Solve for $$\displaystyle \hat{\theta}_{MLE} $$.

Simple Example (Coin Toss): Data: 7 heads in 10 tosses. Likelihood for probability $p$: $$\displaystyle L(p) = p^7(1-p)^3 $$. MLE $$\displaystyle \hat{p} = 7/10 = 0.7 $$.

D. Bayesian Modeling

Incorporates prior knowledge/beliefs about parameters ($P(\theta)$) with data likelihood ($P(D|\theta)$) to get posterior distribution $P(\theta|D)$.

$$ P(\theta|D) = \frac{P(D|\theta) P(\theta)}{P(D)} \propto P(D|\theta) P(\theta) $$

Advantages:

  • Incorporates prior information.

  • Provides full probability distribution (uncertainty quantification).

  • Intuitive interpretation of credible intervals.

Disadvantages:

  • Choice of prior can be subjective.

  • Computationally intensive for complex models (MCMC methods).

  • Can be less efficient with large, informative datasets.

E. Resampling Methods

Non-parametric methods that use repeated samples from the observed data.

  • Bootstrapping:

    1. Create many "bootstrap samples" by sampling with replacement from original data (same size).

    2. Calculate statistic (e.g., mean) for each bootstrap sample.

    3. Use distribution of bootstrap statistics to estimate standard error, confidence intervals.

  • Cross-Validation:

    • k-fold CV: Split data into k folds. Use k-1 folds for training, 1 for testing. Repeat k times. Average performance.

    • Purpose: Assess model generalization, tune hyperparameters, avoid overfitting.


IV. Data Exploration Techniques

A. Univariate Analysis

Explores a single variable.

  • Numerical: Histogram, Box Plot, Summary stats (mean, sd, IQR).

  • Categorical: Bar Chart, Frequency Table, Mode.

B. Bivariate Analysis

Explores relationship between two variables.

  • Num-Num: Scatter Plot, Correlation Coefficient ($r$ or $\rho$).

$$ r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}} $$

  • Cat-Cat: Contingency Table, Stacked Bar Chart, Chi-Square Test.

  • Num-Cat: Box Plot (by category), Z-test/t-test (if two groups).

C. Multivariate Analysis

Explores relationships among multiple variables.

  • Principal Component Analysis (PCA): Dimensionality reduction. Transforms data to new orthogonal axes (Principal Components) capturing maximum variance. Used for feature extraction, noise reduction.

  • Cluster Analysis: Groups similar observations.

    • k-Means: Partitions data into k clusters by minimizing within-cluster sum of squares.

    • Hierarchical: Builds tree (dendrogram) of clusters.

[!TIP] Exam Link: "Multivariate analysis" questions often expect PCA and Cluster Analysis as key examples.


V. Data Management and Wrangling

A. Data Wrangling Process

  1. Collection: Acquire data from sources (APIs, databases, files).

  2. Cleaning: Handle missing values (impute/remove), correct errors, remove duplicates.

  3. Transformation: Normalize/scales features, encode categories, create new features.

  4. Integration: Merge/join datasets from different sources.

B. Data Management and Indexing

  • Database Concepts: Relational (SQL) vs. Non-Relational (NoSQL). Tables, keys (Primary, Foreign).

  • Indexing: Data structure (e.g., B-tree, Hash index) to speed up query retrieval. Trade-off: faster reads, slower writes, extra storage.

C. Data Analyst Ecosystem

Category Tools Primary Use
Querying SQL (MySQL, PostgreSQL) Data extraction/aggregation from databases.
Programming Python (Pandas, NumPy), R Data manipulation, statistical analysis, modeling.
Visualization Tableau, PowerBI, Matplotlib, Seaborn Creating dashboards and plots.
Big Data Spark, Hadoop (Hive) Processing large-scale datasets.

VI. Data Visualization

A. Principles and Types

  • Design Principles: Clarity, accuracy, efficiency. Minimize "chartjunk." Use appropriate color scales (sequential, diverging, categorical). Label axes, add titles.

  • Chart Selection Guide:

    • Comparison: Bar Chart, Column Chart.

    • Distribution: Histogram, Box Plot.

    • Relationship: Scatter Plot, Line Chart.

    • Composition: Pie Chart (few categories), Stacked Bar.

B. Creating Custom Visualizations

For complex datasets: Combine multiple plot types (e.g., scatter matrix, faceted plots). Use interactive libraries (Plotly) for drill-down. Focus on storytelling—highlight key insights.

C. Python Visualization Libraries

Library Strengths Typical Use
Pandas Quick, built-in plotting (.plot()). Fast exploratory plots from DataFrames.
Matplotlib Highly customizable, low-level "canvas." Creating publication-quality, complex figures.
Seaborn Statistical, attractive defaults. Built on Matplotlib. Visualizing distributions, relationships, categorical data.
plotnine (ggplot) Grammar of Graphics paradigm. Declarative syntax. Building plots layer-by-layer (data -> aesthetics -> geometry).

D. PowerBI

  • PowerBI Desktop: Free application for data modeling, transformation (Power Query), and report creation.

  • PowerBI Service: Cloud-based SaaS for sharing, collaboration, scheduling refresh.

  • PowerBI Mobile: Apps for viewing dashboards on devices.

  • Key Tools/Features:

    • DAX (Data Analysis Expressions): Formula language for calculated columns/measures (e.g., Total Sales = SUM(Sales[Amount])).

    • Query Editor (Power Query): GUI for data shaping/cleaning (similar to wrangling).

    • Dashboards: Single-page, tile-based, real-time visualizations.


VII. File Formats

Structured Data Formats

  • CSV (Comma-Separated Values): Plain text, rows/columns. Simple, universal. Delimiter issues, no schema.

  • JSON (JavaScript Object Notation): Hierarchical, key-value pairs. Flexible, web-friendly. More verbose than CSV.

  • XML (eXtensible Markup Language): Markup language with tags. Very verbose, schema-driven (XSD). Used in older systems.

Binary Formats

  • Store data in binary (not human-readable). Efficient for storage/access.

  • Examples: Pickle (Python objects), Parquet (columnar, efficient for analytics), HDF5 (hierarchical, large scientific data).

Image and Video Formats (Contextualized)

Format Type Key Property Use Case
JPEG Lossy High compression, quality loss. Web images, photos.
PNG Lossless Supports transparency, larger files. Graphics, logos, screenshots.
MP4 (H.264/HEVC) Lossy Video High compression, widely supported. Streaming, storage.
AVI Less Compressed Older format, larger files, simple codec. Archival, editing (less common now).
RAW Unprocessed Sensor data, no in-camera processing. Professional photography, max quality.

[!TIP] Context Link: In image/video processing, format choice affects compression artifacts (JPEG blocking), color depth (PNG 16-bit vs JPEG 8-bit), and editing workflow (RAW vs JPEG).


VIII. Big Data Processing

A. Big Data Concepts (4 V's)

  • Volume: Massive scale (TB, PB).

  • Velocity: Speed of data generation/ingestion (real-time streams).

  • Variety: Structured, unstructured (images, video), semi-structured (JSON).

  • Veracity: Uncertainty, quality, trustworthiness of data.

B. Hadoop Ecosystem

  • HDFS (Hadoop Distributed File System):

    • Architecture: Master-Slave. NameNode (master, metadata), DataNodes (slaves, store blocks).

    • Principle: Files split into 128MB blocks, replicated (default 3x) across DataNodes. Fault-tolerant.

    • DiagramCANVAS: Draw NameNode connected to multiple DataNodes in a rack. Show a file split into blocks (A1, A2, A3) distributed across 3 DataNodes with replicas.

  • Hive:

    • Data warehouse infrastructure on Hadoop.

    • Provides HiveQL (SQL-like query language).

    • Translates queries into MapReduce/Tez/Spark jobs. For batch processing of structured data.

C. Other Big Data Tools

  • Apache Spark: In-memory cluster computing. Faster than Hadoop MapReduce for iterative algorithms (ML). Modules: Spark SQL, MLlib, Streaming.

  • HBase: NoSQL, column-family store. Provides real-time, random read/write access to big data (on top of HDFS). Good for sparse data.

  • Apache Kafka: Distributed streaming platform. Handles high-velocity data feeds (e.g., video sensor streams).

[!TIP] Comparison: Use HDFS for storage. Use Hive for SQL-like batch queries on large structured data. Use Spark for faster, complex analytics/ML on the same data. Use Kafka for ingesting real-time data streams.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in