Skip to content
AL-603 (A) · Image and Video Processing/Quick Revision Short Notes

Image and Video Processing (AL-603 (A)) - Unit 4 Short Notes

UNIT 4: Data Analytics & Visualization (for Image/Video Processing Context)

I. Foundational Statistical Concepts

A. Measures of Central Tendency

  • Definition: Single values that describe the center of a data distribution.

  • Mean (Arithmetic Average): Sum of all values divided by count.

$$ \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$

  • Median: Middle value in an ordered dataset. For even n, median = average of two middle values.

  • Mode: Most frequently occurring value(s). Can be unimodal, bimodal, etc.

  • [!TIP] Exam Tip: For skewed distributions (common in image intensity histograms), median is a better central measure than mean as it is robust to outliers.

B. Measures of Dispersion/Location

  • Range: Max - Min. Highly sensitive to outliers.

  • Variance (s²): Average of squared deviations from the mean.

$$ s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} \quad \text{(Sample Variance)} $$

  • Standard Deviation (s): Square root of variance. Same units as data.

$$ s = \sqrt{s^2} $$

  • Interquartile Range (IQR): Spread of middle 50% of data. IQR = Q3 - Q1. Used in box plots to identify outliers (typically, outliers < Q1 - 1.5IQR or > Q3 + 1.5IQR).

C. Levels of Measurement

  1. Nominal: Categorical labels only (e.g., image class: "cat", "dog"). No order.

  2. Ordinal: Categorical with order but no consistent difference (e.g., image quality rating: "poor", "fair", "good").

  3. Interval: Numerical with order and consistent difference, but no true zero (e.g., temperature in Celsius). Ratios meaningless.

  4. Ratio: Numerical with order, consistent difference, and true zero (e.g., pixel intensity [0-255], image dimensions, duration). All mathematical operations valid.

D. Variables and Data Categorization

  • Variable: A characteristic/attribute that can take different values.

  • Categorization:

    • By Type: Categorical (Nominal/Ordinal) vs. Numerical (Interval/Ratio).

    • By Role in Analysis: Independent/Predictor (e.g., filter parameters) vs. Dependent/Response (e.g., processed image quality metric).

    • By Nature: Discrete (countable, e.g., number of objects) vs. Continuous (measurable, e.g., pixel coordinates).


II. Inferential Statistics & Hypothesis Testing

A. Statistical Inferences

  • Point Estimation: Using sample statistic to estimate population parameter (e.g., sample mean x̄ estimates population mean μ).

  • Confidence Interval (CI): Range of values likely to contain the population parameter.

$$ \text{CI} = \text{Point Estimate} \pm (\text{Critical Value} \times \text{Standard Error}) $$

*   For mean (large n/z-dist): `x̄ ± z*(σ/√n)`

*   For mean (small n/t-dist): `x̄ ± t*(s/√n)`

B. Chi-Square Test (χ²)

  • Goodness of Fit: Tests if observed frequency distribution fits an expected theoretical distribution.

$$ \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$

*   `O_i` = Observed frequency, `E_i` = Expected frequency.

*   **df = (number of categories) - 1**
  • Test of Independence: Tests if two categorical variables are associated in a contingency table.

$$ \chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}} \quad \text{for all cells (i,j)} $$

*   `E_{ij} = (row total_i * col total_j) / grand total`

*   **df = (rows - 1) * (cols - 1)**
  • [!TIP] Assumption: Expected frequencies E_i should be ≥5 for >80% of cells.

C. t-test

  • One-Sample t-test: Compares sample mean (x̄) to a known population mean (μ₀).

$$ t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} \quad \text{df} = n-1 $$

> **Example (May 2024):** Testing if potato yield (`X`) is better than standard (`μ=20`). `H₀: μ = 20`, `H₁: μ > 20`. Calculate `x̄`, `s`, then `t`. Compare to critical t-value.
  • Two-Sample Independent t-test: Compares means of two independent groups.

$$ t = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} \quad \text{df via Welch's approximation} $$

  • Two-Sample Paired t-test: Compares means of two related samples (e.g., before/after processing).

$$ t = \frac{\bar{d}}{s_d / \sqrt{n}} \quad \text{df} = n-1 $$

where `d` = difference for each pair.
  • [!TIP] Key Assumption: Data should be approximately normally distributed (check with histogram/Q-Q plot) and samples should have equal variance for standard independent t-test (use Welch's t-test if variances unequal).

D. Resampling Methods

  • Bootstrap: Repeatedly samples with replacement from the original sample to estimate sampling distribution of a statistic (e.g., mean, median). Used for CI estimation when distribution is unknown.

  • Cross-Validation: Resampling method for model evaluation. Data is split into k folds; model trained on k-1 folds and validated on the held-out fold, repeated k times. Common: k-fold CV (k=5 or 10).


III. Regression & Predictive Modeling

A. Regression Analysis

  • Simple Linear Regression: Models relationship between one independent (X) and one dependent (Y) variable.

$$ Y = \beta_0 + \beta_1 X + \epsilon $$

*   `β₀` = intercept, `β₁` = slope, `ε` = error term.

*   Estimated via **Ordinary Least Squares (OLS)**: Minimizes sum of squared residuals `Σ(y_i - ŷ_i)²`.
  • Multiple Linear Regression: Extends to multiple independent variables (X₁, X₂, ..., Xₖ).

$$ Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + ... + \beta_k X_k + \epsilon $$

  • [!TIP] Interpretation: β₁ represents the change in Y for a one-unit change in X₁, holding all other X constant.

B. Types of Variables in Regression Modeling

  • Dependent Variable (Y): The target variable being predicted (e.g., image sharpness score).

  • Independent/Predictor Variables (X): Features used for prediction (e.g., filter kernel size, compression ratio).

  • Dummy Variables: Binary (0/1) variables used to represent categorical predictors (e.g., Image_Type_Color=1 if color, 0 if grayscale). k categories require k-1 dummy variables to avoid multicollinearity (dummy variable trap).

C. Multivariate Analysis

  • Concept: Simultaneous analysis of more than one outcome/dependent variable. (Contrast with Multiple Regression, which has one dependent variable but multiple independents).

  • Applications in Image/Video:

    • Principal Component Analysis (PCA): Dimensionality reduction. Finds orthogonal axes (principal components) of maximum variance in multivariate data (e.g., pixel values across many images).

    • Cluster Analysis: Grouping similar objects (e.g., grouping similar image features, video scenes).

    • Multivariate Analysis of Variance (MANOVA): Extension of ANOVA for multiple dependent variables (e.g., comparing groups on multiple image quality metrics: PSNR, SSIM, VMAF).

  • [!TIP] Key Distinction: Multivariate = Multiple dependent variables. Multiple Regression = One dependent, multiple independent variables.


IV. Advanced Statistical Methods

A. Bayesian Modeling (Principles)

  • Core Idea: Treats parameters as random variables with probability distributions, updating prior beliefs with observed data to get posterior beliefs.

  • Bayes' Theorem:

$$ P(\theta|D) = \frac{P(D|\theta) \cdot P(\theta)}{P(D)} $$

*   `P(θ|D)` = **Posterior** (updated belief about parameter θ given data D).

*   `P(D|θ)` = **Likelihood** (probability of data given θ).

*   `P(θ)` = **Prior** (initial belief about θ).

*   `P(D)` = Marginal likelihood (normalizing constant).
  • Advantages: Incorporates prior knowledge, provides full probability distribution for parameters (not just point estimates), intuitive uncertainty quantification.

  • Disadvantages: Choice of prior can be subjective, computationally intensive for complex models (MCMC methods).

B. Maximum Likelihood Estimation (MLE)

  • Goal: Find parameter values (θ) that maximize the likelihood P(D|θ)—i.e., make the observed data most probable.

  • Steps:

    1. Write down the likelihood function L(θ) = P(D|θ). Often use log-likelihood ℓ(θ) = log L(θ) for easier computation.

    2. Differentiate ℓ(θ) w.r.t. θ.

    3. Set derivative = 0 and solve for θ.

  • Example (Bernoulli/Binomial): For n coin flips with k heads, L(p) = p^k (1-p)^(n-k). MLE for p is p̂ = k/n.

  • For Linear Regression: MLE under normal error assumption yields the same OLS estimates.

  • [!TIP] MLE vs. Bayesian: MLE gives a single "best" parameter value. Bayesian gives a distribution. MLE is a special case of Bayesian with a uniform (flat) prior.


V. Data Handling, Wrangling & Formats

A. Data Wrangling Process

  1. Data Collection: Gathering data from sources (sensors, logs, databases, web scraping).

  2. Data Cleaning: Handling missing values (impute/remove), correcting errors, removing duplicates, ensuring consistency.

  3. Data Transformation: Normalization/scaling, encoding categorical variables, creating new features (feature engineering), aggregating.

  4. Data Enrichment: Adding relevant external data (e.g., adding geographical data to image metadata).

  • [!TIP] Rule of Thumb: Data wrangling often consumes 60-80% of a data analyst's time.

B. File Formats

  • Structured:

    • CSV (Comma-Separated Values): Plain text, rows & columns. Simple, universal. Delimiters can cause issues.

    • JSON (JavaScript Object Notation): Hierarchical, key-value pairs. Flexible for nested data (e.g., image metadata with nested tags).

  • Unstructured:

    • Images: JPEG, PNG, TIFF (lossless), RAW. Contain pixel arrays + metadata (EXIF).

    • Video: MP4 (H.264/AVC, H.265/HEVC), AVI, MKV. Contain encoded video/audio streams + containers.

  • Binary (Optimized for Big Data):

    • Parquet: Columnar storage. Highly efficient for querying specific columns (common in Spark/Hadoop).

    • Avro: Row-based, compact, good for serialization.

C. Pandas in Python

  • Core Structure: DataFrame (2D labeled table) and Series (1D labeled array).

  • Key Operations:

    • pd.read_csv(), pd.read_json()

    • df.head(), df.info(), df.describe()

    • df[col], df.loc[], df.iloc[] (indexing/selection)

    • df.groupby(), df.merge(), df.pivot_table()

    • Handling missing data: df.isnull(), df.dropna(), df.fillna()

  • [!TIP] Performance: Vectorized operations (df['col'].mean()) are vastly faster than Python loops (for index, row in df.iterrows()).

D. Data Management and Indexing

  • Concept: Organizing, storing, and retrieving data efficiently.

  • Database Indexing: Creating a separate data structure (index) to speed up querying on a column(s).

    • B-tree Index: Most common. Allows fast equality/range queries. Slows down INSERT/UPDATE/DELETE.

    • Hash Index: Fast for exact matches only.

    • Bitmap Index: Efficient for low-cardinality columns (e.g., categorical variables like image "scene_type").

  • [!TIP] Trade-off: Indexes improve read speed but add overhead for write operations and consume storage.


VI. Data Visualization & Exploration

A. Data Visualization Fundamentals

  1. Univariate Exploration (Single Variable):

    • Histogram: Shows distribution of a continuous variable. Bins choice affects perception.

    • Box Plot (Whisker Plot): Shows 5-number summary (min, Q1, median, Q3, max) and outliers. Excellent for comparing distributions across categories.

    DiagramCANVAS: A standard box plot with labeled median, quartiles, whiskers, and outliers marked as individual points beyond 1.5*IQR
  2. Bivariate Exploration (Two Variables):

    • Scatter Plot: Relationship between two continuous variables. Reveals correlation, clusters, non-linearity.

    • Heatmap: Visualizes matrix data, e.g., correlation matrix (df.corr()). Color intensity represents value magnitude.

  3. Multivariate Exploration (>2 Variables):

    • Pair Plot (Scatter Matrix): Grid of scatter plots for all pairwise combinations + histograms on diagonal. Use seaborn.pairplot().

    • 3D Plots: Can show 3 continuous variables (x,y,z) or use color/size for 4th/5th variables. Often interactive (Plotly).

    • Faceting (Small Multiples): Creating multiple plots of the same type for different subsets of data (e.g., scatter plot of feature1 vs feature2 for each image_category).

B. Python Visualization Libraries

  • Matplotlib: Foundation library. Low-level, highly customizable. Steeper learning curve.

  • Seaborn: Built on Matplotlib. Statistical visualizations. Simpler syntax, beautiful defaults. Great for distplot, boxplot, heatmap, pairplot.

  • Plotly: Interactive, web-based visualizations. Supports 3D, zooming, hover info. Good for dashboards.

  • Bokeh: Interactive visualizations for web apps. Similar to Plotly, focuses on large/streaming data.

C. ggplot2 in R (Grammar of Graphics)

  • Philosophy: Build plots layer by layer: Data + Aesthetics (aes) + Geometries (geom_*) + Facets + Statistics + Coordinates + Theme.

  • Example: ggplot(data, aes(x=var1, y=var2)) + geom_point() + facet_wrap(~category)

  • Strength: Consistent, logical syntax. Excellent for complex multi-plot figures.

D. Creating Custom Visualizations for Complex Datasets

  1. Understand the Question: What relationship/story are you trying to show?

  2. Choose the Right Chart Type: Match data structure (categorical vs. continuous) and relationship (comparison, distribution, composition, relationship).

  3. Simplify & Declutter: Remove non-essential ink (chartjunk). Use color purposefully (qualitative, sequential, diverging palettes).

  4. Leverage Interactivity: For high-dimensional data, use tools like Plotly to allow brushing, zooming, filtering.

  5. Combine Views: Use dashboards (PowerBI, Dash) or faceted plots to show multiple linked perspectives.


VII. Big Data Ecosystem & Processing

A. Big Data (4 V's) & Tools

  • Characteristics:

    • Volume: Massive scale (TB/PB).

    • Velocity: High speed of data in/out (real-time streams from video feeds).

    • Variety: Structured, semi-structured, unstructured (images, video, logs).

    • Veracity: Uncertainty, quality, trustworthiness of data.

  • Processing Tools:

    • Hadoop: Framework for distributed storage & processing. Core: HDFS (storage) + MapReduce (processing).

    • Spark: Faster in-memory processing engine. Can run on Hadoop YARN or standalone. Better for iterative algorithms (e.g., ML) and streaming. APIs in Scala, Python (PySpark), Java, R.

B. HDFS (Hadoop Distributed File System)

  • Architecture:

    • NameNode: Master server. Manages file system namespace (metadata: file->block mapping). Single Point of Failure (HA solutions exist).

    • DataNode: Slave nodes. Store actual data blocks (default 128MB/256MB). Report status to NameNode.

    • Client: Interface to read/write data. Gets block locations from NameNode, then talks directly to DataNodes.

  • Operations: Files are broken into blocks, replicated (default 3x) across DataNodes for fault tolerance. Write-once-read-many (WORM) model.

  • DiagramSEARCH: "HDFS architecture diagram NameNode DataNode"

C. Hive (Data Warehousing on Hadoop)

  • Concept: Provides SQL-like interface (HiveQL) to query data stored in HDFS.

  • How it works: Hive translates HiveQL queries into MapReduce (or Tez/Spark) jobs.

  • Components:

    • Metastore: Stores table schema, partitions, location (in RDBMS like MySQL).

    • Driver: Compiles, optimizes, and executes queries.

    • Executors: Run the tasks (Map/Reduce).

  • Use Case: Batch processing of large, structured logs or tabular data (e.g., aggregated video view counts). Not for low-latency queries.

D. Data Analyst Ecosystem

  • Roles: Data Analyst, Business Analyst, Data Scientist, Data Engineer.

  • Workflow: Data Ingestion → Storage (Data Lake/Warehouse) → Wrangling (ETL/ELT) → Analysis/Modeling → Visualization → Dashboarding.

  • Tools: SQL (Hive, PostgreSQL), Python/R (Pandas, Scikit-learn), Visualization (PowerBI, Tableau), Big Data (Spark, Hive), Orchestration (Airflow).


VIII. Business Intelligence & Dashboarding

A. PowerBI

  • Components:

    • Power BI Desktop: Free Windows application for building reports & data modeling. Core features: Power Query (data transformation), Data Model (relationships, DAX measures), Report View (visualizations).

    • Power BI Service: Cloud-based SaaS service for publishing, sharing, collaboration. Sets up scheduled refreshes, creates dashboards (single-page pin-based summaries).

    • Power BI Mobile: Apps for iOS/Android to view dashboards/reports on-the-go.

  • Key Tools:

    • Power Query (M Language): Data connection & transformation engine. Used in Desktop for ETL.

    • DAX (Data Analysis Expressions): Formula language for creating custom measures and calculated columns in the data model.

      • Example Measure: Total Sales = SUM(Sales[Amount])

      • Example Calculated Column: Profit = [Sales] - [Cost]

    • Visualizations: Standard (bar, line, pie) + custom visuals from marketplace.

  • Workflow: Connect to data source → Transform in Power Query → Model data (create relationships) → Build report with visuals & DAX → Publish to Service → Share/Dashboard.

  • [!TIP] Best Practice: Perform heavy transformations in Power Query (during load) rather than with DAX (during query) for better performance.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in