UNIT 5: Data and Visual Analytics (for Image and Video Processing)
I. Fundamentals of Data
Variables and Data Categorization
-
Variable: A characteristic or attribute that can be measured or counted (e.g., pixel intensity, frame rate, video duration).
-
Data Categorization:
-
Categorical (Qualitative): Represents categories or groups.
-
Nominal: No inherent order (e.g., video format: MP4, AVI, MOV).
-
Ordinal: Has a meaningful order but no fixed interval (e.g., video quality rating: Low, Medium, High).
-
-
Numerical (Quantitative): Represents measurable quantities.
-
Discrete: Countable values (e.g., number of frames, object count in a video).
-
Continuous: Any value within a range (e.g., pixel brightness, video bitrate).
-
-
Levels of Measurement
| Level | Description | Example (Image/Video Context) | Mathematical Operations |
|---|---|---|---|
| Nominal | Labels/categories only. No order. | Video codec type (H.264, HEVC). | Equality/inequality checks. |
| Ordinal | Ordered categories. Unknown intervals. | Compression level (Low, Medium, High). | Greater/less than comparisons. |
| Interval | Ordered with equal intervals. No true zero. | Timestamp in a video (seconds). | Addition/subtraction. |
| Ratio | All interval properties + a true zero. | Image width/height (pixels), file size (MB). | All arithmetic operations. |
[!TIP] Exam Focus: Be prepared to classify given scenarios into the correct level of measurement and state which operations are valid.
II. Descriptive Statistics
Measures of Central Tendency
- Mean (Arithmetic Average): Sum of all values divided by count.
$$ \bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_i $$
-
Median: Middle value in an ordered dataset. Robust to outliers.
-
Mode: Most frequently occurring value. Can be used for categorical data.
Measures of Dispersion (Variability)
| Measure | Formula (Sample) | Description |
|---|---|---|
| Range | $\text{Max} - \text{Min}$ | Simplest measure, sensitive to outliers. |
| Variance | $$\displaystyle s^2 = \frac{1}{n-1} \sum_{i=1}^{n} (x_i - \bar{x})^2 $$ | Average squared deviation from mean. |
| Standard Deviation | $$\displaystyle s = \sqrt{s^2} $$ | Square root of variance. In same units as data. |
| IQR | $$\displaystyle Q_3 - Q_1 $$ | Spread of middle 50% of data. Robust to outliers. |
[!TIP] Common Pitfall: Remember to use $n-1$ (Bessel's correction) for sample variance/standard deviation when inferring about a population.
III. Inferential Statistics
A. Hypothesis Testing
Core Steps:
-
State Null ($$\displaystyle H_0 $$) and Alternative ($$\displaystyle H_1 $$) hypotheses.
-
Choose significance level ($\alpha$, typically 0.05).
-
Select appropriate test statistic and calculate its value.
-
Determine p-value or critical region.
-
Make a decision: Reject $$\displaystyle H_0 $$ if p-value < $\alpha$.
t-test
Used to compare means when population standard deviation is unknown and/or sample size is small.
| Test Type | Purpose | Test Statistic |
|---|---|---|
| One-sample | Compare sample mean ($\bar{x}$) to a known population mean ($$\displaystyle \mu_0 $$). | $$\displaystyle t = \frac{\bar{x} - \mu_0}{s/\sqrt{n}} $$ |
| Two-sample (Independent) | Compare means of two independent groups. | $$\displaystyle t = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} $$ (Welch's t) |
| Paired (Dependent) | Compare means of two related measurements (before/after). | $$\displaystyle t = \frac{\bar{d}}{s_d/\sqrt{n}} $$ where $d$ is difference. |
Example (From May 2024 Paper): Test if potato yield from 12 farms is significantly better than standard $$\displaystyle \mu=20 $$.
Data: $$\displaystyle X = [21.5, 24.5, 18.5, 17.2, 14.5, 23.2, 22.1, 20.5, 19.4, 18.1, 24.1, 18.5] $$
Solution:
- $$\displaystyle H_0: \mu = 20 $$ (No improvement), $$\displaystyle H_1: \mu > 20 $$ (Better yield - one-tailed).
- $$\displaystyle \alpha = 0.05 $$, df = 11.
- $$\displaystyle \bar{x} = 20.15 $$, $s \approx 3.15$.
- $$\displaystyle t = \frac{20.15 - 20}{3.15/\sqrt{12}} \approx 0.164 $$.
- Critical $$\displaystyle t_{0.05,11} \approx 1.796 $$. Since $$\displaystyle 0.164 < 1.796 $$, Fail to reject $$\displaystyle H_0 $$. No significant evidence yield is better.
Chi-Square ($$\displaystyle \chi^2 $$) Test
Used for categorical data.
- Goodness of Fit: Tests if observed frequencies fit an expected distribution.
$$ \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$
df = (number of categories - 1).
-
Test of Independence: Tests if two categorical variables are associated (contingency table).
df = (rows - 1) * (columns - 1).
Condition: All expected frequencies $$\displaystyle E_i \ge 5 $$.
[!TIP] Key Difference: t-test compares means of numerical data. Chi-square compares frequencies of categorical data.
B. Regression Analysis
Models relationship between a dependent variable (Y) and one or more independent variables (X).
| Type | Equation | Use Case |
|---|---|---|
| Simple Linear | $$\displaystyle Y = \beta_0 + \beta_1 X + \epsilon $$ | One predictor (e.g., predict image quality score from compression ratio). |
| Multiple | $$\displaystyle Y = \beta_0 + \beta_1 X_1 + ... + \beta_k X_k + \epsilon $$ | Multiple predictors (e.g., predict video bitrate from resolution, frame rate, codec). |
-
Dependent Variable (Y): Outcome being predicted.
-
Independent Variable (X): Predictor(s).
-
Dummy Variables: Categorical X converted to binary (0/1) for regression (e.g.,
Is_MP4,Is_AVI).
C. Maximum Likelihood Estimation (MLE)
Method to estimate model parameters by finding the values that maximize the likelihood of observing the given data.
Steps:
-
Assume a probability distribution for the data (e.g., Normal, Bernoulli).
-
Write the Likelihood Function $L(\theta)$: probability of data given parameters $\theta$.
-
Take the Log-Likelihood $$\displaystyle \ell(\theta) = \log L(\theta) $$ (simplifies multiplication to addition).
-
Differentiate $\ell(\theta)$ w.r.t. $\theta$ and set to zero.
-
Solve for $$\displaystyle \hat{\theta}_{MLE} $$.
Simple Example (Coin Toss): Data: 7 heads in 10 tosses. Likelihood for probability $p$: $$\displaystyle L(p) = p^7(1-p)^3 $$. MLE $$\displaystyle \hat{p} = 7/10 = 0.7 $$.
D. Bayesian Modeling
Incorporates prior knowledge/beliefs about parameters ($P(\theta)$) with data likelihood ($P(D|\theta)$) to get posterior distribution $P(\theta|D)$.
$$ P(\theta|D) = \frac{P(D|\theta) P(\theta)}{P(D)} \propto P(D|\theta) P(\theta) $$
Advantages:
-
Incorporates prior information.
-
Provides full probability distribution (uncertainty quantification).
-
Intuitive interpretation of credible intervals.
Disadvantages:
-
Choice of prior can be subjective.
-
Computationally intensive for complex models (MCMC methods).
-
Can be less efficient with large, informative datasets.
E. Resampling Methods
Non-parametric methods that use repeated samples from the observed data.
-
Bootstrapping:
-
Create many "bootstrap samples" by sampling with replacement from original data (same size).
-
Calculate statistic (e.g., mean) for each bootstrap sample.
-
Use distribution of bootstrap statistics to estimate standard error, confidence intervals.
-
-
Cross-Validation:
-
k-fold CV: Split data into k folds. Use k-1 folds for training, 1 for testing. Repeat k times. Average performance.
-
Purpose: Assess model generalization, tune hyperparameters, avoid overfitting.
-
IV. Data Exploration Techniques
A. Univariate Analysis
Explores a single variable.
-
Numerical: Histogram, Box Plot, Summary stats (mean, sd, IQR).
-
Categorical: Bar Chart, Frequency Table, Mode.
B. Bivariate Analysis
Explores relationship between two variables.
- Num-Num: Scatter Plot, Correlation Coefficient ($r$ or $\rho$).
$$ r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}} $$
-
Cat-Cat: Contingency Table, Stacked Bar Chart, Chi-Square Test.
-
Num-Cat: Box Plot (by category), Z-test/t-test (if two groups).
C. Multivariate Analysis
Explores relationships among multiple variables.
-
Principal Component Analysis (PCA): Dimensionality reduction. Transforms data to new orthogonal axes (Principal Components) capturing maximum variance. Used for feature extraction, noise reduction.
-
Cluster Analysis: Groups similar observations.
-
k-Means: Partitions data into k clusters by minimizing within-cluster sum of squares.
-
Hierarchical: Builds tree (dendrogram) of clusters.
-
[!TIP] Exam Link: "Multivariate analysis" questions often expect PCA and Cluster Analysis as key examples.
V. Data Management and Wrangling
A. Data Wrangling Process
-
Collection: Acquire data from sources (APIs, databases, files).
-
Cleaning: Handle missing values (impute/remove), correct errors, remove duplicates.
-
Transformation: Normalize/scales features, encode categories, create new features.
-
Integration: Merge/join datasets from different sources.
B. Data Management and Indexing
-
Database Concepts: Relational (SQL) vs. Non-Relational (NoSQL). Tables, keys (Primary, Foreign).
-
Indexing: Data structure (e.g., B-tree, Hash index) to speed up query retrieval. Trade-off: faster reads, slower writes, extra storage.
C. Data Analyst Ecosystem
| Category | Tools | Primary Use |
|---|---|---|
| Querying | SQL (MySQL, PostgreSQL) | Data extraction/aggregation from databases. |
| Programming | Python (Pandas, NumPy), R | Data manipulation, statistical analysis, modeling. |
| Visualization | Tableau, PowerBI, Matplotlib, Seaborn | Creating dashboards and plots. |
| Big Data | Spark, Hadoop (Hive) | Processing large-scale datasets. |
VI. Data Visualization
A. Principles and Types
-
Design Principles: Clarity, accuracy, efficiency. Minimize "chartjunk." Use appropriate color scales (sequential, diverging, categorical). Label axes, add titles.
-
Chart Selection Guide:
-
Comparison: Bar Chart, Column Chart.
-
Distribution: Histogram, Box Plot.
-
Relationship: Scatter Plot, Line Chart.
-
Composition: Pie Chart (few categories), Stacked Bar.
-
B. Creating Custom Visualizations
For complex datasets: Combine multiple plot types (e.g., scatter matrix, faceted plots). Use interactive libraries (Plotly) for drill-down. Focus on storytelling—highlight key insights.
C. Python Visualization Libraries
| Library | Strengths | Typical Use |
|---|---|---|
| Pandas | Quick, built-in plotting (.plot()). |
Fast exploratory plots from DataFrames. |
| Matplotlib | Highly customizable, low-level "canvas." | Creating publication-quality, complex figures. |
| Seaborn | Statistical, attractive defaults. Built on Matplotlib. | Visualizing distributions, relationships, categorical data. |
| plotnine (ggplot) | Grammar of Graphics paradigm. Declarative syntax. | Building plots layer-by-layer (data -> aesthetics -> geometry). |
D. PowerBI
-
PowerBI Desktop: Free application for data modeling, transformation (Power Query), and report creation.
-
PowerBI Service: Cloud-based SaaS for sharing, collaboration, scheduling refresh.
-
PowerBI Mobile: Apps for viewing dashboards on devices.
-
Key Tools/Features:
-
DAX (Data Analysis Expressions): Formula language for calculated columns/measures (e.g.,
Total Sales = SUM(Sales[Amount])). -
Query Editor (Power Query): GUI for data shaping/cleaning (similar to wrangling).
-
Dashboards: Single-page, tile-based, real-time visualizations.
-
VII. File Formats
Structured Data Formats
-
CSV (Comma-Separated Values): Plain text, rows/columns. Simple, universal. Delimiter issues, no schema.
-
JSON (JavaScript Object Notation): Hierarchical, key-value pairs. Flexible, web-friendly. More verbose than CSV.
-
XML (eXtensible Markup Language): Markup language with tags. Very verbose, schema-driven (XSD). Used in older systems.
Binary Formats
-
Store data in binary (not human-readable). Efficient for storage/access.
-
Examples: Pickle (Python objects), Parquet (columnar, efficient for analytics), HDF5 (hierarchical, large scientific data).
Image and Video Formats (Contextualized)
| Format | Type | Key Property | Use Case |
|---|---|---|---|
| JPEG | Lossy | High compression, quality loss. | Web images, photos. |
| PNG | Lossless | Supports transparency, larger files. | Graphics, logos, screenshots. |
| MP4 (H.264/HEVC) | Lossy Video | High compression, widely supported. | Streaming, storage. |
| AVI | Less Compressed | Older format, larger files, simple codec. | Archival, editing (less common now). |
| RAW | Unprocessed | Sensor data, no in-camera processing. | Professional photography, max quality. |
[!TIP] Context Link: In image/video processing, format choice affects compression artifacts (JPEG blocking), color depth (PNG 16-bit vs JPEG 8-bit), and editing workflow (RAW vs JPEG).
VIII. Big Data Processing
A. Big Data Concepts (4 V's)
-
Volume: Massive scale (TB, PB).
-
Velocity: Speed of data generation/ingestion (real-time streams).
-
Variety: Structured, unstructured (images, video), semi-structured (JSON).
-
Veracity: Uncertainty, quality, trustworthiness of data.
B. Hadoop Ecosystem
-
HDFS (Hadoop Distributed File System):
-
Architecture: Master-Slave. NameNode (master, metadata), DataNodes (slaves, store blocks).
-
Principle: Files split into 128MB blocks, replicated (default 3x) across DataNodes. Fault-tolerant.
-
DiagramCANVAS: Draw NameNode connected to multiple DataNodes in a rack. Show a file split into blocks (A1, A2, A3) distributed across 3 DataNodes with replicas.
-
-
Hive:
-
Data warehouse infrastructure on Hadoop.
-
Provides HiveQL (SQL-like query language).
-
Translates queries into MapReduce/Tez/Spark jobs. For batch processing of structured data.
-
C. Other Big Data Tools
-
Apache Spark: In-memory cluster computing. Faster than Hadoop MapReduce for iterative algorithms (ML). Modules: Spark SQL, MLlib, Streaming.
-
HBase: NoSQL, column-family store. Provides real-time, random read/write access to big data (on top of HDFS). Good for sparse data.
-
Apache Kafka: Distributed streaming platform. Handles high-velocity data feeds (e.g., video sensor streams).
[!TIP] Comparison: Use HDFS for storage. Use Hive for SQL-like batch queries on large structured data. Use Spark for faster, complex analytics/ML on the same data. Use Kafka for ingesting real-time data streams.