UNIT 2: DATA ANALYSIS & VISUALIZATION FOR IMAGE/VIDEO PROCESSING
I. FOUNDATIONS OF DATA ANALYSIS
A. Variables and Data Categorization
-
Variables: Characteristics or quantities that can be measured or counted. They are the building blocks of datasets.
-
Types of Variables:
-
Categorical (Qualitative): Represents categories or groups.
-
Nominal: No inherent order (e.g., Color: Red, Blue, Green).
-
Ordinal: Has a meaningful order but not fixed intervals (e.g., Rating: Poor, Good, Excellent).
-
-
Numerical (Quantitative): Represents measurable quantities.
-
Discrete: Countable, distinct values (e.g., Number of pixels in an image, Count of objects).
-
Continuous: Measurable, infinite possible values within a range (e.g., Pixel intensity values, Video frame rate).
-
-
[!TIP] Exam Focus: Be prepared to classify given examples (e.g., "Image resolution" is discrete numerical; "Pixel brightness" is continuous numerical).
B. Levels of Measurement
| Scale | Properties | Example (Image Context) | Permissible Operations |
|---|---|---|---|
| Nominal | Categories only, no order | Image file format (JPEG, PNG) | Count, Mode, Chi-square |
| Ordinal | Ordered categories, unequal intervals | Image quality rating (1-5 stars) | Median, Percentiles, Non-parametric tests |
| Interval | Ordered, equal intervals, no true zero | Temperature in Celsius, Date | Mean, Standard Deviation, Correlation |
| Ratio | All interval properties + true zero | Pixel intensity (0-255), Video duration, File size | All statistical operations, Geometric mean |
C. Measures of Central Tendency
- Mean (Arithmetic Average): Sum of all values divided by count. Sensitive to outliers.
$$ \bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_i $$
-
Median: Middle value in an ordered dataset. Robust to outliers.
-
Mode: Most frequently occurring value. Can be used for nominal data.
-
Comparative Analysis: For symmetric distributions, Mean ≈ Median. For skewed, Median is better. For categorical data, Mode is the only measure.
D. Measures of Dispersion (Location of Dispersions)
-
Range: Max - Min. Sensitive to outliers.
-
Variance (s²): Average of squared deviations from the mean.
$$ s^2 = \frac{1}{n-1} \sum_{i=1}^{n} (x_i - \bar{x})^2 $$
- Standard Deviation (s): Square root of variance. In same units as data.
$$ s = \sqrt{s^2} $$
-
Interquartile Range (IQR): Range of middle 50% of data (Q3 - Q1). Robust to outliers.
-
Coefficient of Variation (CV): Standardized measure of dispersion (s / Mean). Useful for comparing variability across datasets with different units/means.
$$ CV = \frac{s}{\bar{x}} \times 100\% $$
[!TIP] Common Pitfall: Remember to use n-1 (sample) for variance/standard deviation when estimating population parameters from a sample.
II. STATISTICAL INFERENCE & HYPOTHESIS TESTING
A. Statistical Inferences
-
Population: Entire group of interest.
-
Sample: Subset of the population.
-
Point Estimation: Single value estimate of a population parameter (e.g., sample mean $\bar{x}$ estimates population mean $\mu$).
-
Interval Estimation (Confidence Interval): Range of values likely to contain the population parameter.
$$ \text{CI} = \text{Point Estimate} \pm (\text{Critical Value} \times \text{Standard Error}) $$
*Example:* 95% CI for mean: $$\displaystyle \bar{x} \pm t_{\alpha/2, df} \cdot \frac{s}{\sqrt{n}} $$
B. Hypothesis Testing Framework
-
Null Hypothesis (H₀): Status quo, no effect, no difference (e.g., $$\displaystyle \mu = 20 $$).
-
Alternative Hypothesis (H₁): Research claim, effect exists (e.g., $$\displaystyle \mu > 20 $$).
-
Significance Level (α): Probability of rejecting H₀ when it is true (Type I Error). Common values: 0.05, 0.01.
-
Test Statistic: Calculated from sample data (e.g., t-statistic, χ²).
-
p-value: Probability of observing the test statistic (or more extreme) if H₀ is true.
- Decision: If p-value ≤ α, reject H₀ (result is statistically significant).
-
Type II Error (β): Failing to reject H₀ when it is false.
-
Power (1-β): Probability of correctly rejecting a false H₀.
C. Parametric Tests
1. t-test (Student's t-test)
-
Assumptions: Data is approximately normally distributed, samples are independent (for two-sample), equal variances (for standard two-sample).
-
One-sample t-test: Compares sample mean to a known/hypothesized population mean.
$$ t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} \quad \text{df} = n-1 $$
-
Two-sample t-test:
-
Independent (Unpaired): Two distinct groups.
-
Paired (Dependent): Same subjects measured twice (e.g., before/after processing).
-
$$ t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} \quad \text{df} \approx \text{Welch-Satterthwaite} $$
[!TIP] Exam Application (May 2024): Testing potato yield.
Given: $$\displaystyle \mu_0 = 20 $$, $$\displaystyle n=12 $$, $$\displaystyle X = [21.5, 24.5, ...] $$
Steps:
- Calculate $\bar{x}$ and $s$ from data.
- H₀: $$\displaystyle \mu = 20 $$ (no improvement), H₁: $$\displaystyle \mu > 20 $$ (better yield). One-tailed test.
- Compute $$\displaystyle t = \frac{\bar{x} - 20}{s/\sqrt{12}} $$.
- Find critical t-value for df=11, α=0.05 (one-tailed). $$\displaystyle t_{crit} \approx 1.796 $$.
- Compare calculated t to $$\displaystyle t_{crit} $$. If $$\displaystyle t_{calc} > t_{crit} $$, reject H₀ → yield is significantly better.
2. Chi-Square (χ²) Test
- Chi-Square Test for Independence: Tests association between two categorical variables in a contingency table.
$$ \chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}} \quad \text{df} = (r-1)(c-1) $$
$$\displaystyle O_{ij} $$ = Observed frequency, $$\displaystyle E_{ij} $$ = Expected frequency = $$\displaystyle \frac{(\text{row total}) \times (\text{col total})}{\text{grand total}} $$.
*Assumption:* Expected frequencies should generally be ≥5.
- Chi-Square Goodness-of-Fit Test: Tests if a single categorical variable follows a hypothesized distribution.
$$ \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} \quad \text{df} = k - 1 - m $$
$k$ = number of categories, $m$ = number of estimated parameters.
D. Maximum Likelihood Estimation (MLE)
-
Concept: Finds parameter values ($\theta$) that maximize the likelihood of observing the given sample data.
-
Likelihood Function (L): Probability (or probability density) of the observed data given parameters.
For i.i.d. data: $$\displaystyle L(\theta) = \prod_{i=1}^{n} f(x_i | \theta) $$.
-
Procedure:
-
Write the likelihood function for the assumed distribution (e.g., Normal, Poisson).
-
Take the natural logarithm (Log-Likelihood, $\ell$) for easier differentiation.
-
Differentiate $\ell$ with respect to $\theta$.
-
Set derivative to zero and solve for $$\displaystyle \hat{\theta}_{MLE} $$.
-
-
Properties: Consistent, asymptotically normal, but can be biased in small samples.
-
Example (Normal mean): For $$\displaystyle X \sim N(\mu, \sigma^2) $$, $$\displaystyle \hat{\mu}_{MLE} = \bar{x} $$ (same as sample mean).
E. Re-sampling Methods
-
Bootstrapping:
-
Goal: Estimate sampling distribution of a statistic (e.g., mean, median) and its uncertainty (standard error, CI).
-
Process: Repeatedly sample with replacement from the original sample to create many "bootstrap samples" (size n). Calculate the statistic for each. The distribution of these bootstrap statistics approximates the true sampling distribution.
-
-
Cross-Validation:
-
Goal: Assess model performance and prevent overfitting.
-
Process (k-fold): Split data into k subsets (folds). Train model on k-1 folds, test on the held-out fold. Repeat k times. Average performance metric (e.g., accuracy, MSE) across folds.
-
III. ADVANCED STATISTICAL MODELING
A. Regression Analysis
1. Simple Linear Regression
-
Model: $$\displaystyle Y = \beta_0 + \beta_1 X + \epsilon $$
-
$Y$: Dependent (Response) variable.
-
$X$: Independent (Predictor) variable.
-
$$\displaystyle \beta_0 $$: Intercept (value of Y when X=0).
-
$$\displaystyle \beta_1 $$: Slope (change in Y per unit change in X).
-
$\epsilon$: Random error.
-
-
Least Squares Estimation: Minimizes sum of squared residuals (observed - predicted).
$$ \hat{\beta}_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2}, \quad \hat{\beta}_0 = \bar{y} - \hat{\beta}_1\bar{x} $$
- R-squared (R²): Proportion of variance in Y explained by X.
$$ R^2 = 1 - \frac{SS_{res}}{SS_{tot}} \in [0, 1] $$
2. Multiple Linear Regression
-
Model: $$\displaystyle Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + ... + \beta_p X_p + \epsilon $$
-
Interpretation: $$\displaystyle \beta_j $$ is the effect of $$\displaystyle X_j $$ on Y holding all other X's constant.
-
Adjusted R-squared: Penalizes adding irrelevant predictors. Better for model comparison than R².
-
Dummy Variables: Convert categorical predictors into binary (0/1) variables. One category is the reference baseline.
B. Multivariate Analysis
-
Definition: Simultaneous analysis of more than two variables to understand relationships and structure.
-
Key Techniques:
-
Principal Component Analysis (PCA): Dimensionality reduction. Creates new uncorrelated variables (Principal Components) that capture maximum variance.
-
Cluster Analysis: Groups observations into clusters based on similarity (e.g., image segmentation).
-
Multivariate Regression: Regression with multiple dependent variables.
-
C. Bayesian Modelling
-
Paradigm: Parameters are random variables with probability distributions (beliefs). Updates prior belief with data (likelihood) to get posterior belief.
-
Bayes' Theorem:
$$ P(\theta | \text{data}) = \frac{P(\text{data} | \theta) \cdot P(\theta)}{P(\text{data})} $$
* **Posterior** $\propto$ **Likelihood** $\times$ **Prior**
* $P(\theta)$: Prior distribution (initial belief about $\theta$).
* $P(\text{data} | \theta)$: Likelihood (probability of data given $\theta$).
* $P(\theta | \text{data})$: Posterior distribution (updated belief after seeing data).
-
Advantages:
-
Incorporates prior knowledge/expert opinion.
-
Provides full probability distribution for parameters (not just point estimates).
-
Intuitive interpretation of credible intervals.
-
-
Disadvantages:
-
Choice of prior can be subjective.
-
Computationally intensive for complex models (MCMC methods often needed).
-
Can be less familiar to those trained in frequentist statistics.
-
IV. DATA MANAGEMENT, WRANGLING & BIG DATA TOOLS
A. Data Wrangling (Data Munging)
The process of cleaning, structuring, and enriching raw data into a desired format for analysis. Detailed Process:
-
Data Gathering: Acquiring data from various sources (databases, APIs, files, web scraping).
-
Data Cleaning:
-
Missing Values: Identify (isnull), handle (delete rows/columns, impute with mean/median/mode, or use predictive models).
-
Outliers: Detect (IQR method, Z-score) and decide (remove, transform, or keep).
-
Inconsistencies: Fix formatting (dates, strings), correct errors, handle duplicates.
-
-
Data Transformation:
-
Normalization/Standardization: Scale numeric features (Min-Max, Z-score).
-
Encoding: Convert categorical variables to numeric (Label Encoding, One-Hot Encoding).
-
Aggregation: Summarize data (groupby, pivot tables).
-
-
Data Enrichment: Add new relevant features from existing ones (e.g., derive "aspect ratio" from image width/height).
-
Data Validation: Ensure data quality, consistency, and integrity after transformations.
B. File Formats for Data Storage
| Format | Type | Structure | Pros | Cons | Best For |
|---|---|---|---|---|---|
| CSV | Text | Tabular, rows/columns | Human-readable, simple, universal | No schema, inefficient for large data, no data types | Small-medium tabular data exchange |
| JSON | Text | Hierarchical (key-value) | Flexible, supports nested data, web-friendly | Verbose, slower to parse than binary | Web APIs, configuration, semi-structured data |
| Parquet | Binary | Columnar storage | Highly compressed, fast query (column pruning), efficient for analytics | Not human-readable, write once read many | Big Data analytics (Spark, Hive), large datasets |
| HDF5 | Binary | Hierarchical (groups/datasets) | Fast I/O for large numerical arrays, supports metadata | Complex API, less portable than Parquet | Scientific data (images, simulations, large arrays) |
C. Big Data Processing Ecosystem
1. Hadoop Distributed File System (HDFS)
-
Architecture:
-
NameNode: Master server. Manages file system namespace and metadata (file->block mapping). Single Point of Failure (HA setups exist).
-
DataNode: Slave servers. Store actual data blocks. Report status to NameNode.
-
-
Key Features:
-
Fault Tolerance: Data replicated (default 3x) across multiple DataNodes.
-
Large Blocks: Default 128MB/256MB (vs. 4KB in local FS). Optimized for large sequential reads.
-
Write Once, Read Many (WORM): Files are immutable after write.
-
2. Apache Hive
-
Definition: Data warehouse infrastructure built on Hadoop.
-
Purpose: Provides SQL-like querying (HiveQL) for data summarization, analysis, and ETL on massive datasets stored in HDFS.
-
Components:
-
HiveQL: Compiles queries into MapReduce/Tez/Spark jobs.
-
Metastore: Central repository storing table schema, partition info, and location of data in HDFS.
-
-
Use Case: Batch processing of very large, structured/semi-structured data where latency of seconds/minutes is acceptable.
3. Other Tools Overview
-
Apache Spark: In-memory cluster computing. Much faster than MapReduce (Hive's original engine) for iterative algorithms and interactive queries. Core components: Spark SQL (structured data), MLlib (machine learning), Spark Streaming.
-
NoSQL Databases (e.g., HBase): Provide low-latency, random read/write access to large datasets on top of HDFS. Schema-flexible. Used for real-time queries.
D. Data Management and Indexing
-
Database Concepts:
-
Table: Collection of related data in rows (records) and columns (fields).
-
Primary Key: Unique identifier for each row in a table.
-
Foreign Key: Field in one table that links to the primary key of another table (enforces referential integrity).
-
-
Indexing:
-
Purpose: Dramatically speed up data retrieval (SELECT, WHERE, JOIN) by creating a separate, optimized data structure (like a book's index).
-
Trade-off: Accelerates reads but slows down writes (INSERT/UPDATE/DELETE) as index must be maintained.
-
-
Types of Indexes:
-
Primary Index: Built on primary key. Often clustered (table rows physically ordered by key). One per table.
-
Secondary (Non-Clustered) Index: Built on non-key columns. Points to data rows (can be multiple). Can be clustered or non-clustered.
-
V. DATA VISUALIZATION PRINCIPLES & TOOLS
A. Data Visualization Fundamentals
-
Goals: Exploration (find patterns), Communication (tell a story), Insight Discovery.
-
Exploration by Dimensionality:
-
Univariate (1 variable): Distribution of a single variable.
-
Plots: Histogram, Box Plot, Violin Plot, KDE Plot.
-
Insights: Central tendency, spread, skewness, outliers.
-
-
Bivariate (2 variables): Relationship between two variables.
-
Numeric vs. Numeric: Scatter Plot, Line Chart (time series).
-
Categorical vs. Numeric: Box Plot, Violin Plot, Bar Chart (aggregated).
-
Categorical vs. Categorical:* Heatmap (from contingency table), Stacked Bar Chart.
-
Insights: Correlation, trends, group differences.
-
-
Multivariate (>2 variables): Explore relationships among multiple variables.
-
Encodings: Use Color, Size, Shape to represent additional variables on a bivariate plot.
-
Faceting (Small Multiples): Create multiple sub-plots for different subsets (e.g., by category).
-
Pair Plots: Matrix of scatterplots for all numeric variable pairs.
-
Insights: Interaction effects, complex patterns.
-
-
B. Python Visualization Libraries
| Library | Purpose/Strength | Key Features | Typical Use Case |
|---|---|---|---|
| Pandas | Quick, integrated plotting | .plot() method on DataFrames/Series. Simple line, bar, hist, scatter. |
Initial data exploration directly from DataFrame. |
| Matplotlib | Foundational, full control | Low-level "pyplot" API. Explicit control over Figure (overall window) and Axes (individual plot). Highly customizable. | Creating publication-quality static plots, complex multi-plot figures. |
| Seaborn | Statistical, attractive defaults | High-level interface. Built on Matplotlib. Beautiful styles, color palettes. Simplifies complex plots: distplot, catplot, relplot, lmplot. |
Statistical visualization (distributions, relationships, categorical data). |
| Plotly | Interactive, web-based | Creates interactive, web-ready plots (zoom, pan, hover). plotly.express for simple syntax, graph_objects for complex. |
Dashboards, reports needing interactivity, sharing online. |
ggplot (via plotnine) |
Grammar of Graphics | Implements R's ggplot2 in Python. Layered approach: ggplot(data) + geom_point(aes(x,y)) + .... |
Users familiar with ggplot2, declarative plot building. |
C. Power BI Ecosystem
-
Definition: Microsoft's suite of business analytics tools for data visualization and interactive reporting.
-
Components:
-
Power BI Desktop: Free application for report authoring. Connect to data sources, model data (Power Query, DAX), create visualizations and reports (.pbix files).
-
Power BI Service: Cloud-based (SaaS) platform. Publish, share, and collaborate on reports. Set up data refresh schedules, create dashboards, manage security.
-
Power BI Mobile: Apps for iOS/Android/Windows to consume reports and dashboards on the go.
-
Power BI Report Server: On-premises server for hosting and managing reports internally (for organizations with data residency/compliance needs). Reports developed in Desktop can be published here.
-
D. Creating Custom Visualizations
Process for Complex Datasets:
-
Understand the Data & Question: What story needs to be told? What are the key variables and relationships?
-
Choose Appropriate Encodings: Map variables to visual properties (position, length, color, size, shape) based on data type and perceptual effectiveness.
-
Iterative Design & Prototyping: Start with a simple plot (e.g., scatter). Add encodings (color for category, size for magnitude). Use faceting if needed. Refine for clarity.
-
Leverage Libraries: Use Matplotlib/Seaborn for static custom plots (subclassing
matplotlib.axes.Axes). Use Plotly for interactive custom charts. Combine libraries (e.g., Seaborn style with Matplotlib customization). -
Apply Design Principles:
-
Clarity: Label axes, use legends, avoid clutter.
-
Accuracy: Represent proportions correctly (e.g., bar charts from zero).
-
Efficiency: Minimize "chartjunk," use color meaningfully (sequential for magnitude, diverging for deviation, categorical for groups).
-
VI. DATA ANALYST ECOSYSTEM & SYNTHESIS
A. The Data Analyst Ecosystem
A modern workflow integrates multiple tools:
-
Storage: HDFS (for massive raw data), Relational/NoSQL databases.
-
Processing & Wrangling: Apache Spark (fast processing), Hive/Spark SQL (querying), Python (Pandas) for medium-data wrangling.
-
Analysis & Modeling: Python/R (libraries: statsmodels, scikit-learn, PyMC3 for Bayesian).
-
Visualization & Reporting: Power BI/Tableau (business dashboards), Python libs (Matplotlib/Seaborn/Plotly for custom analysis), Jupyter Notebooks for exploratory analysis.
-
Orchestration: Workflow tools (Apache Airflow) to automate pipelines.
B. Application-Oriented Synthesis
-
Connecting Tests to Questions:
-
"Is the average compression ratio of our new algorithm (sample) significantly different from the old one (known value)?" → One-sample t-test.
-
"Is there an association between video resolution (HD/4K) and user engagement (High/Low)?" → Chi-Square Test for Independence.
-
"How do file size, frame rate, and resolution predict video streaming buffering events?" → Multiple Linear Regression.
-
-
Choosing Visualizations:
-
See distribution of image file sizes? → Histogram or Box Plot (Univariate).
-
See relationship between image resolution and file size? → Scatter Plot (Bivariate).
-
Compare resolution distributions across 3 video codecs? → Faceted Box Plots or Overlaid KDEs (Multivariate via faceting/color).
-
-
End-to-End Pipeline:
Raw Logs/Images→ Wrangle (clean, transform with Pandas/Spark) → Analyze (compute stats, run t-test/regression with Python) → Visualize (create interactive Power BI dashboard or Seaborn report) → Insight/Decision.