UNIT 4: Data Analytics & Visualization (for Image/Video Processing Context)
I. Foundational Statistical Concepts
A. Measures of Central Tendency
-
Definition: Single values that describe the center of a data distribution.
-
Mean (Arithmetic Average): Sum of all values divided by count.
$$ \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$
-
Median: Middle value in an ordered dataset. For even
n, median = average of two middle values. -
Mode: Most frequently occurring value(s). Can be unimodal, bimodal, etc.
-
[!TIP] Exam Tip: For skewed distributions (common in image intensity histograms), median is a better central measure than mean as it is robust to outliers.
B. Measures of Dispersion/Location
-
Range:
Max - Min. Highly sensitive to outliers. -
Variance (
s²): Average of squared deviations from the mean.
$$ s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} \quad \text{(Sample Variance)} $$
- Standard Deviation (
s): Square root of variance. Same units as data.
$$ s = \sqrt{s^2} $$
- Interquartile Range (IQR): Spread of middle 50% of data.
IQR = Q3 - Q1. Used in box plots to identify outliers (typically, outliers < Q1 - 1.5IQR or > Q3 + 1.5IQR).
C. Levels of Measurement
-
Nominal: Categorical labels only (e.g., image class: "cat", "dog"). No order.
-
Ordinal: Categorical with order but no consistent difference (e.g., image quality rating: "poor", "fair", "good").
-
Interval: Numerical with order and consistent difference, but no true zero (e.g., temperature in Celsius). Ratios meaningless.
-
Ratio: Numerical with order, consistent difference, and true zero (e.g., pixel intensity [0-255], image dimensions, duration). All mathematical operations valid.
D. Variables and Data Categorization
-
Variable: A characteristic/attribute that can take different values.
-
Categorization:
-
By Type: Categorical (Nominal/Ordinal) vs. Numerical (Interval/Ratio).
-
By Role in Analysis: Independent/Predictor (e.g., filter parameters) vs. Dependent/Response (e.g., processed image quality metric).
-
By Nature: Discrete (countable, e.g., number of objects) vs. Continuous (measurable, e.g., pixel coordinates).
-
II. Inferential Statistics & Hypothesis Testing
A. Statistical Inferences
-
Point Estimation: Using sample statistic to estimate population parameter (e.g., sample mean
x̄estimates population meanμ). -
Confidence Interval (CI): Range of values likely to contain the population parameter.
$$ \text{CI} = \text{Point Estimate} \pm (\text{Critical Value} \times \text{Standard Error}) $$
* For mean (large n/z-dist): `x̄ ± z*(σ/√n)`
* For mean (small n/t-dist): `x̄ ± t*(s/√n)`
B. Chi-Square Test (χ²)
- Goodness of Fit: Tests if observed frequency distribution fits an expected theoretical distribution.
$$ \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$
* `O_i` = Observed frequency, `E_i` = Expected frequency.
* **df = (number of categories) - 1**
- Test of Independence: Tests if two categorical variables are associated in a contingency table.
$$ \chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}} \quad \text{for all cells (i,j)} $$
* `E_{ij} = (row total_i * col total_j) / grand total`
* **df = (rows - 1) * (cols - 1)**
-
[!TIP] Assumption: Expected frequencies
E_ishould be ≥5 for >80% of cells.
C. t-test
- One-Sample t-test: Compares sample mean (
x̄) to a known population mean (μ₀).
$$ t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} \quad \text{df} = n-1 $$
> **Example (May 2024):** Testing if potato yield (`X`) is better than standard (`μ=20`). `H₀: μ = 20`, `H₁: μ > 20`. Calculate `x̄`, `s`, then `t`. Compare to critical t-value.
- Two-Sample Independent t-test: Compares means of two independent groups.
$$ t = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} \quad \text{df via Welch's approximation} $$
- Two-Sample Paired t-test: Compares means of two related samples (e.g., before/after processing).
$$ t = \frac{\bar{d}}{s_d / \sqrt{n}} \quad \text{df} = n-1 $$
where `d` = difference for each pair.
-
[!TIP] Key Assumption: Data should be approximately normally distributed (check with histogram/Q-Q plot) and samples should have equal variance for standard independent t-test (use Welch's t-test if variances unequal).
D. Resampling Methods
-
Bootstrap: Repeatedly samples with replacement from the original sample to estimate sampling distribution of a statistic (e.g., mean, median). Used for CI estimation when distribution is unknown.
-
Cross-Validation: Resampling method for model evaluation. Data is split into
kfolds; model trained onk-1folds and validated on the held-out fold, repeatedktimes. Common: k-fold CV (k=5 or 10).
III. Regression & Predictive Modeling
A. Regression Analysis
- Simple Linear Regression: Models relationship between one independent (
X) and one dependent (Y) variable.
$$ Y = \beta_0 + \beta_1 X + \epsilon $$
* `β₀` = intercept, `β₁` = slope, `ε` = error term.
* Estimated via **Ordinary Least Squares (OLS)**: Minimizes sum of squared residuals `Σ(y_i - ŷ_i)²`.
- Multiple Linear Regression: Extends to multiple independent variables (
X₁, X₂, ..., Xₖ).
$$ Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + ... + \beta_k X_k + \epsilon $$
-
[!TIP] Interpretation:
β₁represents the change inYfor a one-unit change inX₁, holding all other X constant.
B. Types of Variables in Regression Modeling
-
Dependent Variable (Y): The target variable being predicted (e.g., image sharpness score).
-
Independent/Predictor Variables (X): Features used for prediction (e.g., filter kernel size, compression ratio).
-
Dummy Variables: Binary (0/1) variables used to represent categorical predictors (e.g.,
Image_Type_Color=1if color,0if grayscale).kcategories requirek-1dummy variables to avoid multicollinearity (dummy variable trap).
C. Multivariate Analysis
-
Concept: Simultaneous analysis of more than one outcome/dependent variable. (Contrast with Multiple Regression, which has one dependent variable but multiple independents).
-
Applications in Image/Video:
-
Principal Component Analysis (PCA): Dimensionality reduction. Finds orthogonal axes (principal components) of maximum variance in multivariate data (e.g., pixel values across many images).
-
Cluster Analysis: Grouping similar objects (e.g., grouping similar image features, video scenes).
-
Multivariate Analysis of Variance (MANOVA): Extension of ANOVA for multiple dependent variables (e.g., comparing groups on multiple image quality metrics: PSNR, SSIM, VMAF).
-
-
[!TIP] Key Distinction: Multivariate = Multiple dependent variables. Multiple Regression = One dependent, multiple independent variables.
IV. Advanced Statistical Methods
A. Bayesian Modeling (Principles)
-
Core Idea: Treats parameters as random variables with probability distributions, updating prior beliefs with observed data to get posterior beliefs.
-
Bayes' Theorem:
$$ P(\theta|D) = \frac{P(D|\theta) \cdot P(\theta)}{P(D)} $$
* `P(θ|D)` = **Posterior** (updated belief about parameter θ given data D).
* `P(D|θ)` = **Likelihood** (probability of data given θ).
* `P(θ)` = **Prior** (initial belief about θ).
* `P(D)` = Marginal likelihood (normalizing constant).
-
Advantages: Incorporates prior knowledge, provides full probability distribution for parameters (not just point estimates), intuitive uncertainty quantification.
-
Disadvantages: Choice of prior can be subjective, computationally intensive for complex models (MCMC methods).
B. Maximum Likelihood Estimation (MLE)
-
Goal: Find parameter values (
θ) that maximize the likelihoodP(D|θ)—i.e., make the observed data most probable. -
Steps:
-
Write down the likelihood function
L(θ) = P(D|θ). Often use log-likelihoodℓ(θ) = log L(θ)for easier computation. -
Differentiate
ℓ(θ)w.r.t.θ. -
Set derivative = 0 and solve for
θ.
-
-
Example (Bernoulli/Binomial): For
ncoin flips withkheads,L(p) = p^k (1-p)^(n-k). MLE forpisp̂ = k/n. -
For Linear Regression: MLE under normal error assumption yields the same OLS estimates.
-
[!TIP] MLE vs. Bayesian: MLE gives a single "best" parameter value. Bayesian gives a distribution. MLE is a special case of Bayesian with a uniform (flat) prior.
V. Data Handling, Wrangling & Formats
A. Data Wrangling Process
-
Data Collection: Gathering data from sources (sensors, logs, databases, web scraping).
-
Data Cleaning: Handling missing values (impute/remove), correcting errors, removing duplicates, ensuring consistency.
-
Data Transformation: Normalization/scaling, encoding categorical variables, creating new features (feature engineering), aggregating.
-
Data Enrichment: Adding relevant external data (e.g., adding geographical data to image metadata).
-
[!TIP] Rule of Thumb: Data wrangling often consumes 60-80% of a data analyst's time.
B. File Formats
-
Structured:
-
CSV (Comma-Separated Values): Plain text, rows & columns. Simple, universal. Delimiters can cause issues.
-
JSON (JavaScript Object Notation): Hierarchical, key-value pairs. Flexible for nested data (e.g., image metadata with nested tags).
-
-
Unstructured:
-
Images: JPEG, PNG, TIFF (lossless), RAW. Contain pixel arrays + metadata (EXIF).
-
Video: MP4 (H.264/AVC, H.265/HEVC), AVI, MKV. Contain encoded video/audio streams + containers.
-
-
Binary (Optimized for Big Data):
-
Parquet: Columnar storage. Highly efficient for querying specific columns (common in Spark/Hadoop).
-
Avro: Row-based, compact, good for serialization.
-
C. Pandas in Python
-
Core Structure:
DataFrame(2D labeled table) andSeries(1D labeled array). -
Key Operations:
-
pd.read_csv(),pd.read_json() -
df.head(),df.info(),df.describe() -
df[col],df.loc[],df.iloc[](indexing/selection) -
df.groupby(),df.merge(),df.pivot_table() -
Handling missing data:
df.isnull(),df.dropna(),df.fillna()
-
-
[!TIP] Performance: Vectorized operations (
df['col'].mean()) are vastly faster than Python loops (for index, row in df.iterrows()).
D. Data Management and Indexing
-
Concept: Organizing, storing, and retrieving data efficiently.
-
Database Indexing: Creating a separate data structure (index) to speed up querying on a column(s).
-
B-tree Index: Most common. Allows fast equality/range queries. Slows down
INSERT/UPDATE/DELETE. -
Hash Index: Fast for exact matches only.
-
Bitmap Index: Efficient for low-cardinality columns (e.g., categorical variables like image "scene_type").
-
-
[!TIP] Trade-off: Indexes improve read speed but add overhead for write operations and consume storage.
VI. Data Visualization & Exploration
A. Data Visualization Fundamentals
-
Univariate Exploration (Single Variable):
-
Histogram: Shows distribution of a continuous variable. Bins choice affects perception.
-
Box Plot (Whisker Plot): Shows 5-number summary (min, Q1, median, Q3, max) and outliers. Excellent for comparing distributions across categories.
DiagramCANVAS: A standard box plot with labeled median, quartiles, whiskers, and outliers marked as individual points beyond 1.5*IQR -
-
Bivariate Exploration (Two Variables):
-
Scatter Plot: Relationship between two continuous variables. Reveals correlation, clusters, non-linearity.
-
Heatmap: Visualizes matrix data, e.g., correlation matrix (
df.corr()). Color intensity represents value magnitude.
-
-
Multivariate Exploration (>2 Variables):
-
Pair Plot (Scatter Matrix): Grid of scatter plots for all pairwise combinations + histograms on diagonal. Use
seaborn.pairplot(). -
3D Plots: Can show 3 continuous variables (x,y,z) or use color/size for 4th/5th variables. Often interactive (Plotly).
-
Faceting (Small Multiples): Creating multiple plots of the same type for different subsets of data (e.g., scatter plot of
feature1vsfeature2for eachimage_category).
-
B. Python Visualization Libraries
-
Matplotlib: Foundation library. Low-level, highly customizable. Steeper learning curve.
-
Seaborn: Built on Matplotlib. Statistical visualizations. Simpler syntax, beautiful defaults. Great for
distplot,boxplot,heatmap,pairplot. -
Plotly: Interactive, web-based visualizations. Supports 3D, zooming, hover info. Good for dashboards.
-
Bokeh: Interactive visualizations for web apps. Similar to Plotly, focuses on large/streaming data.
C. ggplot2 in R (Grammar of Graphics)
-
Philosophy: Build plots layer by layer:
Data+Aesthetics (aes)+Geometries (geom_*)+Facets+Statistics+Coordinates+Theme. -
Example:
ggplot(data, aes(x=var1, y=var2)) + geom_point() + facet_wrap(~category) -
Strength: Consistent, logical syntax. Excellent for complex multi-plot figures.
D. Creating Custom Visualizations for Complex Datasets
-
Understand the Question: What relationship/story are you trying to show?
-
Choose the Right Chart Type: Match data structure (categorical vs. continuous) and relationship (comparison, distribution, composition, relationship).
-
Simplify & Declutter: Remove non-essential ink (chartjunk). Use color purposefully (qualitative, sequential, diverging palettes).
-
Leverage Interactivity: For high-dimensional data, use tools like Plotly to allow brushing, zooming, filtering.
-
Combine Views: Use dashboards (PowerBI, Dash) or faceted plots to show multiple linked perspectives.
VII. Big Data Ecosystem & Processing
A. Big Data (4 V's) & Tools
-
Characteristics:
-
Volume: Massive scale (TB/PB).
-
Velocity: High speed of data in/out (real-time streams from video feeds).
-
Variety: Structured, semi-structured, unstructured (images, video, logs).
-
Veracity: Uncertainty, quality, trustworthiness of data.
-
-
Processing Tools:
-
Hadoop: Framework for distributed storage & processing. Core: HDFS (storage) + MapReduce (processing).
-
Spark: Faster in-memory processing engine. Can run on Hadoop YARN or standalone. Better for iterative algorithms (e.g., ML) and streaming. APIs in Scala, Python (PySpark), Java, R.
-
B. HDFS (Hadoop Distributed File System)
-
Architecture:
-
NameNode: Master server. Manages file system namespace (metadata: file->block mapping). Single Point of Failure (HA solutions exist).
-
DataNode: Slave nodes. Store actual data blocks (default 128MB/256MB). Report status to NameNode.
-
Client: Interface to read/write data. Gets block locations from NameNode, then talks directly to DataNodes.
-
-
Operations: Files are broken into blocks, replicated (default 3x) across DataNodes for fault tolerance. Write-once-read-many (WORM) model.
-
DiagramSEARCH: "HDFS architecture diagram NameNode DataNode"
C. Hive (Data Warehousing on Hadoop)
-
Concept: Provides SQL-like interface (HiveQL) to query data stored in HDFS.
-
How it works: Hive translates HiveQL queries into MapReduce (or Tez/Spark) jobs.
-
Components:
-
Metastore: Stores table schema, partitions, location (in RDBMS like MySQL).
-
Driver: Compiles, optimizes, and executes queries.
-
Executors: Run the tasks (Map/Reduce).
-
-
Use Case: Batch processing of large, structured logs or tabular data (e.g., aggregated video view counts). Not for low-latency queries.
D. Data Analyst Ecosystem
-
Roles: Data Analyst, Business Analyst, Data Scientist, Data Engineer.
-
Workflow: Data Ingestion → Storage (Data Lake/Warehouse) → Wrangling (ETL/ELT) → Analysis/Modeling → Visualization → Dashboarding.
-
Tools: SQL (Hive, PostgreSQL), Python/R (Pandas, Scikit-learn), Visualization (PowerBI, Tableau), Big Data (Spark, Hive), Orchestration (Airflow).
VIII. Business Intelligence & Dashboarding
A. PowerBI
-
Components:
-
Power BI Desktop: Free Windows application for building reports & data modeling. Core features: Power Query (data transformation), Data Model (relationships, DAX measures), Report View (visualizations).
-
Power BI Service: Cloud-based SaaS service for publishing, sharing, collaboration. Sets up scheduled refreshes, creates dashboards (single-page pin-based summaries).
-
Power BI Mobile: Apps for iOS/Android to view dashboards/reports on-the-go.
-
-
Key Tools:
-
Power Query (M Language): Data connection & transformation engine. Used in Desktop for ETL.
-
DAX (Data Analysis Expressions): Formula language for creating custom measures and calculated columns in the data model.
-
Example Measure:
Total Sales = SUM(Sales[Amount]) -
Example Calculated Column:
Profit = [Sales] - [Cost]
-
-
Visualizations: Standard (bar, line, pie) + custom visuals from marketplace.
-
-
Workflow: Connect to data source → Transform in Power Query → Model data (create relationships) → Build report with visuals & DAX → Publish to Service → Share/Dashboard.
-
[!TIP] Best Practice: Perform heavy transformations in Power Query (during load) rather than with DAX (during query) for better performance.