Skip to content
AL-603 (A) · Image and Video Processing/Quick Revision Short Notes

Image and Video Processing (AL-603 (A)) - Unit 1 Short Notes

UNIT 1: FOUNDATIONS OF DATA ANALYSIS & VISUALIZATION


I. FOUNDATIONAL STATISTICAL CONCEPTS

Measures of Central Tendency

  • Mean (Arithmetic Average): Sum of all values divided by count. Sensitive to outliers.

$$ \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$

  • Median: Middle value in ordered data. Robust to outliers.

  • Mode: Most frequent value. Can be multi-modal.

[!TIP] For skewed distributions, median is a better measure than mean.

Measures of Dispersion (Spread) & Location

  • Range: Max - Min. Sensitive to outliers.

  • Variance (s²): Average squared deviation from mean.

$$ s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} $$

  • Standard Deviation (s): Square root of variance. Same units as data.

  • Interquartile Range (IQR): Q3 - Q1. Measures spread of middle 50%. Robust.

  • Five-Number Summary: Min, Q1, Median, Q3, Max.

  • Quartiles/Percentiles: Values dividing data into 4/100 equal parts.

Levels of Measurement (Scales of Data)

Scale Characteristics Permissible Operations Example
Nominal Categories only, no order Count, mode, frequency Gender, Color
Ordinal Ordered categories All nominal + median, percentiles Likert scale, Ranks
Interval Ordered, equal intervals, no true zero All ordinal + mean, SD Temperature (°C)
Ratio Interval + absolute zero All operations, geometric mean Height, Weight, Income

Variables and Data Categorization

  • By Role: Independent (predictor) vs. Dependent (outcome).

  • By Type:

    • Qualitative/Categorical: Nominal/Ordinal.

    • Quantitative/Numerical: Discrete (countable) vs. Continuous (measurable).


II. INFERENTIAL STATISTICS & HYPOTHESIS TESTING

Statistical Inferences

  • Core Concept: Drawing conclusions about a population using sample data.

  • Key Components:

    • Population: Entire group of interest.

    • Sample: Subset of population.

    • Parameter: Population value (e.g., μ, σ).

    • Statistic: Sample value (e.g., x̄, s).

    • Sampling Distribution: Distribution of a statistic over many samples.

t-Test (Student's t-test)

  • One-Sample t-test: Tests if sample mean differs from a known/hypothesized population mean (μ₀).

  • Hypotheses:

    • H₀: μ = μ₀ (no difference)

    • H₁: μ ≠ μ₀ (two-tailed) or μ > μ₀ / μ < μ₀ (one-tailed)

  • Assumptions: Random sample, approximate normality (or n>30), unknown population σ.

  • Test Statistic:

$$ t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} \quad \text{df} = n-1 $$

  • Decision: Compare |t| to critical t-value or use p-value (p < α → reject H₀).

[!EXAMPLE] Potato Yield t-Test (May 2024)

Given: μ₀ = 20, X = [21.5, 24.5, 18.5, 17.2, 14.5, 23.2, 22.1, 20.5, 19.4, 18.1, 24.1, 18.5], n=12.

  1. Calculate x̄ = 19.825, s ≈ 3.15.
  1. H₀: μ = 20; H₁: μ > 20 (one-tailed, "significantly better").
  1. t = (19.825 - 20) / (3.15/√12) ≈ -0.192.
  1. df = 11. Critical t(0.05, 11) ≈ 1.796.
  1. Since -0.192 < 1.796, fail to reject H₀. No significant evidence yield is better.

Chi-Square (χ²) Test

  • Test for Independence (Association): Tests relationship between two categorical variables using a contingency table.

    • H₀: Variables are independent.

    • H₁: Variables are associated.

    • Expected Frequency: E = (row total × column total) / grand total.

    • Test Statistic:

$$ \chi^2 = \sum \frac{(O - E)^2}{E} \quad \text{df} = (r-1)(c-1) $$

*   **Assumption:** All E ≥ 5 (for 2x2, all E ≥ 5; larger tables, 80% E≥5, none <1).
  • Goodness-of-Fit Test: Tests if observed frequencies fit a hypothesized distribution.

    • H₀: Data fits the distribution.

    • df = (k - 1 - c) where k = categories, c = estimated parameters.

[!EXAMPLE] Chi-Square Calculation Steps:

  1. State H₀ and H₁.
  1. Create contingency table with observed (O) counts.
  1. Calculate expected (E) counts for each cell.
  1. Compute χ² = Σ(O-E)²/E.
  1. Find df and critical χ² value.
  1. Conclude: if χ²_calc > χ²_crit, reject H₀.

Maximum Likelihood Estimation (MLE)

  • Concept: Find parameter values (θ) that maximize the likelihood of observing the given sample data.

  • Likelihood Function (L(θ)): Probability of data given parameters. For independent observations: L(θ) = Π P(x_i|θ).

  • Procedure:

    1. Write likelihood function.

    2. Take log (log-likelihood, easier to maximize).

    3. Differentiate w.r.t. θ, set derivative = 0.

    4. Solve for θ (MLE estimator, θ̂).

  • Example (Bernoulli, p): Data: k successes, n-k failures.

    L(p) = p^k (1-p)^(n-k) → log L = k ln p + (n-k) ln(1-p).

    d/dp (log L) = k/p - (n-k)/(1-p) = 0 → θ̂ = k/n.

Re-sampling Methods (Briefly)

  • Bootstrapping: Repeatedly sample with replacement from the original sample to estimate sampling distribution.

  • Cross-Validation: Partition data into training/validation sets to assess model performance (e.g., k-fold CV).


III. REGRESSION & MULTIVARIATE ANALYSIS

Regression Analysis

  • Simple Linear Regression (SLR):

    • Model: Y = β₀ + β₁X + ε

    • β₁ (Slope): Change in Y per unit change in X.

    • β₀ (Intercept): Value of Y when X=0.

    • Estimation (Least Squares): Minimize Σ(y_i - ŷ_i)².

    • R-squared (R²): Proportion of variance in Y explained by X. 0 ≤ R² ≤ 1.

  • Multiple Linear Regression (MLR):

    • Model: Y = β₀ + β₁X₁ + β₂X₂ + ... + βₖXₖ + ε

    • Interpretation: βⱼ is effect of Xⱼ on Y holding other X's constant.

    • Assumptions: Linearity, Independence, Homoscedasticity, Normality of errors, No multicollinearity.

  • Variables in Regression Modeling:

    • Dummy Variables: Encode categorical predictors (e.g., 0/1 for "Male"/"Female").

    • Interaction Terms: X₁*X₂ to model effect modification.

Multivariate Analysis

  • Objective: Simultaneously analyze relationships among multiple variables (p > 1).

  • Key Techniques:

    • Principal Component Analysis (PCA): Dimensionality reduction. Finds orthogonal axes (PCs) of maximum variance.

    • Factor Analysis: Models observed variables as linear combinations of latent factors.

    • Cluster Analysis: Groups observations into clusters based on similarity (e.g., K-means).

Multivariate Exploration of Data

  • Visualization Techniques:

    • Scatterplot Matrix: Grid of scatterplots for all variable pairs.

    • Pair Plots: Similar to scatterplot matrix, often with histograms on diagonal.

    • Correlation Heatmap: Visual matrix of correlation coefficients (color-coded).

    • 3D Plots / Parallel Coordinates: For >3 variables.

[!TIP] Use color/shape/size in scatterplots to encode additional categorical/quantitative variables.


IV. BAYESIAN STATISTICS

Bayesian Modelling

  • Core Philosophy: Parameters are random variables with probability distributions.

  • Bayes' Theorem:

$$ P(\theta|data) = \frac{P(data|\theta) \cdot P(\theta)}{P(data)} $$

*   **Prior P(θ):** Belief about θ before seeing data.

*   **Likelihood P(data|θ):** Probability of data given θ.

*   **Posterior P(θ|data):** Updated belief after seeing data.

*   **Evidence P(data):** Normalizing constant (often hard to compute).
  • Process: Start with Prior → Collect Data → Update to Posterior → (Posterior becomes new Prior for next update).

  • Contrast with Frequentist: Frequentist treats θ as fixed, uses confidence intervals; Bayesian gives probability distributions for θ.

  • Advantages:

    • Incorporates prior knowledge/expert opinion.

    • Intuitive probabilistic interpretation of parameters.

    • Naturally handles uncertainty.

  • Disadvantages:

    • Subjectivity of prior choice can influence results.

    • Computationally intensive for complex models (MCMC methods often needed).


V. DATA MANAGEMENT, WRANGLING & FORMATS

Data Wrangling (Data Munging)

  • Process:

    1. Data Gathering: Acquire from databases, APIs, files, web scraping.

    2. Data Cleaning: Handle missing values (impute/drop), correct errors, remove duplicates, fix inconsistencies.

    3. Data Transformation: Normalize/standardize, aggregate (groupby), reshape (pivot/melt), encode categoricals.

    4. Data Enrichment: Add new features/variables from external sources.

  • Tools: Pandas (Python), dplyr (R), SQL.

Types of File Formats

Format Type Structure Pros Cons Use Case
CSV/TSV Structured Plain text, delimited Human-readable, universal No schema, inefficient for large data Simple tabular data exchange
JSON Semi-structured Key-value pairs, nested Flexible, hierarchical, web-friendly Verbose, slower parse APIs, config files, NoSQL
XML Semi-structured Tagged hierarchy Self-describing, validated (XSD) Verbose, complex Legacy systems, documents
Parquet Binary/Columnar Column-oriented storage High compression, fast column queries Not human-readable Big Data (Hadoop/Spark), analytics
Avro Binary/Row-based Row-oriented, schema embedded Compact, fast serialization Less optimized for column reads Streaming, messaging (Kafka)

Big Data & Processing Tools

  • Big Data Characteristics (4 V's): Volume, Velocity, Variety, Veracity.

  • Hadoop Ecosystem:

    • HDFS (Hadoop Distributed File System):

      • Architecture: Master/Slave. NameNode (metadata, namespace) + DataNodes (block storage).

      • Blocks: Files split into 128MB/256MB blocks (default), replicated (default=3) across DataNodes.

      • Fault Tolerance: If a DataNode fails, NameNode re-replicates blocks from other copies.

      • DiagramSEARCH: "HDFS architecture diagram NameNode DataNode"
    • Hive:

      • Data warehouse infrastructure on top of HDFS.

      • HiveQL: SQL-like query language (translates to MapReduce/Tez/Spark jobs).

      • Metastore: Stores table schema and partition info (usually in RDBMS).

  • Context: Apache Spark (in-memory processing, faster than MapReduce), NoSQL databases (HBase, Cassandra for non-relational data).


VI. DATA VISUALIZATION PRINCIPLES & TOOLS

Data Visualization Fundamentals

  • Purpose: Exploration (find patterns), Communication (tell story), Discovery.

  • Univariate Exploration (1 variable):

    • Continuous: Histogram, Density plot, Box plot.

    • Categorical: Bar chart, Pie chart (use sparingly), Frequency table.

  • Bivariate Exploration (2 variables):

    • Continuous-Continuous: Scatter plot, Hexbin plot, Line chart (time series).

    • Categorical-Continuous: Box plot (by category), Violin plot.

    • Categorical-Categorical: Grouped bar chart, Mosaic plot, Heatmap (contingency).

  • Multivariate Exploration (>2 variables): Use aesthetics (color, size, shape, facets) in bivariate plots, or use techniques like scatterplot matrix, parallel coordinates.

Python Visualization Libraries

Library Key Features Best For Level
Pandas .plot() method, quick & dirty Fast exploratory plots from DataFrames Basic
Matplotlib Foundational, full control, object-oriented Custom, publication-quality figures Low-level
Seaborn Statistical plots, beautiful defaults, built on Matplotlib Complex stats viz (regression plots, distributions, categorical) High-level
Plotly Interactive, web-based, hover tooltips Dashboards, interactive web apps Interactive
ggplot (plotnine) Grammar of Graphics (R's ggplot2 port) Layered, declarative plotting Grammar-based

[!TIP] Seaborn is ideal for statistical exploration; Plotly for interactive reports.

Power BI Ecosystem

  • Power BI Desktop: Free application for report authoring. Connect to data, model (DAX), create visuals, design reports.

  • Power BI Service: Cloud platform (SaaS). Publish, share, collaborate, schedule refreshes, create dashboards (single-page, real-time).

  • Power BI Mobile: App for consuming reports/dashboards on phones/tablets.

  • Power BI Report Server: On-premises server for hosting reports (for organizations with strict data governance).

  • Core Components:

    • Datasets: Underlying data model (can be shared).

    • Reports: Multi-page collection of visuals.

    • Dashboards: Single-page, tile-based, pinned from reports/datasets.

    • Apps: Packaged reports/dashboards for distribution.

Creating Custom Visualizations for Complex Datasets

  1. Understand the Story: What question are you answering? What insight is hidden?

  2. Choose Chart Type: Go beyond standard (bar, line, scatter). Consider: Sankey diagram (flows), Network graph (relationships), Heatmap with dendrogram (clustering), Geospatial map, Treemap (hierarchy).

  3. Design for Clarity: Use appropriate color scales (sequential, diverging, categorical), clear labels, annotations, avoid clutter (chartjunk).

  4. Iterative Refinement: Prototype, get feedback, simplify, ensure accessibility (colorblind-friendly).


VII. DATA ANALYST ECOSYSTEM & INDEXING

Data Analyst Ecosystem

  • Overview: Integrated workflow and tools:

    Data Acquisition (SQL, APIs) → Storage (DB, Data Lake) → Processing (Python/R, Spark) → Analysis (Stats, ML) → Visualization (Power BI, Tableau) → Decision Support.

  • Key Roles/Tools Integration: SQL for extraction, Python/R for analysis/ML, Power BI/Tableau for dashboards, Big Data platforms (Hadoop/Spark) for large-scale processing.

Data Management and Indexing

  • Data Management:

    • DBMS (Database Management System): Software to manage databases (e.g., PostgreSQL, MySQL).

    • Data Modeling: Conceptual (ER diagram), Logical (tables, columns), Physical (indexes, partitions).

    • Data Integrity: Accuracy & consistency (constraints: primary key, foreign key, unique, not null).

    • ACID Properties: Atomicity, Consistency, Isolation, Durability (for transactional reliability).

  • Indexing:

    • Purpose: Speed up data retrieval (SELECT, WHERE, JOIN) by creating a lookup structure (like a book index).

    • Trade-off: Faster reads, slower writes (INSERT/UPDATE/DELETE), extra storage.

    • Common Types:

      • B-tree: Balanced tree, sorted, good for range queries and equality. Most common.

      • Hash: Key-value lookup, very fast for equality, not for range.

    • How it works: Index stores column value + pointer to row. DB engine searches index first to find pointers, then fetches rows.

    [!TIP] Index foreign keys and frequently queried columns. Avoid over-indexing on frequently updated tables.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in