UNIT 1: FOUNDATIONS OF DATA ANALYSIS & VISUALIZATION
I. FOUNDATIONAL STATISTICAL CONCEPTS
Measures of Central Tendency
- Mean (Arithmetic Average): Sum of all values divided by count. Sensitive to outliers.
$$ \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$
-
Median: Middle value in ordered data. Robust to outliers.
-
Mode: Most frequent value. Can be multi-modal.
[!TIP] For skewed distributions, median is a better measure than mean.
Measures of Dispersion (Spread) & Location
-
Range: Max - Min. Sensitive to outliers.
-
Variance (s²): Average squared deviation from mean.
$$ s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} $$
-
Standard Deviation (s): Square root of variance. Same units as data.
-
Interquartile Range (IQR): Q3 - Q1. Measures spread of middle 50%. Robust.
-
Five-Number Summary: Min, Q1, Median, Q3, Max.
-
Quartiles/Percentiles: Values dividing data into 4/100 equal parts.
Levels of Measurement (Scales of Data)
| Scale | Characteristics | Permissible Operations | Example |
|---|---|---|---|
| Nominal | Categories only, no order | Count, mode, frequency | Gender, Color |
| Ordinal | Ordered categories | All nominal + median, percentiles | Likert scale, Ranks |
| Interval | Ordered, equal intervals, no true zero | All ordinal + mean, SD | Temperature (°C) |
| Ratio | Interval + absolute zero | All operations, geometric mean | Height, Weight, Income |
Variables and Data Categorization
-
By Role: Independent (predictor) vs. Dependent (outcome).
-
By Type:
-
Qualitative/Categorical: Nominal/Ordinal.
-
Quantitative/Numerical: Discrete (countable) vs. Continuous (measurable).
-
II. INFERENTIAL STATISTICS & HYPOTHESIS TESTING
Statistical Inferences
-
Core Concept: Drawing conclusions about a population using sample data.
-
Key Components:
-
Population: Entire group of interest.
-
Sample: Subset of population.
-
Parameter: Population value (e.g., μ, σ).
-
Statistic: Sample value (e.g., x̄, s).
-
Sampling Distribution: Distribution of a statistic over many samples.
-
t-Test (Student's t-test)
-
One-Sample t-test: Tests if sample mean differs from a known/hypothesized population mean (μ₀).
-
Hypotheses:
-
H₀: μ = μ₀ (no difference)
-
H₁: μ ≠ μ₀ (two-tailed) or μ > μ₀ / μ < μ₀ (one-tailed)
-
-
Assumptions: Random sample, approximate normality (or n>30), unknown population σ.
-
Test Statistic:
$$ t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} \quad \text{df} = n-1 $$
- Decision: Compare |t| to critical t-value or use p-value (p < α → reject H₀).
[!EXAMPLE] Potato Yield t-Test (May 2024)
Given: μ₀ = 20, X = [21.5, 24.5, 18.5, 17.2, 14.5, 23.2, 22.1, 20.5, 19.4, 18.1, 24.1, 18.5], n=12.
- Calculate x̄ = 19.825, s ≈ 3.15.
- H₀: μ = 20; H₁: μ > 20 (one-tailed, "significantly better").
- t = (19.825 - 20) / (3.15/√12) ≈ -0.192.
- df = 11. Critical t(0.05, 11) ≈ 1.796.
- Since -0.192 < 1.796, fail to reject H₀. No significant evidence yield is better.
Chi-Square (χ²) Test
-
Test for Independence (Association): Tests relationship between two categorical variables using a contingency table.
-
H₀: Variables are independent.
-
H₁: Variables are associated.
-
Expected Frequency: E = (row total × column total) / grand total.
-
Test Statistic:
-
$$ \chi^2 = \sum \frac{(O - E)^2}{E} \quad \text{df} = (r-1)(c-1) $$
* **Assumption:** All E ≥ 5 (for 2x2, all E ≥ 5; larger tables, 80% E≥5, none <1).
-
Goodness-of-Fit Test: Tests if observed frequencies fit a hypothesized distribution.
-
H₀: Data fits the distribution.
-
df = (k - 1 - c) where k = categories, c = estimated parameters.
-
[!EXAMPLE] Chi-Square Calculation Steps:
- State H₀ and H₁.
- Create contingency table with observed (O) counts.
- Calculate expected (E) counts for each cell.
- Compute χ² = Σ(O-E)²/E.
- Find df and critical χ² value.
- Conclude: if χ²_calc > χ²_crit, reject H₀.
Maximum Likelihood Estimation (MLE)
-
Concept: Find parameter values (θ) that maximize the likelihood of observing the given sample data.
-
Likelihood Function (L(θ)): Probability of data given parameters. For independent observations: L(θ) = Π P(x_i|θ).
-
Procedure:
-
Write likelihood function.
-
Take log (log-likelihood, easier to maximize).
-
Differentiate w.r.t. θ, set derivative = 0.
-
Solve for θ (MLE estimator, θ̂).
-
-
Example (Bernoulli, p): Data: k successes, n-k failures.
L(p) = p^k (1-p)^(n-k) → log L = k ln p + (n-k) ln(1-p).
d/dp (log L) = k/p - (n-k)/(1-p) = 0 → θ̂ = k/n.
Re-sampling Methods (Briefly)
-
Bootstrapping: Repeatedly sample with replacement from the original sample to estimate sampling distribution.
-
Cross-Validation: Partition data into training/validation sets to assess model performance (e.g., k-fold CV).
III. REGRESSION & MULTIVARIATE ANALYSIS
Regression Analysis
-
Simple Linear Regression (SLR):
-
Model: Y = β₀ + β₁X + ε
-
β₁ (Slope): Change in Y per unit change in X.
-
β₀ (Intercept): Value of Y when X=0.
-
Estimation (Least Squares): Minimize Σ(y_i - ŷ_i)².
-
R-squared (R²): Proportion of variance in Y explained by X. 0 ≤ R² ≤ 1.
-
-
Multiple Linear Regression (MLR):
-
Model: Y = β₀ + β₁X₁ + β₂X₂ + ... + βₖXₖ + ε
-
Interpretation: βⱼ is effect of Xⱼ on Y holding other X's constant.
-
Assumptions: Linearity, Independence, Homoscedasticity, Normality of errors, No multicollinearity.
-
-
Variables in Regression Modeling:
-
Dummy Variables: Encode categorical predictors (e.g., 0/1 for "Male"/"Female").
-
Interaction Terms: X₁*X₂ to model effect modification.
-
Multivariate Analysis
-
Objective: Simultaneously analyze relationships among multiple variables (p > 1).
-
Key Techniques:
-
Principal Component Analysis (PCA): Dimensionality reduction. Finds orthogonal axes (PCs) of maximum variance.
-
Factor Analysis: Models observed variables as linear combinations of latent factors.
-
Cluster Analysis: Groups observations into clusters based on similarity (e.g., K-means).
-
Multivariate Exploration of Data
-
Visualization Techniques:
-
Scatterplot Matrix: Grid of scatterplots for all variable pairs.
-
Pair Plots: Similar to scatterplot matrix, often with histograms on diagonal.
-
Correlation Heatmap: Visual matrix of correlation coefficients (color-coded).
-
3D Plots / Parallel Coordinates: For >3 variables.
-
[!TIP] Use color/shape/size in scatterplots to encode additional categorical/quantitative variables.
IV. BAYESIAN STATISTICS
Bayesian Modelling
-
Core Philosophy: Parameters are random variables with probability distributions.
-
Bayes' Theorem:
$$ P(\theta|data) = \frac{P(data|\theta) \cdot P(\theta)}{P(data)} $$
* **Prior P(θ):** Belief about θ before seeing data.
* **Likelihood P(data|θ):** Probability of data given θ.
* **Posterior P(θ|data):** Updated belief after seeing data.
* **Evidence P(data):** Normalizing constant (often hard to compute).
-
Process: Start with Prior → Collect Data → Update to Posterior → (Posterior becomes new Prior for next update).
-
Contrast with Frequentist: Frequentist treats θ as fixed, uses confidence intervals; Bayesian gives probability distributions for θ.
-
Advantages:
-
Incorporates prior knowledge/expert opinion.
-
Intuitive probabilistic interpretation of parameters.
-
Naturally handles uncertainty.
-
-
Disadvantages:
-
Subjectivity of prior choice can influence results.
-
Computationally intensive for complex models (MCMC methods often needed).
-
V. DATA MANAGEMENT, WRANGLING & FORMATS
Data Wrangling (Data Munging)
-
Process:
-
Data Gathering: Acquire from databases, APIs, files, web scraping.
-
Data Cleaning: Handle missing values (impute/drop), correct errors, remove duplicates, fix inconsistencies.
-
Data Transformation: Normalize/standardize, aggregate (groupby), reshape (pivot/melt), encode categoricals.
-
Data Enrichment: Add new features/variables from external sources.
-
-
Tools: Pandas (Python), dplyr (R), SQL.
Types of File Formats
| Format | Type | Structure | Pros | Cons | Use Case |
|---|---|---|---|---|---|
| CSV/TSV | Structured | Plain text, delimited | Human-readable, universal | No schema, inefficient for large data | Simple tabular data exchange |
| JSON | Semi-structured | Key-value pairs, nested | Flexible, hierarchical, web-friendly | Verbose, slower parse | APIs, config files, NoSQL |
| XML | Semi-structured | Tagged hierarchy | Self-describing, validated (XSD) | Verbose, complex | Legacy systems, documents |
| Parquet | Binary/Columnar | Column-oriented storage | High compression, fast column queries | Not human-readable | Big Data (Hadoop/Spark), analytics |
| Avro | Binary/Row-based | Row-oriented, schema embedded | Compact, fast serialization | Less optimized for column reads | Streaming, messaging (Kafka) |
Big Data & Processing Tools
-
Big Data Characteristics (4 V's): Volume, Velocity, Variety, Veracity.
-
Hadoop Ecosystem:
-
HDFS (Hadoop Distributed File System):
-
Architecture: Master/Slave. NameNode (metadata, namespace) + DataNodes (block storage).
-
Blocks: Files split into 128MB/256MB blocks (default), replicated (default=3) across DataNodes.
-
Fault Tolerance: If a DataNode fails, NameNode re-replicates blocks from other copies.
-
DiagramSEARCH: "HDFS architecture diagram NameNode DataNode"
-
-
Hive:
-
Data warehouse infrastructure on top of HDFS.
-
HiveQL: SQL-like query language (translates to MapReduce/Tez/Spark jobs).
-
Metastore: Stores table schema and partition info (usually in RDBMS).
-
-
-
Context: Apache Spark (in-memory processing, faster than MapReduce), NoSQL databases (HBase, Cassandra for non-relational data).
VI. DATA VISUALIZATION PRINCIPLES & TOOLS
Data Visualization Fundamentals
-
Purpose: Exploration (find patterns), Communication (tell story), Discovery.
-
Univariate Exploration (1 variable):
-
Continuous: Histogram, Density plot, Box plot.
-
Categorical: Bar chart, Pie chart (use sparingly), Frequency table.
-
-
Bivariate Exploration (2 variables):
-
Continuous-Continuous: Scatter plot, Hexbin plot, Line chart (time series).
-
Categorical-Continuous: Box plot (by category), Violin plot.
-
Categorical-Categorical: Grouped bar chart, Mosaic plot, Heatmap (contingency).
-
-
Multivariate Exploration (>2 variables): Use aesthetics (color, size, shape, facets) in bivariate plots, or use techniques like scatterplot matrix, parallel coordinates.
Python Visualization Libraries
| Library | Key Features | Best For | Level |
|---|---|---|---|
| Pandas | .plot() method, quick & dirty |
Fast exploratory plots from DataFrames | Basic |
| Matplotlib | Foundational, full control, object-oriented | Custom, publication-quality figures | Low-level |
| Seaborn | Statistical plots, beautiful defaults, built on Matplotlib | Complex stats viz (regression plots, distributions, categorical) | High-level |
| Plotly | Interactive, web-based, hover tooltips | Dashboards, interactive web apps | Interactive |
| ggplot (plotnine) | Grammar of Graphics (R's ggplot2 port) | Layered, declarative plotting | Grammar-based |
[!TIP] Seaborn is ideal for statistical exploration; Plotly for interactive reports.
Power BI Ecosystem
-
Power BI Desktop: Free application for report authoring. Connect to data, model (DAX), create visuals, design reports.
-
Power BI Service: Cloud platform (SaaS). Publish, share, collaborate, schedule refreshes, create dashboards (single-page, real-time).
-
Power BI Mobile: App for consuming reports/dashboards on phones/tablets.
-
Power BI Report Server: On-premises server for hosting reports (for organizations with strict data governance).
-
Core Components:
-
Datasets: Underlying data model (can be shared).
-
Reports: Multi-page collection of visuals.
-
Dashboards: Single-page, tile-based, pinned from reports/datasets.
-
Apps: Packaged reports/dashboards for distribution.
-
Creating Custom Visualizations for Complex Datasets
-
Understand the Story: What question are you answering? What insight is hidden?
-
Choose Chart Type: Go beyond standard (bar, line, scatter). Consider: Sankey diagram (flows), Network graph (relationships), Heatmap with dendrogram (clustering), Geospatial map, Treemap (hierarchy).
-
Design for Clarity: Use appropriate color scales (sequential, diverging, categorical), clear labels, annotations, avoid clutter (chartjunk).
-
Iterative Refinement: Prototype, get feedback, simplify, ensure accessibility (colorblind-friendly).
VII. DATA ANALYST ECOSYSTEM & INDEXING
Data Analyst Ecosystem
-
Overview: Integrated workflow and tools:
Data Acquisition (SQL, APIs) → Storage (DB, Data Lake) → Processing (Python/R, Spark) → Analysis (Stats, ML) → Visualization (Power BI, Tableau) → Decision Support. -
Key Roles/Tools Integration: SQL for extraction, Python/R for analysis/ML, Power BI/Tableau for dashboards, Big Data platforms (Hadoop/Spark) for large-scale processing.
Data Management and Indexing
-
Data Management:
-
DBMS (Database Management System): Software to manage databases (e.g., PostgreSQL, MySQL).
-
Data Modeling: Conceptual (ER diagram), Logical (tables, columns), Physical (indexes, partitions).
-
Data Integrity: Accuracy & consistency (constraints: primary key, foreign key, unique, not null).
-
ACID Properties: Atomicity, Consistency, Isolation, Durability (for transactional reliability).
-
-
Indexing:
-
Purpose: Speed up data retrieval (SELECT, WHERE, JOIN) by creating a lookup structure (like a book index).
-
Trade-off: Faster reads, slower writes (INSERT/UPDATE/DELETE), extra storage.
-
Common Types:
-
B-tree: Balanced tree, sorted, good for range queries and equality. Most common.
-
Hash: Key-value lookup, very fast for equality, not for range.
-
-
How it works: Index stores column value + pointer to row. DB engine searches index first to find pointers, then fetches rows.
[!TIP] Index foreign keys and frequently queried columns. Avoid over-indexing on frequently updated tables.
-