UNIT 3: STATISTICAL FOUNDATIONS, DATA WRANGLING & VISUALIZATION
I. DESCRIPTIVE STATISTICS & MEASUREMENT
A. Measures of Central Tendency
- Mean ($\bar{x}$): Arithmetic average. Sensitive to outliers.
$$\bar{x} = \frac{\sum_{i=1}^{n} x_i}{n}$$
-
Median: Middle value in ordered data. Robust to outliers.
-
Mode: Most frequent value. Applicable to nominal data.
-
Comparison: Use mean for symmetric distributions; median for skewed; mode for categorical.
B. Measures of Dispersion (Variability)
-
Range: $\text{max} - \text{min}$. Sensitive to extremes.
-
Variance:
-
Sample: $$\displaystyle s^2 = \frac{\sum (x_i - \bar{x})^2}{n-1} $$
-
Population: $$\displaystyle \sigma^2 = \frac{\sum (x_i - \mu)^2}{N} $$
-
-
Standard Deviation (SD): $$\displaystyle s = \sqrt{s^2} $$. Same units as data.
$$\boxed{s = \sqrt{\frac{\sum (x_i - \bar{x})^2}{n-1}}}$$
-
Interquartile Range (IQR): $$\displaystyle Q_3 - Q_1 $$. Robust measure.
-
Coefficient of Variation (CV): $$\displaystyle CV = \frac{s}{\bar{x}} \times 100\% $$. Compares relative variability.
C. Levels of Measurement (Scales of Data)
| Scale | Properties | Permissible Operations | Example |
|---|---|---|---|
| Nominal | Categories only | Count, mode | Gender, color |
| Ordinal | Ordered categories | Median, percentiles | Satisfaction (1-5) |
| Interval | Ordered, equal intervals, no true zero | Mean, SD (but ratios meaningless) | Temperature (°C) |
| Ratio | Interval + true zero | All operations | Height, weight |
D. Variables and Data Categorization
-
Categorical: Nominal (no order), Ordinal (ordered).
-
Numerical: Discrete (countable), Continuous (measurable).
-
Dependent (Response): Outcome variable.
-
Independent (Predictor): Explanatory variable.
[!TIP]
Exam Focus: Differences between scales dictate which statistics are valid. E.g., mean for interval/ratio, median for ordinal.
II. INFERENTIAL STATISTICS & HYPOTHESIS TESTING
A. Statistical Inferences
-
Population vs. Sample: Population = entire group; Sample = subset.
-
Point Estimation: Single value estimate (e.g., $\bar{x}$ for $\mu$).
-
Interval Estimation (Confidence Interval): Range likely to contain parameter.
$$\bar{x} \pm t_{\alpha/2, df} \cdot \frac{s}{\sqrt{n}}$$
B. Parametric Hypothesis Testing
1. t-Test
-
Assumptions: Normality, independence, equal variances (for two-sample).
-
One-Sample t-Test: Test sample mean vs. known $$\displaystyle \mu_0 $$.
$$\boxed{t = \frac{\bar{x} - \mu_0}{s/\sqrt{n}}}$$
$$\displaystyle df = n-1 $$.
Example (Potato Yield):
Given $$\displaystyle \mu_0 = 20 $$, $$\displaystyle n=12 $$, $$\displaystyle \bar{x}=20.175 $$, $$\displaystyle s=3.021 $$.
$$\displaystyle H_0: \mu = 20 $$, $$\displaystyle H_1: \mu > 20 $$ (one-tailed).
$$\displaystyle t = \frac{20.175-20}{3.021/\sqrt{12}} = 0.201 $$, $$\displaystyle df=11 $$.
Critical $$\displaystyle t_{0.05,11} \approx 1.796 $$. Since $$\displaystyle 0.201 < 1.796 $$, fail to reject $$\displaystyle H_0 $$. Yield not significantly better.
-
Two-Sample Independent t-Test: Compare means of two groups.
-
Equal variances: $$\displaystyle t = \frac{\bar{x}_1 - \bar{x}_2}{s_p \sqrt{\frac{1}{n_1} + \frac{1}{n_2}}} $$, where $$\displaystyle s_p^2 = \frac{(n_1-1)s_1^2 + (n_2-1)s_2^2}{n_1+n_2-2} $$.
-
Unequal variances (Welch’s): $$\displaystyle t = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} $$.
-
-
Paired t-Test: For matched pairs. Use differences $$\displaystyle d_i = x_{i1} - x_{i2} $$, then $$\displaystyle t = \frac{\bar{d}}{s_d/\sqrt{n}} $$.
2. Chi-Square Test
- Chi-Square Test for Independence: Test association between two categorical variables in a contingency table.
$$\boxed{\chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}}}$$
$$\displaystyle df = (r-1)(c-1) $$, where $$\displaystyle E_{ij} = \frac{(\text{row total}) \times (\text{column total})}{\text{grand total}} $$.
-
Goodness-of-Fit Test: Compare observed frequencies to expected theoretical distribution.
$$\displaystyle df = k-1 $$ (k = categories).
-
Steps: Compute expected counts, calculate $$\displaystyle \chi^2 $$, compare to critical value from $$\displaystyle \chi^2 $$ table.
-
Assumption: Expected counts $\geq 5$ for reliability.
C. Non-Parametric Tests (Context: Re-Sampling)
-
Need: When data violates normality or sample size small.
-
Bootstrapping: Resample with replacement to estimate sampling distribution (e.g., confidence interval for median).
-
Permutation Tests: Shuffle group labels to generate null distribution; compute p-value as proportion of permuted statistics more extreme than observed.
[!TIP]
Common Pitfall: Chi-square requires independent observations and adequate expected counts. Use Fisher’s exact test for small samples.
III. REGRESSION & ADVANCED MODELING
A. Regression Analysis
1. Simple Linear Regression
-
Model: $$\displaystyle y = \beta_0 + \beta_1 x + \epsilon $$
-
Coefficients:
$$\beta_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2}, \quad \beta_0 = \bar{y} - \beta_1 \bar{x}$$
-
Assumptions: Linearity, independence, homoscedasticity, normal errors.
-
Interpretation: $$\displaystyle \beta_1 $$ = change in $y$ per unit increase in $x$.
2. Multiple Linear Regression
-
Model: $$\displaystyle y = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + ... + \beta_p x_p + \epsilon $$
-
Interpretation: Coefficients hold other predictors constant.
-
Model Fit: $$\displaystyle R^2 $$ = proportion of variance explained.
$$R^2 = 1 - \frac{SS_{res}}{SS_{tot}}, \quad SS_{res} = \sum (y_i - \hat{y}_i)^2, \quad SS_{tot} = \sum (y_i - \bar{y})^2$$
Adjusted $$\displaystyle R^2 $$ penalizes extra predictors.
3. Types of Variables in Regression Modeling
-
Dummy Variables: Encode categorical predictors (e.g., gender: 0/1). For $k$ categories, use $k-1$ dummies.
-
Interaction Terms: Product of predictors to capture effect modification.
e.g., $$\displaystyle y = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + \beta_3 (x_1 \times x_2) + \epsilon $$
B. Maximum Likelihood Estimation (MLE)
-
Philosophy: Choose parameters that maximize probability of observed data.
-
Procedure:
-
Likelihood function: $$\displaystyle L(\theta) = \prod_{i=1}^{n} f(x_i|\theta) $$
-
Log-likelihood: $$\displaystyle \ell(\theta) = \log L(\theta) = \sum \log f(x_i|\theta) $$
-
Maximization: Solve $$\displaystyle \frac{d\ell}{d\theta} = 0 $$.
-
-
Example: For normal $$\displaystyle N(\mu,\sigma^2) $$, MLE: $$\displaystyle \hat{\mu} = \bar{x} $$, $$\displaystyle \hat{\sigma}^2 = \frac{\sum (x_i - \bar{x})^2}{n} $$ (uses $n$, not $n-1$).
C. Bayesian Modeling
-
Bayesian vs. Frequentist: Parameters have probability distributions (Bayesian) vs. fixed (Frequentist).
-
Core Concepts:
-
Prior $P(\theta)$: Initial belief about parameters.
-
Likelihood $P(\text{data}|\theta)$: Probability of data given parameters.
-
Posterior $P(\theta|\text{data})$: Updated belief after seeing data.
-
-
Bayes’ Theorem:
$$\boxed{P(\theta|D) = \frac{P(D|\theta) P(\theta)}{P(D)}}$$
-
Process: Start with prior, update with likelihood to get posterior. Posterior can become prior for new data.
-
Advantages: Incorporates prior knowledge, full uncertainty quantification.
-
Disadvantages: Prior choice subjective, computationally intensive (MCMC).
[!TIP]
Exam Focus: Contrast MLE (point estimate) with Bayesian (distribution). Know normal distribution MLE results.
IV. MULTIVARIATE ANALYSIS & DATA EXPLORATION
A. Multivariate Analysis (General)
-
Purpose: Analyze relationships among >2 variables simultaneously.
-
Techniques:
-
PCA (Principal Component Analysis): Dimensionality reduction, orthogonal components.
-
Factor Analysis: Identify latent factors underlying observed variables.
-
Cluster Analysis: Group similar observations (e.g., k-means).
-
B. Exploratory Data Analysis (EDA)
1. Univariate Exploration
-
Goal: Understand distribution of single variable.
-
Techniques:
-
Histogram: Shape, skewness, modality.
-
Box Plot: Five-number summary (min, Q1, median, Q3, max), outliers (IQR method).
-
Summary Statistics: Mean, median, SD, IQR.
-
2. Bivariate Exploration
-
Goal: Relationship between two variables.
-
Techniques:
-
Scatter Plot: Pattern, correlation, outliers.
-
Correlation Coefficient (Pearson’s $r$): Linear association strength/direction.
-
$$r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}}$$
- Cross-Tabulation: For two categorical variables.
3. Multivariate Exploration
-
Goal: Patterns across multiple variables.
-
Techniques:
-
Scatterplot Matrix: Pairwise scatter plots for all variable pairs.
-
Colored/Scaled Scatter Plots: Use color/size for third (or fourth) variable.
-
Small Multiples: Series of similar plots for different subsets (faceting).
-
[!TIP]
Common Pitfall: Relying solely on correlation; always visualize to detect non-linear relationships.
V. DATA MANAGEMENT, WRANGLING & FORMATS
A. Data Wrangling (Data Munging)
-
Definition: Process of cleaning, transforming, and structuring raw data for analysis.
-
Importance: Raw data is often messy; quality data is prerequisite for accurate analysis.
-
Detailed Process:
-
Data Acquisition: Collect from databases, APIs, files, web scraping.
-
Data Cleaning:
-
Missing Values: Delete (if few), impute (mean/median/mode for numerical, mode for categorical), or model-based.
-
Outliers: Detect via IQR (values < Q1–1.5IQR or > Q3+1.5IQR) or Z-score (>3). Investigate cause; correct, transform, or cap.
-
Inconsistencies: Fix typos, standardize formats (dates, units), resolve duplicates.
-
-
Data Transformation:
-
Normalization: Min-Max ($$\displaystyle x' = \frac{x - \min}{\max - \min} $$), Z-score ($$\displaystyle z = \frac{x - \mu}{\sigma} $$).
-
Binning: Discretize continuous variables (e.g., age groups).
-
-
Data Structuring: Reshape (pivot/melt), merge/join datasets, aggregate.
-
B. File Formats & Data Storage
1. Types of File Formats
| Format | Type | Use Case | Pros | Cons |
|---|---|---|---|---|
| CSV | Structured | Simple tabular data | Human-readable, universal | No schema, no data types |
| JSON | Semi-structured | Web APIs, nested data | Flexible, hierarchical | Larger size, parsing overhead |
| XML | Semi-structured | Document storage, configs | Self-describing, validatable | Verbose |
| Parquet | Binary | Big data, columnar queries | Efficient compression, columnar | Not human-readable |
| Avro | Binary | Schema evolution, serialization | Compact, schema with data | Less tool support than Parquet |
2. Data Management and Indexing
-
Indexing: Create auxiliary data structures (e.g., B-trees) on columns to accelerate query operations (WHERE, JOIN).
-
Trade-off: Faster read queries, slower writes, additional storage overhead.
-
Example: Index on
customer_idin a database table speeds up lookups.
C. Big Data Ecosystem & Tools
1. What is Big Data?
-
4 V’s:
-
Volume: Massive scale (TB/PB).
-
Velocity: High generation/processing speed (real-time).
-
Variety: Structured, semi-structured, unstructured.
-
Veracity: Uncertainty, quality issues.
-
2. Big Data Processing Frameworks & Tools
-
Hadoop Ecosystem:
-
HDFS (Hadoop Distributed File System):
-
Architecture: Master (NameNode) manages metadata; Slaves (DataNodes) store blocks.
-
Features: Fault-tolerant (replication, default 3x), block size (default 128 MB), write-once-read-many.
-
-
MapReduce: Programming model for batch processing.
-
Map: Process input key-value pairs → intermediate key-value pairs.
-
Shuffle: Group values by key.
-
Reduce: Aggregate values per key.
-
-
Hive: Data warehouse on Hadoop. Provides SQL-like querying (HiveQL), translates to MapReduce jobs. For structured data.
-
-
Other Tools:
-
Apache Spark: In-memory processing, faster than MapReduce. Uses RDDs/DataFrames. Supports batch, stream, ML, graph processing.
-
NoSQL Databases: Non-relational, scalable, flexible schema.
-
Key-Value: Redis (caching).
-
Document: MongoDB (JSON-like).
-
Column-Family: Cassandra (wide-column).
-
Graph: Neo4j (relationships).
-
-
[!TIP]
Exam Focus: Distinguish HDFS (storage) vs. MapReduce (processing) vs. Hive (querying). Spark’s advantage: in-memory.
VI. DATA VISUALIZATION & TOOLS
A. Data Visualization Fundamentals
-
Definition: Graphical representation of data to communicate insights.
-
Goals: Explore data patterns, convey stories, support decision-making.
-
Principles:
-
Clarity: Simple, avoid clutter; use labels, legends.
-
Accuracy: Truthful representation; avoid distortion (e.g., truncated axes).
-
Efficiency: Quick comprehension; appropriate chart type.
-
B. Python Visualization Libraries
-
Matplotlib: Foundation, low-level, highly customizable. Basic plots:
plt.plot(),plt.bar(),plt.scatter(). Steeper learning curve. -
Seaborn: Built on Matplotlib, statistical graphics, attractive defaults. Less code:
sns.histplot(),sns.boxplot(),sns.heatmap(). -
Plotly: Interactive, web-based visualizations.
plotly.expressfor quick plots;graph_objectsfor customization. Supports zoom, hover, animation. -
Pandas Integration:
DataFrame.plot()method for quick, inline plots (uses Matplotlib backend).
C. ggplot2 (in R context)
-
Grammar of Graphics: Plot built as layers.
-
Layers:
-
Data: The dataset.
-
Aesthetics (
aes()): Map variables to x, y, color, size. -
Geometries (
geom_*): Points, lines, bars. -
Facets (
facet_*): Subplots for categories. -
Statistics (
stat_*): Smoothing, binning. -
Themes (
theme_*): Non-data ink (background, fonts).
-
-
Example:
ggplot(data, aes(x=var1, y=var2)) + geom_point() + theme_minimal()
D. Power BI Ecosystem
-
What is Power BI?: Microsoft’s business analytics service for interactive visualizations and business intelligence.
-
Tools:
-
Power BI Desktop: Free application for report creation. Data modeling (relationships, DAX), transformations (Power Query), visualizations.
-
Power BI Service: Cloud platform (app.powerbi.com) for publishing, sharing, collaboration, scheduling refreshes.
-
Power BI Mobile: iOS/Android apps for on-the-go access.
-
-
Key Components:
-
Reports: Multi-page collections of visuals (charts, tables, maps).
-
Dashboards: Single-page, real-time “cockpits” with pinned tiles from reports.
-
Datasets: Underlying data sources; can be reused across reports.
-
E. Creating Custom Visualizations
-
Process:
-
Understand Data & Story: What question are you answering? What message?
-
Choose Chart Type: Bar (comparison), line (trend), scatter (relationship), heatmap (correlation), etc.
-
Design: Use color strategically (categorical vs. sequential), label axes, add title, ensure accessibility (colorblind-friendly).
-
Iterate: Test with audience, refine for clarity.
-
-
Design Considerations: Avoid chartjunk, maintain aspect ratio, use appropriate scales, highlight key data points.
[!TIP]
Common Pitfall: Overcomplicating visuals; prioritize simplicity. In Power BI, use bookmarks for custom navigation.
VII. DATA ANALYST ECOSYSTEM
A. The Data Analyst Ecosystem
-
Roles:
-
Data Analyst: Interpret data, create reports/visualizations, answer business questions.
-
Data Scientist: Advanced modeling, machine learning, predictive analytics.
-
Data Engineer: Build/maintain data pipelines, storage, ETL processes.
-
-
Tools:
-
Analysis: SQL (databases), Python/R (pandas, NumPy, scikit-learn), Excel.
-
Visualization: Power BI, Tableau, Matplotlib/Seaborn.
-
Big Data: Hadoop, Spark, Hive.
-
Version Control: Git.
-
-
Processes: CRISP-DM (Cross-Industry Standard Process for Data Mining):
-
Business Understanding
-
Data Understanding
-
Data Preparation (wrangling)
-
Modeling
-
Evaluation
-
Deployment
-
-
Interaction Flow: Data Sources → Storage (DB/Data Lake) → Processing (ETL/Spark) → Analysis (Python/R) → Visualization (Power BI) → Decision Support.
[!TIP]
Exam Focus: Know the CRISP-DM phases and differentiate roles (analyst vs. scientist vs. engineer).