UNIT 2: FOUNDATIONAL STATISTICAL CONCEPTS
Measures of Central Tendency
Definition: Single values that describe the center of a data distribution.
| Measure | Definition & Formula | Properties | Applications |
|---|---|---|---|
| Mean ($\bar{x}$) | Sum of all values divided by count: $$\displaystyle \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$ | Sensitive to outliers; uses all data points. | Most common for interval/ratio data. |
| Median | Middle value when data is sorted. For even n: avg of two middle values. | Robust to outliers; divides data into halves. | Ordinal data or skewed distributions. |
| Mode | Most frequently occurring value(s). Can be multi-modal. | Applicable to all data types; may not be unique. | Nominal data; identifying common categories. |
[!TIP] Exam Focus: Be prepared to calculate each from a small dataset and justify which is most appropriate for a given data type (e.g., median for income data).
Measures of Dispersion (Spread)
Definition: Quantify the variability or spread of data points around the central value.
| Measure | Formula (Sample) | Interpretation |
|---|---|---|
| Range | $\text{Max} - \text{Min}$ | Sensitive to extremes; simple but crude. |
| Variance ($$\displaystyle s^2 $$) | $$\displaystyle s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} $$ | Average squared deviation; units are squared. |
| Standard Deviation ($s$) | $$\displaystyle s = \sqrt{s^2} = \sqrt{\frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1}} $$ | Most common; in original units. |
| IQR | $$\displaystyle Q_3 - Q_1 $$ (75th - 25th percentile) | Spread of middle 50%; robust to outliers. |
[!TIP] Key Point: Variance and SD use $n-1$ (degrees of freedom) for sample estimation to provide an unbiased estimate of population variance.
Levels of Measurement
| Scale | Characteristics | Permissible Statistics | Example |
|---|---|---|---|
| Nominal | Categories only; no order. | Mode, frequency, chi-square. | Gender, blood type. |
| Ordinal | Ordered categories; unequal intervals. | Median, percentiles, non-parametric tests. | Likert scale, education level. |
| Interval | Ordered, equal intervals, no true zero. | Mean, SD, correlation. | Temperature (°C), IQ score. |
| Ratio | All interval properties + absolute zero. | All statistics, including geometric mean. | Height, weight, income. |
Variables and Data Categorization
-
Categorical Variables:
-
Nominal: No intrinsic order (e.g., color, country).
-
Ordinal: Has order but not fixed intervals (e.g., small/medium/large).
-
-
Numerical (Quantitative) Variables:
-
Discrete: Countable values (integers), gaps (e.g., number of children).
-
Continuous: Any value in a range, measurable (e.g., time, weight).
-
-
Coding Schemes: Converting categorical data to numerical for analysis.
-
Dummy Variables (0/1): For nominal/ordinal in regression.
-
Label Encoding: Assigns integers (can imply false order—use cautiously).
-
One-Hot Encoding: Creates binary column for each category (avoids ordinal assumption).
-
UNIT 2: INFERENTIAL STATISTICS
Statistical Inference
-
Estimation:
-
Point Estimation: Single value (e.g., $\bar{x}$ estimates $\mu$).
-
Interval Estimation (Confidence Interval): Range likely to contain population parameter.
-
For mean (known $\sigma$): $$\displaystyle \bar{x} \pm z_{\alpha/2} \frac{\sigma}{\sqrt{n}} $$
-
For mean (unknown $\sigma$): $$\displaystyle \bar{x} \pm t_{\alpha/2, df} \frac{s}{\sqrt{n}} $$
-
-
-
Hypothesis Testing:
-
Null Hypothesis ($$\displaystyle H_0 $$): Status quo, no effect.
-
Alternative Hypothesis ($$\displaystyle H_1 $$ or $$\displaystyle H_a $$): Research claim.
-
p-value: Probability of observing data given $$\displaystyle H_0 $$ is true. Reject $$\displaystyle H_0 $$ if p-value < $\alpha$ (significance level).
-
Type I Error ($\alpha$): Rejecting true $$\displaystyle H_0 $$.
-
Type II Error ($\beta$): Failing to reject false $$\displaystyle H_0 $$. Power = $1-\beta$.
-
Parametric Tests
t-test
Assumptions: Random sample, normality (or large n), homogeneity of variance (for two-sample).
| Type | Purpose | Test Statistic | Degrees of Freedom |
|---|---|---|---|
| One-sample | Compare sample mean to known $$\displaystyle \mu_0 $$. | $$\displaystyle t = \frac{\bar{x} - \mu_0}{s/\sqrt{n}} $$ | $$\displaystyle df = n-1 $$ |
| Independent two-sample | Compare means of two independent groups. | $$\displaystyle t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} $$ | Welch's approx. or pooled $df$ |
| Paired | Compare two measurements on same subjects. | $$\displaystyle t = \frac{\bar{d} - \mu_d}{s_d/\sqrt{n}} $$ | $$\displaystyle df = n-1 $$ (where $d$ = differences) |
Example (Potato Yield - One-sample t-test from May 2024):
$$\displaystyle H_0: \mu = 20 $$ (standard yield), $$\displaystyle H_1: \mu > 20 $$ (better).
Data: $$\displaystyle X = [21.5, 24.5, ..., 18.5] $$, $$\displaystyle n=12 $$.
- Compute $\bar{x}$ and $s$.
- Calculate $$\displaystyle t = \frac{\bar{x} - 20}{s/\sqrt{12}} $$.
- Compare to $$\displaystyle t_{critical, \alpha, df=11} $$ (one-tailed). Reject $$\displaystyle H_0 $$ if $$\displaystyle t_{calc} > t_{crit} $$.
Chi-Square ($$\displaystyle \chi^2 $$) Test
Assumptions: Independent observations, expected frequencies ≥ 5 (mostly), categorical data.
| Test | Purpose | Statistic | Degrees of Freedom |
|---|---|---|---|
| Goodness-of-Fit | Compare observed freq. to expected freq. (one variable). | $$\displaystyle \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$ | $$\displaystyle df = k - 1 - m $$ (k categories, m estimated params) |
| Test of Independence | Test association between two categorical variables (contingency table). | Same formula, using row/column totals to find $$\displaystyle E_{ij} = \frac{(row_i \ total) \times (col_j \ total)}{n} $$ | $$\displaystyle df = (r-1)(c-1) $$ |
[!TIP] Calculation Steps: 1) State $$\displaystyle H_0 $$ (no association/good fit). 2) Build table, compute $$\displaystyle E_i $$. 3) Calculate $$\displaystyle \chi^2_{calc} $$. 4) Find critical $$\displaystyle \chi^2 $$ from table or compute p-value. 5) Conclude.
Maximum Likelihood Estimation (MLE)
Goal: Find parameter values ($\theta$) that maximize the likelihood of observed data.
-
Likelihood Function: $$\displaystyle L(\theta | data) = \prod_{i=1}^{n} f(x_i; \theta) $$ (joint PDF/PMF).
-
Log-Likelihood: $$\displaystyle \ell(\theta) = \ln L(\theta) = \sum \ln f(x_i; \theta) $$ (easier to maximize).
-
Estimation: Solve $$\displaystyle \frac{d\ell}{d\theta} = 0 $$ for $$\displaystyle \hat{\theta}_{MLE} $$.
-
Properties:
-
Consistency: Converges to true $\theta$ as $n \to \infty$.
-
Efficiency: Achieves Cramér-Rao lower bound asymptotically.
-
Invariance: If $\hat{\theta}$ is MLE of $\theta$, then $g(\hat{\theta})$ is MLE of $g(\theta)$.
-
Example: For Normal($$\displaystyle \mu, \sigma^2 $$), MLE for $\mu$ is $\bar{x}$, for $$\displaystyle \sigma^2 $$ is $$\displaystyle \frac{1}{n}\sum (x_i - \bar{x})^2 $$ (note: biased, uses n not n-1).
Resampling Methods
-
Bootstrap: Sample with replacement from original data to create many "bootstrap samples." Estimate sampling distribution (e.g., CI for median).
-
Cross-Validation:
-
k-fold: Split data into k subsets; train on k-1, test on 1; repeat k times. Average performance.
-
Leave-One-Out (LOO): k = n. High variance, computationally expensive.
-
-
Permutation Test: Randomly shuffle labels/group assignments to generate null distribution of test statistic. Non-parametric alternative.
UNIT 2: MODELING TECHNIQUES
Regression Analysis
Simple Linear Regression (SLR):
-
Model: $$\displaystyle y = \beta_0 + \beta_1 x + \epsilon $$, $$\displaystyle \epsilon \sim N(0, \sigma^2) $$.
-
Least Squares Estimation: Minimize $$\displaystyle \sum (y_i - \hat{y}_i)^2 $$.
-
$$\displaystyle \hat{\beta}_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2} = \frac{S_{xy}}{S_{xx}} $$
-
$$\displaystyle \hat{\beta}_0 = \bar{y} - \hat{\beta}_1 \bar{x} $$
-
-
Interpretation: $$\displaystyle \hat{\beta}_1 $$ = change in y per unit change in x.
Multiple Linear Regression (MLR):
-
Model: $$\displaystyle y = \beta_0 + \beta_1 x_1 + ... + \beta_p x_p + \epsilon $$.
-
Assumptions: Linearity, independence, homoscedasticity, normality of errors, no perfect multicollinearity.
-
Multicollinearity: High correlation among predictors. Diagnose with VIF (Variance Inflation Factor). VIF > 5-10 indicates problem.
-
Variables:
-
Dependent (y): Outcome.
-
Independent (x): Predictors.
-
Dummy Variables: 0/1 coding for categorical predictors (e.g., gender).
-
Interaction Terms: $$\displaystyle x_1 \times x_2 $$ to model effect modification.
-
Multivariate Analysis
Purpose: Analyze multiple variables simultaneously to understand relationships, reduce dimensionality, or group observations.
| Technique | Goal | Key Idea |
|---|---|---|
| PCA (Principal Component Analysis) | Dimensionality reduction. | Find orthogonal axes (PCs) maximizing variance. |
| Factor Analysis | Identify latent factors. | Model observed vars as linear combos of unobserved factors. |
| Cluster Analysis | Group similar observations. | Distance metrics (Euclidean), algorithms (k-means, hierarchical). |
| Discriminant Analysis | Classify into known groups. | Find linear combinations (LDs) maximizing between/within-group variance. |
| Multivariate Visualization | Explore high-dim data. | Pair plots, 3D scatter, parallel coordinates, glyphs. |
Bayesian Modeling
Core: Bayes' Theorem: $$\displaystyle P(\theta | D) = \frac{P(D | \theta) P(\theta)}{P(D)} $$
-
Posterior $P(\theta | D)$: Updated belief after seeing data.
-
Likelihood $P(D | \theta)$: Probability of data given parameter.
-
Prior $P(\theta)$: Initial belief about parameter.
-
Process:
-
Specify prior distribution (informative, weakly informative, flat/uniform).
-
Write likelihood based on data model.
-
Compute posterior (analytically for conjugate priors, or via MCMC sampling like Gibbs, HMC).
-
-
Advantages:
-
Incorporates prior knowledge/expert opinion.
-
Provides full probability distribution (uncertainty quantification).
-
Natural for sequential/online learning.
-
-
Disadvantages:
-
Computationally intensive for complex models.
-
Choice of prior can influence results (sensitivity analysis needed).
-
Can be subjective.
-
UNIT 2: DATA HANDLING AND VISUALIZATION
Data Wrangling (Munging)
Process:
-
Data Cleaning:
-
Missing Values: Delete, impute (mean/median/mode, model-based), or flag.
-
Outliers: Detect via IQR (points beyond $Q1-1.5IQR$, $Q3+1.5IQR$) or Z-score (>3). Investigate cause; cap/transform or remove if erroneous.
-
Inconsistencies: Fix formatting (dates, units), correct typos.
-
-
Data Transformation:
-
Normalization (Min-Max): $$\displaystyle x' = \frac{x - \min}{\max - \min} $$ → [0,1].
-
Standardization (Z-score): $$\displaystyle x' = \frac{x - \mu}{\sigma} $$ → mean=0, SD=1.
-
Encoding: One-hot for nominal, label/ordinal for ordinal.
-
-
Data Integration: Merge/join datasets (SQL-like), concatenate, reshape (pivot/melt).
-
Tools: Python (Pandas:
df.dropna(),df.merge(),pd.get_dummies()), R (dplyr, tidyr).
Data Visualization
| Exploration | Goal | Common Plots |
|---|---|---|
| Univariate | Distribution of single variable. | Histogram, box plot, bar chart (categorical), density plot, violin plot. |
| Bivariate | Relationship between two variables. | Scatter plot (num-num), line chart (time series), box plot (cat-num), heatmap (correlation), contingency table (cat-cat). |
| Multivariate | Explore >2 variables. | Pair plot (scatter matrix), 3D scatter plot, faceted/trellis plots (small multiples), parallel coordinates, glyphs (star plots), bubble charts (size as 3rd dim). |
Custom Visualization Design
Principles (C.A.S.E.):
-
Clarity: Message is immediately understandable. Avoid clutter.
-
Accuracy: Represent data truthfully; no distorted scales.
-
Simplicity: Minimal ink for maximum insight (Tufte's data-ink ratio).
-
Storytelling: Guide viewer with annotations, logical flow, title. Chart Selection: Match plot to data type and question (e.g., trend over time → line chart; part-to-whole → stacked bar/pie [use sparingly]; distribution → histogram/box). Tools: D3.js (web, highly customizable), Tableau (drag-and-drop BI), Python: Matplotlib (low-level), Seaborn (statistical, high-level), Plotly (interactive). R: ggplot2 (grammar of graphics).
UNIT 2: TOOLS, ECOSYSTEMS, AND DATA MANAGEMENT
Visualization Libraries
| Library (Python) | Key Features | Use Case |
|---|---|---|
| Matplotlib | Foundation; low-level, highly customizable. | Static, publication-quality plots; full control. |
| Seaborn | Built on Matplotlib; statistical plots, nice defaults. | Quick exploration of distributions, relationships (e.g., relplot, catplot). |
| Pandas .plot() | Integrated with DataFrame; simple syntax. | Rapid, basic plots directly from data. |
| Plotly | Interactive, web-based plots (D3.js backend). | Dashboards, hover info, zoom, shareable HTML. |
| plotnine | ggplot2 interface for Python. | Grammar-of-graphics approach for R users. |
| R: ggplot2 | Grammar of graphics; layered, consistent syntax. | Complex, multi-step plot building; industry standard in R. |
Business Intelligence Tools: Power BI
-
Power BI Desktop: Free application for data connection, modeling, visualization creation.
-
Power BI Service: Cloud-based SaaS for publishing, sharing, collaboration, scheduling refreshes.
-
Power BI Mobile: Apps for iOS/Android to view reports on-the-go.
-
Key Features:
-
Data Modeling: Create relationships between tables; define calculated columns/measures.
-
DAX (Data Analysis Expressions): Formula language for custom calculations (e.g.,
CALCULATE,FILTER). -
Interactive Dashboards: Drill-down, cross-filtering, bookmarks.
-
Report Sharing: Publish to web, share within organization via workspaces.
-
Big Data Processing Tools
| Tool | Core Concept | Key Component/Feature |
|---|---|---|
| Hadoop Ecosystem | Distributed storage & processing of big data. | HDFS: Master-Slave (NameNode, DataNode). Files split into blocks (default 128MB), replicated (default 3) for fault tolerance. MapReduce: Programming model (Map → Shuffle/Sort → Reduce). |
| Hive | Data warehousing on Hadoop. | HiveQL: SQL-like query language. Metastore: Stores schema info (RDBMS). Converts HiveQL to MapReduce/Tez/Spark jobs. |
| Apache Spark | In-memory cluster computing. | RDDs (Resilient Distributed Datasets): Immutable, partitioned collections. DataFrames: Distributed tables (optimized via Catalyst). Spark SQL: Structured data processing. Faster than MapReduce for iterative/multi-pass algorithms. |
Data Management and Indexing
-
SQL (Relational): Structured data, fixed schema, ACID transactions. (e.g., PostgreSQL, MySQL).
-
NoSQL (Non-relational): Flexible schema, horizontal scaling. Types: Document (MongoDB), Key-Value (Redis), Column-family (Cassandra), Graph (Neo4j).
-
Indexing: Data structure to speed up queries.
-
B-tree: Balanced tree; sorted keys; good for range queries. Default in most RDBMS.
-
Hash Index: Key-value lookup; fast equality; not for ranges.
-
Bitmap Index: Bit vectors for low-cardinality columns; efficient for multi-condition AND/OR (e.g., data warehouses).
-
Data Analyst Ecosystem
Roles: Collect, clean, analyze, visualize data; translate business questions to data questions; create reports/dashboards. Common Toolchain:
-
SQL: Extract/transform data from databases.
-
Python/R: Advanced analysis, statistics, machine learning (Pandas, NumPy, SciPy, scikit-learn, tidyverse).
-
Excel/Sheets: Quick analysis, prototyping.
-
BI Tools (Power BI, Tableau): Build interactive dashboards.
-
Version Control (Git): Track code changes, collaborate. End-to-End Workflow: $$\displaystyle \text{Data Collection} \rightarrow \text{Wrangling} \rightarrow \text{Analysis} \rightarrow \text{Visualization} \rightarrow \text{Reporting} $$.
File Formats
| Format | Type | Structure | Pros | Cons | Best For |
|---|---|---|---|---|---|
| CSV/TSV | Structured | Plain text, delimiter-separated. | Human-readable, universal. | No schema, no types, inefficient storage. | Small-medium data exchange, simple import/export. |
| JSON | Semi-structured | Hierarchical (key-value), nested. | Flexible schema, web-friendly. | Verbose, slower to parse. | APIs, config files, NoSQL (MongoDB). |
| XML | Semi-structured | Tag-based, hierarchical. | Self-describing, validation (XSD). | Very verbose, complex. | Legacy systems, documents (e.g., SOAP). |
| Parquet | Binary/Columnar | Column-oriented storage. | Highly compressed, fast column queries, schema embedded. | Not human-readable, write-optimized. | Big data analytics (Spark, Hive), data lakes. |
| Avro | Binary/Row-based | Compact binary, schema stored separately. | Fast serialization, schema evolution. | Less optimized for column scans than Parquet. | Streaming, Kafka, row-based processing. |
| ORC | Binary/Columnar | Optimized for Hive (like Parquet). | High compression, ACID properties in Hive. | Less ecosystem support than Parquet. | Hive/Impala workloads. |
[!TIP] Rule of Thumb: Use Parquet for analytical queries in big data (column reads). Use CSV/JSON for interchange or small data. Use Avro for row-based streaming or when schema evolves frequently.