UNIT 1: FOUNDATIONS OF DATA ANALYSIS & VISUALIZATION
1.0 FUNDAMENTAL DATA CONCEPTS & MEASUREMENT
1.1 Variables and Data Categorization
-
Variable: A characteristic or attribute that can be measured or observed and can vary among individuals or over time.
-
Data Categorization:
-
Categorical (Qualitative): Represents categories or groups.
-
Nominal: No natural order (e.g., Gender, Color, Country).
-
Ordinal: Natural order but not fixed intervals (e.g., Education Level, Satisfaction Rating: Low/Medium/High).
-
-
Numerical (Quantitative): Represents measurable quantities.
-
Discrete: Countable, distinct values (e.g., Number of students, Shoe size).
-
Continuous: Any value within a range, measurable (e.g., Height, Weight, Temperature).
-
-
[!TIP] Exam Focus: Distinguish Ordinal (ordered categories) from Nominal (no order). Discrete is countable; continuous is measurable.
1.2 Levels of Measurement (Scale of Measurement)
| Scale | Properties | Examples | Permissible Operations |
|---|---|---|---|
| Nominal | Classification only | Gender, ID numbers | Equality/Inequality (=, ≠) |
| Ordinal | Classification + Order | Rankings, Grades (A,B,C) | Equality/Inequality, Greater/Less (<, >) |
| Interval | Order + Fixed intervals, no true zero | Temperature (°C, °F), Calendar years | All above + Addition/Subtraction (+,-) |
| Ratio | All properties + true zero | Height, Weight, Age, Income | All arithmetic operations (+, -, ×, ÷) |
[!TIP] Key Difference: Interval has arbitrary zero (0°C doesn't mean no heat); Ratio has absolute zero (0 kg means no mass). Only ratio scale allows meaningful ratios (e.g., 20kg is twice 10kg).
1.3 Measures of Central Tendency
Summarizes the center of a dataset.
-
Mean (Arithmetic Average): $$\displaystyle \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$. Sensitive to outliers.
-
Median: Middle value when sorted. Robust to outliers.
-
Mode: Most frequent value. Can be multimodal. Applicable to nominal data.
-
Comparison:
-
Symmetric distribution: Mean ≈ Median ≈ Mode.
-
Right-skewed (positive): Mean > Median > Mode.
-
Left-skewed (negative): Mean < Median < Mode.
-
[!TIP] For nominal data, only Mode is meaningful. For ordinal, Median is often best.
1.4 Measures of Location and Dispersion (Spread)
-
Location (Percentiles):
-
Quartiles: Q1 (25th), Q2=Median (50th), Q3 (75th).
-
Five-Number Summary: Min, Q1, Median, Q3, Max.
-
Interquartile Range (IQR): $$\displaystyle IQR = Q3 - Q1 $$. Measures spread of middle 50%, robust to outliers.
-
-
Dispersion:
-
Range: $Max - Min$. Sensitive to outliers.
-
Variance ($$\displaystyle s^2 $$ for sample, $$\displaystyle \sigma^2 $$ for population): $$\displaystyle s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} $$. Average squared deviation.
-
Standard Deviation ($s$ or $\sigma$): $$\displaystyle s = \sqrt{s^2} $$. In original units.
-
Coefficient of Variation (CV): $$\displaystyle CV = \frac{s}{\bar{x}} \times 100\% $$. Unitless measure of relative variability. Useful for comparing datasets with different units/means.
-
[!FORMULA] Box Plot Whiskers: Typically extend to
Q1 - 1.5*IQRandQ3 + 1.5*IQR. Points beyond are potential outliers.
1.5 Introduction to Multivariate Analysis
-
Definition: Simultaneous analysis of more than two variables to understand relationships and structures.
-
Need: Real-world problems involve multiple interacting factors. Univariate/bivariate analysis is insufficient.
-
Key Methods (Conceptual):
-
Multiple Regression: Models relationship between one dependent variable and multiple independent variables.
-
Factor Analysis: Reduces many variables into fewer underlying latent factors.
-
Cluster Analysis: Groups observations into clusters based on similarity.
-
Principal Component Analysis (PCA): Reduces dimensionality by transforming variables into uncorrelated principal components.
-
2.0 DATA MANAGEMENT, WRANGLING & FORMATS
2.1 Data Wrangling (Data Munging) Process
The process of cleaning, structuring, and enriching raw data into a desired format for analysis.
-
Data Acquisition: Gathering data from sources (databases, APIs, files, web scraping).
-
Data Cleaning:
-
Handling Missing Values: Deletion, Imputation (mean/median/mode), Prediction.
-
Handling Outliers: Detection (IQR, Z-score), Capping, Transformation, or Investigation.
-
Correcting Errors: Inconsistencies, typos, invalid entries.
-
-
Data Transformation:
-
Normalization/Standardization (scaling).
-
Encoding categorical variables (One-Hot, Label Encoding).
-
Aggregation (summarizing).
-
Pivoting/Melting.
-
-
Data Integration: Combining data from multiple sources (merging, joining, concatenating).
-
Data Enrichment: Adding new derived features or external data to improve analysis.
[!TIP] Golden Rule: Data wrling takes ~60-80% of a data scientist's time. "Garbage in, garbage out."
2.2 File Formats and Data Storage
| Format | Type | Structure | Key Characteristics | Best For |
|---|---|---|---|---|
| CSV | Structured | Tabular, plain text | Human-readable, universal, no schema, no type info. | Small-medium datasets, simple exchange. |
| JSON | Semi-structured | Key-value pairs, nested | Hierarchical, flexible schema, web-friendly. | APIs, config files, NoSQL (MongoDB). |
| XML | Semi-structured | Tagged elements, tree | Verbose, strict schema (XSD), self-descriptive. | Legacy systems, documents, configs. |
| Parquet | Structured | Columnar storage | Highly compressed, efficient for queries (read only needed columns), schema stored. | Big Data (Hadoop/Spark), analytics workloads. |
| Avro | Structured | Row-based, binary | Compact, fast serialization, schema evolution support. | Streaming data (Kafka), Hadoop. |
| Excel (.xlsx) | Structured | Tabular, multiple sheets | GUI-based, formulas, formatting. | Business users, small datasets, reports. |
[!TIP] Columnar (Parquet) is faster for analytical queries (read specific columns). Row-based (Avro, CSV) is better for transactional/row-wise operations.
2.3 Data Management and Indexing (Conceptual)
-
Database Fundamentals:
-
Relational (SQL): Tables with fixed schemas, ACID transactions (e.g., MySQL, PostgreSQL). Good for structured data, complex queries.
-
NoSQL: Flexible schema, horizontal scaling. Types: Document (MongoDB), Key-Value (Redis), Column-family (Cassandra), Graph (Neo4j).
-
-
Indexing: A database object that improves the speed of data retrieval operations.
-
Analogy: Like a book's index—allows finding a topic (data) without scanning every page (full table scan).
-
Trade-off: Speeds up
SELECT/WHERE/JOINbut slows downINSERT/UPDATE/DELETE(index must be updated).
-
2.4 Big Data Ecosystem Overview
-
Characteristics (4 V's):
-
Volume: Massive scale (TB, PB, EB).
-
Velocity: High speed of data generation/need for processing (real-time streams).
-
Variety: Structured, semi-structured, unstructured.
-
Veracity: Uncertainty, inconsistency, missing values.
-
-
Core Processing Frameworks:
-
Hadoop:
-
HDFS (Hadoop Distributed File System): Master-Slave architecture. Stores large files across clusters. Fault-tolerant.
-
MapReduce: Programming model for processing.
Map(filter/sort) →Shuffle→Reduce(aggregate). Disk-based, high latency.
-
-
Apache Spark:
-
In-memory processing. DAG (Directed Acyclic Graph) execution engine.
-
Resilient Distributed Datasets (RDDs): Immutable, partitioned collections.
-
Faster than MapReduce for iterative algorithms (ML) and interactive queries. Supports Spark SQL, Streaming, MLlib.
-
-
-
Query/Processing Tools:
-
Hive: Data warehouse on Hadoop. Uses HiveQL (SQL-like) which compiles to MapReduce/Tez/Spark. For batch processing.
-
Pig: High-level scripting language (Pig Latin) for data flows. Compiles to MapReduce. Good for ETL pipelines.
-
[!TIP] Spark vs. MapReduce: Spark is in-memory and uses DAGs, making it ~100x faster for many workloads. MapReduce is disk-based, more stable for very large, non-iterative jobs.
3.0 STATISTICAL INFERENCE & HYPOTHESIS TESTING
3.1 Statistical Inferences
-
Population: Entire group of interest (parameter: $\mu, \sigma$).
-
Sample: Subset of population (statistic: $\bar{x}, s$).
-
Estimation:
-
Point Estimation: Single value estimate (e.g., $\bar{x}$ estimates $\mu$).
-
Interval Estimation (Confidence Interval - CI): Range of values likely to contain population parameter.
-
95% CI for $\mu$ (known $\sigma$): $$\displaystyle \bar{x} \pm z_{\alpha/2} \frac{\sigma}{\sqrt{n}} $$
-
95% CI for $\mu$ (unknown $\sigma$): $$\displaystyle \bar{x} \pm t_{\alpha/2, n-1} \frac{s}{\sqrt{n}} $$
-
-
3.2 Parametric Hypothesis Testing
General Steps:
-
State Hypotheses:
-
$$\displaystyle H_0 $$ (Null): Status quo, no effect, status to be rejected.
-
$$\displaystyle H_1 $$ (Alternative): Research claim, effect exists.
-
-
Choose Significance Level ($\alpha$): Probability of Type I error (rejecting true $$\displaystyle H_0 $$). Common: 0.05, 0.01.
-
Select Appropriate Test & Compute Test Statistic.
-
Determine p-value or Critical Region.
-
Decision: If p-value ≤ $\alpha$, reject $$\displaystyle H_0 $$. Or if test stat in critical region, reject $$\displaystyle H_0 $$.
-
Conclusion in context.
3.2.1 t-Test
Compares means. Assumes (approximately) normally distributed population.
-
One-Sample t-Test: Tests if sample mean differs from a known/hypothesized population mean $$\displaystyle \mu_0 $$.
-
Test Statistic: $$\displaystyle t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} $$, df = $n-1$.
-
Example (Potato Yield - May 2024):
$$\displaystyle X = [21.5, 24.5, 18.5, 17.2, 14.5, 23.2, 22.1, 20.5, 19.4, 18.1, 24.1, 18.5] $$, $$\displaystyle \mu_0 = 20 $$, $$\displaystyle n=12 $$.
$$\displaystyle \bar{x} = 19.9 $$, $s \approx 2.94$.
$$\displaystyle t = \frac{19.9 - 20}{2.94 / \sqrt{12}} \approx -0.118 $$.
df=11. For one-tailed test (better than standard), p-value > 0.05. Fail to reject $$\displaystyle H_0 $$. No significant evidence yield is better.
-
-
Two-Sample t-Test:
-
Independent: Two separate groups. Assumes equal variances (pooled) or not (Welch's).
-
Paired (Dependent): Same subjects measured twice (before/after). Uses differences $$\displaystyle d_i = x_{1i} - x_{2i} $$, then one-sample t on $d$.
-
3.2.2 Chi-Square (χ²) Test
For categorical data. Compares observed frequencies to expected frequencies.
-
Goodness-of-Fit Test: Tests if a single variable's distribution fits a hypothesized distribution.
- $$\displaystyle \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$, df = $k-1$ (k = categories).
-
Test of Independence (Contingency Table): Tests if two categorical variables are independent.
-
Contingency Table: Rows = categories of Var A, Columns = categories of Var B.
-
$$\displaystyle E_{ij} = \frac{(Row_i \ Total) \times (Col_j \ Total)}{Grand \ Total} $$.
-
$$\displaystyle \chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}} $$, df = $(r-1)(c-1)$.
-
Assumption: All $$\displaystyle E_{ij} \geq 5 $$ (otherwise use Fisher's Exact Test).
-
-
Decision: Reject $$\displaystyle H_0 $$ (variables are independent) if $$\displaystyle \chi^2_{calc} > \chi^2_{crit} $$ or p-value < $\alpha$.
[!TIP] Chi-Square: $$\displaystyle H_0 $$: Variables are independent (no association). Rejecting $$\displaystyle H_0 $$ implies association.
3.3 Regression Analysis
Models relationship between variables.
-
Dependent (Response) Variable $Y$: Variable being predicted/explained.
-
Independent (Predictor) Variable(s) $X$: Variable(s) used for prediction.
-
Simple Linear Regression (SLR):
-
Model: $$\displaystyle Y = \beta_0 + \beta_1 X + \epsilon $$
-
$$\displaystyle \beta_0 $$: Intercept (value of Y when X=0).
-
$$\displaystyle \beta_1 $$: Slope (change in Y per unit change in X).
-
Least Squares Estimation: Minimizes $$\displaystyle \sum (y_i - \hat{y}_i)^2 $$.
-
$$\displaystyle \hat{\beta}_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2} $$, $$\displaystyle \hat{\beta}_0 = \bar{y} - \hat{\beta}_1 \bar{x} $$.
-
-
Multiple Linear Regression (MLR):
-
Model: $$\displaystyle Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + ... + \beta_p X_p + \epsilon $$
-
Each $$\displaystyle \beta_j $$ represents the effect of $$\displaystyle X_j $$ holding all other X's constant.
-
-
Categorical Predictors (Dummy Variables): Convert categories into binary (0/1) variables. For k categories, use k-1 dummies to avoid multicollinearity (dummy variable trap).
3.4 Maximum Likelihood Estimation (MLE)
-
Philosophy: Find parameter values ($\theta$) that make the observed data most probable.
-
Steps:
-
Write down the likelihood function $$\displaystyle L(\theta | data) = P(data | \theta) = \prod_{i=1}^{n} f(x_i | \theta) $$ (for i.i.d. data).
-
Take the log-likelihood: $$\displaystyle \ell(\theta) = \log L(\theta) $$ (simplifies product to sum).
-
Differentiate $\ell(\theta)$ w.r.t. $\theta$ and set to zero: $$\displaystyle \frac{d\ell}{d\theta} = 0 $$.
-
Solve for $\theta$. This $$\displaystyle \hat{\theta}_{MLE} $$ maximizes likelihood.
-
-
Example (Bernoulli): Data: k successes in n trials. $$\displaystyle L(p) = p^k (1-p)^{n-k} $$. $$\displaystyle \ell(p) = k \log p + (n-k) \log(1-p) $$. $$\displaystyle \frac{d\ell}{dp} = \frac{k}{p} - \frac{n-k}{1-p} = 0 $$ → $$\displaystyle \hat{p} = \frac{k}{n} $$ (sample proportion).
[!TIP] MLE is a general principle, not a specific test. Used to estimate parameters for many models (regression, logistic, etc.).
3.5 Bayesian Modelling (Conceptual)
-
Bayes' Theorem: $$\displaystyle P(\theta | D) = \frac{P(D | \theta) P(\theta)}{P(D)} $$
-
$P(\theta | D)$: Posterior (updated belief about $\theta$ after seeing data D).
-
$P(D | \theta)$: Likelihood (probability of data given $\theta$).
-
$P(\theta)$: Prior (belief about $\theta$ before seeing data).
-
$P(D)$: Marginal likelihood (normalizing constant).
-
-
Bayesian vs. Frequentist:
-
Bayesian: Parameters are random variables with probability distributions. Incorporates prior knowledge. Results are probability statements about parameters (e.g., "There's a 95% probability $\theta$ is in this interval").
-
Frequentist: Parameters are fixed but unknown. Relies on repeated sampling concept. Confidence intervals are about long-run coverage.
-
-
Advantages: Incorporates prior info, intuitive probabilistic interpretation, handles small data well, full posterior distribution.
-
Disadvantages: Computationally intensive (MCMC), choice of prior can be subjective/influential.
3.6 Re-sampling Methods (Conceptual)
Estimate uncertainty/performance without strict distributional assumptions.
-
Cross-Validation:
-
k-Fold CV: Split data into k folds. Train on k-1, test on 1. Repeat k times. Average performance.
-
Purpose: Model selection, hyperparameter tuning, estimating generalization error.
-
-
Bootstrap:
-
Sample with replacement from original data to create many "bootstrap samples" (same size n).
-
Compute statistic (e.g., mean) for each sample → bootstrap distribution.
-
Use distribution to estimate standard error, confidence intervals (percentile method).
-
Purpose: Estimating variability of almost any statistic.
-
4.0 DATA VISUALIZATION PRINCIPLES & EXPLORATION
4.1 Introduction to Data Visualization
-
Goals:
-
Exploration: Discover patterns, outliers, relationships (for yourself).
-
Explanation/Communication: Present findings to others clearly.
-
-
Principles of Effective Visualization:
-
Data-Ink Ratio (Tufte): Maximize ink used for data, minimize non-data ink (decoration, 3D effects).
-
Chart Selection: Match plot type to question (distribution? relationship? comparison?).
-
Color Theory: Use sequential for ordered, diverging for deviation from midpoint, categorical for distinct groups. Avoid rainbow scales. Consider colorblindness.
-
Avoid Distortion: Truncated axes, inappropriate scales can mislead.
-
Label Clearly: Axes, titles, legends.
-
4.2 Univariate Exploration (One Variable)
| Data Type | Plots | Summarizes |
|---|---|---|
| Numerical | Histogram (bins), Density Plot (smoothed), Box Plot, Stem-and-Leaf | Shape (symmetry, peaks), Center, Spread, Outliers |
| Categorical | Bar Chart (counts/frequencies), Pie Chart (use sparingly) | Frequency/Proportion of each category |
[!TIP] Box Plot shows 5-number summary and outliers. Histogram bin width choice greatly affects perceived shape.
4.3 Bivariate Exploration (Two Variables)
| X (Predictor) | Y (Response) | Recommended Plots | Measure of Association |
|---|---|---|---|
| Numerical | Numerical | Scatter Plot, Line Chart (time series) | Pearson's r (linear), Spearman's ρ (rank) |
| Categorical | Numerical | Box Plot (by category), Violin Plot, Bar Chart (with summary stat) | ANOVA (F-test) for mean differences |
| Categorical | Categorical | Grouped/Mosaic Bar Chart, Heatmap (contingency table) | Chi-Square (χ²) test of independence, Cramer's V |
4.4 Multivariate Exploration (>2 Variables)
-
Using 2D Plot Aesthetics (add dimensions to scatter/bar):
-
Color: Categorical or continuous (sequential palette).
-
Size: Numerical (area proportional to value).
-
Shape: Categorical (distinct markers).
-
Faceting (Small Multiples): Create multiple sub-plots for levels of a categorical variable. Best practice for adding dimensions without clutter.
-
-
3D Plots: Often hard to interpret; avoid unless truly necessary and interactive.
-
Specialized Plots:
-
Scatterplot Matrix (Pairs Plot): Grid of scatterplots for all numerical variable pairs.
-
Parallel Coordinates: Each variable is an axis; observations are lines. Good for high-dimensional patterns.
-
4.5 Creating Custom Visualizations for Complex Datasets
Process:
-
Understand the Question & Data: What story? What variables? What relationships?
-
Choose the Right Encodings: Map variables to visual channels (x, y, color, size, shape) based on data type and importance.
-
Iterative Design:
-
Start simple (univariate, bivariate).
-
Add dimensions cautiously (faceting > color > size).
-
Refine: Improve labels, titles, legends, remove clutter.
-
Test: Does it convey the intended message quickly?
-
-
Leverage Interactivity (for exploration): Tooltips, zoom, filter, linked views ( brushing & linking).
5.0 TOOLS & TECHNOLOGIES FOR ANALYSIS & VISUALIZATION
5.1 Python Ecosystem
5.1.1 Pandas
-
Core Data Structures:
-
Series: 1D labeled array. -
DataFrame: 2D labeled table (rows: index, columns: columns). Workhorse.
-
-
Core Operations:
-
Indexing/Selection:
df[col],df.loc[](label-based),df.iloc[](position-based). -
Filtering: Boolean indexing
df[df['col'] > 5]. -
Grouping & Aggregation:
df.groupby('cat_col').agg({'num_col': ['mean', 'sum']}). -
Merging/Joining:
pd.merge(df1, df2, on='key')(SQL-like joins),pd.concat(). -
Handling Missing:
isna(),fillna(),dropna().
-
5.1.2 Core Visualization Libraries
| Library | Philosophy | Strengths | Use Case |
|---|---|---|---|
| Matplotlib | Foundation, low-level | Full control, highly customizable, object-oriented API (Figure, Axes). |
Complex, publication-quality static plots; building blocks for others. |
| Seaborn | Statistical graphics, high-level | Built on Matplotlib. Simplifies complex plots (distribution, regression, categorical). Beautiful defaults/themes. | Quick statistical exploration (distributions, relationships, categorical summaries). |
| Plotly | Interactive, web-based | Interactive by default (zoom, pan, hover). Web-based (HTML). plotly.express for simple syntax. |
Dashboards, web apps, presentations, exploratory analysis needing interactivity. |
[!TIP] Typical Workflow: Use Pandas for data manipulation → Seaborn for quick statistical plots → Matplotlib for fine-tuning → Plotly for interactive needs.
5.2 Specialized Visualization Tools
-
ggplot2 (R): Based on Grammar of Graphics.
-
Layered Approach:
ggplot(data) + geom_point(mapping=aes(x,y)) + labs(...). -
Components: Data, Aesthetics (
aes()), Geometries (geom_*), Facets, Statistics, Scales, Coordinates, Themes. -
Strength: Consistent, logical framework for building complex plots layer by layer.
-
5.3 Business Intelligence & Dashboarding Tools
-
Power BI Ecosystem:
-
Power BI Desktop: Free Windows application for data modeling (using DAX - Data Analysis Expressions), report creation, and visualization.
-
Power BI Service: Cloud-based SaaS for publishing, sharing, collaborating, and scheduling refreshes.
-
Power BI Mobile: Apps for viewing reports on mobile devices.
-
Key Features: Drag-and-drop interface, rich visualizations, real-time dashboards, natural language Q&A, integration with Azure/Office 365.
-
5.4 The Data Analyst Ecosystem (Holistic View)
A typical pipeline integrates multiple tools:
-
Storage: SQL databases (structured), HDFS/Data Lakes (big, raw).
-
Wrangling & Analysis: SQL (initial extraction), Python (Pandas, NumPy) / R for cleaning, transformation, statistical analysis, ML.
-
Visualization & Exploration: Python libs (Matplotlib/Seaborn/Plotly), ggplot2, or BI tools for initial exploration.
-
Dashboarding & Sharing: Power BI, Tableau, Plotly Dash, Streamlit for interactive dashboards and storytelling.
-
Deployment: Embed reports in websites/apps, schedule refreshes, set up alerts.
[!TIP] Role Clarity: Data Analyst focuses on descriptive/predictive analytics, SQL, BI tools, business metrics. Data Scientist adds advanced ML, Python/R coding, experimental design. Data Engineer builds/maintains data pipelines (Spark, Airflow, cloud services).