UNIT 3: Data Analysis and Visualization
1.0 Fundamental Statistical Concepts
1.1 Measures of Central Tendency
Definition: Single values that represent the center or typical value of a dataset.
-
Mean (Arithmetic Average):
-
Formula: $$\displaystyle \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$
-
Sensitive to outliers.
-
-
Median:
-
Middle value when data is sorted.
-
Robust to outliers.
-
-
Mode:
-
Most frequently occurring value.
-
Can be used for nominal data.
-
[!TIP] For skewed distributions, median is often a better measure than mean.
1.2 Measures of Dispersion
Definition: Quantify the spread or variability in a dataset.
-
Range: $Max - Min$
-
Variance: $$\displaystyle \sigma^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n} $$ (population), $$\displaystyle s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} $$ (sample)
-
Standard Deviation: $$\displaystyle \sigma = \sqrt{\sigma^2} $$, $$\displaystyle s = \sqrt{s^2} $$
-
Interquartile Range (IQR): $Q3 - Q1$ (resistant to outliers)
-
Coefficient of Variation (CV): $$\displaystyle \text{CV} = \frac{s}{\bar{x}} \times 100\% $$ (unitless, compares relative variability)
1.3 Levels of Measurement
| Scale | Properties | Examples | Allowed Operations |
|---|---|---|---|
| Nominal | Categories only | Gender, Color | Count, Mode |
| Ordinal | Ordered categories | Likert scale, Ranks | Count, Median, Percentiles |
| Interval | Ordered, equal intervals, no true zero | Temperature (°C), IQ | All above + Mean, SD |
| Ratio | Interval + true zero | Height, Weight, Income | All statistical operations |
1.4 Variables and Data Categorization
-
Categorical (Qualitative):
-
Nominal: No order (e.g., country).
-
Ordinal: Ordered categories (e.g., education level).
-
-
Numerical (Quantitative):
-
Discrete: Countable values (e.g., number of students).
-
Continuous: Measurable values (e.g., weight).
-
-
Dependent Variable: Outcome variable (Y).
-
Independent Variable: Predictor variable (X).
-
Confounding Variable: External variable affecting both X and Y, causing spurious association.
2.0 Inferential Statistical Methods
2.1 Statistical Inferences
-
Point Estimation: Single value estimate of a population parameter (e.g., $\bar{x}$ estimates $\mu$).
-
Interval Estimation (Confidence Interval): Range of values likely containing the parameter.
- Example (mean, known $\sigma$): $$\displaystyle \bar{x} \pm z_{\alpha/2} \frac{\sigma}{\sqrt{n}} $$
-
Hypothesis Testing Framework:
-
State $$\displaystyle H_0 $$ (null) and $$\displaystyle H_1 $$ (alternative).
-
Choose significance level $\alpha$ (e.g., 0.05).
-
Compute test statistic.
-
Find p-value or compare with critical value.
-
Reject $$\displaystyle H_0 $$ if p-value $$\displaystyle < \alpha $$ or statistic in rejection region.
-
2.2 t-test
One-Sample t-test: Compares sample mean $\bar{x}$ to known population mean $$\displaystyle \mu_0 $$.
-
Test statistic: $$\displaystyle t = \frac{\bar{x} - \mu_0}{s/\sqrt{n}} $$, df = $n-1$
-
Assumptions: Random sample, normality (or large n), scale data.
[!EXAMPLE] Potato Yield Problem (May 2024)
Given: $$\displaystyle X = [21.5, 24.5, 18.5, 17.2, 14.5, 23.2, 22.1, 20.5, 19.4, 18.1, 24.1, 18.5] $$, $$\displaystyle \mu_0 = 20 $$, $$\displaystyle n=12 $$.
Steps:
- $$\displaystyle \bar{x} = 20.15 $$, $s \approx 3.16$
- $$\displaystyle t = \frac{20.15 - 20}{3.16/\sqrt{12}} \approx 0.164 $$
- df=11, one-tailed critical $$\displaystyle t_{0.05,11} \approx 1.796 $$
- Since $$\displaystyle 0.164 < 1.796 $$, fail to reject $$\displaystyle H_0 $$. No significant evidence yield is better.
Two-Sample Independent t-test: Compares means of two independent groups.
-
Pooled variance (equal variances assumed): $$\displaystyle s_p^2 = \frac{(n_1-1)s_1^2 + (n_2-1)s_2^2}{n_1+n_2-2} $$
-
$$\displaystyle t = \frac{\bar{x}_1 - \bar{x}_2}{s_p \sqrt{\frac{1}{n_1} + \frac{1}{n_2}}} $$, df = $$\displaystyle n_1+n_2-2 $$
Paired t-test: For dependent samples (e.g., before-after).
- $$\displaystyle t = \frac{\bar{d}}{s_d/\sqrt{n}} $$, df = $n-1$, where $d$ = differences.
2.3 Chi-Square Test
-
Goodness of Fit: Tests if observed frequencies fit expected distribution.
- $$\displaystyle \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$, df = $k-1$ (k categories)
-
Test of Independence: Tests association between two categorical variables in a contingency table.
-
$$\displaystyle \chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}} $$, df = $(r-1)(c-1)$
-
$$\displaystyle E_{ij} = \frac{(\text{row total}) \times (\text{col total})}{\text{grand total}} $$
-
-
Decision: Reject $$\displaystyle H_0 $$ (independence) if $$\displaystyle \chi^2_{calc} > \chi^2_{\alpha, df} $$.
2.4 Regression Analysis
-
Simple Linear Regression: One predictor $X$.
-
Model: $$\displaystyle Y = \beta_0 + \beta_1 X + \epsilon $$
-
$$\displaystyle \beta_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2} $$, $$\displaystyle \beta_0 = \bar{y} - \beta_1 \bar{x} $$
-
-
Multiple Linear Regression: $$\displaystyle Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + ... + \epsilon $$
-
R-squared: Proportion of variance explained. $$\displaystyle R^2 = \frac{SS_{reg}}{SS_{tot}} = 1 - \frac{SS_{res}}{SS_{tot}} $$
-
Adjusted R-squared: Adjusts for number of predictors.
-
Assumptions: Linearity, independence, homoscedasticity, normality of residuals.
2.5 Bayesian Modeling
-
Bayes' Theorem: $$\displaystyle P(A|B) = \frac{P(B|A) P(A)}{P(B)} $$
-
Components:
-
Prior $P(\theta)$: Belief about parameter $\theta$ before data.
-
Likelihood $P(D|\theta)$: Probability of data given $\theta$.
-
Posterior $P(\theta|D)$: Updated belief after data.
-
-
Advantages: Incorporates prior knowledge, full probability distribution, intuitive.
-
Disadvantages: Subjective prior choice, computationally intensive for complex models.
2.6 Maximum Likelihood Estimation (MLE)
-
Objective: Find parameter $\theta$ that maximizes likelihood $$\displaystyle L(\theta) = P(D|\theta) $$.
-
Steps:
-
Write likelihood function for data.
-
Take log-likelihood: $$\displaystyle \ell(\theta) = \log L(\theta) $$ (simplifies multiplication to sum).
-
Differentiate w.r.t. $\theta$, set to zero, solve.
-
Verify maximum (second derivative).
-
-
Properties: Consistent, asymptotically normal, efficient (under regularity conditions).
-
Example: For normal distribution $$\displaystyle N(\mu, \sigma^2) $$, MLE of $\mu$ is $\bar{x}$, $$\displaystyle \sigma^2 $$ is $$\displaystyle \frac{1}{n}\sum (x_i - \bar{x})^2 $$.
2.7 Resampling Methods
-
Bootstrapping:
-
Sample with replacement from original data to create many "bootstrap samples".
-
Compute statistic (e.g., mean) for each sample.
-
Use distribution of bootstrap statistics to estimate standard error, confidence intervals.
-
-
Cross-Validation:
-
k-fold: Split data into k subsets; train on k-1, test on 1; repeat k times.
-
Leave-One-Out (LOO): k = n (extreme case).
-
Used for model evaluation, hyperparameter tuning, preventing overfitting.
-
3.0 Multivariate Analysis
3.1 Definition and Scope
Analysis of multiple variables simultaneously to understand relationships, patterns, and structures. Differs from univariate (one variable) and bivariate (two variables) by handling complexity and interdependencies.
3.2 Key Techniques
-
Principal Component Analysis (PCA):
-
Goal: Dimensionality reduction; transform correlated variables into uncorrelated principal components (PCs).
-
Process: Compute covariance matrix, eigenvectors/values; PCs = eigenvectors ordered by eigenvalue magnitude.
-
DiagramCANVAS: Scatter plot of 2D data with original axes and rotated principal component axes showing maximum variance direction
-
-
Cluster Analysis:
-
k-means: Partition data into k clusters by minimizing within-cluster sum of squares.
- Steps: Initialize centroids, assign points, update centroids, repeat until convergence.
-
Hierarchical: Build tree (dendrogram) of clusters via agglomerative (bottom-up) or divisive (top-down).
-
-
Multivariate Analysis of Variance (MANOVA): Extension of ANOVA for multiple dependent variables; tests if group means differ across a combination of DVs.
-
Factor Analysis: Identifies underlying latent factors explaining correlations among observed variables.
3.3 Applications and Interpretation
-
Pattern Recognition: Group similar images/features (clustering).
-
Data Reduction: PCA for feature extraction before classification.
-
Visualization: Reduce high-dimensional data to 2D/3D for plotting.
-
Interpretation: Loadings (PCA/factor) show variable contributions; cluster assignments reveal natural groupings.
4.0 Data Management and Processing
4.1 Data Wrangling
Process of cleaning, transforming, and organizing raw data for analysis.
-
Gathering: Collect from APIs, web scraping, databases.
-
Cleaning: Handle missing values (impute/remove), outliers (cap/remove), inconsistencies (standardize formats).
-
Transformation: Normalization (scale to [0,1]), standardization (z-score: $$\displaystyle z = \frac{x - \mu}{\sigma} $$), aggregation (summarize).
-
Integration: Merge/join datasets from different sources.
4.2 Big Data Processing Tools
-
HDFS (Hadoop Distributed File System):
-
Architecture: Master-Slave (NameNode manages metadata, DataNodes store blocks).
-
Features: Fault-tolerant, scalable, stores large files across clusters, block size (default 128MB).
-
-
Hive:
-
Data warehousing on Hadoop; provides SQL-like querying (HiveQL).
-
Converts queries to MapReduce/Tez/Spark jobs.
-
-
Spark: In-memory processing, faster than MapReduce for iterative tasks.
-
HBase: NoSQL database on HDFS for real-time read/write access.
4.3 File Formats
| Format | Type | Readability | Size | Speed | Schema | Best For |
|---|---|---|---|---|---|---|
| CSV | Structured | High | Large | Slow | No | Simple exchange |
| JSON | Semi-structured | High | Medium | Medium | Flexible | Web APIs, configs |
| XML | Semi-structured | Medium | Large | Slow | Strict | Document-centric |
| Parquet | Binary columnar | Low | Small | Fast | Yes | Big data analytics (column queries) |
| Avro | Binary row-based | Low | Small | Fast | Yes | Serialization, row-based ops |
[!TIP] Parquet is optimal for read-heavy analytical queries on Hadoop/Spark due to columnar storage and compression.
4.4 Data Management and Indexing
-
DBMS Concepts: Organizes data storage, retrieval, security (e.g., relational vs. NoSQL).
-
Indexing: Speeds up query performance.
-
B-tree: Balanced tree, efficient for range queries, inserts/updates (used in databases).
-
Hash: Key-value lookup, constant time O(1), not for ranges.
-
-
Storage Strategies: Partitioning (horizontal), sharding (distributed), columnar vs. row-oriented.
4.5 Data Analyst Ecosystem
-
Roles:
-
Data Analyst: Focus on descriptive analytics, reporting, visualization.
-
Data Scientist: Predictive modeling, machine learning.
-
Data Engineer: Build/maintain data pipelines, infrastructure.
-
-
Toolchain: Collection (APIs, sensors) → Storage (DB, HDFS) → Processing (SQL, Spark) → Visualization (PowerBI, Python).
-
Workflow: Raw data → Clean/transform → Analyze/Model → Visualize/Report → Insights/Decisions.
5.0 Data Visualization
5.1 Principles of Effective Visualization
-
Chart Selection: Match chart to data story (comparison: bar; distribution: histogram; relationship: scatter).
-
Design Elements:
-
Color: Use sequential for ordered, diverging for deviation, categorical for groups. Avoid rainbow scales.
-
Labeling: Clear axes, titles, legends.
-
Clutter Reduction: Remove non-essential ink (Tufte's data-ink ratio).
-
-
Storytelling: Guide viewer with annotations, logical flow, highlight key insights.
5.2 Univariate Exploration
Visualize single variable distribution.
-
Histogram: Binned frequency counts (choose bin width carefully).
-
Box Plot: Shows median, IQR, outliers (five-number summary).
-
Bar Chart: For categorical frequencies.
-
Density Plot: Smoothed histogram (KDE).
5.3 Bivariate Exploration
Relationships between two variables.
-
Scatter Plot: Numeric vs. numeric; reveals correlation, clusters, non-linearity.
-
Line Chart: Time series (time on x-axis).
-
Heatmap: Correlation matrix; color intensity shows strength/direction.
5.4 Multivariate Exploration
Techniques for >2 variables.
-
Scatterplot Matrix: Grid of scatter plots for all pairs; diagonal often shows histograms.
-
3D Plots: Can be hard to interpret; use cautiously.
-
Faceting (Small Multiples): Split by a categorical variable into multiple plots with shared axes.
-
Parallel Coordinates: Each variable is an axis; lines represent observations. Good for high-D.
-
Encoding: Use color, size, shape, animation to represent additional dimensions.
5.5 Python Visualization Libraries
| Library | Level | Strengths | Typical Use |
|---|---|---|---|
| Pandas | High | Integrated with DataFrames (df.plot(kind='...')) |
Quick exploratory plots |
| ggplot | High | Grammar of graphics (R's ggplot2 port); declarative | Complex, layered plots |
| Matplotlib | Low | Highly customizable; "mother of all libraries" | Full control, publication quality |
| Seaborn | Mid | Statistical plots, beautiful defaults, built on Matplotlib | Visualizing distributions, relationships, categorical data |
[!TIP] Use Seaborn for quick statistical visuals (e.g.,
sns.regplot(),sns.boxplot()); use Matplotlib for fine-tuning.
5.6 PowerBI and Its Tools
-
Power BI Desktop: Free application for report authoring (data modeling, DAX queries, visual design).
-
Power BI Service: Cloud platform for sharing, collaboration, scheduling refreshes.
-
Power BI Report Server: On-premises deployment for organizations with data residency requirements.
-
Key Features: Interactive dashboards, natural language Q&A, AI visuals, robust data connectivity.
5.7 Creating Custom Visualizations
-
When Needed: Standard charts insufficient for unique data structures or storytelling.
-
Tools:
-
D3.js: JavaScript library for web-based, dynamic, interactive visualizations (steep learning curve).
-
Custom Python/Matplotlib: Programmatically create novel plots.
-
-
Process:
-
Design: Sketch, identify encodings, user needs.
-
Prototype: Build simple version (e.g., with Matplotlib).
-
Implement: Refine, add interactivity (with D3.js or Plotly).
-
Validate: Usability testing, accuracy check.
-