UNIT 1: Data and Visual Analytics - Short Notes
1.0 Foundational Concepts & Data Types
Data Analyst Ecosystem
The ecosystem comprises roles, tools, and processes for extracting insights from data.
-
Key Roles: Data Analyst, Data Scientist, Data Engineer, Business Analyst.
-
Core Skills: Statistics, Programming (Python/R/SQL), Domain Knowledge, Visualization, Communication.
-
Workflow: Data Collection → Wrangling → Exploration → Modeling → Visualization → Interpretation.
Variables and Data Categorization
-
Variable: A characteristic, number, or quantity that can be measured or counted.
-
Categorical (Qualitative): Represents categories or groups.
-
Nominal: No inherent order (e.g., Gender, Color).
-
Ordinal: Ordered categories (e.g., Satisfaction: Low, Medium, High).
-
-
Numerical (Quantitative): Represents measurable quantities.
-
Discrete: Countable integers (e.g., Number of students, Defects).
-
Continuous: Measurable on a scale (e.g., Height, Weight, Time).
-
Levels of Measurement
Scales defining the nature of information within variable values.
| Scale | Properties | Operations Allowed | Example |
|---|---|---|---|
| Nominal | Labels only, no order | Equality/inequality, mode | Gender, Zip Code |
| Ordinal | Ordered categories | + above + ranking, median | Likert Scale, Grades |
| Interval | Ordered, equal intervals, no true zero | + above + mean/SD, addition | Temperature (°C), Year |
| Ratio | Has a true zero point | All arithmetic operations | Height, Weight, Income |
[!TIP] Exam Focus: Distinguish Interval (no true zero, e.g., 20°C is not "twice" 10°C) from Ratio (true zero, e.g., 0kg means no weight).
2.0 Descriptive Statistics & Data Exploration
Measures of Central Tendency
Summarizes the center of a dataset.
-
Mean ($\bar{x}$ or $\mu$):
-
Formula (Sample): $$\displaystyle \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$
-
Property: Uses all data, sensitive to outliers.
-
-
Median (M):
-
Middle value when sorted. For even n: avg of two middle values.
-
Property: Robust to outliers, used for ordinal/ratio data.
-
-
Mode:
-
Most frequent value(s).
-
Property: Can be used for nominal data, dataset can be multimodal.
-
Measures of Dispersion (Spread)
Quantifies variability in data.
-
Range: $Max - Min$. Sensitive to outliers.
-
Variance ($$\displaystyle s^2 $$ or $$\displaystyle \sigma^2 $$):
-
Average squared deviation from mean.
-
Sample: $$\displaystyle s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} $$
-
-
Standard Deviation ($s$ or $\sigma$):
- $$\displaystyle \boxed{s = \sqrt{s^2}} $$. Same units as data.
-
Interquartile Range (IQR):
-
$$\displaystyle IQR = Q_3 - Q_1 $$ (spread of middle 50%).
-
Robust to outliers. Used to detect outliers: $$\displaystyle [Q_1 - 1.5 \times IQR,\; Q_3 + 1.5 \times IQR] $$.
-
Data Exploration Techniques
| Technique | Variables Analyzed | Purpose | Common Tools |
|---|---|---|---|
| Univariate | One | Describe distribution (shape, center, spread) | Histogram, Box plot, Summary stats |
| Bivariate | Two | Explore relationship/correlation | Scatter plot, Cross-tab, Correlation coeff. |
| Multivariate Exploration | Multiple (>2) | Initial view of relationships across many variables | Pair plots, Correlation matrix, Heatmap |
| Multivariate Analysis | Multiple | Advanced modeling to understand complex structures (e.g., PCA, Cluster Analysis) | Dimensionality reduction, Clustering algos |
[!TIP] Key Distinction: Multivariate Exploration is descriptive (visualizing many variables). Multivariate Analysis is inferential/modeling (e.g., Multiple Regression, MANOVA).
3.0 Inferential Statistics & Statistical Modeling
Statistical Inferences
Drawing conclusions about a population from a sample.
-
Estimation: Point estimate (e.g., $\bar{x}$) & Interval estimate (Confidence Interval).
-
Hypothesis Testing: Formal procedure to reject/retain a null hypothesis ($$\displaystyle H_0 $$).
Hypothesis Testing Framework
-
State $$\displaystyle H_0 $$ (null) & $$\displaystyle H_1 $$ (alternative).
-
Choose significance level ($\alpha$, usually 0.05).
-
Select appropriate test & compute test statistic.
-
Find p-value or critical value.
-
Decision: Reject $$\displaystyle H_0 $$ if p-value < $\alpha$ (or statistic in rejection region).
Parametric Tests
t-test (compares means, assumes normality & equal variance for two-sample).
-
One-sample: Test if sample mean differs from known $\mu$.
- $$\displaystyle t = \frac{\bar{x} - \mu}{s / \sqrt{n}} $$, df = n-1
-
Independent two-sample: Compare means of two unrelated groups.
- $$\displaystyle t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} $$ (Welch's for unequal variance)
-
Paired: Compare two measurements on same subjects.
- $$\displaystyle t = \frac{\bar{d} - \mu_d}{s_d / \sqrt{n}} $$, where $d$ = differences.
Example (Potato Yield - One-sample t-test from May 2024):
Given: $$\displaystyle \mu = 20 $$, $$\displaystyle X = [21.5, 24.5, ..., 18.5] $$, n=12.
- $$\displaystyle \bar{x} = 20.03 $$, $s \approx 3.22$
- $$\displaystyle H_0: \mu = 20 $$ vs $$\displaystyle H_1: \mu > 20 $$ (one-tailed, "better than")
- $$\displaystyle t = \frac{20.03 - 20}{3.22 / \sqrt{12}} \approx 0.028 $$
- df=11, critical t(0.05, 11) ≈ 1.796.
- Since 0.028 < 1.796, Fail to Reject $$\displaystyle H_0 $$. No significant evidence yield is better.
Chi-Square ($$\displaystyle \chi^2 $$) Test
-
Goodness of Fit: Tests if observed frequencies fit expected distribution.
- $$\displaystyle \boxed{\chi^2 = \sum \frac{(O_i - E_i)^2}{E_i}} $$, df = k-1 (k=categories)
-
Test of Independence: Tests association between two categorical variables in a contingency table.
-
$$\displaystyle \chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}} $$, df = (r-1)(c-1)
-
$$\displaystyle E_{ij} = \frac{(Row_i \; Total) \times (Col_j \; Total)}{Grand \; Total} $$
-
Regression Analysis
Models relationship between dependent (Y) and independent (X) variables.
-
Simple Linear Regression (SLR):
-
Model: $$\displaystyle \boxed{Y = \beta_0 + \beta_1 X + \epsilon} $$
-
$$\displaystyle \beta_1 $$ (slope): Change in Y per unit change in X.
-
$$\displaystyle \beta_0 $$ (intercept): Value of Y when X=0.
-
Estimated via Least Squares: Minimizes $$\displaystyle \sum (y_i - \hat{y}_i)^2 $$.
-
-
Multiple Linear Regression (MLR):
-
Model: $$\displaystyle Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + ... + \beta_p X_p + \epsilon $$
-
Assumptions: Linearity, Independence, Homoscedasticity, Normality of residuals, No multicollinearity.
-
-
Variables in Regression:
-
Dependent (Y): Outcome variable.
-
Independent (X): Predictor variable(s).
-
Dummy Variable: Categorical X converted to binary (0/1) for modeling (e.g., Gender: Male=0, Female=1).
-
Bayesian Modeling
-
Core Principle: Updates probability for a hypothesis as evidence/data becomes available. Uses prior knowledge.
-
Bayes' Theorem:
$$\boxed{P(H|D) = \frac{P(D|H) \cdot P(H)}{P(D)}}$$
* $P(H|D)$: Posterior (updated belief after data D).
* $P(D|H)$: Likelihood (probability of data given hypothesis).
* $P(H)$: Prior (initial belief).
* $P(D)$: Marginal likelihood (normalizing constant).
-
How it works: Start with Prior → Collect Data → Calculate Posterior → New Prior for next iteration.
-
Advantages: Incorporates prior info, intuitive probability statements, good for small data.
-
Disadvantages: Prior selection can be subjective, computationally intensive for complex models.
Maximum Likelihood Estimation (MLE)
-
Concept: Finds parameter values ($\theta$) that maximize the likelihood of observing the given sample data.
-
Procedure:
-
Write likelihood function $$\displaystyle L(\theta) = P(data|\theta) $$ (joint probability of sample).
-
Often use log-likelihood: $$\displaystyle \ell(\theta) = \log L(\theta) $$ (simpler to maximize).
-
Solve $$\displaystyle \frac{d\ell}{d\theta} = 0 $$ for $\theta$.
-
-
Properties: Consistent, Asymptotically normal, Efficient (under regularity conditions).
-
Example (Bernoulli): For n coin flips with k heads, $$\displaystyle L(p) = p^k (1-p)^{n-k} $$. MLE $$\displaystyle \hat{p} = k/n $$.
Resampling Methods
-
Concept: Estimate sampling distribution by repeatedly drawing samples from the observed data (with replacement).
-
Bootstrapping:
-
Create B bootstrap samples (resample n observations with replacement).
-
Calculate statistic (e.g., mean) for each sample.
-
Use distribution of B statistics for CI/SE.
-
-
Cross-Validation:
-
Assess model performance & prevent overfitting.
-
k-fold CV: Split data into k folds, train on k-1, test on 1, repeat k times.
-
4.0 Data Handling, Wrangling & Storage
Data Wrangling (Data Munging)
Process of cleaning, structuring, and enriching raw data.
-
Gathering: Collect data from sources (DBs, APIs, files).
-
Cleaning: Handle missing values, correct errors, remove duplicates.
-
Transformation: Normalize, aggregate, reshape (pivot/melt), encode.
-
Enrichment: Add relevant external data (e.g., geocoding).
-
Validation: Ensure data quality, consistency, and integrity.
File Formats
| Format | Type | Structure | Use Case |
|---|---|---|---|
| CSV | Structured | Tabular, plain text, delimiter-separated | Simple data exchange, spreadsheets |
| Excel (.xlsx) | Structured | Tabular, multiple sheets, formulas, formatting | Business reports, user-friendly analysis |
| JSON | Semi-structured | Key-value pairs, nested objects (tree-like) | Web APIs, NoSQL DBs (MongoDB), configs |
| XML | Semi-structured | Tags/elements, hierarchical, verbose | Legacy systems, document markup, SOAP APIs |
| Unstructured | Unstructured | No fixed schema (text, images, video) | Text mining, image recognition, logs |
Data Management & Indexing
-
Purpose: Efficient storage, retrieval, and update of data.
-
Indexing: Data structure (e.g., B-tree, Hash) that improves speed of data retrieval operations on a table.
-
Similar to a book's index—points to data location without scanning entire table.
-
Trade-off: Faster reads, slower writes, extra storage.
-
Big Data Fundamentals
-
Definition: Datasets too large/complex for traditional DBMS/tools.
-
Characteristics (3V/5V):
-
Volume (size), Velocity (speed of generation), Variety (structured/semi/unstructured).
-
Extended: Veracity (uncertainty), Value (usefulness).
-
-
Challenges: Storage, Processing, Analysis, Visualization, Security.
Big Data Processing Tools & Frameworks
Hadoop Ecosystem
-
HDFS (Hadoop Distributed File System):
-
Architecture: Master-Slave. NameNode (master, metadata), DataNode (slave, stores blocks).
-
Components: Block (default 128MB), Replication (default 3), Rack Awareness.
-
Operations:
hadoop fs -ls,-put,-get,-cat. -
DiagramCANVAS: Show NameNode managing metadata for file split into blocks across 3 DataNodes, with replication to another rack's DataNode
-
-
Hive:
-
Architecture: Metastore (schema/table metadata), HiveServer2 (query execution), Driver/Compiler/Executor.
-
HiveQL (HQL): SQL-like query language, translates to MapReduce/Tez/Spark jobs.
-
Use Cases: Data warehousing, ETL on Hadoop, batch processing of large structured data.
-
-
Other Tools:
-
MapReduce: Programming model (Map → Shuffle/Sort → Reduce) for parallel processing.
-
Spark: In-memory processing, faster than MapReduce for iterative tasks (ML, streaming).
-
5.0 Data Visualization Tools & Libraries
Python Visualization Ecosystem
| Library | Purpose | Key Features | Best For |
|---|---|---|---|
| Pandas | Data manipulation | .plot() method (wraps Matplotlib), quick exploratory plots |
Quick univariate/bivariate plots from DataFrame |
| Matplotlib | Foundation library | Low-level, highly customizable, object-oriented (Figure, Axes) | Static, publication-quality plots, full control |
| Seaborn | Statistical graphics | Built on Matplotlib, attractive defaults, complex plots (violin, heatmap) | Statistical visualizations, distribution relationships |
| Plotly | Interactive, web-based | Interactive (zoom, hover), Dash for web apps, 3D plots | Interactive dashboards, web deployment, complex interactivity |
R Visualization: ggplot2
-
Grammar of Graphics: Build plots layer by layer.
-
Core Components:
-
data: Dataset. -
aes(): Aesthetics (x, y, color, size). -
geom_*(): Geometric objects (points, lines, bars). -
facet_*(): Small multiples. -
theme(): Non-data ink (labels, grid).
-
-
Example:
ggplot(data, aes(x=var1, y=var2)) + geom_point() + geom_smooth()
Business Intelligence Tools: Power BI
-
Components:
-
Power BI Desktop: Free app for data connection, modeling, report creation.
-
Power BI Service: Cloud-based SaaS for sharing, collaboration, dashboards.
-
Power BI Mobile: Apps for iOS/Android to view reports.
-
-
Key Features:
-
Data Modeling: Create relationships between tables, define hierarchies.
-
DAX (Data Analysis Expressions): Formula language for custom calculations & measures (e.g.,
TOTALYTD(),RELATED()). -
Interactive Dashboards: Drill-down, slicers, cross-filtering between visuals.
-
6.0 Synthesis & Application
Creating Custom Visualizations for Complex Datasets
-
Understand Data & Goal: What question are you answering? Know variables, types, relationships.
-
Choose Chart Types:
-
Comparison: Bar, Column.
-
Distribution: Histogram, Box, Violin.
-
Relationship: Scatter, Heatmap, Bubble.
-
Composition: Stacked Bar, Pie (use sparingly), Treemap.
-
Trend: Line, Area.
-
-
Apply Design Principles:
-
Clarity: Minimize clutter (chartjunk).
-
Accuracy: Represent proportions correctly (e.g., area vs. height for circles).
-
Color: Use purposefully (categorical vs. sequential), ensure accessibility.
-
Labeling: Clear axes, titles, legends.
-
-
Tool Selection: Match tool to need (Matplotlib for static, Plotly/Dash for interactive, Power BI for business dashboards).
-
Iterate: Prototype, get feedback, refine.
Integrating Statistical Analysis with Visualization
-
Purpose: Visuals make statistical results intuitive and communicate uncertainty.
-
Methods:
-
Overlay model predictions (regression line, confidence band) on scatter plots.
-
Use box plots/violin plots to show group differences from t-tests/ANOVA.
-
Visualize correlation matrices with heatmaps.
-
Plot residuals to check regression assumptions (linearity, homoscedasticity).
-
Show confidence intervals as error bars on bar charts for group means.
-
-
Example: A scatter plot with a best-fit line (SLR result) and a shaded 95% confidence interval around the line visually communicates both the relationship and its uncertainty.