UNIT 4: Data and Visual Analytics
I. Foundational Data Concepts & Statistics
A. Nature of Data & Variables
-
Variable: A characteristic or attribute that can be measured or quantified.
-
Types of Variables:
-
Categorical (Qualitative):
-
Nominal: No inherent order (e.g., Gender, Color, Country).
-
Ordinal: Has a meaningful order but not fixed intervals (e.g., Education Level, Satisfaction Rating: Low/Medium/High).
-
-
Numerical (Quantitative):
-
Discrete: Countable, finite/integer values (e.g., Number of students, Shoe size).
-
Continuous: Measurable, infinite values within a range (e.g., Height, Weight, Temperature).
-
-
-
Levels of Measurement (Scale of Measurement):
-
Nominal: Classification only.
-
Ordinal: Rank order.
-
Interval: Ordered with equal intervals, but no true zero (e.g., Celsius temperature, IQ score).
-
Ratio: All properties of interval plus a meaningful, non-arbitrary zero (e.g., Weight, Height, Income). Allows for meaningful ratios.
-
B. Descriptive Statistics
-
Measures of Central Tendency (Location): Describe the center of a data distribution.
- Mean ($\bar{x}$ or $\mu$): Sum of all values divided by count. Sensitive to outliers.
$$\bar{x} = \frac{\sum_{i=1}^{n} x_i}{n}$$
* **Median ($M$):** Middle value when sorted. Robust to outliers.
* **Mode:** Most frequently occurring value. Can be multimodal.
> [!TIP] For symmetric distributions, Mean ≈ Median ≈ Mode. For skewed distributions, order is Mode < Median < Mean (positive skew).
-
Measures of Dispersion (Variability): Describe the spread of data.
-
Range: Max - Min. Sensitive to outliers.
-
Variance ($$\displaystyle s^2 $$ for sample, $$\displaystyle \sigma^2 $$ for population): Average squared deviation from the mean.
-
$$s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1}$$
* **Standard Deviation ($s$ or $\sigma$):** Square root of variance. In same units as data.
$$s = \sqrt{s^2} = \sqrt{\frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1}}$$
* **Interquartile Range (IQR):** Spread of middle 50% of data. $$\displaystyle IQR = Q_3 - Q_1 $$. Robust to outliers.
-
Measures of Location:
-
Percentiles: Value below which a given percentage of data falls (e.g., 90th percentile).
-
Quartiles: Specific percentiles: Q1 (25th), Q2 (50th = Median), Q3 (75th).
-
C. Exploratory Data Analysis (EDA)
-
Goal: Summarize main characteristics of data, often with visual methods, to discover patterns, spot anomalies, test hypotheses.
-
Univariate Exploration: Analysis of a single variable.
-
Techniques: Histograms, box plots, density plots, frequency tables.
-
Focus: Distribution shape (symmetry, skewness, modality), central tendency, spread, outliers.
-
-
Bivariate Exploration: Analysis of two variables to explore relationships.
-
Techniques:
-
Numerical vs. Numerical: Scatter plot, correlation coefficient (Pearson's r).
-
Categorical vs. Categorical: Cross-tabulation (contingency table), stacked/mosaic plots.
-
Numerical vs. Categorical: Box plots (by category), violin plots.
-
-
-
Multivariate Exploration: Analysis of three or more variables.
- Techniques: Pair plots (scatter matrix), 3D plots, faceting (small multiples), color/shape/size encoding in 2D plots, dimensionality reduction (PCA for visualization).
[!TIP] EDA is iterative and non-prescriptive. Visualization is its most powerful tool. Always ask: "What story is this data telling?"
II. Statistical Inference & Hypothesis Testing
A. Statistical Inference
-
Definition: The process of using sample data to draw conclusions (make inferences) about an underlying population.
-
Two Main Types:
-
Estimation:
-
Point Estimation: Provide a single value as an estimate of a population parameter (e.g., sample mean $\bar{x}$ estimates population mean $\mu$).
-
Interval Estimation (Confidence Interval): Provide a range of values within which the parameter is likely to lie, with a stated confidence level (e.g., 95% CI for $\mu$).
-
-
Hypothesis Testing: A formal procedure for assessing whether sample evidence is strong enough to reject a null hypothesis ($$\displaystyle H_0 $$) about a population.
-
B. Parametric Tests
-
Assumptions: Data follows a specific distribution (usually normal), homogeneity of variance, independent observations.
-
1. t-test: Compares means. Used when population standard deviation ($\sigma$) is unknown and sample size is small, assuming normality.
- One-sample t-test: Tests if a sample mean ($\bar{x}$) differs from a known/hypothesized population mean ($$\displaystyle \mu_0 $$).
$$t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}}$$
*Degrees of Freedom (df) = n-1.*
> **Example (Potato Yield):** Given $$\displaystyle \mu_0=20 $$, $$\displaystyle n=12 $$, data $X$. Calculate $\bar{x}$, $s$, then $$\displaystyle t_{calc} $$. Compare with $$\displaystyle t_{critical} $$ (from t-table at $$\displaystyle \alpha=0.05 $$, df=11, one-tailed). If $$\displaystyle t_{calc} > t_{crit} $$, reject $$\displaystyle H_0 $$ (yield is significantly better).
* **Two-sample (Independent) t-test:** Compares means of **two independent groups**.
$$t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}}$$
*Use Welch's approximation if variances are unequal.*
* **Paired t-test:** Compares means of **two related samples** (e.g., before/after, matched pairs). Uses differences ($$\displaystyle d_i = x_{1i} - x_{2i} $$).
$$t = \frac{\bar{d}}{s_d / \sqrt{n}}$$
*df = n-1, where n = number of pairs.*
-
2. Chi-Square ($$\displaystyle \chi^2 $$) Test: Used for categorical data. Tests for association or goodness-of-fit.
-
Chi-Square Test for Independence (Association): Tests if two categorical variables are independent.
Steps:
-
State $$\displaystyle H_0 $$: Variables are independent. $$\displaystyle H_1 $$: They are associated.
-
Create contingency table (Observed frequencies, $$\displaystyle O_{ij} $$).
-
Calculate Expected frequencies: $$\displaystyle E_{ij} = \frac{(Row\ Total_i) \times (Column\ Total_j)}{Grand\ Total} $$.
-
Compute test statistic:
-
-
$$\chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}}$$
5. df = $(r-1)(c-1)$ where r=rows, c=columns.
6. Compare $$\displaystyle \chi^2_{calc} $$ with critical value from $$\displaystyle \chi^2 $$ table at chosen $\alpha$.
* **Chi-Square Goodness-of-Fit Test:** Tests if a single categorical variable follows a hypothesized distribution.
$$\chi^2 = \sum \frac{(O_i - E_i)^2}{E_i}$$
*df = (number of categories - 1 - number of estimated parameters).*
C. Non-Parametric & Resampling Methods
-
Re-sampling: Makes no strict assumptions about population distribution.
-
Bootstrap: Repeatedly samples with replacement from the original sample to estimate sampling distribution of a statistic (e.g., mean, median). Used for confidence intervals.
-
Permutation Test: Repeatedly shuffles/permutes group labels to create a null distribution of the test statistic. Compares observed statistic to this distribution.
-
D. Regression Analysis
-
Goal: Model the relationship between a dependent (response) variable and one or more independent (predictor) variables.
-
Simple Linear Regression (SLR): One predictor ($x$).
-
Model: $$\displaystyle y = \beta_0 + \beta_1 x + \epsilon $$
-
$$\displaystyle \beta_0 $$: Intercept (value of y when x=0).
-
$$\displaystyle \beta_1 $$: Slope (change in y for a one-unit change in x).
-
$\epsilon$: Random error.
-
-
Estimation (Ordinary Least Squares - OLS): Finds $$\displaystyle \hat{\beta}_0, \hat{\beta}_1 $$ that minimize sum of squared residuals ($$\displaystyle \sum e_i^2 $$).
-
Interpretation: Sign of $$\displaystyle \hat{\beta}_1 $$ indicates direction. Magnitude indicates strength (in units of y per unit x).
-
-
Multiple Linear Regression (MLR): Two or more predictors.
-
Model: $$\displaystyle y = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + ... + \beta_p x_p + \epsilon $$
-
Interpretation: Holding other variables constant, $$\displaystyle \hat{\beta}_j $$ is the change in y for a one-unit change in $$\displaystyle x_j $$. (Ceteris paribus).
-
Key Assumptions: Linearity, Independence, Homoscedasticity (constant variance), Normality of errors, No multicollinearity.
-
-
Types of Variables in Regression:
-
Dependent (Y): Variable being predicted.
-
Independent (X): Predictor variables.
-
Dummy Variables (Indicator Variables): Binary (0/1) variables used to represent categorical predictors in the model (e.g., Gender: Male=0, Female=1).
-
E. Advanced Inference Techniques
-
1. Maximum Likelihood Estimation (MLE):
-
Concept: A method for estimating model parameters by finding the parameter values that maximize the likelihood function—the probability of observing the given sample data.
-
Steps:
-
Write down the likelihood function $L(\theta|data)$ (joint probability of all observations, often product of PDFs/PMFs).
-
Take the natural logarithm to get Log-Likelihood ($\ell$) for easier differentiation.
-
Differentiate $\ell$ with respect to the parameter(s) $\theta$.
-
Set derivatives to zero and solve for $$\displaystyle \hat{\theta}_{MLE} $$.
-
-
Example (Bernoulli): For $n$ Bernoulli trials with $k$ successes, $$\displaystyle L(p) = p^k (1-p)^{n-k} $$. $$\displaystyle \hat{p}_{MLE} = k/n $$.
[!TIP] MLE is a general method, not a specific test. It's the foundation for many parametric models (including logistic regression).
-
-
2. Bayesian Modelling:
- Core Concept (Bayes' Theorem): Updates prior beliefs about a parameter ($\theta$) with observed data (likelihood) to obtain a posterior belief.
$$P(\theta|Data) = \frac{P(Data|\theta) \cdot P(\theta)}{P(Data)}$$
* **Prior $P(\theta)$:** Belief about $\theta$ before seeing data.
* **Likelihood $P(Data|\theta)$:** Probability of data given $\theta$.
* **Posterior $P(\theta|Data)$:** Updated belief after seeing data.
* **How it Works (vs. Frequentist):**
* **Frequentist:** Parameters are fixed but unknown. Probability is long-run frequency of data.
* **Bayesian:** Parameters are **random variables** with probability distributions. Probability measures **degree of belief**.
* **Advantages:** Incorporates prior knowledge, provides full probability distribution for parameters (credible intervals), intuitive interpretation of intervals.
* **Disadvantages:** Choice of prior can be subjective, computationally intensive for complex models.
III. Data Management & Big Data Ecosystem
A. Data Wrangling / Munging
-
Definition: The process of cleaning, structuring, and enriching raw data into a desired format for analysis. Often takes 60-80% of an analyst's time.
-
Process Steps:
-
Data Acquisition: Gathering data from sources (APIs, databases, files, web scraping).
-
Data Cleaning:
-
Handling missing values (delete, impute with mean/median/mode, model-based).
-
Handling outliers (detect via IQR/Z-score, decide to cap, transform, or remove).
-
Correcting data types, fixing inconsistencies, removing duplicates.
-
-
Data Transformation: Normalization, standardization, binning, encoding categorical variables, creating new features (feature engineering).
-
Data Enrichment: Adding relevant external data or calculated columns.
-
Data Validation: Ensuring data quality, consistency, and integrity for analysis.
-
B. File Formats & Storage
| Category | Formats | Key Characteristics |
|---|---|---|
| Structured | CSV, Excel, JSON (tabular) | Fixed schema, rows & columns. Easy to query/process. CSV: plain text, no types. |
| Semi-structured | JSON, XML | Tags/markers separate elements. Flexible schema. Self-describing. Good for web data. |
| Unstructured | Text files, Images, Video, Audio | No predefined model. Requires NLP/computer vision for analysis. |
| Binary (Columnar) | Parquet, Avro, ORC | Optimized for read/write performance in distributed systems. Columnar storage (Parquet) enables efficient querying of specific columns. |
C. Big Data Processing Tools & Frameworks
-
Hadoop Ecosystem: Framework for distributed storage and processing.
-
HDFS (Hadoop Distributed File System):
-
Architecture: Master-Slave. NameNode (master, metadata) manages DataNodes (slaves, store blocks).
-
Key Concept: Files split into blocks (default 128MB/256MB) and distributed across cluster. Replication (default 3x) for fault tolerance.
-
Advantage: Cost-effective storage of massive datasets on commodity hardware.
-
-
Hive (Data Warehouse Infrastructure):
-
Provides SQL-like interface (HiveQL) to query data stored in HDFS.
-
Translates HiveQL queries into MapReduce/Tez/Spark jobs.
-
Schema-on-read: Apply structure when reading data, not when storing.
-
Use Case: Data summarization, ad-hoc analysis, ETL on large, static datasets.
-
-
-
Other Paradigms:
-
MapReduce: Programming model for processing large datasets with parallel, distributed algorithms. Map (filter/sort) -> Shuffle -> Reduce (aggregate).
-
Apache Spark: In-memory cluster computing. Faster than MapReduce for iterative algorithms. Supports batch, stream, machine learning (MLlib), graph processing. Core abstraction: Resilient Distributed Dataset (RDD).
-
IV. Data Visualization: Tools & Techniques
A. Principles of Data Visualization
-
Definition: The graphical representation of information and data. Purpose is to communicate insights clearly, efficiently, and accurately.
-
Grammar of Graphics (ggplot2 paradigm): A framework that breaks down a graphic into distinct components:
-
Data: The dataset being plotted.
-
Aesthetics (aes): Mappings of data variables to visual properties (x, y, color, size, shape).
-
Geometries (geom): The geometric objects that represent data points (points, lines, bars, boxplots).
-
Facets: Small multiples for conditioning on a variable.
-
Statistics: Statistical transformations (e.g., smoothing, binning for histograms).
-
Coordinates: Coordinate system (cartesian, polar).
-
Themes: Non-data ink (fonts, grids, backgrounds).
-
B. Python Visualization Libraries
| Library | Primary Use Case | Key Features |
|---|---|---|
| Matplotlib | Foundational, low-level plotting. | Highly customizable, object-oriented API. Basis for many other libraries. |
| Seaborn | Statistical visualizations. | Built on Matplotlib. Simplified syntax for complex plots (heatmaps, violin plots, pair plots). Integrates with Pandas DataFrames. |
| Plotly | Interactive, web-based visualizations. | Creates interactive charts (zoom, hover, click). Dash framework for web apps. |
| Pandas | Quick, integrated plotting. | .plot() method on DataFrames/Series. Simple wrapper around Matplotlib. Good for EDA. |
| ggplot | Implementation of Grammar of Graphics. | Based on R's ggplot2. Uses a declarative syntax: ggplot(data, aes(x,y)) + geom_point(). |
C. Business Intelligence Tools
-
Power BI:
-
Definition: Microsoft's suite of business analytics tools.
-
Components:
-
Power BI Desktop: Free Windows application for data connection, transformation (Power Query), modeling, and report creation.
-
Power BI Service: Cloud-based SaaS service for publishing, sharing, and collaborative viewing of reports/dashboards.
-
Power BI Mobile: Apps for iOS/Android to view reports on-the-go.
-
Power BI Report Server: On-premises server for hosting Power BI reports (for organizations with cloud restrictions).
-
-
Key Tools within Ecosystem:
-
Power Query: Data connection and transformation engine (Get & Transform).
-
DAX (Data Analysis Expressions): Formula language for creating custom calculations and measures in data models.
-
Rich Visualizations: Custom visuals marketplace, drill-through, bookmarks.
-
-
D. Creating Custom Visualizations
-
Process for Complex Datasets:
-
Understand the Data & Question: What story needs to be told? What is the key insight?
-
Choose Appropriate Chart Type: Match visualization to relationship (comparison, distribution, composition, relationship).
DiagramCANVAS: Chart selection flowchart based on data question -
Apply Design Principles:
-
Maximize Data-Ink Ratio: Remove non-essential ink (chartjunk).
-
Use Color Strategically: For categorization or sequential data, not decoration. Ensure accessibility.
-
Label Clearly: Axes, titles, data labels. Avoid ambiguity.
-
Tell a Story: Use titles to state the insight, not just "Chart of X vs Y".
-
-
Incorporate Interactivity (if needed): Tooltips, filtering, zooming, linking views (dashboard).
-
Iterate & Test: Get feedback. Is the message clear?
-
V. The Data Analyst Ecosystem
A. Data Analyst Ecosystem
-
Definition: The interconnected set of technologies, processes, and roles involved in turning raw data into actionable insights.
-
Key Components (Data Pipeline):
-
Data Sources: Databases, APIs, flat files, IoT sensors.
-
Storage: Data Lakes (raw), Data Warehouses (structured), Databases (SQL/NoSQL).
-
Processing (ETL/ELT): Extract, Transform, Load (or Extract, Load, Transform). Tools: Apache Airflow, dbt, SSIS.
-
Analysis: Statistical analysis, machine learning, data mining (using Python/R/SQL).
-
Visualization & BI: Creating dashboards/reports (Power BI, Tableau, Python libraries).
-
Consumption: End-users, stakeholders, applications.
DiagramSEARCH: data analytics pipeline diagram etl bi -
B. Data Management & Indexing
-
Database Management Systems (DBMS): Software for storing, retrieving, and managing data.
- Types: Relational (SQL: PostgreSQL, MySQL), Non-Relational (NoSQL: MongoDB, Cassandra), NewSQL.
-
Indexing:
-
Purpose: Dramatically speed up data retrieval (query) operations on a table, at the cost of slower writes and additional storage.
-
How it Works: An index is a separate data structure (often a B-tree or Hash table) that stores the values of a column (or columns) in sorted order, along with pointers to the actual data rows.
-
Analogy: Like an index in a textbook—find a topic quickly without reading every page.
-
Trade-off: More indexes → faster reads, slower inserts/updates/deletes, more disk space. Choose indexes on frequently queried (
WHERE,JOIN,ORDER BY) columns.
-