UNIT 5: Data and Visual Analytics - Short Notes
I. FOUNDATIONAL STATISTICAL CONCEPTS
Measures of Central Tendency
Definition: Single values that represent the center or typical value of a dataset.
- Mean (Arithmetic Average): Sum of all values divided by count.
$$ \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$
*Sensitive to outliers.*
-
Median: Middle value when data is sorted. For even
n, median = average of two middle values.Robust to outliers.
-
Mode: Most frequently occurring value. Can be multimodal.
Applicable to nominal data.
| Measure | Best For | Sensitive to Outliers? |
|---|---|---|
| Mean | Symmetric, numerical data | Yes |
| Median | Skewed data, ordinal data | No |
| Mode | Categorical (nominal/ordinal) data | No |
Measures of Dispersion / Location
-
Range:
Max - Min. Highly sensitive to outliers. -
Variance (Sample): Average squared deviation from mean.
$$ s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} $$
- Standard Deviation (s): Square root of variance. Same units as data.
$$ s = \sqrt{s^2} $$
- Interquartile Range (IQR): Spread of middle 50% of data.
$$ IQR = Q_3 - Q_1 $$
*Robust measure.*
-
Quartiles (Q1, Q2, Q3): Values at 25th, 50th (median), 75th percentiles.
-
Percentiles: Value below which a given percentage of data falls.
Levels of Measurement (Scales of Data)
| Scale | Characteristics | Examples | Permissible Operations |
|---|---|---|---|
| Nominal | Categories only, no order | Gender, Blood Type | Count, Mode, Chi-Square |
| Ordinal | Ordered categories, unequal intervals | Likert Scale, Ranks | All Nominal + Median, Percentiles |
| Interval | Ordered, equal intervals, no true zero | Celsius Temp, IQ | All Ordinal + Mean, SD |
| Ratio | All Interval properties + true zero | Height, Weight, Income | All statistical operations |
Variables and Data Categorization
-
By Nature:
-
Categorical: Nominal (no order), Ordinal (ordered).
-
Numerical: Discrete (countable), Continuous (measurable).
-
-
By Role in Analysis:
-
Independent Variable (Predictor/Feature): Manipulated or used to explain.
-
Dependent Variable (Response/Target): Outcome being measured/predicted.
-
[!TIP] Exam Focus: Be prepared to classify given variables into correct types and scales. Remember: Ratio scale is required for meaningful ratios (e.g., "twice as heavy").
II. INFERENTIAL STATISTICS & HYPOTHESIS TESTING
Statistical Inferences
-
Definition: Using sample data to make conclusions about a larger population.
-
Estimation:
-
Point Estimation: Single value estimate (e.g., sample mean
\bar{x}estimates population mean\mu). -
Interval Estimation (Confidence Interval): Range of values likely to contain population parameter.
-
$$ \text{CI} = \text{Point Estimate} \pm (\text{Critical Value} \times \text{Standard Error}) $$
Parametric Hypothesis Tests
General Steps: 1. State H₀ & H₁. 2. Choose significance level (α). 3. Calculate test statistic. 4. Find p-value/critical value. 5. Decision & Interpretation.
1. t-Test
- One-Sample t-Test: Compares sample mean to a known population mean (μ).
$$ t = \frac{\bar{x} - \mu}{s / \sqrt{n}} $$
*Degrees of Freedom (df) = n-1.*
> **May 2024 Example:** Given potato yield data `X` and `μ=20`.
> 1. Calculate `\bar{x}` and `s`.
> 2. Compute `t` using formula above.
> 3. Compare to critical `t` (df=11, α=0.05, one-tailed) or find p-value.
> 4. **Interpret:** "If p < α, reject H₀. There is significant evidence that yield is better than standard."
-
Two-Sample (Independent) t-Test: Compares means of two independent groups.
Assumes equal variances (pooled) or not (Welch's).
-
Paired t-Test: Compares means of two related samples (e.g., before-after).
$$ t = \frac{\bar{d}}{s_d / \sqrt{n}} $$
where `d` is the difference for each pair.
2. Chi-Square (χ²) Test
- Test for Independence: Tests association between two categorical variables in a contingency table.
$$ \chi^2 = \sum \frac{(O - E)^2}{E} $$
*df = (rows-1)*(cols-1)*
*E = (row total * col total) / grand total*
-
Goodness-of-Fit Test: Tests if observed frequencies fit expected distribution.
df = (categories - 1)
[!TIP] Key Assumption for χ²: Expected frequency
E >= 5for most cells. If violated, use Fisher's Exact Test or combine categories.
Maximum Likelihood Estimation (MLE)
-
Philosophy: Find parameter values (θ) that make the observed sample most probable.
-
Procedure:
-
Write likelihood function
L(θ) = P(data|θ). -
Often work with log-likelihood
ℓ(θ) = ln(L(θ)). -
Maximize
ℓ(θ)by taking derivative w.r.t θ, setting to 0, solving.
-
-
Simple Example (Bernoulli): For
ncoin flips withkheads,L(p) = p^k (1-p)^(n-k). Maximizing gives\hat{p} = k/n.
Re-Sampling Methods
-
Bootstrapping: Resample with replacement from the original sample to create many "bootstrap samples." Use distribution of a statistic (e.g., mean) from these samples to estimate its sampling distribution and confidence intervals.
-
Permutation Test: Resample without replacement by randomly shuffling group labels. Recalculates test statistic to build null distribution. Advantage: Makes fewer distributional assumptions (e.g., normality).
[!TIP] Common Pitfall: MLE can be biased for small samples. Bootstrapping requires the original sample to be representative of the population.
III. REGRESSION & PREDICTIVE MODELING
Regression Analysis
- Simple Linear Regression (SLR): Models linear relationship between one independent (X) and one dependent (Y) variable.
$$ Y = \beta_0 + \beta_1 X + \epsilon $$
* `\beta_0` = intercept, `\beta_1` = slope.
* **Assumptions:** Linearity, Independence, Homoscedasticity, Normality of residuals.
* **Interpretation:** `\beta_1` = change in Y for a 1-unit change in X.
- Multiple Linear Regression (MLR): Extends SLR to multiple predictors.
$$ Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + ... + \beta_p X_p + \epsilon $$
* **Interpretation:** `\beta_j` = effect of `X_j` *holding all other X's constant*.
* **Key Metric:** Adjusted R² (penalizes for extra variables).
Variables in Regression Modeling
-
Dummy Variables (Indicator Variables): Convert categorical variables into binary (0/1) for regression.
- For
kcategories, usek-1dummy variables to avoid dummy variable trap.
- For
-
Interaction Terms: Product of two or more predictors (e.g.,
X1*X2). Captures effect of one predictor depending on the level of another.
Multivariate Analysis
-
Definition: Simultaneous analysis of more than two variables to understand relationships.
-
Distinction from MLR: MLR has one dependent variable. Multivariate analysis often has multiple dependent variables.
-
Common Techniques:
-
MANOVA (Multivariate ANOVA): Extension of ANOVA for multiple DVs.
-
Factor Analysis: Reduces many variables into fewer underlying factors.
-
Cluster Analysis: Groups observations into similar clusters.
-
Primarily conceptual for this unit.
-
Bayesian Modelling
- Core Principle: Bayes' Theorem updates prior belief with data to get posterior belief.
$$ P(\theta|data) = \frac{P(data|\theta) \cdot P(\theta)}{P(data)} $$
* `P(θ)` = Prior (belief before data).
* `P(θ|data)` = Posterior (updated belief).
* `P(data|θ)` = Likelihood.
-
Workflow: 1. Specify Prior
P(θ). 2. Collect Data & define Likelihood. 3. Compute Posterior (often via MCMC sampling). -
Advantages:
-
Incorporates existing knowledge (prior).
-
Provides full probability distribution for parameters (not just point estimates).
-
Intuitive probabilistic interpretation.
-
-
Disadvantages:
-
Computationally intensive (especially for complex models).
-
Choice of prior can be subjective and influence results.
-
Can be slow for large datasets.
-
[!TIP] Key Difference: Frequentist (e.g., MLE, t-test) treats parameters as fixed. Bayesian treats parameters as random variables with distributions.
IV. DATA VISUALIZATION & EXPLORATORY DATA ANALYSIS (EDA)
Data Visualization Fundamentals
-
Purpose: Discover patterns, trends, relationships, outliers; communicate findings.
-
Principles of Effective Viz:
-
Clarity: Message should be immediately understandable.
-
Accuracy: Represent data truthfully (no distortion).
-
Efficiency: Minimize "chartjunk," maximize data-ink ratio (Tufte).
-
Exploration of Data (EDA)
-
Univariate Exploration (1 variable):
-
Categorical: Bar chart, Pie chart (use sparingly), Frequency table.
-
Numerical: Histogram, Box plot (shows IQR, median, outliers), Stem-and-leaf.
-
-
Bivariate Exploration (2 variables):
-
Num-Num: Scatter plot (primary), Correlation coefficient (r).
-
Cat-Cat: Stacked/Grouped bar chart, Mosaic plot.
-
Num-Cat: Box plot (by category), Violin plot.
-
-
Multivariate Exploration (>2 variables):
-
Encodings: Use color, shape, size, facet (small multiples) to add dimensions to 2D plots.
-
Techniques: Pair plot (scatter matrix), 3D plots (use cautiously), Heatmaps (for correlation matrices).
-
Creating Custom Visualizations for Complex Datasets
Process:
-
Define Question: What relationship/story are you investigating?
-
Choose Chart Type: Match question to visual encoding (e.g., trend over time → line chart; part-to-whole → stacked bar).
-
Design for Clarity: Label axes, add title, use appropriate color scales (sequential for ordered, diverging for deviation from center, categorical for distinct groups).
-
Iterate & Refine: Simplify, remove non-data ink, ensure accessibility (colorblind-friendly palettes).
[!TIP] Common Pitfall: Using 3D pie charts or excessive colors that distort perception. Always prefer 2D and test if the visualization answers the question without explanation.
V. PROGRAMMATIC VISUALIZATION & TOOLS
Python Visualization Libraries
-
Matplotlib: Foundation library. Highly customizable but verbose. Good for static, publication-quality plots.
plt.plot(x, y); plt.xlabel(); plt.title() -
Seaborn: Built on Matplotlib. Simplifies statistical plots with beautiful defaults.
-
relplot()(scatter/line),catplot()(categorical),displot()(distributions). -
Key Plots:
heatmap(),violinplot(),pairplot().
-
-
Plotly/Bokeh: Interactive visualizations (zoom, hover, click). Ideal for web dashboards.
- Plotly (
plotly.express) is more user-friendly; Bokeh offers more control.
- Plotly (
-
Pandas Integration: Quick plotting via
.plot()method on DataFrames/Series (uses Matplotlib backend).df['column'].plot(kind='hist')
Power BI Ecosystem
-
Power BI Desktop: Primary tool. Free application for:
-
Data connection & transformation (Power Query Editor).
-
Data modeling (relationships, DAX formulas).
-
Report creation (drag-and-drop visuals).
-
-
Power BI Service: Cloud platform (
app.powerbi.com).-
Publish, share, collaborate on reports/dashboards.
-
Schedule data refreshes, set up row-level security.
-
-
Power BI Mobile: Consume reports on iOS/Android devices.
-
Power BI Report Server: On-premises server for hosting reports internally (for organizations with strict data governance).
ggplot2 (R)
-
Philosophy: Grammar of Graphics. Build plots layer by layer.
-
Core Syntax:
ggplot(data = df, aes(x = var1, y = var2)) + # 1. Data & Aesthetics geom_point() + # 2. Geometric object (layer) facet_wrap(~category) + # 3. Facets (subplots) theme_minimal() # 4. Theme (styling) -
Key Components:
ggplot()(init),aes()(mapping),geom_*()(marks),facet_*()(panels),scale_*()(adjust axes/color),theme()(non-data ink).
[!TIP] Python vs. R: Matplotlib/Seaborn = imperative (plot step-by-step). ggplot2 = declarative (describe the final plot components). Plotly is the closest Python equivalent to ggplot2's layered approach.
VI. BIG DATA TECHNOLOGIES & ECOSYSTEM
Big Data Fundamentals
-
Definition: Datasets too large, fast, or complex for traditional DBMS.
-
Vs (Extended):
-
Volume: Size (TB/PB).
-
Velocity: Speed of generation/processing.
-
Variety: Structured, semi-structured, unstructured.
-
Veracity: Uncertainty, quality, trustworthiness.
-
Value: Ultimate goal—extracting useful insights.
-
Big Data Processing Tools
Hadoop Ecosystem:
-
HDFS (Hadoop Distributed File System):
-
Architecture: Master-Slave. NameNode (master, metadata), DataNodes (slaves, store blocks).
-
File Storage: Split into blocks (default 128MB/256MB), replicated (default 3x) across nodes.
-
Advantage: Fault-tolerant, scalable, cost-effective (commodity hardware).
-
-
MapReduce:
-
Programming Model:
Map(filter/sort → key-value pairs) →Shuffle & Sort→Reduce(aggregate). -
Limitation: High latency due to disk I/O between stages.
-
-
Hive:
-
Purpose: Data warehouse infrastructure on Hadoop.
-
HiveQL: SQL-like query language. Translates queries to MapReduce/Tez/Spark jobs.
-
Metastore: Stores schema information (table definitions).
-
Apache Spark:
-
Core Innovation: In-memory processing. Caches data in RAM across cluster.
-
Key Abstractions:
-
RDD (Resilient Distributed Dataset): Immutable, partitioned collection. Fault-tolerant via lineage.
-
DataFrame/Dataset: Higher-level API (like Pandas/SQL tables). Optimized with Catalyst optimizer.
-
-
Speed: 10-100x faster than MapReduce for iterative algorithms (ML, graph) due to reduced disk I/O.
-
Components: Spark SQL, MLlib, Spark Streaming, GraphX.
[!TIP] Hadoop vs. Spark: Hadoop (HDFS+MapReduce) is disk-based, good for huge batch jobs. Spark is in-memory, good for iterative processing, streaming, and SQL. Often used together: Spark reads from HDFS.
VII. DATA WRANGLING & MANAGEMENT
Data Wrangling (Data Munging)
Critical, time-consuming (~60-80% of analyst time) process of making raw data usable.
Detailed Process:
-
Data Acquisition:
-
Sources: CSV/Excel, SQL databases, APIs (REST), Web scraping, Sensors.
-
Tools:
pandas.read_csv(),pd.read_sql(),requestslibrary.
-
-
Data Cleaning:
-
Missing Values: Identify (
isnull()), Handle (delete, impute with mean/median/mode, or advanced methods). -
Outliers: Detect (IQR method:
Q1 - 1.5*IQR,Q3 + 1.5*IQR; Z-score), Treat (cap, transform, or keep if valid). -
Data Types: Correct (
astype()), parse dates (pd.to_datetime()). -
Duplicates: Identify & remove (
drop_duplicates()).
-
-
Data Transformation:
- Normalization (Min-Max): Scale to [0,1].
$$ x' = \frac{x - x_{min}}{x_{max} - x_{min}} $$
* **Standardization (Z-score):** Mean=0, SD=1.
$$ z = \frac{x - \mu}{\sigma} $$
* **Binning:** Convert continuous to categorical (e.g., age groups).
* **Derived Variables:** Create new columns from existing ones (e.g., `profit = revenue - cost`).
* **Reshaping:** `pivot()` (wide to long), `melt()` (long to wide).
-
Data Integration:
-
Merging/Joining: Combine datasets from different sources.
pd.merge(df1, df2, on='key')(SQL-like joins: inner, left, right, outer).
-
Concatenation:
pd.concat([df1, df2])(stack vertically/horizontally).
-
Data Management and Indexing
-
DBMS: Software to store, retrieve, manage data (e.g., MySQL, PostgreSQL).
-
Indexing: Data structure (like B-tree) that speeds up query retrieval at the cost of extra storage and slower writes.
-
Primary Index: On primary key (unique, not null).
-
Secondary (Clustered) Index: Determines physical order of data (one per table).
-
Secondary (Non-Clustered) Index: Separate structure pointing to data rows (multiple allowed).
-
File Formats
| Format | Type | Schema? | Human-Readable? | Best For | Big Data? |
|---|---|---|---|---|---|
| CSV | Structured | No (implied) | Yes | Simple exchange, small data | No (no compression) |
| JSON | Semi-structured | Flexible | Yes | Web APIs, nested data | No (text-based, verbose) |
| Parquet | Structured | Yes (defined) | No | Columnar storage in Hadoop/Spark. High compression, efficient querying (read only needed columns). | Yes |
| Avro | Structured | Yes (in file) | No | Row-based, serialization in Hadoop. Schema evolution. | Yes |
| XML | Semi-structured | Flexible | Yes (verbose) | Legacy systems, documents | No |
[!TIP] Rule of Thumb: Use Parquet for analytical queries in big data (columnar). Use CSV/JSON for small, simple, or web-based interchange. Use Avro for row-based serialization in streaming.
VIII. DATA ANALYST ECOSYSTEM & WORKFLOW
Data Analyst Ecosystem Overview
A toolchain supporting the full data lifecycle:
-
Storage: SQL (PostgreSQL), NoSQL (MongoDB), HDFS (big data).
-
Processing & Analysis: Python (pandas, NumPy, SciPy), R, Spark (PySpark).
-
Visualization & Reporting: Power BI, Tableau, Python libs (Seaborn, Plotly), ggplot2.
-
Collaboration & Reproducibility: Git (version control), Jupyter Notebooks/Lab, R Markdown.
End-to-End Workflow Integration
-
Acquire: Connect to source (SQL DB, CSV, API →
pandas.read_sql(),requests). -
Wrangle: Clean & transform (Python/R:
pandas,dplyr). For big data, use Spark DataFrames. -
Analyze: Exploratory analysis (visualization, summary stats) → Inferential stats (t-test, χ²) → Modeling (regression, ML).
-
Visualize & Report: Create static/interactive dashboards (Power BI, Plotly Dash) or notebooks for narrative.
-
Deploy & Share: Publish to Power BI Service, schedule refreshes, or share notebooks via GitHub.
[!TIP] Exam Focus: Be able to map tools to workflow stages. Example: "For a 50GB CSV file, which tool for wrangling?" → Spark (PySpark) because pandas would run out of memory. "For creating an interactive sales dashboard for management?" → Power BI or Plotly Dash.
\boxed{\text{End of Unit 5 Notes}} Based on RGPV May 2024 & May 2023 exam analysis. Focus on definitions, formulas, step-by-step procedures (t-test, χ²), tool comparisons, and workflow mapping.