UNIT 2: Data and Visual Analytics - Short Notes
I. Foundational Statistics and Data Types
Measures of Central Tendency
-
Definition: Single value representing the center point of a dataset.
-
Types:
-
Mean (Arithmetic Average): $$\displaystyle \bar{x} = \frac{\sum x_i}{n} $$. Sensitive to outliers.
-
Median: Middle value in sorted data. Robust to outliers.
-
Mode: Most frequent value. Used for nominal data.
-
-
Properties & Applications:
-
Mean: Best for symmetric, interval/ratio data.
-
Median: Preferred for skewed distributions or ordinal data.
-
Mode: Only measure for nominal data; can be multimodal.
-
Measures of Dispersion (Spread)
-
Definition: Quantify variability or spread in data.
-
Types:
-
Range: $Max - Min$. Sensitive to outliers.
-
Variance ($$\displaystyle s^2 $$ for sample, $$\displaystyle \sigma^2 $$ for population): $$\displaystyle s^2 = \frac{\sum (x_i - \bar{x})^2}{n-1} $$. Measures average squared deviation.
-
Standard Deviation ($s$ or $\sigma$): $\sqrt{Variance}$. Same units as data.
-
Interquartile Range (IQR): $Q3 - Q1$. Measures spread of middle 50%; resistant to outliers.
-
-
Applications: Standard deviation used in z-scores, confidence intervals. IQR for box plots.
Levels of Measurement
| Level | Characteristics | Examples | Operations Allowed |
|---|---|---|---|
| Nominal | Categories only, no order | Gender, Blood type | Count, Mode |
| Ordinal | Ordered categories, unequal intervals | Likert scale, Education level | Count, Median, Percentiles |
| Interval | Ordered, equal intervals, no true zero | Temperature (°C), IQ score | Mean, Std Dev (but ratios meaningless) |
| Ratio | Interval + true zero, meaningful ratios | Height, Weight, Age | All operations including geometric mean |
[!TIP] Exam Focus: Distinguish Interval vs Ratio (true zero). Nominal/Ordinal limit statistical tests.
Variables and Data Categorization
-
Types of Variables:
-
Categorical (Qualitative): Nominal or Ordinal.
-
Numerical (Quantitative):
-
Discrete: Countable integers (e.g., number of children).
-
Continuous: Measurable on a scale (e.g., weight, time).
-
-
-
Data Classification Methods:
-
By Nature: Primary (collected firsthand) vs Secondary (existing sources).
-
By Structure: Structured (tabular), Semi-structured (JSON/XML), Unstructured (text/images).
-
II. Exploratory Data Analysis (EDA)
Univariate Exploration
-
Goal: Understand distribution of a single variable.
-
Techniques:
-
Summary Statistics: Mean, median, std dev, IQR, min/max.
-
Visualizations:
-
Histogram: Shows frequency distribution; choose bin size carefully.
-
Box Plot: Displays median, quartiles, outliers; ideal for comparing groups.
-
KDE Plot: Smoothed version of histogram.
-
Bar Chart: For categorical data frequencies.
-
-
Bivariate Exploration
-
Goal: Explore relationship between two variables.
-
Techniques:
-
Scatter Plot: For two numerical variables; reveals correlation, patterns, outliers.
-
Correlation Coefficient ($r$): $-1 \leq r \leq 1$. Linear relationship strength/direction.
-
Cross-Tabulation (Contingency Table): For two categorical variables; shows frequencies.
-
Line Plot: For time-series data (time vs metric).
-
Multivariate Exploration
-
Goal: Analyze interactions among three or more variables.
-
Techniques:
-
Pair Plots (Scatter Matrix): Grid of scatter plots for all numerical variable pairs; diagonal shows histograms/KDEs.
-
Heatmap: Visualizes correlation matrix; colors indicate strength/direction.
-
Dimensionality Reduction:
-
PCA (Principal Component Analysis): Reduces dimensions to 2D/3D while preserving variance; used in scatter plots.
-
t-SNE: Non-linear reduction for high-dimensional data visualization.
-
-
3D Plots/Animation: Add third variable via color, size, or motion.
-
[!TIP] Common Pitfall: Correlation ≠ Causation. Always check for confounding variables in bivariate analysis.
III. Inferential Statistics and Hypothesis Testing
Statistical Inferences
-
Definition: Drawing conclusions about a population from a sample.
-
Types:
-
Estimation:
-
Point Estimate: Single value (e.g., $\bar{x}$ for $\mu$).
-
Confidence Interval (CI): Range $[\bar{x} - E, \bar{x} + E]$, where $$\displaystyle E = t_{\alpha/2} \cdot \frac{s}{\sqrt{n}} $$ (for means). Confidence level (e.g., 95%) indicates long-run success rate.
-
-
Hypothesis Testing:
-
Null Hypothesis ($$\displaystyle H_0 $$): Status quo (e.g., $$\displaystyle \mu = \mu_0 $$).
-
Alternative Hypothesis ($$\displaystyle H_a $$): Claim to test (one-tailed or two-tailed).
-
p-value: Probability of observing sample result (or more extreme) if $$\displaystyle H_0 $$ true. Reject $$\displaystyle H_0 $$ if p-value < $\alpha$ (significance level, e.g., 0.05).
-
-
t-Test
-
Used when: Population std dev unknown, sample size small ($$\displaystyle n < 30 $$), data approximately normal.
-
Types:
-
One-Sample t-test: Compare sample mean to known $$\displaystyle \mu_0 $$.
- $$\displaystyle t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} $$, df = $n-1$.
-
Independent Two-Sample t-test: Compare means of two independent groups.
-
Equal variances: Pooled variance $$\displaystyle s_p^2 $$, df = $$\displaystyle n_1 + n_2 - 2 $$.
-
Unequal variances (Welch's): Separate variances, approximate df.
-
-
Paired t-test: Compare two measurements on same subjects (e.g., before/after).
- $$\displaystyle t = \frac{\bar{d}}{s_d / \sqrt{n}} $$, where $d$ = differences, df = $n-1$.
-
-
Assumptions: Normality (or large n), independence, homogeneity of variance (for two-sample equal var).
-
Example: Potato Yield Problem (One-sample, one-tailed):
-
$$\displaystyle H_0: \mu = 20 $$, $$\displaystyle H_a: \mu > 20 $$.
-
Data: $$\displaystyle X = [21.5, 24.5, ..., 18.5] $$, $$\displaystyle n=12 $$.
-
$$\displaystyle \bar{x} = 19.79 $$, $$\displaystyle s = 2.98 $$.
-
$$\displaystyle t = \frac{19.79 - 20}{2.98 / \sqrt{12}} = -0.25 $$.
-
df = 11, critical $$\displaystyle t_{0.05,11} = 1.796 $$ (one-tailed). Since $$\displaystyle |t| < 1.796 $$, fail to reject $$\displaystyle H_0 $$. Not significantly better.
-
Chi-Square Test ($$\displaystyle \chi^2 $$)
-
Used for: Categorical data.
-
Types:
-
Goodness-of-Fit: Compare observed frequencies to expected frequencies from a known distribution.
- $$\displaystyle \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$, df = $k-1$ (k = categories).
-
Test for Independence: Assess association between two categorical variables in a contingency table.
-
$$\displaystyle \chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}} $$, df = $(r-1)(c-1)$.
-
$$\displaystyle E_{ij} = \frac{(row \ total_i) \times (col \ total_j)}{grand \ total} $$.
-
-
-
Assumptions: Independent observations, expected frequencies $\geq 5$ (for 80%+ cells).
-
Steps: State hypotheses, compute $$\displaystyle \chi^2 $$, find critical value/p-value, conclude.
Re-sampling Methods
-
Goal: Estimate sampling distribution without parametric assumptions.
-
Bootstrap:
-
Sample with replacement from original data to create many "bootstrap samples".
-
Compute statistic (e.g., mean) for each.
-
Use distribution of bootstrap statistics for CI or SE.
-
-
Cross-Validation:
-
k-fold CV: Split data into k folds; train on k-1, test on 1; repeat k times. Average performance.
-
Used for model evaluation and hyperparameter tuning.
-
-
Applications: Bootstrap for CI on medians; CV for preventing overfitting.
Maximum Likelihood Estimation (MLE)
-
Concept: Find parameter values that maximize the likelihood of observing the given sample.
-
Steps:
-
Write likelihood function $$\displaystyle L(\theta) = \prod f(x_i|\theta) $$ (joint PDF/PMF).
-
Take log: $$\displaystyle \ell(\theta) = \log L(\theta) = \sum \log f(x_i|\theta) $$.
-
Differentiate w.r.t. $\theta$, set to 0, solve.
-
-
Example: For i.i.d. Normal($$\displaystyle \mu, \sigma^2 $$), MLE for $\mu$ is $\bar{x}$, for $$\displaystyle \sigma^2 $$ is $$\displaystyle \frac{\sum (x_i - \bar{x})^2}{n} $$ (biased).
-
Properties: Consistent, asymptotically normal, but may be biased in small samples.
IV. Predictive and Advanced Analytics
Regression Analysis
-
Goal: Model relationship between dependent variable (Y) and independent variable(s) (X).
-
Simple Linear Regression (SLR): $$\displaystyle Y = \beta_0 + \beta_1 X + \epsilon $$.
-
$$\displaystyle \hat{\beta}_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2} $$, $$\displaystyle \hat{\beta}_0 = \bar{y} - \hat{\beta}_1 \bar{x} $$.
-
$$\displaystyle R^2 $$: Proportion of variance in Y explained by X.
-
-
Multiple Linear Regression (MLR): $$\displaystyle Y = \beta_0 + \beta_1 X_1 + ... + \beta_p X_p + \epsilon $$.
-
Variables:
-
Dependent: Target variable (continuous).
-
Independent: Predictors (continuous or categorical).
-
Dummy Variables: Binary (0/1) for categorical predictors (e.g., Gender: Male=1, Female=0). Reference category omitted to avoid multicollinearity.
-
-
Interpretation: $$\displaystyle \beta_j $$ = change in Y per unit change in $$\displaystyle X_j $$, holding others constant.
-
-
Assumptions: Linearity, independence, homoscedasticity, normality of residuals, no multicollinearity.
Multivariate Analysis
-
Goal: Simultaneously analyze multiple outcome variables.
-
Techniques:
-
MANOVA (Multivariate ANOVA): Extension of ANOVA for multiple dependent variables; tests if group means differ on a combination of DVs.
-
Factor Analysis: Reduce many variables to fewer latent factors (e.g., "socioeconomic status" from income, education).
-
Cluster Analysis: Group similar observations (e.g., customer segmentation). Methods: K-means (distance-based), hierarchical.
-
-
Applications: Market research, psychology, bioinformatics.
Bayesian Modeling
-
Core: Bayes' Theorem: $$\displaystyle P(A|B) = \frac{P(B|A) \cdot P(A)}{P(B)} $$.
-
$P(A)$: Prior belief about A.
-
$P(A|B)$: Posterior belief after observing B.
-
$P(B|A)$: Likelihood.
-
-
How Bayesian Inference Works:
-
Specify prior distribution for parameter(s).
-
Collect data, compute likelihood.
-
Update prior with likelihood to get posterior distribution (using MCMC or conjugate priors).
-
Posterior becomes new prior for sequential updating.
-
-
Advantages:
-
Incorporates prior knowledge.
-
Provides full probability distribution (uncertainty quantification).
-
Intuitive interpretation of credible intervals.
-
-
Disadvantages:
-
Choice of prior can be subjective.
-
Computationally intensive for complex models.
-
Less familiar to traditional frequentist audiences.
-
V. Data Management and Engineering
Data Wrangling Process
-
Gathering: Collect from databases, APIs, files, web scraping.
-
Cleaning: Handle missing values (impute/remove), correct errors, remove duplicates.
-
Transformation: Normalize, aggregate, pivot, create new features.
-
Enrichment: Merge with external data sources.
-
Validation: Ensure data quality, consistency, and integrity.
- Tools: Python (Pandas, NumPy), R (dplyr), OpenRefine, SQL.
File Formats
| Category | Formats | Characteristics | Use Cases |
|---|---|---|---|
| Structured | CSV, Excel | Tabular, fixed schema, rows/columns | Small-medium datasets, spreadsheets |
| Semi-structured | JSON, XML, Parquet | Tags/markers separate elements; schema flexible | Web data, configs, big data (Parquet) |
| Unstructured | Text, Images, Videos | No predefined schema | Social media, logs, multimedia |
| Columnar | Parquet, ORC | Store by column; efficient for analytics | Big data processing (Hive/Spark) |
[!TIP] Exam Key: Parquet is columnar & compressed → faster queries on specific columns vs row-based CSV.
Data Management and Indexing
-
Database Concepts:
-
Relational (SQL): Tables, ACID transactions, primary/foreign keys.
-
NoSQL: Document (MongoDB), Key-Value (Redis), Column-family (Cassandra), Graph (Neo4j). Scale horizontally, flexible schema.
-
-
Indexing: Data structure (e.g., B-tree) to speed up queries.
-
Clustered Index: Determines physical order of data (one per table).
-
Non-Clustered Index: Separate structure pointing to data rows.
-
Trade-off: Faster reads, slower writes, extra storage.
-
-
Query Optimization: Use indexes, avoid
SELECT *, join order, analyze execution plans.
Big Data Processing Tools
-
Hadoop Ecosystem:
-
HDFS (Hadoop Distributed File System):
-
Architecture: Master-Slave. NameNode (master, metadata), DataNodes (slaves, store blocks). Files split into 128MB blocks, replicated (default 3).
-
Function: Fault-tolerant, scalable storage for large files.
-
-
Hive:
-
Data warehousing on Hadoop. Provides SQL-like language (HiveQL).
-
Translates queries to MapReduce/Tez/Spark jobs. Schema-on-read (vs schema-on-write).
-
-
-
Other Tools:
-
Spark: In-memory processing; faster than MapReduce for iterative tasks. Supports SQL, Streaming, MLlib.
-
NoSQL Databases: MongoDB (document), Cassandra (wide-column) for high write throughput and scalability.
-
Data Analyst Ecosystem
-
Components:
-
Data Sources: Databases, APIs, logs, IoT.
-
Ingestion/ETL: Tools like Apache NiFi, Talend, Stitch.
-
Storage: Data lakes (S3, HDFS), warehouses (Redshift, Snowflake), databases.
-
Processing/Transformation: Spark, Pandas, SQL.
-
Analysis/ML: Python/R (scikit-learn, statsmodels), SAS.
-
Visualization: Tableau, Power BI, Matplotlib, Seaborn.
-
Orchestration: Airflow, Luigi.
-
-
Roles:
-
Data Analyst: Focus on descriptive/predictive analytics, reporting, visualization.
-
Data Engineer: Build/maintain data pipelines, infrastructure.
-
Data Scientist: Advanced modeling, ML, statistical inference.
-
-
Workflow Integration: Analysts often use cleaned data from engineers; visualize outputs for stakeholders.
VI. Visualization Tools and Implementation
Python Visualization Libraries
| Library | Primary Use | Key Features | Best For |
|---|---|---|---|
| Matplotlib | Foundational 2D plotting | Highly customizable, object-oriented API | Static, publication-quality plots |
| Seaborn | Statistical visuals | Built on Matplotlib, default themes, simplifies complex plots (heatmaps, violin plots) | Exploratory data analysis |
| Plotly | Interactive visuals | Web-based, zoom/pan, hover, Dash for apps | Dashboards, web apps |
| Pandas | Integrated plotting | .plot() method uses Matplotlib backend |
Quick EDA on DataFrames |
| ggplot | Grammar of Graphics (R port) | Layered approach: data → aesthetics → geometries | Consistent, complex multi-plot figures |
Power BI Ecosystem
-
Power BI Desktop: Free Windows application for report creation. Connect to data sources, model data (Power Query, DAX), design interactive visuals.
-
Power BI Service: Cloud platform (SaaS) for publishing, sharing, collaboration. Apps, dashboards, scheduled refresh.
-
Power BI Tools:
-
Power BI Report Server: On-premises deployment for organizations with cloud restrictions.
-
Power BI Gateway: Bridge between on-premises data sources and Power BI Service.
-
Power BI Mobile: iOS/Android apps for viewing reports.
-
Power BI Embedded: Integrate visuals into custom apps (developer-focused).
-
-
Workflow: Desktop (develop) → Publish to Service (share) → Mobile/Gateway (access).
Designing Custom Visualizations
-
Principles for Complex Datasets:
-
Know the audience: Tailor complexity.
-
Choose appropriate chart type: Match data story (comparison, distribution, relationship, composition).
-
Simplify: Remove clutter (chartjunk), use color purposefully.
-
Ensure clarity: Label axes, add titles, use legends.
-
Enable interaction: Tooltips, filters, drill-down for large datasets.
-
-
Tool Selection:
-
Static/Publication: Matplotlib, ggplot.
-
Interactive/Dashboards: Plotly Dash, Power BI, Tableau.
-
Geospatial: Folium, Plotly Express.
-
-
Design Process:
-
Define question/goal.
-
Explore data (EDA).
-
Sketch visualization concepts.
-
Prototype in chosen tool.
-
Iterate based on feedback.
-
Add storytelling: annotations, logical flow, narrative.
-
-
Storytelling with Data: Guide viewer with title, annotations, sequential slides. Emphasize insights, not just data.
[!TIP] Exam Tip: For "designing custom visuals", emphasize clarity, audience, and narrative flow over technical complexity. Mention specific tools (e.g., "use Plotly for interactive web-based storytelling").
These notes consolidate key definitions, formulas, and concepts from the UNIT 2 syllabus, aligned with past RGPV exam patterns. Focus on understanding applications and distinctions (e.g., measurement levels, regression types, tool comparisons).