UNIT 4: Data Exploration and Visualization
This unit focuses on techniques to understand data structure, patterns, and relationships through statistical summaries and graphical representations. It is a critical step before any modeling.
IV.1 Univariate Exploration
Analysis of a single variable to understand its distribution and central tendencies.
| Technique | Purpose & Key Features | Exam Insight |
|---|---|---|
| Histogram | - Visualizes frequency distribution of numerical data.<br>- Bins divide data range; bar height = count/frequency in bin.<br>- Reveals shape: normal, skewed, bimodal. | > [!TIP] Use for continuous numerical data. Choice of bin size significantly impacts interpretation. |
| Box Plot (Whisker Plot) | - Shows five-number summary: min, Q1, median (Q2), Q3, max.<br>- IQR = Q3 - Q1. Whiskers extend to 1.5*IQR; points beyond are outliers.<br>- Excellent for comparing distributions across categories. | Formula for Outliers:<br>Lower fence = Q1 - 1.5IQR<br>Upper fence = Q3 + 1.5IQR |
| Summary Statistics | - Central Tendency: Mean ($\mu$), Median, Mode.<br>- Dispersion: Range, Variance ($$\displaystyle \sigma^2 $$), Standard Deviation ($\sigma$), IQR.<br>- Shape: Skewness, Kurtosis. | > [!CAUTION] Mean is sensitive to outliers; median is robust. Use IQR for skewed data. |
Diagram Suggestion:
DiagramSEARCH: histogram vs box plot comparison
IV.2 Bivariate Exploration
Analysis of two variables to identify relationships.
| Technique | Purpose & Key Features | Exam Insight |
|---|---|---|
| Scatter Plot | - Plots numerical vs. numerical variables.<br>- Reveals relationship direction (positive/negative), strength (tight/loose cluster), and form (linear/non-linear).<br>- Can encode a 3rd variable via color/size. | Correlation Coefficient (Pearson's r):<br> |
$$ r = \frac{\sum_{i=1}^n (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum_{i=1}^n (x_i - \bar{x})^2 \sum_{i=1}^n (y_i - \bar{y})^2}} $$
<br>Range: $r \in [-1, 1]$. |
| Cross-Tabulation (Contingency Table) | - Analyzes categorical vs. categorical variables.<br>- Displays frequency counts/percentages in a matrix.<br>- Foundation for Chi-Square Test of Independence. | > [!TIP] Use row/column percentages to better understand association patterns. | | Heatmap | - Visualizes magnitude of values in a matrix using color intensity.<br>- Commonly used for correlation matrices (r values). | > [!CAUTION] Correlation ≠ Causation! A high r does not imply one variable causes the other. |
Diagram Suggestion:
DiagramSEARCH: scatter plot with correlation examples
IV.3 Multivariate Exploration
Analysis of three or more variables to uncover complex interactions.
| Technique | Purpose & Key Features | Exam Insight |
|---|---|---|
| Pair Plot (Scatter Matrix) | - Grid of scatter plots for all numerical variable pairs.<br>- Diagonal often shows histograms/KDEs of single variables.<br>- Quick overview of pairwise relationships and potential multicollinearity. | Best for datasets with <10 numerical variables. Becomes cluttered otherwise. |
| Heatmap (Extended) | - Visualizes matrix data beyond correlation (e.g., missing values, feature covariance).<br>- Color scale must be chosen carefully (sequential vs. diverging). | |
| Dimensionality Reduction (Intro) | - Techniques to reduce # of variables while preserving structure.<br>- Principal Component Analysis (PCA): Linear transform to orthogonal components (max variance).<br>- t-SNE / UMAP: Non-linear, better for visualization of clusters. | Goal: Combat the "curse of dimensionality" and enable 2D/3D plotting. PCA components are linear combinations of original features. |
Diagram Suggestion:
DiagramCANVAS: A pair plot grid showing histograms on diagonal and scatter plots off-diagonal
IV.4 Visualization Principles
Guidelines for creating effective, truthful visualizations.
-
Chart Selection:
-
Comparison: Bar chart (categories), Line chart (time series).
-
Distribution: Histogram, Box plot, Density plot.
-
Relationship: Scatter plot, Heatmap.
-
Composition: Pie chart (simple parts-of-whole), Stacked bar.
[!TIP] Avoid pie charts for many categories or precise comparisons. Use bar charts instead.
-
-
Storytelling & Design:
-
Clarity over decoration: Remove unnecessary ink ("chartjunk").
-
Label clearly: Axes, titles, units.
-
Use color intentionally: Highlight important data, ensure accessibility (colorblind-friendly palettes).
-
Guide the viewer: Use annotations to point out key insights.
-
IV.5 Custom Visualization for Complex Datasets
Adapting standard techniques or creating new ones for specific data challenges.
-
Geospatial Data: Use choropleth maps (color regions by value) or point maps.
-
Network/Graph Data: Use node-link diagrams (force-directed layouts) or adjacency matrices.
-
Hierarchical Data: Use treemaps or sunburst charts.
-
Time-Series with Multiple Variables: Use small multiples (faceting) or stacked area charts.
-
Interactive Dashboards: Tools like Plotly Dash or Streamlit allow users to filter, zoom, and drill down.
Process: 1) Understand the question you want the viz to answer. 2) Profile the data structure (types, dimensions). 3) Prototype standard charts. 4) Iterate for clarity and insight. 5) Add interactivity if needed for exploration.
Quick Reference: Common Plots & Use Cases
| Plot Type | Primary Data Type | Reveals |
|---|---|---|
| Histogram | 1 Numerical | Distribution shape, skewness, modality |
| Box Plot | 1 Numerical + 1 Categorical (optional) | Summary stats, outliers, comparison across groups |
| Scatter Plot | 2 Numerical | Relationship (correlation, pattern, clusters) |
| Bar Chart | 1 Categorical | Frequency / comparison of categories |
| Line Chart | 1 Numerical + 1 Temporal (ordered) | Trend over time |
| Heatmap | Matrix (e.g., correlation) | Magnitude of values via color intensity |
| Pair Plot | Multiple Numerical | All pairwise relationships at a glance |
\boxed{\text{Core Principle: Always let the data and the analytical question dictate the visualization choice.}}