Skip to content
AL-603 (B) · Data and Visual Analytics/Quick Revision Short Notes

Data and Visual Analytics (AL-603 (B)) - Unit 1 Short Notes

UNIT 1: Data and Visual Analytics - Short Notes

1.0 Foundational Concepts & Data Types

Data Analyst Ecosystem

The ecosystem comprises roles, tools, and processes for extracting insights from data.

  • Key Roles: Data Analyst, Data Scientist, Data Engineer, Business Analyst.

  • Core Skills: Statistics, Programming (Python/R/SQL), Domain Knowledge, Visualization, Communication.

  • Workflow: Data Collection → Wrangling → Exploration → Modeling → Visualization → Interpretation.

Variables and Data Categorization

  • Variable: A characteristic, number, or quantity that can be measured or counted.

  • Categorical (Qualitative): Represents categories or groups.

    • Nominal: No inherent order (e.g., Gender, Color).

    • Ordinal: Ordered categories (e.g., Satisfaction: Low, Medium, High).

  • Numerical (Quantitative): Represents measurable quantities.

    • Discrete: Countable integers (e.g., Number of students, Defects).

    • Continuous: Measurable on a scale (e.g., Height, Weight, Time).

Levels of Measurement

Scales defining the nature of information within variable values.

Scale Properties Operations Allowed Example
Nominal Labels only, no order Equality/inequality, mode Gender, Zip Code
Ordinal Ordered categories + above + ranking, median Likert Scale, Grades
Interval Ordered, equal intervals, no true zero + above + mean/SD, addition Temperature (°C), Year
Ratio Has a true zero point All arithmetic operations Height, Weight, Income

[!TIP] Exam Focus: Distinguish Interval (no true zero, e.g., 20°C is not "twice" 10°C) from Ratio (true zero, e.g., 0kg means no weight).


2.0 Descriptive Statistics & Data Exploration

Measures of Central Tendency

Summarizes the center of a dataset.

  • Mean ($\bar{x}$ or $\mu$):

    • Formula (Sample): $$\displaystyle \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$

    • Property: Uses all data, sensitive to outliers.

  • Median (M):

    • Middle value when sorted. For even n: avg of two middle values.

    • Property: Robust to outliers, used for ordinal/ratio data.

  • Mode:

    • Most frequent value(s).

    • Property: Can be used for nominal data, dataset can be multimodal.

Measures of Dispersion (Spread)

Quantifies variability in data.

  • Range: $Max - Min$. Sensitive to outliers.

  • Variance ($$\displaystyle s^2 $$ or $$\displaystyle \sigma^2 $$):

    • Average squared deviation from mean.

    • Sample: $$\displaystyle s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} $$

  • Standard Deviation ($s$ or $\sigma$):

    • $$\displaystyle \boxed{s = \sqrt{s^2}} $$. Same units as data.
  • Interquartile Range (IQR):

    • $$\displaystyle IQR = Q_3 - Q_1 $$ (spread of middle 50%).

    • Robust to outliers. Used to detect outliers: $$\displaystyle [Q_1 - 1.5 \times IQR,\; Q_3 + 1.5 \times IQR] $$.

Data Exploration Techniques

Technique Variables Analyzed Purpose Common Tools
Univariate One Describe distribution (shape, center, spread) Histogram, Box plot, Summary stats
Bivariate Two Explore relationship/correlation Scatter plot, Cross-tab, Correlation coeff.
Multivariate Exploration Multiple (>2) Initial view of relationships across many variables Pair plots, Correlation matrix, Heatmap
Multivariate Analysis Multiple Advanced modeling to understand complex structures (e.g., PCA, Cluster Analysis) Dimensionality reduction, Clustering algos

[!TIP] Key Distinction: Multivariate Exploration is descriptive (visualizing many variables). Multivariate Analysis is inferential/modeling (e.g., Multiple Regression, MANOVA).


3.0 Inferential Statistics & Statistical Modeling

Statistical Inferences

Drawing conclusions about a population from a sample.

  • Estimation: Point estimate (e.g., $\bar{x}$) & Interval estimate (Confidence Interval).

  • Hypothesis Testing: Formal procedure to reject/retain a null hypothesis ($$\displaystyle H_0 $$).

Hypothesis Testing Framework

  1. State $$\displaystyle H_0 $$ (null) & $$\displaystyle H_1 $$ (alternative).

  2. Choose significance level ($\alpha$, usually 0.05).

  3. Select appropriate test & compute test statistic.

  4. Find p-value or critical value.

  5. Decision: Reject $$\displaystyle H_0 $$ if p-value < $\alpha$ (or statistic in rejection region).

Parametric Tests

t-test (compares means, assumes normality & equal variance for two-sample).

  • One-sample: Test if sample mean differs from known $\mu$.

    • $$\displaystyle t = \frac{\bar{x} - \mu}{s / \sqrt{n}} $$, df = n-1
  • Independent two-sample: Compare means of two unrelated groups.

    • $$\displaystyle t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} $$ (Welch's for unequal variance)
  • Paired: Compare two measurements on same subjects.

    • $$\displaystyle t = \frac{\bar{d} - \mu_d}{s_d / \sqrt{n}} $$, where $d$ = differences.

Example (Potato Yield - One-sample t-test from May 2024):

Given: $$\displaystyle \mu = 20 $$, $$\displaystyle X = [21.5, 24.5, ..., 18.5] $$, n=12.

  1. $$\displaystyle \bar{x} = 20.03 $$, $s \approx 3.22$
  1. $$\displaystyle H_0: \mu = 20 $$ vs $$\displaystyle H_1: \mu > 20 $$ (one-tailed, "better than")
  1. $$\displaystyle t = \frac{20.03 - 20}{3.22 / \sqrt{12}} \approx 0.028 $$
  1. df=11, critical t(0.05, 11) ≈ 1.796.
  1. Since 0.028 < 1.796, Fail to Reject $$\displaystyle H_0 $$. No significant evidence yield is better.

Chi-Square ($$\displaystyle \chi^2 $$) Test

  • Goodness of Fit: Tests if observed frequencies fit expected distribution.

    • $$\displaystyle \boxed{\chi^2 = \sum \frac{(O_i - E_i)^2}{E_i}} $$, df = k-1 (k=categories)
  • Test of Independence: Tests association between two categorical variables in a contingency table.

    • $$\displaystyle \chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}} $$, df = (r-1)(c-1)

    • $$\displaystyle E_{ij} = \frac{(Row_i \; Total) \times (Col_j \; Total)}{Grand \; Total} $$

Regression Analysis

Models relationship between dependent (Y) and independent (X) variables.

  • Simple Linear Regression (SLR):

    • Model: $$\displaystyle \boxed{Y = \beta_0 + \beta_1 X + \epsilon} $$

    • $$\displaystyle \beta_1 $$ (slope): Change in Y per unit change in X.

    • $$\displaystyle \beta_0 $$ (intercept): Value of Y when X=0.

    • Estimated via Least Squares: Minimizes $$\displaystyle \sum (y_i - \hat{y}_i)^2 $$.

  • Multiple Linear Regression (MLR):

    • Model: $$\displaystyle Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + ... + \beta_p X_p + \epsilon $$

    • Assumptions: Linearity, Independence, Homoscedasticity, Normality of residuals, No multicollinearity.

  • Variables in Regression:

    • Dependent (Y): Outcome variable.

    • Independent (X): Predictor variable(s).

    • Dummy Variable: Categorical X converted to binary (0/1) for modeling (e.g., Gender: Male=0, Female=1).

Bayesian Modeling

  • Core Principle: Updates probability for a hypothesis as evidence/data becomes available. Uses prior knowledge.

  • Bayes' Theorem:

$$\boxed{P(H|D) = \frac{P(D|H) \cdot P(H)}{P(D)}}$$

*   $P(H|D)$: Posterior (updated belief after data D).

*   $P(D|H)$: Likelihood (probability of data given hypothesis).

*   $P(H)$: Prior (initial belief).

*   $P(D)$: Marginal likelihood (normalizing constant).
  • How it works: Start with Prior → Collect Data → Calculate Posterior → New Prior for next iteration.

  • Advantages: Incorporates prior info, intuitive probability statements, good for small data.

  • Disadvantages: Prior selection can be subjective, computationally intensive for complex models.

Maximum Likelihood Estimation (MLE)

  • Concept: Finds parameter values ($\theta$) that maximize the likelihood of observing the given sample data.

  • Procedure:

    1. Write likelihood function $$\displaystyle L(\theta) = P(data|\theta) $$ (joint probability of sample).

    2. Often use log-likelihood: $$\displaystyle \ell(\theta) = \log L(\theta) $$ (simpler to maximize).

    3. Solve $$\displaystyle \frac{d\ell}{d\theta} = 0 $$ for $\theta$.

  • Properties: Consistent, Asymptotically normal, Efficient (under regularity conditions).

  • Example (Bernoulli): For n coin flips with k heads, $$\displaystyle L(p) = p^k (1-p)^{n-k} $$. MLE $$\displaystyle \hat{p} = k/n $$.

Resampling Methods

  • Concept: Estimate sampling distribution by repeatedly drawing samples from the observed data (with replacement).

  • Bootstrapping:

    • Create B bootstrap samples (resample n observations with replacement).

    • Calculate statistic (e.g., mean) for each sample.

    • Use distribution of B statistics for CI/SE.

  • Cross-Validation:

    • Assess model performance & prevent overfitting.

    • k-fold CV: Split data into k folds, train on k-1, test on 1, repeat k times.


4.0 Data Handling, Wrangling & Storage

Data Wrangling (Data Munging)

Process of cleaning, structuring, and enriching raw data.

  1. Gathering: Collect data from sources (DBs, APIs, files).

  2. Cleaning: Handle missing values, correct errors, remove duplicates.

  3. Transformation: Normalize, aggregate, reshape (pivot/melt), encode.

  4. Enrichment: Add relevant external data (e.g., geocoding).

  5. Validation: Ensure data quality, consistency, and integrity.

File Formats

Format Type Structure Use Case
CSV Structured Tabular, plain text, delimiter-separated Simple data exchange, spreadsheets
Excel (.xlsx) Structured Tabular, multiple sheets, formulas, formatting Business reports, user-friendly analysis
JSON Semi-structured Key-value pairs, nested objects (tree-like) Web APIs, NoSQL DBs (MongoDB), configs
XML Semi-structured Tags/elements, hierarchical, verbose Legacy systems, document markup, SOAP APIs
Unstructured Unstructured No fixed schema (text, images, video) Text mining, image recognition, logs

Data Management & Indexing

  • Purpose: Efficient storage, retrieval, and update of data.

  • Indexing: Data structure (e.g., B-tree, Hash) that improves speed of data retrieval operations on a table.

    • Similar to a book's index—points to data location without scanning entire table.

    • Trade-off: Faster reads, slower writes, extra storage.

Big Data Fundamentals

  • Definition: Datasets too large/complex for traditional DBMS/tools.

  • Characteristics (3V/5V):

    • Volume (size), Velocity (speed of generation), Variety (structured/semi/unstructured).

    • Extended: Veracity (uncertainty), Value (usefulness).

  • Challenges: Storage, Processing, Analysis, Visualization, Security.

Big Data Processing Tools & Frameworks

Hadoop Ecosystem

  • HDFS (Hadoop Distributed File System):

    • Architecture: Master-Slave. NameNode (master, metadata), DataNode (slave, stores blocks).

    • Components: Block (default 128MB), Replication (default 3), Rack Awareness.

    • Operations: hadoop fs -ls, -put, -get, -cat.

    • DiagramCANVAS: Show NameNode managing metadata for file split into blocks across 3 DataNodes, with replication to another rack's DataNode
  • Hive:

    • Architecture: Metastore (schema/table metadata), HiveServer2 (query execution), Driver/Compiler/Executor.

    • HiveQL (HQL): SQL-like query language, translates to MapReduce/Tez/Spark jobs.

    • Use Cases: Data warehousing, ETL on Hadoop, batch processing of large structured data.

  • Other Tools:

    • MapReduce: Programming model (Map → Shuffle/Sort → Reduce) for parallel processing.

    • Spark: In-memory processing, faster than MapReduce for iterative tasks (ML, streaming).


5.0 Data Visualization Tools & Libraries

Python Visualization Ecosystem

Library Purpose Key Features Best For
Pandas Data manipulation .plot() method (wraps Matplotlib), quick exploratory plots Quick univariate/bivariate plots from DataFrame
Matplotlib Foundation library Low-level, highly customizable, object-oriented (Figure, Axes) Static, publication-quality plots, full control
Seaborn Statistical graphics Built on Matplotlib, attractive defaults, complex plots (violin, heatmap) Statistical visualizations, distribution relationships
Plotly Interactive, web-based Interactive (zoom, hover), Dash for web apps, 3D plots Interactive dashboards, web deployment, complex interactivity

R Visualization: ggplot2

  • Grammar of Graphics: Build plots layer by layer.

  • Core Components:

    • data: Dataset.

    • aes(): Aesthetics (x, y, color, size).

    • geom_*(): Geometric objects (points, lines, bars).

    • facet_*(): Small multiples.

    • theme(): Non-data ink (labels, grid).

  • Example: ggplot(data, aes(x=var1, y=var2)) + geom_point() + geom_smooth()

Business Intelligence Tools: Power BI

  • Components:

    • Power BI Desktop: Free app for data connection, modeling, report creation.

    • Power BI Service: Cloud-based SaaS for sharing, collaboration, dashboards.

    • Power BI Mobile: Apps for iOS/Android to view reports.

  • Key Features:

    • Data Modeling: Create relationships between tables, define hierarchies.

    • DAX (Data Analysis Expressions): Formula language for custom calculations & measures (e.g., TOTALYTD(), RELATED()).

    • Interactive Dashboards: Drill-down, slicers, cross-filtering between visuals.


6.0 Synthesis & Application

Creating Custom Visualizations for Complex Datasets

  1. Understand Data & Goal: What question are you answering? Know variables, types, relationships.

  2. Choose Chart Types:

    • Comparison: Bar, Column.

    • Distribution: Histogram, Box, Violin.

    • Relationship: Scatter, Heatmap, Bubble.

    • Composition: Stacked Bar, Pie (use sparingly), Treemap.

    • Trend: Line, Area.

  3. Apply Design Principles:

    • Clarity: Minimize clutter (chartjunk).

    • Accuracy: Represent proportions correctly (e.g., area vs. height for circles).

    • Color: Use purposefully (categorical vs. sequential), ensure accessibility.

    • Labeling: Clear axes, titles, legends.

  4. Tool Selection: Match tool to need (Matplotlib for static, Plotly/Dash for interactive, Power BI for business dashboards).

  5. Iterate: Prototype, get feedback, refine.

Integrating Statistical Analysis with Visualization

  • Purpose: Visuals make statistical results intuitive and communicate uncertainty.

  • Methods:

    • Overlay model predictions (regression line, confidence band) on scatter plots.

    • Use box plots/violin plots to show group differences from t-tests/ANOVA.

    • Visualize correlation matrices with heatmaps.

    • Plot residuals to check regression assumptions (linearity, homoscedasticity).

    • Show confidence intervals as error bars on bar charts for group means.

  • Example: A scatter plot with a best-fit line (SLR result) and a shaded 95% confidence interval around the line visually communicates both the relationship and its uncertainty.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in