Skip to content
AL-603 (B) · Data and Visual Analytics/Quick Revision Short Notes

Data and Visual Analytics (AL-603 (B)) - Unit 5 Short Notes

UNIT 5: Data and Visual Analytics - Short Notes

I. FOUNDATIONAL STATISTICAL CONCEPTS

Measures of Central Tendency

Definition: Single values that represent the center or typical value of a dataset.

  • Mean (Arithmetic Average): Sum of all values divided by count.

$$ \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$

*Sensitive to outliers.*
  • Median: Middle value when data is sorted. For even n, median = average of two middle values.

    Robust to outliers.

  • Mode: Most frequently occurring value. Can be multimodal.

    Applicable to nominal data.

Measure Best For Sensitive to Outliers?
Mean Symmetric, numerical data Yes
Median Skewed data, ordinal data No
Mode Categorical (nominal/ordinal) data No

Measures of Dispersion / Location

  • Range: Max - Min. Highly sensitive to outliers.

  • Variance (Sample): Average squared deviation from mean.

$$ s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} $$

  • Standard Deviation (s): Square root of variance. Same units as data.

$$ s = \sqrt{s^2} $$

  • Interquartile Range (IQR): Spread of middle 50% of data.

$$ IQR = Q_3 - Q_1 $$

*Robust measure.*
  • Quartiles (Q1, Q2, Q3): Values at 25th, 50th (median), 75th percentiles.

  • Percentiles: Value below which a given percentage of data falls.

Levels of Measurement (Scales of Data)

Scale Characteristics Examples Permissible Operations
Nominal Categories only, no order Gender, Blood Type Count, Mode, Chi-Square
Ordinal Ordered categories, unequal intervals Likert Scale, Ranks All Nominal + Median, Percentiles
Interval Ordered, equal intervals, no true zero Celsius Temp, IQ All Ordinal + Mean, SD
Ratio All Interval properties + true zero Height, Weight, Income All statistical operations

Variables and Data Categorization

  • By Nature:

    • Categorical: Nominal (no order), Ordinal (ordered).

    • Numerical: Discrete (countable), Continuous (measurable).

  • By Role in Analysis:

    • Independent Variable (Predictor/Feature): Manipulated or used to explain.

    • Dependent Variable (Response/Target): Outcome being measured/predicted.

[!TIP] Exam Focus: Be prepared to classify given variables into correct types and scales. Remember: Ratio scale is required for meaningful ratios (e.g., "twice as heavy").


II. INFERENTIAL STATISTICS & HYPOTHESIS TESTING

Statistical Inferences

  • Definition: Using sample data to make conclusions about a larger population.

  • Estimation:

    • Point Estimation: Single value estimate (e.g., sample mean \bar{x} estimates population mean \mu).

    • Interval Estimation (Confidence Interval): Range of values likely to contain population parameter.

$$ \text{CI} = \text{Point Estimate} \pm (\text{Critical Value} \times \text{Standard Error}) $$

Parametric Hypothesis Tests

General Steps: 1. State H₀ & H₁. 2. Choose significance level (α). 3. Calculate test statistic. 4. Find p-value/critical value. 5. Decision & Interpretation.

1. t-Test

  • One-Sample t-Test: Compares sample mean to a known population mean (μ).

$$ t = \frac{\bar{x} - \mu}{s / \sqrt{n}} $$

*Degrees of Freedom (df) = n-1.*

> **May 2024 Example:** Given potato yield data `X` and `μ=20`.

> 1. Calculate `\bar{x}` and `s`.

> 2. Compute `t` using formula above.

> 3. Compare to critical `t` (df=11, α=0.05, one-tailed) or find p-value.

> 4. **Interpret:** "If p < α, reject H₀. There is significant evidence that yield is better than standard."
  • Two-Sample (Independent) t-Test: Compares means of two independent groups.

    Assumes equal variances (pooled) or not (Welch's).

  • Paired t-Test: Compares means of two related samples (e.g., before-after).

$$ t = \frac{\bar{d}}{s_d / \sqrt{n}} $$

where `d` is the difference for each pair.

2. Chi-Square (χ²) Test

  • Test for Independence: Tests association between two categorical variables in a contingency table.

$$ \chi^2 = \sum \frac{(O - E)^2}{E} $$

*df = (rows-1)*(cols-1)*

*E = (row total * col total) / grand total*
  • Goodness-of-Fit Test: Tests if observed frequencies fit expected distribution.

    df = (categories - 1)

[!TIP] Key Assumption for χ²: Expected frequency E >= 5 for most cells. If violated, use Fisher's Exact Test or combine categories.

Maximum Likelihood Estimation (MLE)

  • Philosophy: Find parameter values (θ) that make the observed sample most probable.

  • Procedure:

    1. Write likelihood function L(θ) = P(data|θ).

    2. Often work with log-likelihood ℓ(θ) = ln(L(θ)).

    3. Maximize ℓ(θ) by taking derivative w.r.t θ, setting to 0, solving.

  • Simple Example (Bernoulli): For n coin flips with k heads, L(p) = p^k (1-p)^(n-k). Maximizing gives \hat{p} = k/n.

Re-Sampling Methods

  • Bootstrapping: Resample with replacement from the original sample to create many "bootstrap samples." Use distribution of a statistic (e.g., mean) from these samples to estimate its sampling distribution and confidence intervals.

  • Permutation Test: Resample without replacement by randomly shuffling group labels. Recalculates test statistic to build null distribution. Advantage: Makes fewer distributional assumptions (e.g., normality).

[!TIP] Common Pitfall: MLE can be biased for small samples. Bootstrapping requires the original sample to be representative of the population.


III. REGRESSION & PREDICTIVE MODELING

Regression Analysis

  • Simple Linear Regression (SLR): Models linear relationship between one independent (X) and one dependent (Y) variable.

$$ Y = \beta_0 + \beta_1 X + \epsilon $$

*   `\beta_0` = intercept, `\beta_1` = slope.

*   **Assumptions:** Linearity, Independence, Homoscedasticity, Normality of residuals.

*   **Interpretation:** `\beta_1` = change in Y for a 1-unit change in X.
  • Multiple Linear Regression (MLR): Extends SLR to multiple predictors.

$$ Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + ... + \beta_p X_p + \epsilon $$

*   **Interpretation:** `\beta_j` = effect of `X_j` *holding all other X's constant*.

*   **Key Metric:** Adjusted R² (penalizes for extra variables).

Variables in Regression Modeling

  • Dummy Variables (Indicator Variables): Convert categorical variables into binary (0/1) for regression.

    • For k categories, use k-1 dummy variables to avoid dummy variable trap.
  • Interaction Terms: Product of two or more predictors (e.g., X1*X2). Captures effect of one predictor depending on the level of another.

Multivariate Analysis

  • Definition: Simultaneous analysis of more than two variables to understand relationships.

  • Distinction from MLR: MLR has one dependent variable. Multivariate analysis often has multiple dependent variables.

  • Common Techniques:

    • MANOVA (Multivariate ANOVA): Extension of ANOVA for multiple DVs.

    • Factor Analysis: Reduces many variables into fewer underlying factors.

    • Cluster Analysis: Groups observations into similar clusters.

    • Primarily conceptual for this unit.

Bayesian Modelling

  • Core Principle: Bayes' Theorem updates prior belief with data to get posterior belief.

$$ P(\theta|data) = \frac{P(data|\theta) \cdot P(\theta)}{P(data)} $$

*   `P(θ)` = Prior (belief before data).

*   `P(θ|data)` = Posterior (updated belief).

*   `P(data|θ)` = Likelihood.
  • Workflow: 1. Specify Prior P(θ). 2. Collect Data & define Likelihood. 3. Compute Posterior (often via MCMC sampling).

  • Advantages:

    • Incorporates existing knowledge (prior).

    • Provides full probability distribution for parameters (not just point estimates).

    • Intuitive probabilistic interpretation.

  • Disadvantages:

    • Computationally intensive (especially for complex models).

    • Choice of prior can be subjective and influence results.

    • Can be slow for large datasets.

[!TIP] Key Difference: Frequentist (e.g., MLE, t-test) treats parameters as fixed. Bayesian treats parameters as random variables with distributions.


IV. DATA VISUALIZATION & EXPLORATORY DATA ANALYSIS (EDA)

Data Visualization Fundamentals

  • Purpose: Discover patterns, trends, relationships, outliers; communicate findings.

  • Principles of Effective Viz:

    1. Clarity: Message should be immediately understandable.

    2. Accuracy: Represent data truthfully (no distortion).

    3. Efficiency: Minimize "chartjunk," maximize data-ink ratio (Tufte).

Exploration of Data (EDA)

  • Univariate Exploration (1 variable):

    • Categorical: Bar chart, Pie chart (use sparingly), Frequency table.

    • Numerical: Histogram, Box plot (shows IQR, median, outliers), Stem-and-leaf.

  • Bivariate Exploration (2 variables):

    • Num-Num: Scatter plot (primary), Correlation coefficient (r).

    • Cat-Cat: Stacked/Grouped bar chart, Mosaic plot.

    • Num-Cat: Box plot (by category), Violin plot.

  • Multivariate Exploration (>2 variables):

    • Encodings: Use color, shape, size, facet (small multiples) to add dimensions to 2D plots.

    • Techniques: Pair plot (scatter matrix), 3D plots (use cautiously), Heatmaps (for correlation matrices).

Creating Custom Visualizations for Complex Datasets

Process:

  1. Define Question: What relationship/story are you investigating?

  2. Choose Chart Type: Match question to visual encoding (e.g., trend over time → line chart; part-to-whole → stacked bar).

  3. Design for Clarity: Label axes, add title, use appropriate color scales (sequential for ordered, diverging for deviation from center, categorical for distinct groups).

  4. Iterate & Refine: Simplify, remove non-data ink, ensure accessibility (colorblind-friendly palettes).

[!TIP] Common Pitfall: Using 3D pie charts or excessive colors that distort perception. Always prefer 2D and test if the visualization answers the question without explanation.


V. PROGRAMMATIC VISUALIZATION & TOOLS

Python Visualization Libraries

  • Matplotlib: Foundation library. Highly customizable but verbose. Good for static, publication-quality plots.

    
    plt.plot(x, y); plt.xlabel(); plt.title()
    
    
  • Seaborn: Built on Matplotlib. Simplifies statistical plots with beautiful defaults.

    • relplot() (scatter/line), catplot() (categorical), displot() (distributions).

    • Key Plots: heatmap(), violinplot(), pairplot().

  • Plotly/Bokeh: Interactive visualizations (zoom, hover, click). Ideal for web dashboards.

    • Plotly (plotly.express) is more user-friendly; Bokeh offers more control.
  • Pandas Integration: Quick plotting via .plot() method on DataFrames/Series (uses Matplotlib backend).

    
    df['column'].plot(kind='hist')
    
    

Power BI Ecosystem

  • Power BI Desktop: Primary tool. Free application for:

    • Data connection & transformation (Power Query Editor).

    • Data modeling (relationships, DAX formulas).

    • Report creation (drag-and-drop visuals).

  • Power BI Service: Cloud platform (app.powerbi.com).

    • Publish, share, collaborate on reports/dashboards.

    • Schedule data refreshes, set up row-level security.

  • Power BI Mobile: Consume reports on iOS/Android devices.

  • Power BI Report Server: On-premises server for hosting reports internally (for organizations with strict data governance).

ggplot2 (R)

  • Philosophy: Grammar of Graphics. Build plots layer by layer.

  • Core Syntax:

    
    ggplot(data = df, aes(x = var1, y = var2)) +  # 1. Data & Aesthetics
    
      geom_point() +                             # 2. Geometric object (layer)
    
      facet_wrap(~category) +                    # 3. Facets (subplots)
    
      theme_minimal()                            # 4. Theme (styling)
    
    
  • Key Components: ggplot() (init), aes() (mapping), geom_*() (marks), facet_*() (panels), scale_*() (adjust axes/color), theme() (non-data ink).

[!TIP] Python vs. R: Matplotlib/Seaborn = imperative (plot step-by-step). ggplot2 = declarative (describe the final plot components). Plotly is the closest Python equivalent to ggplot2's layered approach.


VI. BIG DATA TECHNOLOGIES & ECOSYSTEM

Big Data Fundamentals

  • Definition: Datasets too large, fast, or complex for traditional DBMS.

  • Vs (Extended):

    • Volume: Size (TB/PB).

    • Velocity: Speed of generation/processing.

    • Variety: Structured, semi-structured, unstructured.

    • Veracity: Uncertainty, quality, trustworthiness.

    • Value: Ultimate goal—extracting useful insights.

Big Data Processing Tools

Hadoop Ecosystem:

  • HDFS (Hadoop Distributed File System):

    • Architecture: Master-Slave. NameNode (master, metadata), DataNodes (slaves, store blocks).

    • File Storage: Split into blocks (default 128MB/256MB), replicated (default 3x) across nodes.

    • Advantage: Fault-tolerant, scalable, cost-effective (commodity hardware).

  • MapReduce:

    • Programming Model: Map (filter/sort → key-value pairs) → Shuffle & Sort → Reduce (aggregate).

    • Limitation: High latency due to disk I/O between stages.

  • Hive:

    • Purpose: Data warehouse infrastructure on Hadoop.

    • HiveQL: SQL-like query language. Translates queries to MapReduce/Tez/Spark jobs.

    • Metastore: Stores schema information (table definitions).

Apache Spark:

  • Core Innovation: In-memory processing. Caches data in RAM across cluster.

  • Key Abstractions:

    • RDD (Resilient Distributed Dataset): Immutable, partitioned collection. Fault-tolerant via lineage.

    • DataFrame/Dataset: Higher-level API (like Pandas/SQL tables). Optimized with Catalyst optimizer.

  • Speed: 10-100x faster than MapReduce for iterative algorithms (ML, graph) due to reduced disk I/O.

  • Components: Spark SQL, MLlib, Spark Streaming, GraphX.

[!TIP] Hadoop vs. Spark: Hadoop (HDFS+MapReduce) is disk-based, good for huge batch jobs. Spark is in-memory, good for iterative processing, streaming, and SQL. Often used together: Spark reads from HDFS.


VII. DATA WRANGLING & MANAGEMENT

Data Wrangling (Data Munging)

Critical, time-consuming (~60-80% of analyst time) process of making raw data usable.

Detailed Process:

  1. Data Acquisition:

    • Sources: CSV/Excel, SQL databases, APIs (REST), Web scraping, Sensors.

    • Tools: pandas.read_csv(), pd.read_sql(), requests library.

  2. Data Cleaning:

    • Missing Values: Identify (isnull()), Handle (delete, impute with mean/median/mode, or advanced methods).

    • Outliers: Detect (IQR method: Q1 - 1.5*IQR, Q3 + 1.5*IQR; Z-score), Treat (cap, transform, or keep if valid).

    • Data Types: Correct (astype()), parse dates (pd.to_datetime()).

    • Duplicates: Identify & remove (drop_duplicates()).

  3. Data Transformation:

    • Normalization (Min-Max): Scale to [0,1].

$$ x' = \frac{x - x_{min}}{x_{max} - x_{min}} $$

*   **Standardization (Z-score):** Mean=0, SD=1.

$$ z = \frac{x - \mu}{\sigma} $$

*   **Binning:** Convert continuous to categorical (e.g., age groups).

*   **Derived Variables:** Create new columns from existing ones (e.g., `profit = revenue - cost`).

*   **Reshaping:** `pivot()` (wide to long), `melt()` (long to wide).
  1. Data Integration:

    • Merging/Joining: Combine datasets from different sources.

      • pd.merge(df1, df2, on='key') (SQL-like joins: inner, left, right, outer).
    • Concatenation: pd.concat([df1, df2]) (stack vertically/horizontally).

Data Management and Indexing

  • DBMS: Software to store, retrieve, manage data (e.g., MySQL, PostgreSQL).

  • Indexing: Data structure (like B-tree) that speeds up query retrieval at the cost of extra storage and slower writes.

    • Primary Index: On primary key (unique, not null).

    • Secondary (Clustered) Index: Determines physical order of data (one per table).

    • Secondary (Non-Clustered) Index: Separate structure pointing to data rows (multiple allowed).

File Formats

Format Type Schema? Human-Readable? Best For Big Data?
CSV Structured No (implied) Yes Simple exchange, small data No (no compression)
JSON Semi-structured Flexible Yes Web APIs, nested data No (text-based, verbose)
Parquet Structured Yes (defined) No Columnar storage in Hadoop/Spark. High compression, efficient querying (read only needed columns). Yes
Avro Structured Yes (in file) No Row-based, serialization in Hadoop. Schema evolution. Yes
XML Semi-structured Flexible Yes (verbose) Legacy systems, documents No

[!TIP] Rule of Thumb: Use Parquet for analytical queries in big data (columnar). Use CSV/JSON for small, simple, or web-based interchange. Use Avro for row-based serialization in streaming.


VIII. DATA ANALYST ECOSYSTEM & WORKFLOW

Data Analyst Ecosystem Overview

A toolchain supporting the full data lifecycle:

  1. Storage: SQL (PostgreSQL), NoSQL (MongoDB), HDFS (big data).

  2. Processing & Analysis: Python (pandas, NumPy, SciPy), R, Spark (PySpark).

  3. Visualization & Reporting: Power BI, Tableau, Python libs (Seaborn, Plotly), ggplot2.

  4. Collaboration & Reproducibility: Git (version control), Jupyter Notebooks/Lab, R Markdown.

End-to-End Workflow Integration

  1. Acquire: Connect to source (SQL DB, CSV, API → pandas.read_sql(), requests).

  2. Wrangle: Clean & transform (Python/R: pandas, dplyr). For big data, use Spark DataFrames.

  3. Analyze: Exploratory analysis (visualization, summary stats) → Inferential stats (t-test, χ²) → Modeling (regression, ML).

  4. Visualize & Report: Create static/interactive dashboards (Power BI, Plotly Dash) or notebooks for narrative.

  5. Deploy & Share: Publish to Power BI Service, schedule refreshes, or share notebooks via GitHub.

[!TIP] Exam Focus: Be able to map tools to workflow stages. Example: "For a 50GB CSV file, which tool for wrangling?" → Spark (PySpark) because pandas would run out of memory. "For creating an interactive sales dashboard for management?" → Power BI or Plotly Dash.


\boxed{\text{End of Unit 5 Notes}} Based on RGPV May 2024 & May 2023 exam analysis. Focus on definitions, formulas, step-by-step procedures (t-test, χ²), tool comparisons, and workflow mapping.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in