Skip to content
AL-603 (B) · Data and Visual Analytics/Quick Revision Short Notes

Data and Visual Analytics (AL-603 (B)) - Unit 4 Short Notes

UNIT 4: Data and Visual Analytics


I. Foundational Data Concepts & Statistics

A. Nature of Data & Variables

  • Variable: A characteristic or attribute that can be measured or quantified.

  • Types of Variables:

    • Categorical (Qualitative):

      • Nominal: No inherent order (e.g., Gender, Color, Country).

      • Ordinal: Has a meaningful order but not fixed intervals (e.g., Education Level, Satisfaction Rating: Low/Medium/High).

    • Numerical (Quantitative):

      • Discrete: Countable, finite/integer values (e.g., Number of students, Shoe size).

      • Continuous: Measurable, infinite values within a range (e.g., Height, Weight, Temperature).

  • Levels of Measurement (Scale of Measurement):

    • Nominal: Classification only.

    • Ordinal: Rank order.

    • Interval: Ordered with equal intervals, but no true zero (e.g., Celsius temperature, IQ score).

    • Ratio: All properties of interval plus a meaningful, non-arbitrary zero (e.g., Weight, Height, Income). Allows for meaningful ratios.

B. Descriptive Statistics

  • Measures of Central Tendency (Location): Describe the center of a data distribution.

    • Mean ($\bar{x}$ or $\mu$): Sum of all values divided by count. Sensitive to outliers.

$$\bar{x} = \frac{\sum_{i=1}^{n} x_i}{n}$$

*   **Median ($M$):** Middle value when sorted. Robust to outliers.

*   **Mode:** Most frequently occurring value. Can be multimodal.

> [!TIP] For symmetric distributions, Mean ≈ Median ≈ Mode. For skewed distributions, order is Mode < Median < Mean (positive skew).
  • Measures of Dispersion (Variability): Describe the spread of data.

    • Range: Max - Min. Sensitive to outliers.

    • Variance ($$\displaystyle s^2 $$ for sample, $$\displaystyle \sigma^2 $$ for population): Average squared deviation from the mean.

$$s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1}$$

*   **Standard Deviation ($s$ or $\sigma$):** Square root of variance. In same units as data.

$$s = \sqrt{s^2} = \sqrt{\frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1}}$$

*   **Interquartile Range (IQR):** Spread of middle 50% of data. $$\displaystyle IQR = Q_3 - Q_1 $$. Robust to outliers.
  • Measures of Location:

    • Percentiles: Value below which a given percentage of data falls (e.g., 90th percentile).

    • Quartiles: Specific percentiles: Q1 (25th), Q2 (50th = Median), Q3 (75th).

C. Exploratory Data Analysis (EDA)

  • Goal: Summarize main characteristics of data, often with visual methods, to discover patterns, spot anomalies, test hypotheses.

  • Univariate Exploration: Analysis of a single variable.

    • Techniques: Histograms, box plots, density plots, frequency tables.

    • Focus: Distribution shape (symmetry, skewness, modality), central tendency, spread, outliers.

  • Bivariate Exploration: Analysis of two variables to explore relationships.

    • Techniques:

      • Numerical vs. Numerical: Scatter plot, correlation coefficient (Pearson's r).

      • Categorical vs. Categorical: Cross-tabulation (contingency table), stacked/mosaic plots.

      • Numerical vs. Categorical: Box plots (by category), violin plots.

  • Multivariate Exploration: Analysis of three or more variables.

    • Techniques: Pair plots (scatter matrix), 3D plots, faceting (small multiples), color/shape/size encoding in 2D plots, dimensionality reduction (PCA for visualization).

    [!TIP] EDA is iterative and non-prescriptive. Visualization is its most powerful tool. Always ask: "What story is this data telling?"


II. Statistical Inference & Hypothesis Testing

A. Statistical Inference

  • Definition: The process of using sample data to draw conclusions (make inferences) about an underlying population.

  • Two Main Types:

    1. Estimation:

      • Point Estimation: Provide a single value as an estimate of a population parameter (e.g., sample mean $\bar{x}$ estimates population mean $\mu$).

      • Interval Estimation (Confidence Interval): Provide a range of values within which the parameter is likely to lie, with a stated confidence level (e.g., 95% CI for $\mu$).

    2. Hypothesis Testing: A formal procedure for assessing whether sample evidence is strong enough to reject a null hypothesis ($$\displaystyle H_0 $$) about a population.

B. Parametric Tests

  • Assumptions: Data follows a specific distribution (usually normal), homogeneity of variance, independent observations.

  • 1. t-test: Compares means. Used when population standard deviation ($\sigma$) is unknown and sample size is small, assuming normality.

    • One-sample t-test: Tests if a sample mean ($\bar{x}$) differs from a known/hypothesized population mean ($$\displaystyle \mu_0 $$).

$$t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}}$$

    *Degrees of Freedom (df) = n-1.*

    > **Example (Potato Yield):** Given $$\displaystyle \mu_0=20 $$, $$\displaystyle n=12 $$, data $X$. Calculate $\bar{x}$, $s$, then $$\displaystyle t_{calc} $$. Compare with $$\displaystyle t_{critical} $$ (from t-table at $$\displaystyle \alpha=0.05 $$, df=11, one-tailed). If $$\displaystyle t_{calc} > t_{crit} $$, reject $$\displaystyle H_0 $$ (yield is significantly better).

*   **Two-sample (Independent) t-test:** Compares means of **two independent groups**.

$$t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}}$$

    *Use Welch's approximation if variances are unequal.*

*   **Paired t-test:** Compares means of **two related samples** (e.g., before/after, matched pairs). Uses differences ($$\displaystyle d_i = x_{1i} - x_{2i} $$).

$$t = \frac{\bar{d}}{s_d / \sqrt{n}}$$

    *df = n-1, where n = number of pairs.*
  • 2. Chi-Square ($$\displaystyle \chi^2 $$) Test: Used for categorical data. Tests for association or goodness-of-fit.

    • Chi-Square Test for Independence (Association): Tests if two categorical variables are independent.

      Steps:

      1. State $$\displaystyle H_0 $$: Variables are independent. $$\displaystyle H_1 $$: They are associated.

      2. Create contingency table (Observed frequencies, $$\displaystyle O_{ij} $$).

      3. Calculate Expected frequencies: $$\displaystyle E_{ij} = \frac{(Row\ Total_i) \times (Column\ Total_j)}{Grand\ Total} $$.

      4. Compute test statistic:

$$\chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}}$$

    5.  df = $(r-1)(c-1)$ where r=rows, c=columns.

    6.  Compare $$\displaystyle \chi^2_{calc} $$ with critical value from $$\displaystyle \chi^2 $$ table at chosen $\alpha$.

*   **Chi-Square Goodness-of-Fit Test:** Tests if a single categorical variable follows a hypothesized distribution.

$$\chi^2 = \sum \frac{(O_i - E_i)^2}{E_i}$$

    *df = (number of categories - 1 - number of estimated parameters).*

C. Non-Parametric & Resampling Methods

  • Re-sampling: Makes no strict assumptions about population distribution.

    • Bootstrap: Repeatedly samples with replacement from the original sample to estimate sampling distribution of a statistic (e.g., mean, median). Used for confidence intervals.

    • Permutation Test: Repeatedly shuffles/permutes group labels to create a null distribution of the test statistic. Compares observed statistic to this distribution.

D. Regression Analysis

  • Goal: Model the relationship between a dependent (response) variable and one or more independent (predictor) variables.

  • Simple Linear Regression (SLR): One predictor ($x$).

    • Model: $$\displaystyle y = \beta_0 + \beta_1 x + \epsilon $$

      • $$\displaystyle \beta_0 $$: Intercept (value of y when x=0).

      • $$\displaystyle \beta_1 $$: Slope (change in y for a one-unit change in x).

      • $\epsilon$: Random error.

    • Estimation (Ordinary Least Squares - OLS): Finds $$\displaystyle \hat{\beta}_0, \hat{\beta}_1 $$ that minimize sum of squared residuals ($$\displaystyle \sum e_i^2 $$).

    • Interpretation: Sign of $$\displaystyle \hat{\beta}_1 $$ indicates direction. Magnitude indicates strength (in units of y per unit x).

  • Multiple Linear Regression (MLR): Two or more predictors.

    • Model: $$\displaystyle y = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + ... + \beta_p x_p + \epsilon $$

    • Interpretation: Holding other variables constant, $$\displaystyle \hat{\beta}_j $$ is the change in y for a one-unit change in $$\displaystyle x_j $$. (Ceteris paribus).

    • Key Assumptions: Linearity, Independence, Homoscedasticity (constant variance), Normality of errors, No multicollinearity.

  • Types of Variables in Regression:

    • Dependent (Y): Variable being predicted.

    • Independent (X): Predictor variables.

    • Dummy Variables (Indicator Variables): Binary (0/1) variables used to represent categorical predictors in the model (e.g., Gender: Male=0, Female=1).

E. Advanced Inference Techniques

  • 1. Maximum Likelihood Estimation (MLE):

    • Concept: A method for estimating model parameters by finding the parameter values that maximize the likelihood function—the probability of observing the given sample data.

    • Steps:

      1. Write down the likelihood function $L(\theta|data)$ (joint probability of all observations, often product of PDFs/PMFs).

      2. Take the natural logarithm to get Log-Likelihood ($\ell$) for easier differentiation.

      3. Differentiate $\ell$ with respect to the parameter(s) $\theta$.

      4. Set derivatives to zero and solve for $$\displaystyle \hat{\theta}_{MLE} $$.

    • Example (Bernoulli): For $n$ Bernoulli trials with $k$ successes, $$\displaystyle L(p) = p^k (1-p)^{n-k} $$. $$\displaystyle \hat{p}_{MLE} = k/n $$.

    [!TIP] MLE is a general method, not a specific test. It's the foundation for many parametric models (including logistic regression).

  • 2. Bayesian Modelling:

    • Core Concept (Bayes' Theorem): Updates prior beliefs about a parameter ($\theta$) with observed data (likelihood) to obtain a posterior belief.

$$P(\theta|Data) = \frac{P(Data|\theta) \cdot P(\theta)}{P(Data)}$$

    *   **Prior $P(\theta)$:** Belief about $\theta$ before seeing data.

    *   **Likelihood $P(Data|\theta)$:** Probability of data given $\theta$.

    *   **Posterior $P(\theta|Data)$:** Updated belief after seeing data.

*   **How it Works (vs. Frequentist):**

    *   **Frequentist:** Parameters are fixed but unknown. Probability is long-run frequency of data.

    *   **Bayesian:** Parameters are **random variables** with probability distributions. Probability measures **degree of belief**.

*   **Advantages:** Incorporates prior knowledge, provides full probability distribution for parameters (credible intervals), intuitive interpretation of intervals.

*   **Disadvantages:** Choice of prior can be subjective, computationally intensive for complex models.

III. Data Management & Big Data Ecosystem

A. Data Wrangling / Munging

  • Definition: The process of cleaning, structuring, and enriching raw data into a desired format for analysis. Often takes 60-80% of an analyst's time.

  • Process Steps:

    1. Data Acquisition: Gathering data from sources (APIs, databases, files, web scraping).

    2. Data Cleaning:

      • Handling missing values (delete, impute with mean/median/mode, model-based).

      • Handling outliers (detect via IQR/Z-score, decide to cap, transform, or remove).

      • Correcting data types, fixing inconsistencies, removing duplicates.

    3. Data Transformation: Normalization, standardization, binning, encoding categorical variables, creating new features (feature engineering).

    4. Data Enrichment: Adding relevant external data or calculated columns.

    5. Data Validation: Ensuring data quality, consistency, and integrity for analysis.

B. File Formats & Storage

Category Formats Key Characteristics
Structured CSV, Excel, JSON (tabular) Fixed schema, rows & columns. Easy to query/process. CSV: plain text, no types.
Semi-structured JSON, XML Tags/markers separate elements. Flexible schema. Self-describing. Good for web data.
Unstructured Text files, Images, Video, Audio No predefined model. Requires NLP/computer vision for analysis.
Binary (Columnar) Parquet, Avro, ORC Optimized for read/write performance in distributed systems. Columnar storage (Parquet) enables efficient querying of specific columns.

C. Big Data Processing Tools & Frameworks

  • Hadoop Ecosystem: Framework for distributed storage and processing.

    • HDFS (Hadoop Distributed File System):

      • Architecture: Master-Slave. NameNode (master, metadata) manages DataNodes (slaves, store blocks).

      • Key Concept: Files split into blocks (default 128MB/256MB) and distributed across cluster. Replication (default 3x) for fault tolerance.

      • Advantage: Cost-effective storage of massive datasets on commodity hardware.

    • Hive (Data Warehouse Infrastructure):

      • Provides SQL-like interface (HiveQL) to query data stored in HDFS.

      • Translates HiveQL queries into MapReduce/Tez/Spark jobs.

      • Schema-on-read: Apply structure when reading data, not when storing.

      • Use Case: Data summarization, ad-hoc analysis, ETL on large, static datasets.

  • Other Paradigms:

    • MapReduce: Programming model for processing large datasets with parallel, distributed algorithms. Map (filter/sort) -> Shuffle -> Reduce (aggregate).

    • Apache Spark: In-memory cluster computing. Faster than MapReduce for iterative algorithms. Supports batch, stream, machine learning (MLlib), graph processing. Core abstraction: Resilient Distributed Dataset (RDD).


IV. Data Visualization: Tools & Techniques

A. Principles of Data Visualization

  • Definition: The graphical representation of information and data. Purpose is to communicate insights clearly, efficiently, and accurately.

  • Grammar of Graphics (ggplot2 paradigm): A framework that breaks down a graphic into distinct components:

    1. Data: The dataset being plotted.

    2. Aesthetics (aes): Mappings of data variables to visual properties (x, y, color, size, shape).

    3. Geometries (geom): The geometric objects that represent data points (points, lines, bars, boxplots).

    4. Facets: Small multiples for conditioning on a variable.

    5. Statistics: Statistical transformations (e.g., smoothing, binning for histograms).

    6. Coordinates: Coordinate system (cartesian, polar).

    7. Themes: Non-data ink (fonts, grids, backgrounds).

B. Python Visualization Libraries

Library Primary Use Case Key Features
Matplotlib Foundational, low-level plotting. Highly customizable, object-oriented API. Basis for many other libraries.
Seaborn Statistical visualizations. Built on Matplotlib. Simplified syntax for complex plots (heatmaps, violin plots, pair plots). Integrates with Pandas DataFrames.
Plotly Interactive, web-based visualizations. Creates interactive charts (zoom, hover, click). Dash framework for web apps.
Pandas Quick, integrated plotting. .plot() method on DataFrames/Series. Simple wrapper around Matplotlib. Good for EDA.
ggplot Implementation of Grammar of Graphics. Based on R's ggplot2. Uses a declarative syntax: ggplot(data, aes(x,y)) + geom_point().

C. Business Intelligence Tools

  • Power BI:

    • Definition: Microsoft's suite of business analytics tools.

    • Components:

      1. Power BI Desktop: Free Windows application for data connection, transformation (Power Query), modeling, and report creation.

      2. Power BI Service: Cloud-based SaaS service for publishing, sharing, and collaborative viewing of reports/dashboards.

      3. Power BI Mobile: Apps for iOS/Android to view reports on-the-go.

      4. Power BI Report Server: On-premises server for hosting Power BI reports (for organizations with cloud restrictions).

    • Key Tools within Ecosystem:

      • Power Query: Data connection and transformation engine (Get & Transform).

      • DAX (Data Analysis Expressions): Formula language for creating custom calculations and measures in data models.

      • Rich Visualizations: Custom visuals marketplace, drill-through, bookmarks.

D. Creating Custom Visualizations

  • Process for Complex Datasets:

    1. Understand the Data & Question: What story needs to be told? What is the key insight?

    2. Choose Appropriate Chart Type: Match visualization to relationship (comparison, distribution, composition, relationship).

      DiagramCANVAS: Chart selection flowchart based on data question

    3. Apply Design Principles:

      • Maximize Data-Ink Ratio: Remove non-essential ink (chartjunk).

      • Use Color Strategically: For categorization or sequential data, not decoration. Ensure accessibility.

      • Label Clearly: Axes, titles, data labels. Avoid ambiguity.

      • Tell a Story: Use titles to state the insight, not just "Chart of X vs Y".

    4. Incorporate Interactivity (if needed): Tooltips, filtering, zooming, linking views (dashboard).

    5. Iterate & Test: Get feedback. Is the message clear?


V. The Data Analyst Ecosystem

A. Data Analyst Ecosystem

  • Definition: The interconnected set of technologies, processes, and roles involved in turning raw data into actionable insights.

  • Key Components (Data Pipeline):

    1. Data Sources: Databases, APIs, flat files, IoT sensors.

    2. Storage: Data Lakes (raw), Data Warehouses (structured), Databases (SQL/NoSQL).

    3. Processing (ETL/ELT): Extract, Transform, Load (or Extract, Load, Transform). Tools: Apache Airflow, dbt, SSIS.

    4. Analysis: Statistical analysis, machine learning, data mining (using Python/R/SQL).

    5. Visualization & BI: Creating dashboards/reports (Power BI, Tableau, Python libraries).

    6. Consumption: End-users, stakeholders, applications.

    DiagramSEARCH: data analytics pipeline diagram etl bi

B. Data Management & Indexing

  • Database Management Systems (DBMS): Software for storing, retrieving, and managing data.

    • Types: Relational (SQL: PostgreSQL, MySQL), Non-Relational (NoSQL: MongoDB, Cassandra), NewSQL.
  • Indexing:

    • Purpose: Dramatically speed up data retrieval (query) operations on a table, at the cost of slower writes and additional storage.

    • How it Works: An index is a separate data structure (often a B-tree or Hash table) that stores the values of a column (or columns) in sorted order, along with pointers to the actual data rows.

    • Analogy: Like an index in a textbook—find a topic quickly without reading every page.

    • Trade-off: More indexes → faster reads, slower inserts/updates/deletes, more disk space. Choose indexes on frequently queried (WHERE, JOIN, ORDER BY) columns.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in