Skip to content
AL-603 (B) · Data and Visual Analytics/Quick Revision Short Notes

Data and Visual Analytics (AL-603 (B)) - Unit 2 Short Notes

UNIT 2: Data and Visual Analytics - Short Notes


I. Foundational Statistics and Data Types

Measures of Central Tendency

  • Definition: Single value representing the center point of a dataset.

  • Types:

    • Mean (Arithmetic Average): $$\displaystyle \bar{x} = \frac{\sum x_i}{n} $$. Sensitive to outliers.

    • Median: Middle value in sorted data. Robust to outliers.

    • Mode: Most frequent value. Used for nominal data.

  • Properties & Applications:

    • Mean: Best for symmetric, interval/ratio data.

    • Median: Preferred for skewed distributions or ordinal data.

    • Mode: Only measure for nominal data; can be multimodal.

Measures of Dispersion (Spread)

  • Definition: Quantify variability or spread in data.

  • Types:

    • Range: $Max - Min$. Sensitive to outliers.

    • Variance ($$\displaystyle s^2 $$ for sample, $$\displaystyle \sigma^2 $$ for population): $$\displaystyle s^2 = \frac{\sum (x_i - \bar{x})^2}{n-1} $$. Measures average squared deviation.

    • Standard Deviation ($s$ or $\sigma$): $\sqrt{Variance}$. Same units as data.

    • Interquartile Range (IQR): $Q3 - Q1$. Measures spread of middle 50%; resistant to outliers.

  • Applications: Standard deviation used in z-scores, confidence intervals. IQR for box plots.

Levels of Measurement

Level Characteristics Examples Operations Allowed
Nominal Categories only, no order Gender, Blood type Count, Mode
Ordinal Ordered categories, unequal intervals Likert scale, Education level Count, Median, Percentiles
Interval Ordered, equal intervals, no true zero Temperature (°C), IQ score Mean, Std Dev (but ratios meaningless)
Ratio Interval + true zero, meaningful ratios Height, Weight, Age All operations including geometric mean

[!TIP] Exam Focus: Distinguish Interval vs Ratio (true zero). Nominal/Ordinal limit statistical tests.

Variables and Data Categorization

  • Types of Variables:

    • Categorical (Qualitative): Nominal or Ordinal.

    • Numerical (Quantitative):

      • Discrete: Countable integers (e.g., number of children).

      • Continuous: Measurable on a scale (e.g., weight, time).

  • Data Classification Methods:

    • By Nature: Primary (collected firsthand) vs Secondary (existing sources).

    • By Structure: Structured (tabular), Semi-structured (JSON/XML), Unstructured (text/images).


II. Exploratory Data Analysis (EDA)

Univariate Exploration

  • Goal: Understand distribution of a single variable.

  • Techniques:

    • Summary Statistics: Mean, median, std dev, IQR, min/max.

    • Visualizations:

      • Histogram: Shows frequency distribution; choose bin size carefully.

      • Box Plot: Displays median, quartiles, outliers; ideal for comparing groups.

      • KDE Plot: Smoothed version of histogram.

      • Bar Chart: For categorical data frequencies.

Bivariate Exploration

  • Goal: Explore relationship between two variables.

  • Techniques:

    • Scatter Plot: For two numerical variables; reveals correlation, patterns, outliers.

    • Correlation Coefficient ($r$): $-1 \leq r \leq 1$. Linear relationship strength/direction.

    • Cross-Tabulation (Contingency Table): For two categorical variables; shows frequencies.

    • Line Plot: For time-series data (time vs metric).

Multivariate Exploration

  • Goal: Analyze interactions among three or more variables.

  • Techniques:

    • Pair Plots (Scatter Matrix): Grid of scatter plots for all numerical variable pairs; diagonal shows histograms/KDEs.

    • Heatmap: Visualizes correlation matrix; colors indicate strength/direction.

    • Dimensionality Reduction:

      • PCA (Principal Component Analysis): Reduces dimensions to 2D/3D while preserving variance; used in scatter plots.

      • t-SNE: Non-linear reduction for high-dimensional data visualization.

    • 3D Plots/Animation: Add third variable via color, size, or motion.

[!TIP] Common Pitfall: Correlation ≠ Causation. Always check for confounding variables in bivariate analysis.


III. Inferential Statistics and Hypothesis Testing

Statistical Inferences

  • Definition: Drawing conclusions about a population from a sample.

  • Types:

    1. Estimation:

      • Point Estimate: Single value (e.g., $\bar{x}$ for $\mu$).

      • Confidence Interval (CI): Range $[\bar{x} - E, \bar{x} + E]$, where $$\displaystyle E = t_{\alpha/2} \cdot \frac{s}{\sqrt{n}} $$ (for means). Confidence level (e.g., 95%) indicates long-run success rate.

    2. Hypothesis Testing:

      • Null Hypothesis ($$\displaystyle H_0 $$): Status quo (e.g., $$\displaystyle \mu = \mu_0 $$).

      • Alternative Hypothesis ($$\displaystyle H_a $$): Claim to test (one-tailed or two-tailed).

      • p-value: Probability of observing sample result (or more extreme) if $$\displaystyle H_0 $$ true. Reject $$\displaystyle H_0 $$ if p-value < $\alpha$ (significance level, e.g., 0.05).

t-Test

  • Used when: Population std dev unknown, sample size small ($$\displaystyle n < 30 $$), data approximately normal.

  • Types:

    1. One-Sample t-test: Compare sample mean to known $$\displaystyle \mu_0 $$.

      • $$\displaystyle t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} $$, df = $n-1$.
    2. Independent Two-Sample t-test: Compare means of two independent groups.

      • Equal variances: Pooled variance $$\displaystyle s_p^2 $$, df = $$\displaystyle n_1 + n_2 - 2 $$.

      • Unequal variances (Welch's): Separate variances, approximate df.

    3. Paired t-test: Compare two measurements on same subjects (e.g., before/after).

      • $$\displaystyle t = \frac{\bar{d}}{s_d / \sqrt{n}} $$, where $d$ = differences, df = $n-1$.
  • Assumptions: Normality (or large n), independence, homogeneity of variance (for two-sample equal var).

  • Example: Potato Yield Problem (One-sample, one-tailed):

    • $$\displaystyle H_0: \mu = 20 $$, $$\displaystyle H_a: \mu > 20 $$.

    • Data: $$\displaystyle X = [21.5, 24.5, ..., 18.5] $$, $$\displaystyle n=12 $$.

    • $$\displaystyle \bar{x} = 19.79 $$, $$\displaystyle s = 2.98 $$.

    • $$\displaystyle t = \frac{19.79 - 20}{2.98 / \sqrt{12}} = -0.25 $$.

    • df = 11, critical $$\displaystyle t_{0.05,11} = 1.796 $$ (one-tailed). Since $$\displaystyle |t| < 1.796 $$, fail to reject $$\displaystyle H_0 $$. Not significantly better.

Chi-Square Test ($$\displaystyle \chi^2 $$)

  • Used for: Categorical data.

  • Types:

    1. Goodness-of-Fit: Compare observed frequencies to expected frequencies from a known distribution.

      • $$\displaystyle \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$, df = $k-1$ (k = categories).
    2. Test for Independence: Assess association between two categorical variables in a contingency table.

      • $$\displaystyle \chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}} $$, df = $(r-1)(c-1)$.

      • $$\displaystyle E_{ij} = \frac{(row \ total_i) \times (col \ total_j)}{grand \ total} $$.

  • Assumptions: Independent observations, expected frequencies $\geq 5$ (for 80%+ cells).

  • Steps: State hypotheses, compute $$\displaystyle \chi^2 $$, find critical value/p-value, conclude.

Re-sampling Methods

  • Goal: Estimate sampling distribution without parametric assumptions.

  • Bootstrap:

    1. Sample with replacement from original data to create many "bootstrap samples".

    2. Compute statistic (e.g., mean) for each.

    3. Use distribution of bootstrap statistics for CI or SE.

  • Cross-Validation:

    • k-fold CV: Split data into k folds; train on k-1, test on 1; repeat k times. Average performance.

    • Used for model evaluation and hyperparameter tuning.

  • Applications: Bootstrap for CI on medians; CV for preventing overfitting.

Maximum Likelihood Estimation (MLE)

  • Concept: Find parameter values that maximize the likelihood of observing the given sample.

  • Steps:

    1. Write likelihood function $$\displaystyle L(\theta) = \prod f(x_i|\theta) $$ (joint PDF/PMF).

    2. Take log: $$\displaystyle \ell(\theta) = \log L(\theta) = \sum \log f(x_i|\theta) $$.

    3. Differentiate w.r.t. $\theta$, set to 0, solve.

  • Example: For i.i.d. Normal($$\displaystyle \mu, \sigma^2 $$), MLE for $\mu$ is $\bar{x}$, for $$\displaystyle \sigma^2 $$ is $$\displaystyle \frac{\sum (x_i - \bar{x})^2}{n} $$ (biased).

  • Properties: Consistent, asymptotically normal, but may be biased in small samples.


IV. Predictive and Advanced Analytics

Regression Analysis

  • Goal: Model relationship between dependent variable (Y) and independent variable(s) (X).

  • Simple Linear Regression (SLR): $$\displaystyle Y = \beta_0 + \beta_1 X + \epsilon $$.

    • $$\displaystyle \hat{\beta}_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2} $$, $$\displaystyle \hat{\beta}_0 = \bar{y} - \hat{\beta}_1 \bar{x} $$.

    • $$\displaystyle R^2 $$: Proportion of variance in Y explained by X.

  • Multiple Linear Regression (MLR): $$\displaystyle Y = \beta_0 + \beta_1 X_1 + ... + \beta_p X_p + \epsilon $$.

    • Variables:

      • Dependent: Target variable (continuous).

      • Independent: Predictors (continuous or categorical).

      • Dummy Variables: Binary (0/1) for categorical predictors (e.g., Gender: Male=1, Female=0). Reference category omitted to avoid multicollinearity.

    • Interpretation: $$\displaystyle \beta_j $$ = change in Y per unit change in $$\displaystyle X_j $$, holding others constant.

  • Assumptions: Linearity, independence, homoscedasticity, normality of residuals, no multicollinearity.

Multivariate Analysis

  • Goal: Simultaneously analyze multiple outcome variables.

  • Techniques:

    • MANOVA (Multivariate ANOVA): Extension of ANOVA for multiple dependent variables; tests if group means differ on a combination of DVs.

    • Factor Analysis: Reduce many variables to fewer latent factors (e.g., "socioeconomic status" from income, education).

    • Cluster Analysis: Group similar observations (e.g., customer segmentation). Methods: K-means (distance-based), hierarchical.

  • Applications: Market research, psychology, bioinformatics.

Bayesian Modeling

  • Core: Bayes' Theorem: $$\displaystyle P(A|B) = \frac{P(B|A) \cdot P(A)}{P(B)} $$.

    • $P(A)$: Prior belief about A.

    • $P(A|B)$: Posterior belief after observing B.

    • $P(B|A)$: Likelihood.

  • How Bayesian Inference Works:

    1. Specify prior distribution for parameter(s).

    2. Collect data, compute likelihood.

    3. Update prior with likelihood to get posterior distribution (using MCMC or conjugate priors).

    4. Posterior becomes new prior for sequential updating.

  • Advantages:

    • Incorporates prior knowledge.

    • Provides full probability distribution (uncertainty quantification).

    • Intuitive interpretation of credible intervals.

  • Disadvantages:

    • Choice of prior can be subjective.

    • Computationally intensive for complex models.

    • Less familiar to traditional frequentist audiences.


V. Data Management and Engineering

Data Wrangling Process

  1. Gathering: Collect from databases, APIs, files, web scraping.

  2. Cleaning: Handle missing values (impute/remove), correct errors, remove duplicates.

  3. Transformation: Normalize, aggregate, pivot, create new features.

  4. Enrichment: Merge with external data sources.

  5. Validation: Ensure data quality, consistency, and integrity.

  • Tools: Python (Pandas, NumPy), R (dplyr), OpenRefine, SQL.

File Formats

Category Formats Characteristics Use Cases
Structured CSV, Excel Tabular, fixed schema, rows/columns Small-medium datasets, spreadsheets
Semi-structured JSON, XML, Parquet Tags/markers separate elements; schema flexible Web data, configs, big data (Parquet)
Unstructured Text, Images, Videos No predefined schema Social media, logs, multimedia
Columnar Parquet, ORC Store by column; efficient for analytics Big data processing (Hive/Spark)

[!TIP] Exam Key: Parquet is columnar & compressed → faster queries on specific columns vs row-based CSV.

Data Management and Indexing

  • Database Concepts:

    • Relational (SQL): Tables, ACID transactions, primary/foreign keys.

    • NoSQL: Document (MongoDB), Key-Value (Redis), Column-family (Cassandra), Graph (Neo4j). Scale horizontally, flexible schema.

  • Indexing: Data structure (e.g., B-tree) to speed up queries.

    • Clustered Index: Determines physical order of data (one per table).

    • Non-Clustered Index: Separate structure pointing to data rows.

    • Trade-off: Faster reads, slower writes, extra storage.

  • Query Optimization: Use indexes, avoid SELECT *, join order, analyze execution plans.

Big Data Processing Tools

  • Hadoop Ecosystem:

    • HDFS (Hadoop Distributed File System):

      • Architecture: Master-Slave. NameNode (master, metadata), DataNodes (slaves, store blocks). Files split into 128MB blocks, replicated (default 3).

      • Function: Fault-tolerant, scalable storage for large files.

    • Hive:

      • Data warehousing on Hadoop. Provides SQL-like language (HiveQL).

      • Translates queries to MapReduce/Tez/Spark jobs. Schema-on-read (vs schema-on-write).

  • Other Tools:

    • Spark: In-memory processing; faster than MapReduce for iterative tasks. Supports SQL, Streaming, MLlib.

    • NoSQL Databases: MongoDB (document), Cassandra (wide-column) for high write throughput and scalability.

Data Analyst Ecosystem

  • Components:

    1. Data Sources: Databases, APIs, logs, IoT.

    2. Ingestion/ETL: Tools like Apache NiFi, Talend, Stitch.

    3. Storage: Data lakes (S3, HDFS), warehouses (Redshift, Snowflake), databases.

    4. Processing/Transformation: Spark, Pandas, SQL.

    5. Analysis/ML: Python/R (scikit-learn, statsmodels), SAS.

    6. Visualization: Tableau, Power BI, Matplotlib, Seaborn.

    7. Orchestration: Airflow, Luigi.

  • Roles:

    • Data Analyst: Focus on descriptive/predictive analytics, reporting, visualization.

    • Data Engineer: Build/maintain data pipelines, infrastructure.

    • Data Scientist: Advanced modeling, ML, statistical inference.

  • Workflow Integration: Analysts often use cleaned data from engineers; visualize outputs for stakeholders.


VI. Visualization Tools and Implementation

Python Visualization Libraries

Library Primary Use Key Features Best For
Matplotlib Foundational 2D plotting Highly customizable, object-oriented API Static, publication-quality plots
Seaborn Statistical visuals Built on Matplotlib, default themes, simplifies complex plots (heatmaps, violin plots) Exploratory data analysis
Plotly Interactive visuals Web-based, zoom/pan, hover, Dash for apps Dashboards, web apps
Pandas Integrated plotting .plot() method uses Matplotlib backend Quick EDA on DataFrames
ggplot Grammar of Graphics (R port) Layered approach: data → aesthetics → geometries Consistent, complex multi-plot figures

Power BI Ecosystem

  • Power BI Desktop: Free Windows application for report creation. Connect to data sources, model data (Power Query, DAX), design interactive visuals.

  • Power BI Service: Cloud platform (SaaS) for publishing, sharing, collaboration. Apps, dashboards, scheduled refresh.

  • Power BI Tools:

    • Power BI Report Server: On-premises deployment for organizations with cloud restrictions.

    • Power BI Gateway: Bridge between on-premises data sources and Power BI Service.

    • Power BI Mobile: iOS/Android apps for viewing reports.

    • Power BI Embedded: Integrate visuals into custom apps (developer-focused).

  • Workflow: Desktop (develop) → Publish to Service (share) → Mobile/Gateway (access).

Designing Custom Visualizations

  • Principles for Complex Datasets:

    1. Know the audience: Tailor complexity.

    2. Choose appropriate chart type: Match data story (comparison, distribution, relationship, composition).

    3. Simplify: Remove clutter (chartjunk), use color purposefully.

    4. Ensure clarity: Label axes, add titles, use legends.

    5. Enable interaction: Tooltips, filters, drill-down for large datasets.

  • Tool Selection:

    • Static/Publication: Matplotlib, ggplot.

    • Interactive/Dashboards: Plotly Dash, Power BI, Tableau.

    • Geospatial: Folium, Plotly Express.

  • Design Process:

    1. Define question/goal.

    2. Explore data (EDA).

    3. Sketch visualization concepts.

    4. Prototype in chosen tool.

    5. Iterate based on feedback.

    6. Add storytelling: annotations, logical flow, narrative.

  • Storytelling with Data: Guide viewer with title, annotations, sequential slides. Emphasize insights, not just data.

[!TIP] Exam Tip: For "designing custom visuals", emphasize clarity, audience, and narrative flow over technical complexity. Mention specific tools (e.g., "use Plotly for interactive web-based storytelling").


These notes consolidate key definitions, formulas, and concepts from the UNIT 2 syllabus, aligned with past RGPV exam patterns. Focus on understanding applications and distinctions (e.g., measurement levels, regression types, tool comparisons).

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in