Skip to content
AL-603 (B) ยท Data and Visual Analytics/Quick Revision Short Notes

Data and Visual Analytics (AL-603 (B)) - Unit 3 Short Notes

UNIT 3: Data Analysis, Visualization, and Big Data Ecosystem


I. Foundational Statistical Concepts

A. Variables and Data Categorization (2023)

  • Variable: A characteristic or attribute that can be measured or observed.

  • Data Categorization:

    • Categorical (Qualitative): Represents categories or groups.

      • Nominal: No inherent order (e.g., Gender, Color).

      • Ordinal: Has a meaningful order but not fixed intervals (e.g., Satisfaction: Low, Medium, High).

    • Numerical (Quantitative): Represents measurable quantities.

      • Discrete: Countable values (e.g., Number of students).

      • Continuous: Infinitely many possible values within a range (e.g., Height, Weight).

B. Levels of Measurement (2024)

The scale on which a variable is measured, determining the types of statistical operations permissible.

Level Description Example Operations Allowed
Nominal Labels/names only; no order. Blood Type (A, B, AB, O) Count, Mode
Ordinal Ordered categories; intervals not equal. Likert Scale (1-5) Count, Median, Percentiles
Interval Ordered, equal intervals, no true zero. Celsius Temperature Mean, Std Dev (but ratios meaningless)
Ratio All interval properties + a true, non-arbitrary zero. Height, Weight, Income All statistical measures

[!TIP] Exam Focus: Be prepared to classify given data examples into the correct level of measurement. The presence of a true zero is the key differentiator between Interval and Ratio.

C. Measures of Central Tendency (2023, 2024) โ˜…โ˜…โ˜…

Summarize the center of a dataset.

  1. Mean (Arithmetic Average):

    • Formula (Population): $$\displaystyle \mu = \frac{\sum_{i=1}^{N} x_i}{N} $$

    • Formula (Sample): $$\displaystyle \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$

    • Property: Sensitive to outliers.

  2. Median:

    • The middle value when data is sorted.

    • For even n: Median = average of two middle values.

    • Property: Robust to outliers.

  3. Mode:

    • The most frequently occurring value(s).

    • Can be used for any data type, including nominal.

    • A dataset can be unimodal, bimodal, or multimodal.

D. Measures of Dispersion/Location (2024) โ˜…โ˜…โ˜…

Describe the spread or variability in the data.

Measure Formula (Sample) Description
Range $\text{Max} - \text{Min}$ Simplest measure; highly sensitive to outliers.
Variance $$\displaystyle s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} $$ Average squared deviation from the mean.
Standard Deviation $$\displaystyle s = \sqrt{s^2} $$ Square root of variance; in original units.
Interquartile Range (IQR) $$\displaystyle Q_3 - Q_1 $$ Spread of the middle 50% of data; robust.
Quartiles Q1 (25th), Q2 (Median, 50th), Q3 (75th) Values dividing data into four equal parts.

E. Data Management and Indexing (2023)

  • Goal: Optimize data storage and retrieval for efficient querying.

  • Indexing: A database object that speeds up data retrieval operations (like a book's index).

    • Types:

      • Primary Index: On the primary key; unique, ordered.

      • Secondary Index: On non-key attributes; may have duplicates.

      • Clustered Index: Determines physical order of data rows (only one per table).

      • Non-Clustered Index: Separate structure pointing to data rows.

  • Impact: Proper indexing dramatically reduces I/O cost and query time but adds overhead on data modification (INSERT/UPDATE/DELETE).


II. Statistical Inference and Modeling Techniques

A. Statistical Inferences (2024)

The process of drawing conclusions about a population based on a sample.

  1. Estimation:

    • Point Estimation: Single value estimate of a population parameter (e.g., $\bar{x}$ for $\mu$).

    • Interval Estimation (Confidence Interval): Range of values likely to contain the parameter, with a confidence level (e.g., 95%).

      • CI for mean (known $\sigma$): $$\displaystyle \bar{x} \pm z_{\alpha/2} \frac{\sigma}{\sqrt{n}} $$

      • CI for mean (unknown $\sigma$): $$\displaystyle \bar{x} \pm t_{\alpha/2, n-1} \frac{s}{\sqrt{n}} $$

  2. Hypothesis Testing: Formal procedure to test a claim about a population parameter.

B. Hypothesis Testing Framework

  1. State Hypotheses:

    • $$\displaystyle H_0 $$ (Null Hypothesis): Status quo, assumed true.

    • $$\displaystyle H_1 $$ or $$\displaystyle H_a $$ (Alternative Hypothesis): What we aim to support.

  2. Choose Significance Level ($\alpha$): Probability of rejecting $$\displaystyle H_0 $$ when it is true (Type I Error). Common: 0.05, 0.01.

  3. Compute Test Statistic: Based on sample data and assumed distribution.

  4. Determine p-value or Critical Region: Probability of observing the sample (or more extreme) if $$\displaystyle H_0 $$ is true.

  5. Make Decision: Reject $$\displaystyle H_0 $$ if p-value $\le \alpha$ (or test stat in critical region).

  6. State Conclusion: In context of the problem.

[!TIP] Common Pitfall: Failing to check test assumptions (e.g., normality for t-test) before proceeding. Always validate assumptions first.

C. Parametric Tests

1. t-test (2023, 2024) โ˜…โ˜…โ˜…

Used when population variance is unknown and sample size is small, assuming approximately normal data.

  • One-Sample t-test: Compares sample mean to a known population mean $\mu$.

    • Test Statistic: $$\displaystyle t = \frac{\bar{x} - \mu}{s/\sqrt{n}} $$, df = $n-1$
  • Independent Two-Sample t-test: Compares means of two independent groups.

    • Equal variances assumed: $$\displaystyle t = \frac{\bar{x}_1 - \bar{x}_2}{s_p \sqrt{\frac{1}{n_1} + \frac{1}{n_2}}} $$, where $$\displaystyle s_p^2 = \frac{(n_1-1)s_1^2 + (n_2-1)s_2^2}{n_1+n_2-2} $$, df = $$\displaystyle n_1+n_2-2 $$

    • Unequal variances (Welch's t-test): $$\displaystyle t = \frac{\bar{x}_1 - \bar{x}_2}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} $$, df approximated by Welchโ€“Satterthwaite equation.

  • Paired Sample t-test: Compares two related measurements (e.g., before/after).

    • Compute differences $$\displaystyle d_i = x_{i1} - x_{i2} $$, then use one-sample t-test on $\bar{d}$.

    • Test Statistic: $$\displaystyle t = \frac{\bar{d}}{s_d/\sqrt{n}} $$, df = $n-1$

Example from May 2024 (One-Sample t-test):

Test if potato yield is significantly better than standard $$\displaystyle \mu=20 $$.

Given: $$\displaystyle n=12 $$, $$\displaystyle \bar{x} = \frac{\sum X}{12} = 19.65 $$, $s \approx 3.19$ (calculated).

$$\displaystyle H_0: \mu = 20 $$, $$\displaystyle H_1: \mu > 20 $$ (one-tailed, "better").

$$\displaystyle t = \frac{19.65 - 20}{3.19/\sqrt{12}} = \frac{-0.35}{0.921} = -0.38 $$

For $$\displaystyle \alpha=0.05 $$, df=11, critical $$\displaystyle t_{0.05,11} = 1.796 $$. Since $$\displaystyle |t| < 1.796 $$, fail to reject $$\displaystyle H_0 $$. No significant evidence yield is better.

2. Chi-Square Test (2023, 2024) โ˜…โ˜…โ˜…

  • Test of Independence: Determines if two categorical variables are associated in a population.

    • Used on a contingency table (r x c).

    • Test Statistic: $$\displaystyle \chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}} $$

    • $$\displaystyle E_{ij} = \frac{(\text{Row } i \text{ Total}) \times (\text{Column } j \text{ Total})}{\text{Grand Total}} $$

    • df = $(r-1)(c-1)$

  • Goodness-of-Fit Test: Determines if a single categorical variable follows a hypothesized distribution.

    • Test Statistic: $$\displaystyle \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$

    • df = $k - 1 - m$, where $k$ = number of categories, $m$ = number of estimated parameters.

  • Assumption: Expected frequencies $$\displaystyle E_{ij} \ge 5 $$ for at least 80% of cells, and all $$\displaystyle E_{ij} \ge 1 $$.

Example Calculation (Conceptual): For a 2x2 table of Gender (M/F) vs Preference (A/B), compute row/column totals, find expected counts for each cell, plug into formula, compare $$\displaystyle \chi^2_{calc} $$ to $$\displaystyle \chi^2_{crit} $$ at df=1, $$\displaystyle \alpha=0.05 $$ (critical โ‰ˆ 3.84).

D. Regression Analysis (2023, 2024) โ˜…โ˜…โ˜…

Models the relationship between a dependent variable and one or more independent variables.

  1. Simple Linear Regression (SLR): One independent variable (X).

    • Model: $$\displaystyle Y = \beta_0 + \beta_1 X + \epsilon $$

    • $$\displaystyle \beta_0 $$: Intercept, $$\displaystyle \beta_1 $$: Slope (change in Y per unit change in X).

    • Estimated via Ordinary Least Squares (OLS): Minimizes $$\displaystyle \sum (y_i - \hat{y}_i)^2 $$.

    • Key Output: $$\displaystyle R^2 $$ (proportion of variance in Y explained by X).

  2. Multiple Linear Regression (MLR): Two or more independent variables.

    • Model: $$\displaystyle Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + ... + \beta_p X_p + \epsilon $$

    • Interpretation: $$\displaystyle \beta_j $$ is the effect of $$\displaystyle X_j $$ on Y holding all other X constant.

  3. Types of Variables in Regression Modeling (2023):

    • Dependent/Response Variable (Y): The outcome we predict.

    • Independent/Explanatory/Predictor Variables (X): The inputs used for prediction.

    • Confounding Variable: A variable that influences both X and Y, creating a spurious association.

    • Dummy Variable: Categorical variable encoded numerically (e.g., 0/1 for Gender) for inclusion in regression.

E. Maximum Likelihood Estimation (MLE) (2023, 2024) โ˜…โ˜…โ˜…

  • Concept: A method to estimate model parameters by finding the parameter values that make the observed sample data most probable.

  • Steps:

    1. Write down the likelihood function $$\displaystyle L(\theta | \text{data}) = P(\text{data} | \theta) = \prod_{i=1}^{n} f(x_i; \theta) $$ for independent observations.

    2. Take the log-likelihood: $$\displaystyle \ell(\theta) = \log L(\theta) $$ (simplifies multiplication to addition).

    3. Differentiate $\ell(\theta)$ with respect to $\theta$.

    4. Set derivative(s) to zero and solve for $$\displaystyle \hat{\theta}_{MLE} $$.

  • Example (Bernoulli Distribution): For coin toss data (x successes in n trials), $$\displaystyle L(p) = p^x (1-p)^{n-x} $$. $$\displaystyle \ell(p) = x\log p + (n-x)\log(1-p) $$. $$\displaystyle \frac{d\ell}{dp} = \frac{x}{p} - \frac{n-x}{1-p} = 0 \Rightarrow \hat{p} = \frac{x}{n} $$ (sample proportion).

[!TIP] MLE is the foundation for many statistical models (including logistic regression). It provides consistent and asymptotically normal estimators under regularity conditions.

F. Bayesian Modeling (2024) โ˜…โ˜…

  • Principles: Treats parameters as random variables and updates beliefs about them using Bayes' Theorem.

    • Prior Probability $P(\theta)$: Belief about parameter $\theta$ before seeing data.

    • Likelihood $P(D|\theta)$: Probability of observed data D given parameter $\theta$.

    • Posterior Probability $P(\theta|D)$: Updated belief after observing data.

    • Formula: $$\displaystyle P(\theta|D) = \frac{P(D|\theta) P(\theta)}{P(D)} $$

  • How it Works: Start with a prior distribution. Use the data (likelihood) to update this prior, resulting in a posterior distribution. This posterior becomes the prior for the next data batch (sequential learning).

  • Advantages:

    • Incorporates prior knowledge/expert opinion.

    • Provides full probability distribution for parameters (not just point estimates).

    • Intuitive interpretation of credible intervals.

  • Disadvantages:

    • Choice of prior can be subjective and influential (especially with small data).

    • Computationally intensive for complex models (often requires MCMC sampling).

G. Multivariate Analysis (2024) โ˜…โ˜…

Analyzes datasets with multiple variables simultaneously to understand relationships and structures.

  • Goals: Reduce dimensionality, identify clusters, find latent factors.

  • Common Techniques:

    • Principal Component Analysis (PCA): Unsupervised dimensionality reduction. Transforms data to new orthogonal axes (principal components) capturing maximum variance.

    • Cluster Analysis: Groups observations into clusters based on similarity (e.g., K-Means, Hierarchical).

    • Factor Analysis: Models observed variables as linear combinations of unobserved latent factors.

    • Multivariate Regression: Extension of MLR with multiple dependent variables.

H. Resampling Methods (2023) โ˜…โ˜…

Non-parametric methods that use repeated samples from the original data to estimate uncertainty or validate models.

  1. Bootstrapping:

    • Create many "bootstrap samples" by sampling with replacement from the original sample (same size).

    • Compute the statistic of interest (e.g., mean) for each bootstrap sample.

    • The distribution of these bootstrap statistics estimates the sampling distribution. Used for CIs, standard errors.

  2. Cross-Validation (CV):

    • Primarily for model evaluation/selection.

    • k-fold CV: Split data into k folds. Train on k-1 folds, test on the held-out fold. Repeat k times. Average performance metric.

    • Prevents overfitting and gives a robust estimate of model performance on unseen data.


III. Data Visualization and Exploratory Data Analysis (EDA)

A. Data Visualization: Concepts and Goals (2024)

  • Concepts: The graphical representation of data and information.

  • Primary Goals:

    1. Exploration: Discover patterns, trends, anomalies, and relationships in data (EDA).

    2. Explanation: Communicate findings and tell a story to others.

    3. Monitoring: Track key metrics and performance over time (dashboards).

  • Principles: Choose the right chart for the question, ensure clarity, avoid distortion, provide context (labels, titles).

B. Exploratory Data Analysis (EDA)

Systematic approach to summarizing, visualizing, and understanding data before formal modeling.

  1. Univariate Exploration (2023, 2024) โ˜…โ˜…โ˜…

    • Focus: Single variable.

    • Tools: Histograms, box plots, density plots, bar charts (for categorical), summary statistics.

    • Purpose: Understand distribution, central tendency, spread, identify outliers.

  2. Bivariate Exploration (2024) โ˜…โ˜…โ˜…

    • Focus: Relationship between two variables.

    • Tools:

      • Numerical vs. Numerical: Scatter plot, line plot.

      • Categorical vs. Numerical: Box plot, violin plot, bar chart (with summary stats).

      • Categorical vs. Categorical: Stacked bar chart, heatmap (contingency table).

    • Purpose: Identify correlations, associations, trends.

  3. Multivariate Exploration (2023, 2024) โ˜…โ˜…โ˜…

    • Focus: Relationships among three or more variables.

    • Tools:

      • Color/Size/Shape: Encode additional variables in a 2D scatter plot using aesthetics.

      • Small Multiples (Faceting): Create multiple sub-plots for different subsets.

      • Pair Plots: Matrix of scatterplots for all numerical variable pairs.

      • Heatmaps: For correlation matrices or 3D data (x, y, color).

    • Purpose: Uncover complex interactions, control for confounding variables.

C. Visualization Tools and Libraries

1. Python Visualization Ecosystem (2024) โ˜…โ˜…โ˜…

Library Primary Use Case Key Features
Matplotlib Foundational, low-level plotting. Highly customizable, "grandfather" of Python plotting. Steeper learning curve.
Seaborn Statistical data visualization. Built on Matplotlib. Simplified syntax for complex plots (heatmaps, violin plots, pair plots). Integrates with Pandas.
Plotly Interactive, web-based visualizations. Creates zoomable, hoverable charts. Supports Dash for building analytical web apps.
Bokeh Interactive visualizations for web apps. Similar to Plotly. Strong focus on streaming data and large datasets.

2. ggplot2 (2023) โ˜…โ˜…

  • Based on: The Grammar of Graphics (Wilkinson, 2005). Decomposes plots into layers: Data + Aesthetics (aes) + Geometries (geom) + Facets + Statistics (stat) + Scales + Themes.

  • Syntax: ggplot(data, aes(x=var1, y=var2)) + geom_point() + theme_minimal()

  • Philosophy: Consistent, layered approach produces elegant and complex plots with relatively simple code. Popular in R ecosystem.

3. Power BI (2024) โ˜…โ˜…โ˜…

  • A Business Analytics Service by Microsoft for visualizing data and sharing insights.

  • Tools:

    • Power BI Desktop: Free Windows application for data connection, transformation (Power Query), modeling (DAX), and report creation.

    • Power BI Service: Cloud-based SaaS platform for publishing, sharing, and collaborating on dashboards/reports. Supports scheduling refreshes.

    • Power BI Mobile: Apps for iOS/Android to view dashboards on the go.

  • Dashboard Creation: Drag-and-drop interface. Connects to diverse data sources. Uses DAX (Data Analysis Expressions) for custom measures and calculated columns. Key components: Reports (multiple pages), Dashboards (single-page, real-time tiles), Datasets.

4. Creating Custom Visualizations for Complex Datasets (2023)

  • When standard charts (bar, line, scatter) are insufficient.

  • Approach:

    1. Understand the Data Structure: What are the dimensions and metrics? What relationships are hidden?

    2. Choose an Appropriate Visual Form:

      • Network/Graph Data: Node-link diagrams, force-directed graphs.

      • Hierarchical Data: Treemaps, sunburst diagrams, circle packing.

      • Geospatial Data: Choropleth maps, point maps, flow maps.

      • Time Series with Multiple Series: Streamgraphs, horizon charts.

    3. Leverage Libraries: Use specialized libraries (e.g., networkx/igraph for networks, squarify for treemaps, folium/geopandas for maps) or extend existing ones (Plotly, Bokeh).

    4. Prioritize Interactivity: Tooltips, zooming, filtering to manage complexity.


IV. Data Wrangling, Formats, and Big Data Infrastructure

A. Data Wrangling Process (2023, 2024) โ˜…โ˜…โ˜…

The iterative process of cleaning, structuring, and enriching raw data into a desired format for analysis. Steps:

  1. Gathering: Collect data from various sources (databases, APIs, files, web scraping).

  2. Assessing: Examine data quality. Identify issues: missing values, duplicates, incorrect data types, outliers, structural problems.

  3. Cleaning: Fix identified issues. Tasks include: handling missing data (drop/impute), correcting errors, removing duplicates, standardizing formats.

  4. Transformation: Modify data to suit analysis. Tasks include: filtering, sorting, merging/joining datasets, pivoting, aggregating, creating new features (feature engineering).

  • Tool (Pandas - 2023): The primary Python library for data wrangling.

    • Key objects: DataFrame (2D table), Series (1D column).

    • Key functions: read_csv(), dropna(), fillna(), merge(), groupby(), pivot_table(), apply().

B. File Formats (2023, 2024) โ˜…โ˜…โ˜…

Category Format Description Pros Cons
Structured CSV Plain text, comma-separated values. Human-readable, universal support. No schema, no data types, inefficient for large data.
Excel (.xlsx) Spreadsheet format with multiple sheets. Widespread, supports formulas, formatting. Proprietary, size limits, slower for big data.
Semi-structured JSON JavaScript Object Notation. Hierarchical, key-value pairs. Flexible schema, web-friendly, readable. Verbose, can be large, parsing overhead.
XML eXtensible Markup Language. Tag-based hierarchy. Self-describing, validated by DTD/XSD. Very verbose, complex parsing.
Binary/Serialized Parquet Columnar storage format. Highly efficient for queries (column pruning), compressed, good for big data (Hadoop/Spark). Not human-readable.
Avro Row-based binary format. Compact, fast serialization/deserialization, schema evolution support. Less efficient for columnar queries than Parquet.

[!TIP] For Big Data processing (Hadoop/Spark), Parquet is often the preferred format due to its columnar nature and high compression.

C. Big Data Fundamentals (2024)

  • Definition: Data that exceeds the processing capacity of traditional relational databases and tools, characterized by the 3 Vs (or more):

    • Volume: Massive scale (Terabytes, Petabytes).

    • Velocity: High speed of generation and need for processing (real-time/streaming).

    • Variety: Multiple formats (structured, unstructured, text, images, logs).

    • (Sometimes added: Veracity, Value).

  • Core Challenge: Storing, processing, and extracting value from such data requires distributed, parallel computing architectures.

D. Big Data Processing Tools (2024) โ˜…โ˜…โ˜…

  • Hadoop: The foundational open-source framework for distributed storage and processing.

    • Core Components:

      • HDFS (Hadoop Distributed File System): Distributed, fault-tolerant file system. Splits files into blocks (default 128MB/256MB) and distributes across cluster nodes. Replication (default 3x) ensures fault tolerance.

      • MapReduce: Programming model for processing. Map function processes key-value pairs to produce intermediate results. Reduce function aggregates intermediate results.

    • Limitation: High latency due to disk I/O in MapReduce; not ideal for iterative/real-time processing.

  • Apache Spark: Modern, faster in-memory data processing engine.

    • Key Feature: Uses Resilient Distributed Datasets (RDDs) and DataFrames/Datasets that can be cached in memory across cluster nodes, enabling iterative algorithms (ML, graph) and interactive queries ~100x faster than Hadoop MapReduce.

    • Ecosystem: Spark SQL (structured data), MLlib (machine learning), Spark Streaming (micro-batch streaming), GraphX (graph processing).

  • Other Tools: Apache Flink (true streaming, low-latency), Apache Kafka (distributed streaming platform for ingest), NoSQL Databases (Cassandra, MongoDB for unstructured/semi-structured data).

E. Hadoop Ecosystem 1. HDFS - Hadoop Distributed File System (2023) โ˜…โ˜…

  • Architecture: Master-Slave.

    • NameNode (Master): Manages file system namespace (metadata: file names, permissions, block locations). Single Point of Failure (mitigated by High Availability setups).

    • DataNode (Slave): Stores actual data blocks. Performs block operations (read/write, replication, deletion).

  • Key Concepts: Files are split into blocks (large, e.g., 128MB). Blocks are replicated across DataNodes for fault tolerance. Client reads/writes directly with DataNodes after contacting NameNode for metadata.

2. Hive (2023) โ˜…โ˜…

  • Data Warehouse Infrastructure built on top of Hadoop.

  • Purpose: Provides SQL-like querying (HiveQL) for data stored in HDFS and other storage systems.

  • How it Works: Hive translates HiveQL queries into MapReduce (or Tez/Spark) jobs. Not suitable for low-latency queries (OLTP), but excellent for batch processing of large datasets (OLAP).

  • Components: Metastore (stores table schemas and metadata), HiveServer2 (serves JDBC/ODBC clients), Executor (runs the jobs).

F. Data Analyst Ecosystem (2023)

The interconnected set of tools, technologies, and roles involved in turning raw data into insights.

  • Data Sources: Databases (SQL/NoSQL), Data Lakes (HDFS, S3), Streams (Kafka), Applications.

  • Storage & Processing: Data Warehouses (Redshift, Snowflake), Hadoop/Spark clusters, Cloud platforms (AWS, GCP, Azure).

  • Analysis & Modeling: Languages (Python, R, SQL), Libraries (Pandas, Scikit-learn, TensorFlow), Statistical software (SAS, SPSS).

  • Visualization & BI: Tools (Power BI, Tableau, Looker), Python/R libraries (Matplotlib, ggplot2).

  • Orchestration & Deployment: Workflow tools (Airflow), Containerization (Docker), Orchestration (Kubernetes), MLOps platforms.

  • Roles: Data Engineer (builds pipelines), Data Analyst (explores, reports), Data Scientist (models), BI Developer (dashboards).


V. Synthesis and Application

A. Integrating Statistical Tests with Real-World Data (e.g., t-test application)

  • Workflow:

    1. Define Question: "Is the average processing time after optimization lower than before?"

    2. Wrangle Data: Collect paired measurements (before/after) for the same processes. Clean missing values.

    3. Check Assumptions: Plot differences (histogram, Q-Q plot) for normality. Use Shapiro-Wilk test if needed.

    4. Choose Test: Paired t-test if normality holds; otherwise, use non-parametric Wilcoxon signed-rank test.

    5. Execute & Interpret: Calculate test statistic and p-value. State conclusion in business context (e.g., "With p=0.02 < 0.05, we conclude the optimization significantly reduced processing time.").

B. Selecting Appropriate Visualization and Analysis Tools for Data Context

Data Context Recommended Analysis Recommended Visualization
Small, clean tabular data Pandas, Statsmodels, Scikit-learn Matplotlib, Seaborn (static)
Interactive web reports/dashboards SQL, DAX (Power BI) Power BI, Tableau
Large-scale distributed data PySpark, HiveQL Use Spark to aggregate, then plot aggregated results with standard tools.
Complex, multi-dimensional relationships PCA, Clustering (Scikit-learn) Pair plots, PCA biplots, interactive scatter (Plotly) with color/faceting.
Geospatial data GeoPandas, spatial joins Choropleth maps, point maps (Folium, Plotly Express).
Time-series forecasting Statsmodels (ARIMA), Prophet Line plots with confidence intervals, decomposition plots.

C. End-to-End Workflow: From Raw Data (Wrangling) to Insight (Visualization/Inference)

  1. Acquire: Connect to source (database, API, file) using appropriate tool (Pandas read_sql, requests).

  2. Assess & Clean: Profile data (.info(), .describe(), .isnull().sum()). Handle missing data, correct types, remove duplicates.

  3. Transform & Enrich: Filter relevant rows/columns. Merge with other datasets. Create calculated features (e.g., df['profit'] = df['revenue'] - df['cost']). Aggregate if needed.

  4. Explore (EDA):

    • Univariate: Histograms of key metrics.

    • Bivariate: Scatter plot of Sales vs. Marketing Spend.

    • Multivariate: Heatmap of correlation matrix; color points in scatter by Region.

  5. Model/Test: Based on EDA insights.

    • If comparing groups: Conduct t-test/ANOVA/Chi-square.

    • If predicting: Build regression/classification model.

  6. Visualize & Communicate:

    • Create polished, focused visualizations for key findings.

    • Build an interactive dashboard (Power BI/Plotly Dash) for stakeholders to explore.

    • Summarize statistical test results with clear effect sizes and confidence intervals.

  7. Iterate: Insights from visualization/modeling may lead to new data questions, requiring further wrangling or analysis.

[!TIP] Exam Application: Be ready to describe this workflow for a given scenario (e.g., "Analyze customer churn"). Mention specific tools (Pandas for wrangling, Seaborn for EDA, logistic regression for modeling, Power BI for dashboard).

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in