Skip to content
AL-603 (C) · Pattern Recognition/Quick Revision Short Notes

Pattern Recognition (AL-603 (C)) - Unit 1 Short Notes

UNIT 1: FOUNDATIONS OF DATA ANALYSIS & VISUALIZATION


1.0 FUNDAMENTAL DATA CONCEPTS & MEASUREMENT

1.1 Variables and Data Categorization

  • Variable: A characteristic or attribute that can be measured or observed and can vary among individuals or over time.

  • Data Categorization:

    • Categorical (Qualitative): Represents categories or groups.

      • Nominal: No natural order (e.g., Gender, Color, Country).

      • Ordinal: Natural order but not fixed intervals (e.g., Education Level, Satisfaction Rating: Low/Medium/High).

    • Numerical (Quantitative): Represents measurable quantities.

      • Discrete: Countable, distinct values (e.g., Number of students, Shoe size).

      • Continuous: Any value within a range, measurable (e.g., Height, Weight, Temperature).

[!TIP] Exam Focus: Distinguish Ordinal (ordered categories) from Nominal (no order). Discrete is countable; continuous is measurable.

1.2 Levels of Measurement (Scale of Measurement)

Scale Properties Examples Permissible Operations
Nominal Classification only Gender, ID numbers Equality/Inequality (=, ≠)
Ordinal Classification + Order Rankings, Grades (A,B,C) Equality/Inequality, Greater/Less (<, >)
Interval Order + Fixed intervals, no true zero Temperature (°C, °F), Calendar years All above + Addition/Subtraction (+,-)
Ratio All properties + true zero Height, Weight, Age, Income All arithmetic operations (+, -, ×, ÷)

[!TIP] Key Difference: Interval has arbitrary zero (0°C doesn't mean no heat); Ratio has absolute zero (0 kg means no mass). Only ratio scale allows meaningful ratios (e.g., 20kg is twice 10kg).

1.3 Measures of Central Tendency

Summarizes the center of a dataset.

  • Mean (Arithmetic Average): $$\displaystyle \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$. Sensitive to outliers.

  • Median: Middle value when sorted. Robust to outliers.

  • Mode: Most frequent value. Can be multimodal. Applicable to nominal data.

  • Comparison:

    • Symmetric distribution: Mean ≈ Median ≈ Mode.

    • Right-skewed (positive): Mean > Median > Mode.

    • Left-skewed (negative): Mean < Median < Mode.

[!TIP] For nominal data, only Mode is meaningful. For ordinal, Median is often best.

1.4 Measures of Location and Dispersion (Spread)

  • Location (Percentiles):

    • Quartiles: Q1 (25th), Q2=Median (50th), Q3 (75th).

    • Five-Number Summary: Min, Q1, Median, Q3, Max.

    • Interquartile Range (IQR): $$\displaystyle IQR = Q3 - Q1 $$. Measures spread of middle 50%, robust to outliers.

  • Dispersion:

    • Range: $Max - Min$. Sensitive to outliers.

    • Variance ($$\displaystyle s^2 $$ for sample, $$\displaystyle \sigma^2 $$ for population): $$\displaystyle s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} $$. Average squared deviation.

    • Standard Deviation ($s$ or $\sigma$): $$\displaystyle s = \sqrt{s^2} $$. In original units.

    • Coefficient of Variation (CV): $$\displaystyle CV = \frac{s}{\bar{x}} \times 100\% $$. Unitless measure of relative variability. Useful for comparing datasets with different units/means.

[!FORMULA] Box Plot Whiskers: Typically extend to Q1 - 1.5*IQR and Q3 + 1.5*IQR. Points beyond are potential outliers.

1.5 Introduction to Multivariate Analysis

  • Definition: Simultaneous analysis of more than two variables to understand relationships and structures.

  • Need: Real-world problems involve multiple interacting factors. Univariate/bivariate analysis is insufficient.

  • Key Methods (Conceptual):

    • Multiple Regression: Models relationship between one dependent variable and multiple independent variables.

    • Factor Analysis: Reduces many variables into fewer underlying latent factors.

    • Cluster Analysis: Groups observations into clusters based on similarity.

    • Principal Component Analysis (PCA): Reduces dimensionality by transforming variables into uncorrelated principal components.


2.0 DATA MANAGEMENT, WRANGLING & FORMATS

2.1 Data Wrangling (Data Munging) Process

The process of cleaning, structuring, and enriching raw data into a desired format for analysis.

  1. Data Acquisition: Gathering data from sources (databases, APIs, files, web scraping).

  2. Data Cleaning:

    • Handling Missing Values: Deletion, Imputation (mean/median/mode), Prediction.

    • Handling Outliers: Detection (IQR, Z-score), Capping, Transformation, or Investigation.

    • Correcting Errors: Inconsistencies, typos, invalid entries.

  3. Data Transformation:

    • Normalization/Standardization (scaling).

    • Encoding categorical variables (One-Hot, Label Encoding).

    • Aggregation (summarizing).

    • Pivoting/Melting.

  4. Data Integration: Combining data from multiple sources (merging, joining, concatenating).

  5. Data Enrichment: Adding new derived features or external data to improve analysis.

[!TIP] Golden Rule: Data wrling takes ~60-80% of a data scientist's time. "Garbage in, garbage out."

2.2 File Formats and Data Storage

Format Type Structure Key Characteristics Best For
CSV Structured Tabular, plain text Human-readable, universal, no schema, no type info. Small-medium datasets, simple exchange.
JSON Semi-structured Key-value pairs, nested Hierarchical, flexible schema, web-friendly. APIs, config files, NoSQL (MongoDB).
XML Semi-structured Tagged elements, tree Verbose, strict schema (XSD), self-descriptive. Legacy systems, documents, configs.
Parquet Structured Columnar storage Highly compressed, efficient for queries (read only needed columns), schema stored. Big Data (Hadoop/Spark), analytics workloads.
Avro Structured Row-based, binary Compact, fast serialization, schema evolution support. Streaming data (Kafka), Hadoop.
Excel (.xlsx) Structured Tabular, multiple sheets GUI-based, formulas, formatting. Business users, small datasets, reports.

[!TIP] Columnar (Parquet) is faster for analytical queries (read specific columns). Row-based (Avro, CSV) is better for transactional/row-wise operations.

2.3 Data Management and Indexing (Conceptual)

  • Database Fundamentals:

    • Relational (SQL): Tables with fixed schemas, ACID transactions (e.g., MySQL, PostgreSQL). Good for structured data, complex queries.

    • NoSQL: Flexible schema, horizontal scaling. Types: Document (MongoDB), Key-Value (Redis), Column-family (Cassandra), Graph (Neo4j).

  • Indexing: A database object that improves the speed of data retrieval operations.

    • Analogy: Like a book's index—allows finding a topic (data) without scanning every page (full table scan).

    • Trade-off: Speeds up SELECT/WHERE/JOIN but slows down INSERT/UPDATE/DELETE (index must be updated).

2.4 Big Data Ecosystem Overview

  • Characteristics (4 V's):

    1. Volume: Massive scale (TB, PB, EB).

    2. Velocity: High speed of data generation/need for processing (real-time streams).

    3. Variety: Structured, semi-structured, unstructured.

    4. Veracity: Uncertainty, inconsistency, missing values.

  • Core Processing Frameworks:

    • Hadoop:

      • HDFS (Hadoop Distributed File System): Master-Slave architecture. Stores large files across clusters. Fault-tolerant.

      • MapReduce: Programming model for processing. Map (filter/sort) → Shuffle → Reduce (aggregate). Disk-based, high latency.

    • Apache Spark:

      • In-memory processing. DAG (Directed Acyclic Graph) execution engine.

      • Resilient Distributed Datasets (RDDs): Immutable, partitioned collections.

      • Faster than MapReduce for iterative algorithms (ML) and interactive queries. Supports Spark SQL, Streaming, MLlib.

  • Query/Processing Tools:

    • Hive: Data warehouse on Hadoop. Uses HiveQL (SQL-like) which compiles to MapReduce/Tez/Spark. For batch processing.

    • Pig: High-level scripting language (Pig Latin) for data flows. Compiles to MapReduce. Good for ETL pipelines.

[!TIP] Spark vs. MapReduce: Spark is in-memory and uses DAGs, making it ~100x faster for many workloads. MapReduce is disk-based, more stable for very large, non-iterative jobs.


3.0 STATISTICAL INFERENCE & HYPOTHESIS TESTING

3.1 Statistical Inferences

  • Population: Entire group of interest (parameter: $\mu, \sigma$).

  • Sample: Subset of population (statistic: $\bar{x}, s$).

  • Estimation:

    • Point Estimation: Single value estimate (e.g., $\bar{x}$ estimates $\mu$).

    • Interval Estimation (Confidence Interval - CI): Range of values likely to contain population parameter.

      • 95% CI for $\mu$ (known $\sigma$): $$\displaystyle \bar{x} \pm z_{\alpha/2} \frac{\sigma}{\sqrt{n}} $$

      • 95% CI for $\mu$ (unknown $\sigma$): $$\displaystyle \bar{x} \pm t_{\alpha/2, n-1} \frac{s}{\sqrt{n}} $$

3.2 Parametric Hypothesis Testing

General Steps:

  1. State Hypotheses:

    • $$\displaystyle H_0 $$ (Null): Status quo, no effect, status to be rejected.

    • $$\displaystyle H_1 $$ (Alternative): Research claim, effect exists.

  2. Choose Significance Level ($\alpha$): Probability of Type I error (rejecting true $$\displaystyle H_0 $$). Common: 0.05, 0.01.

  3. Select Appropriate Test & Compute Test Statistic.

  4. Determine p-value or Critical Region.

  5. Decision: If p-value ≤ $\alpha$, reject $$\displaystyle H_0 $$. Or if test stat in critical region, reject $$\displaystyle H_0 $$.

  6. Conclusion in context.

3.2.1 t-Test

Compares means. Assumes (approximately) normally distributed population.

  • One-Sample t-Test: Tests if sample mean differs from a known/hypothesized population mean $$\displaystyle \mu_0 $$.

    • Test Statistic: $$\displaystyle t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} $$, df = $n-1$.

    • Example (Potato Yield - May 2024):

      $$\displaystyle X = [21.5, 24.5, 18.5, 17.2, 14.5, 23.2, 22.1, 20.5, 19.4, 18.1, 24.1, 18.5] $$, $$\displaystyle \mu_0 = 20 $$, $$\displaystyle n=12 $$.

      $$\displaystyle \bar{x} = 19.9 $$, $s \approx 2.94$.

      $$\displaystyle t = \frac{19.9 - 20}{2.94 / \sqrt{12}} \approx -0.118 $$.

      df=11. For one-tailed test (better than standard), p-value > 0.05. Fail to reject $$\displaystyle H_0 $$. No significant evidence yield is better.

  • Two-Sample t-Test:

    • Independent: Two separate groups. Assumes equal variances (pooled) or not (Welch's).

    • Paired (Dependent): Same subjects measured twice (before/after). Uses differences $$\displaystyle d_i = x_{1i} - x_{2i} $$, then one-sample t on $d$.

3.2.2 Chi-Square (χ²) Test

For categorical data. Compares observed frequencies to expected frequencies.

  • Goodness-of-Fit Test: Tests if a single variable's distribution fits a hypothesized distribution.

    • $$\displaystyle \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$, df = $k-1$ (k = categories).
  • Test of Independence (Contingency Table): Tests if two categorical variables are independent.

    • Contingency Table: Rows = categories of Var A, Columns = categories of Var B.

    • $$\displaystyle E_{ij} = \frac{(Row_i \ Total) \times (Col_j \ Total)}{Grand \ Total} $$.

    • $$\displaystyle \chi^2 = \sum \frac{(O_{ij} - E_{ij})^2}{E_{ij}} $$, df = $(r-1)(c-1)$.

    • Assumption: All $$\displaystyle E_{ij} \geq 5 $$ (otherwise use Fisher's Exact Test).

  • Decision: Reject $$\displaystyle H_0 $$ (variables are independent) if $$\displaystyle \chi^2_{calc} > \chi^2_{crit} $$ or p-value < $\alpha$.

[!TIP] Chi-Square: $$\displaystyle H_0 $$: Variables are independent (no association). Rejecting $$\displaystyle H_0 $$ implies association.

3.3 Regression Analysis

Models relationship between variables.

  • Dependent (Response) Variable $Y$: Variable being predicted/explained.

  • Independent (Predictor) Variable(s) $X$: Variable(s) used for prediction.

  • Simple Linear Regression (SLR):

    • Model: $$\displaystyle Y = \beta_0 + \beta_1 X + \epsilon $$

    • $$\displaystyle \beta_0 $$: Intercept (value of Y when X=0).

    • $$\displaystyle \beta_1 $$: Slope (change in Y per unit change in X).

    • Least Squares Estimation: Minimizes $$\displaystyle \sum (y_i - \hat{y}_i)^2 $$.

    • $$\displaystyle \hat{\beta}_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2} $$, $$\displaystyle \hat{\beta}_0 = \bar{y} - \hat{\beta}_1 \bar{x} $$.

  • Multiple Linear Regression (MLR):

    • Model: $$\displaystyle Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + ... + \beta_p X_p + \epsilon $$

    • Each $$\displaystyle \beta_j $$ represents the effect of $$\displaystyle X_j $$ holding all other X's constant.

  • Categorical Predictors (Dummy Variables): Convert categories into binary (0/1) variables. For k categories, use k-1 dummies to avoid multicollinearity (dummy variable trap).

3.4 Maximum Likelihood Estimation (MLE)

  • Philosophy: Find parameter values ($\theta$) that make the observed data most probable.

  • Steps:

    1. Write down the likelihood function $$\displaystyle L(\theta | data) = P(data | \theta) = \prod_{i=1}^{n} f(x_i | \theta) $$ (for i.i.d. data).

    2. Take the log-likelihood: $$\displaystyle \ell(\theta) = \log L(\theta) $$ (simplifies product to sum).

    3. Differentiate $\ell(\theta)$ w.r.t. $\theta$ and set to zero: $$\displaystyle \frac{d\ell}{d\theta} = 0 $$.

    4. Solve for $\theta$. This $$\displaystyle \hat{\theta}_{MLE} $$ maximizes likelihood.

  • Example (Bernoulli): Data: k successes in n trials. $$\displaystyle L(p) = p^k (1-p)^{n-k} $$. $$\displaystyle \ell(p) = k \log p + (n-k) \log(1-p) $$. $$\displaystyle \frac{d\ell}{dp} = \frac{k}{p} - \frac{n-k}{1-p} = 0 $$ → $$\displaystyle \hat{p} = \frac{k}{n} $$ (sample proportion).

[!TIP] MLE is a general principle, not a specific test. Used to estimate parameters for many models (regression, logistic, etc.).

3.5 Bayesian Modelling (Conceptual)

  • Bayes' Theorem: $$\displaystyle P(\theta | D) = \frac{P(D | \theta) P(\theta)}{P(D)} $$

    • $P(\theta | D)$: Posterior (updated belief about $\theta$ after seeing data D).

    • $P(D | \theta)$: Likelihood (probability of data given $\theta$).

    • $P(\theta)$: Prior (belief about $\theta$ before seeing data).

    • $P(D)$: Marginal likelihood (normalizing constant).

  • Bayesian vs. Frequentist:

    • Bayesian: Parameters are random variables with probability distributions. Incorporates prior knowledge. Results are probability statements about parameters (e.g., "There's a 95% probability $\theta$ is in this interval").

    • Frequentist: Parameters are fixed but unknown. Relies on repeated sampling concept. Confidence intervals are about long-run coverage.

  • Advantages: Incorporates prior info, intuitive probabilistic interpretation, handles small data well, full posterior distribution.

  • Disadvantages: Computationally intensive (MCMC), choice of prior can be subjective/influential.

3.6 Re-sampling Methods (Conceptual)

Estimate uncertainty/performance without strict distributional assumptions.

  • Cross-Validation:

    • k-Fold CV: Split data into k folds. Train on k-1, test on 1. Repeat k times. Average performance.

    • Purpose: Model selection, hyperparameter tuning, estimating generalization error.

  • Bootstrap:

    • Sample with replacement from original data to create many "bootstrap samples" (same size n).

    • Compute statistic (e.g., mean) for each sample → bootstrap distribution.

    • Use distribution to estimate standard error, confidence intervals (percentile method).

    • Purpose: Estimating variability of almost any statistic.


4.0 DATA VISUALIZATION PRINCIPLES & EXPLORATION

4.1 Introduction to Data Visualization

  • Goals:

    • Exploration: Discover patterns, outliers, relationships (for yourself).

    • Explanation/Communication: Present findings to others clearly.

  • Principles of Effective Visualization:

    • Data-Ink Ratio (Tufte): Maximize ink used for data, minimize non-data ink (decoration, 3D effects).

    • Chart Selection: Match plot type to question (distribution? relationship? comparison?).

    • Color Theory: Use sequential for ordered, diverging for deviation from midpoint, categorical for distinct groups. Avoid rainbow scales. Consider colorblindness.

    • Avoid Distortion: Truncated axes, inappropriate scales can mislead.

    • Label Clearly: Axes, titles, legends.

4.2 Univariate Exploration (One Variable)

Data Type Plots Summarizes
Numerical Histogram (bins), Density Plot (smoothed), Box Plot, Stem-and-Leaf Shape (symmetry, peaks), Center, Spread, Outliers
Categorical Bar Chart (counts/frequencies), Pie Chart (use sparingly) Frequency/Proportion of each category

[!TIP] Box Plot shows 5-number summary and outliers. Histogram bin width choice greatly affects perceived shape.

4.3 Bivariate Exploration (Two Variables)

X (Predictor) Y (Response) Recommended Plots Measure of Association
Numerical Numerical Scatter Plot, Line Chart (time series) Pearson's r (linear), Spearman's ρ (rank)
Categorical Numerical Box Plot (by category), Violin Plot, Bar Chart (with summary stat) ANOVA (F-test) for mean differences
Categorical Categorical Grouped/Mosaic Bar Chart, Heatmap (contingency table) Chi-Square (χ²) test of independence, Cramer's V

4.4 Multivariate Exploration (>2 Variables)

  • Using 2D Plot Aesthetics (add dimensions to scatter/bar):

    • Color: Categorical or continuous (sequential palette).

    • Size: Numerical (area proportional to value).

    • Shape: Categorical (distinct markers).

    • Faceting (Small Multiples): Create multiple sub-plots for levels of a categorical variable. Best practice for adding dimensions without clutter.

  • 3D Plots: Often hard to interpret; avoid unless truly necessary and interactive.

  • Specialized Plots:

    • Scatterplot Matrix (Pairs Plot): Grid of scatterplots for all numerical variable pairs.

    • Parallel Coordinates: Each variable is an axis; observations are lines. Good for high-dimensional patterns.

4.5 Creating Custom Visualizations for Complex Datasets

Process:

  1. Understand the Question & Data: What story? What variables? What relationships?

  2. Choose the Right Encodings: Map variables to visual channels (x, y, color, size, shape) based on data type and importance.

  3. Iterative Design:

    • Start simple (univariate, bivariate).

    • Add dimensions cautiously (faceting > color > size).

    • Refine: Improve labels, titles, legends, remove clutter.

    • Test: Does it convey the intended message quickly?

  4. Leverage Interactivity (for exploration): Tooltips, zoom, filter, linked views ( brushing & linking).


5.0 TOOLS & TECHNOLOGIES FOR ANALYSIS & VISUALIZATION

5.1 Python Ecosystem

5.1.1 Pandas

  • Core Data Structures:

    • Series: 1D labeled array.

    • DataFrame: 2D labeled table (rows: index, columns: columns). Workhorse.

  • Core Operations:

    • Indexing/Selection: df[col], df.loc[] (label-based), df.iloc[] (position-based).

    • Filtering: Boolean indexing df[df['col'] > 5].

    • Grouping & Aggregation: df.groupby('cat_col').agg({'num_col': ['mean', 'sum']}).

    • Merging/Joining: pd.merge(df1, df2, on='key') (SQL-like joins), pd.concat().

    • Handling Missing: isna(), fillna(), dropna().

5.1.2 Core Visualization Libraries

Library Philosophy Strengths Use Case
Matplotlib Foundation, low-level Full control, highly customizable, object-oriented API (Figure, Axes). Complex, publication-quality static plots; building blocks for others.
Seaborn Statistical graphics, high-level Built on Matplotlib. Simplifies complex plots (distribution, regression, categorical). Beautiful defaults/themes. Quick statistical exploration (distributions, relationships, categorical summaries).
Plotly Interactive, web-based Interactive by default (zoom, pan, hover). Web-based (HTML). plotly.express for simple syntax. Dashboards, web apps, presentations, exploratory analysis needing interactivity.

[!TIP] Typical Workflow: Use Pandas for data manipulation → Seaborn for quick statistical plots → Matplotlib for fine-tuning → Plotly for interactive needs.

5.2 Specialized Visualization Tools

  • ggplot2 (R): Based on Grammar of Graphics.

    • Layered Approach: ggplot(data) + geom_point(mapping=aes(x,y)) + labs(...).

    • Components: Data, Aesthetics (aes()), Geometries (geom_*), Facets, Statistics, Scales, Coordinates, Themes.

    • Strength: Consistent, logical framework for building complex plots layer by layer.

5.3 Business Intelligence & Dashboarding Tools

  • Power BI Ecosystem:

    • Power BI Desktop: Free Windows application for data modeling (using DAX - Data Analysis Expressions), report creation, and visualization.

    • Power BI Service: Cloud-based SaaS for publishing, sharing, collaborating, and scheduling refreshes.

    • Power BI Mobile: Apps for viewing reports on mobile devices.

    • Key Features: Drag-and-drop interface, rich visualizations, real-time dashboards, natural language Q&A, integration with Azure/Office 365.

5.4 The Data Analyst Ecosystem (Holistic View)

A typical pipeline integrates multiple tools:

  1. Storage: SQL databases (structured), HDFS/Data Lakes (big, raw).

  2. Wrangling & Analysis: SQL (initial extraction), Python (Pandas, NumPy) / R for cleaning, transformation, statistical analysis, ML.

  3. Visualization & Exploration: Python libs (Matplotlib/Seaborn/Plotly), ggplot2, or BI tools for initial exploration.

  4. Dashboarding & Sharing: Power BI, Tableau, Plotly Dash, Streamlit for interactive dashboards and storytelling.

  5. Deployment: Embed reports in websites/apps, schedule refreshes, set up alerts.

[!TIP] Role Clarity: Data Analyst focuses on descriptive/predictive analytics, SQL, BI tools, business metrics. Data Scientist adds advanced ML, Python/R coding, experimental design. Data Engineer builds/maintains data pipelines (Spark, Airflow, cloud services).

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in