Skip to content
AL-603 (C) · Pattern Recognition/Quick Revision Short Notes

Pattern Recognition (AL-603 (C)) - Unit 2 Short Notes

UNIT 2: FOUNDATIONAL STATISTICAL CONCEPTS

Measures of Central Tendency

Definition: Single values that describe the center of a data distribution.

Measure Definition & Formula Properties Applications
Mean ($\bar{x}$) Sum of all values divided by count: $$\displaystyle \bar{x} = \frac{\sum_{i=1}^{n} x_i}{n} $$ Sensitive to outliers; uses all data points. Most common for interval/ratio data.
Median Middle value when data is sorted. For even n: avg of two middle values. Robust to outliers; divides data into halves. Ordinal data or skewed distributions.
Mode Most frequently occurring value(s). Can be multi-modal. Applicable to all data types; may not be unique. Nominal data; identifying common categories.

[!TIP] Exam Focus: Be prepared to calculate each from a small dataset and justify which is most appropriate for a given data type (e.g., median for income data).

Measures of Dispersion (Spread)

Definition: Quantify the variability or spread of data points around the central value.

Measure Formula (Sample) Interpretation
Range $\text{Max} - \text{Min}$ Sensitive to extremes; simple but crude.
Variance ($$\displaystyle s^2 $$) $$\displaystyle s^2 = \frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1} $$ Average squared deviation; units are squared.
Standard Deviation ($s$) $$\displaystyle s = \sqrt{s^2} = \sqrt{\frac{\sum_{i=1}^{n} (x_i - \bar{x})^2}{n-1}} $$ Most common; in original units.
IQR $$\displaystyle Q_3 - Q_1 $$ (75th - 25th percentile) Spread of middle 50%; robust to outliers.

[!TIP] Key Point: Variance and SD use $n-1$ (degrees of freedom) for sample estimation to provide an unbiased estimate of population variance.

Levels of Measurement

Scale Characteristics Permissible Statistics Example
Nominal Categories only; no order. Mode, frequency, chi-square. Gender, blood type.
Ordinal Ordered categories; unequal intervals. Median, percentiles, non-parametric tests. Likert scale, education level.
Interval Ordered, equal intervals, no true zero. Mean, SD, correlation. Temperature (°C), IQ score.
Ratio All interval properties + absolute zero. All statistics, including geometric mean. Height, weight, income.

Variables and Data Categorization

  • Categorical Variables:

    • Nominal: No intrinsic order (e.g., color, country).

    • Ordinal: Has order but not fixed intervals (e.g., small/medium/large).

  • Numerical (Quantitative) Variables:

    • Discrete: Countable values (integers), gaps (e.g., number of children).

    • Continuous: Any value in a range, measurable (e.g., time, weight).

  • Coding Schemes: Converting categorical data to numerical for analysis.

    • Dummy Variables (0/1): For nominal/ordinal in regression.

    • Label Encoding: Assigns integers (can imply false order—use cautiously).

    • One-Hot Encoding: Creates binary column for each category (avoids ordinal assumption).


UNIT 2: INFERENTIAL STATISTICS

Statistical Inference

  • Estimation:

    • Point Estimation: Single value (e.g., $\bar{x}$ estimates $\mu$).

    • Interval Estimation (Confidence Interval): Range likely to contain population parameter.

      • For mean (known $\sigma$): $$\displaystyle \bar{x} \pm z_{\alpha/2} \frac{\sigma}{\sqrt{n}} $$

      • For mean (unknown $\sigma$): $$\displaystyle \bar{x} \pm t_{\alpha/2, df} \frac{s}{\sqrt{n}} $$

  • Hypothesis Testing:

    • Null Hypothesis ($$\displaystyle H_0 $$): Status quo, no effect.

    • Alternative Hypothesis ($$\displaystyle H_1 $$ or $$\displaystyle H_a $$): Research claim.

    • p-value: Probability of observing data given $$\displaystyle H_0 $$ is true. Reject $$\displaystyle H_0 $$ if p-value < $\alpha$ (significance level).

    • Type I Error ($\alpha$): Rejecting true $$\displaystyle H_0 $$.

    • Type II Error ($\beta$): Failing to reject false $$\displaystyle H_0 $$. Power = $1-\beta$.

Parametric Tests

t-test

Assumptions: Random sample, normality (or large n), homogeneity of variance (for two-sample).

Type Purpose Test Statistic Degrees of Freedom
One-sample Compare sample mean to known $$\displaystyle \mu_0 $$. $$\displaystyle t = \frac{\bar{x} - \mu_0}{s/\sqrt{n}} $$ $$\displaystyle df = n-1 $$
Independent two-sample Compare means of two independent groups. $$\displaystyle t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}} $$ Welch's approx. or pooled $df$
Paired Compare two measurements on same subjects. $$\displaystyle t = \frac{\bar{d} - \mu_d}{s_d/\sqrt{n}} $$ $$\displaystyle df = n-1 $$ (where $d$ = differences)

Example (Potato Yield - One-sample t-test from May 2024):

$$\displaystyle H_0: \mu = 20 $$ (standard yield), $$\displaystyle H_1: \mu > 20 $$ (better).

Data: $$\displaystyle X = [21.5, 24.5, ..., 18.5] $$, $$\displaystyle n=12 $$.

  1. Compute $\bar{x}$ and $s$.
  1. Calculate $$\displaystyle t = \frac{\bar{x} - 20}{s/\sqrt{12}} $$.
  1. Compare to $$\displaystyle t_{critical, \alpha, df=11} $$ (one-tailed). Reject $$\displaystyle H_0 $$ if $$\displaystyle t_{calc} > t_{crit} $$.
Chi-Square ($$\displaystyle \chi^2 $$) Test

Assumptions: Independent observations, expected frequencies ≥ 5 (mostly), categorical data.

Test Purpose Statistic Degrees of Freedom
Goodness-of-Fit Compare observed freq. to expected freq. (one variable). $$\displaystyle \chi^2 = \sum \frac{(O_i - E_i)^2}{E_i} $$ $$\displaystyle df = k - 1 - m $$ (k categories, m estimated params)
Test of Independence Test association between two categorical variables (contingency table). Same formula, using row/column totals to find $$\displaystyle E_{ij} = \frac{(row_i \ total) \times (col_j \ total)}{n} $$ $$\displaystyle df = (r-1)(c-1) $$

[!TIP] Calculation Steps: 1) State $$\displaystyle H_0 $$ (no association/good fit). 2) Build table, compute $$\displaystyle E_i $$. 3) Calculate $$\displaystyle \chi^2_{calc} $$. 4) Find critical $$\displaystyle \chi^2 $$ from table or compute p-value. 5) Conclude.

Maximum Likelihood Estimation (MLE)

Goal: Find parameter values ($\theta$) that maximize the likelihood of observed data.

  1. Likelihood Function: $$\displaystyle L(\theta | data) = \prod_{i=1}^{n} f(x_i; \theta) $$ (joint PDF/PMF).

  2. Log-Likelihood: $$\displaystyle \ell(\theta) = \ln L(\theta) = \sum \ln f(x_i; \theta) $$ (easier to maximize).

  3. Estimation: Solve $$\displaystyle \frac{d\ell}{d\theta} = 0 $$ for $$\displaystyle \hat{\theta}_{MLE} $$.

  4. Properties:

    • Consistency: Converges to true $\theta$ as $n \to \infty$.

    • Efficiency: Achieves Cramér-Rao lower bound asymptotically.

    • Invariance: If $\hat{\theta}$ is MLE of $\theta$, then $g(\hat{\theta})$ is MLE of $g(\theta)$.

Example: For Normal($$\displaystyle \mu, \sigma^2 $$), MLE for $\mu$ is $\bar{x}$, for $$\displaystyle \sigma^2 $$ is $$\displaystyle \frac{1}{n}\sum (x_i - \bar{x})^2 $$ (note: biased, uses n not n-1).

Resampling Methods

  • Bootstrap: Sample with replacement from original data to create many "bootstrap samples." Estimate sampling distribution (e.g., CI for median).

  • Cross-Validation:

    • k-fold: Split data into k subsets; train on k-1, test on 1; repeat k times. Average performance.

    • Leave-One-Out (LOO): k = n. High variance, computationally expensive.

  • Permutation Test: Randomly shuffle labels/group assignments to generate null distribution of test statistic. Non-parametric alternative.


UNIT 2: MODELING TECHNIQUES

Regression Analysis

Simple Linear Regression (SLR):

  • Model: $$\displaystyle y = \beta_0 + \beta_1 x + \epsilon $$, $$\displaystyle \epsilon \sim N(0, \sigma^2) $$.

  • Least Squares Estimation: Minimize $$\displaystyle \sum (y_i - \hat{y}_i)^2 $$.

    • $$\displaystyle \hat{\beta}_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2} = \frac{S_{xy}}{S_{xx}} $$

    • $$\displaystyle \hat{\beta}_0 = \bar{y} - \hat{\beta}_1 \bar{x} $$

  • Interpretation: $$\displaystyle \hat{\beta}_1 $$ = change in y per unit change in x.

Multiple Linear Regression (MLR):

  • Model: $$\displaystyle y = \beta_0 + \beta_1 x_1 + ... + \beta_p x_p + \epsilon $$.

  • Assumptions: Linearity, independence, homoscedasticity, normality of errors, no perfect multicollinearity.

  • Multicollinearity: High correlation among predictors. Diagnose with VIF (Variance Inflation Factor). VIF > 5-10 indicates problem.

  • Variables:

    • Dependent (y): Outcome.

    • Independent (x): Predictors.

    • Dummy Variables: 0/1 coding for categorical predictors (e.g., gender).

    • Interaction Terms: $$\displaystyle x_1 \times x_2 $$ to model effect modification.

Multivariate Analysis

Purpose: Analyze multiple variables simultaneously to understand relationships, reduce dimensionality, or group observations.

Technique Goal Key Idea
PCA (Principal Component Analysis) Dimensionality reduction. Find orthogonal axes (PCs) maximizing variance.
Factor Analysis Identify latent factors. Model observed vars as linear combos of unobserved factors.
Cluster Analysis Group similar observations. Distance metrics (Euclidean), algorithms (k-means, hierarchical).
Discriminant Analysis Classify into known groups. Find linear combinations (LDs) maximizing between/within-group variance.
Multivariate Visualization Explore high-dim data. Pair plots, 3D scatter, parallel coordinates, glyphs.

Bayesian Modeling

Core: Bayes' Theorem: $$\displaystyle P(\theta | D) = \frac{P(D | \theta) P(\theta)}{P(D)} $$

  • Posterior $P(\theta | D)$: Updated belief after seeing data.

  • Likelihood $P(D | \theta)$: Probability of data given parameter.

  • Prior $P(\theta)$: Initial belief about parameter.

  • Process:

    1. Specify prior distribution (informative, weakly informative, flat/uniform).

    2. Write likelihood based on data model.

    3. Compute posterior (analytically for conjugate priors, or via MCMC sampling like Gibbs, HMC).

  • Advantages:

    • Incorporates prior knowledge/expert opinion.

    • Provides full probability distribution (uncertainty quantification).

    • Natural for sequential/online learning.

  • Disadvantages:

    • Computationally intensive for complex models.

    • Choice of prior can influence results (sensitivity analysis needed).

    • Can be subjective.


UNIT 2: DATA HANDLING AND VISUALIZATION

Data Wrangling (Munging)

Process:

  1. Data Cleaning:

    • Missing Values: Delete, impute (mean/median/mode, model-based), or flag.

    • Outliers: Detect via IQR (points beyond $Q1-1.5IQR$, $Q3+1.5IQR$) or Z-score (>3). Investigate cause; cap/transform or remove if erroneous.

    • Inconsistencies: Fix formatting (dates, units), correct typos.

  2. Data Transformation:

    • Normalization (Min-Max): $$\displaystyle x' = \frac{x - \min}{\max - \min} $$ → [0,1].

    • Standardization (Z-score): $$\displaystyle x' = \frac{x - \mu}{\sigma} $$ → mean=0, SD=1.

    • Encoding: One-hot for nominal, label/ordinal for ordinal.

  3. Data Integration: Merge/join datasets (SQL-like), concatenate, reshape (pivot/melt).

  4. Tools: Python (Pandas: df.dropna(), df.merge(), pd.get_dummies()), R (dplyr, tidyr).

Data Visualization

Exploration Goal Common Plots
Univariate Distribution of single variable. Histogram, box plot, bar chart (categorical), density plot, violin plot.
Bivariate Relationship between two variables. Scatter plot (num-num), line chart (time series), box plot (cat-num), heatmap (correlation), contingency table (cat-cat).
Multivariate Explore >2 variables. Pair plot (scatter matrix), 3D scatter plot, faceted/trellis plots (small multiples), parallel coordinates, glyphs (star plots), bubble charts (size as 3rd dim).

Custom Visualization Design

Principles (C.A.S.E.):

  • Clarity: Message is immediately understandable. Avoid clutter.

  • Accuracy: Represent data truthfully; no distorted scales.

  • Simplicity: Minimal ink for maximum insight (Tufte's data-ink ratio).

  • Storytelling: Guide viewer with annotations, logical flow, title. Chart Selection: Match plot to data type and question (e.g., trend over time → line chart; part-to-whole → stacked bar/pie [use sparingly]; distribution → histogram/box). Tools: D3.js (web, highly customizable), Tableau (drag-and-drop BI), Python: Matplotlib (low-level), Seaborn (statistical, high-level), Plotly (interactive). R: ggplot2 (grammar of graphics).


UNIT 2: TOOLS, ECOSYSTEMS, AND DATA MANAGEMENT

Visualization Libraries

Library (Python) Key Features Use Case
Matplotlib Foundation; low-level, highly customizable. Static, publication-quality plots; full control.
Seaborn Built on Matplotlib; statistical plots, nice defaults. Quick exploration of distributions, relationships (e.g., relplot, catplot).
Pandas .plot() Integrated with DataFrame; simple syntax. Rapid, basic plots directly from data.
Plotly Interactive, web-based plots (D3.js backend). Dashboards, hover info, zoom, shareable HTML.
plotnine ggplot2 interface for Python. Grammar-of-graphics approach for R users.
R: ggplot2 Grammar of graphics; layered, consistent syntax. Complex, multi-step plot building; industry standard in R.

Business Intelligence Tools: Power BI

  • Power BI Desktop: Free application for data connection, modeling, visualization creation.

  • Power BI Service: Cloud-based SaaS for publishing, sharing, collaboration, scheduling refreshes.

  • Power BI Mobile: Apps for iOS/Android to view reports on-the-go.

  • Key Features:

    • Data Modeling: Create relationships between tables; define calculated columns/measures.

    • DAX (Data Analysis Expressions): Formula language for custom calculations (e.g., CALCULATE, FILTER).

    • Interactive Dashboards: Drill-down, cross-filtering, bookmarks.

    • Report Sharing: Publish to web, share within organization via workspaces.

Big Data Processing Tools

Tool Core Concept Key Component/Feature
Hadoop Ecosystem Distributed storage & processing of big data. HDFS: Master-Slave (NameNode, DataNode). Files split into blocks (default 128MB), replicated (default 3) for fault tolerance. MapReduce: Programming model (Map → Shuffle/Sort → Reduce).
Hive Data warehousing on Hadoop. HiveQL: SQL-like query language. Metastore: Stores schema info (RDBMS). Converts HiveQL to MapReduce/Tez/Spark jobs.
Apache Spark In-memory cluster computing. RDDs (Resilient Distributed Datasets): Immutable, partitioned collections. DataFrames: Distributed tables (optimized via Catalyst). Spark SQL: Structured data processing. Faster than MapReduce for iterative/multi-pass algorithms.

Data Management and Indexing

  • SQL (Relational): Structured data, fixed schema, ACID transactions. (e.g., PostgreSQL, MySQL).

  • NoSQL (Non-relational): Flexible schema, horizontal scaling. Types: Document (MongoDB), Key-Value (Redis), Column-family (Cassandra), Graph (Neo4j).

  • Indexing: Data structure to speed up queries.

    • B-tree: Balanced tree; sorted keys; good for range queries. Default in most RDBMS.

    • Hash Index: Key-value lookup; fast equality; not for ranges.

    • Bitmap Index: Bit vectors for low-cardinality columns; efficient for multi-condition AND/OR (e.g., data warehouses).

Data Analyst Ecosystem

Roles: Collect, clean, analyze, visualize data; translate business questions to data questions; create reports/dashboards. Common Toolchain:

  1. SQL: Extract/transform data from databases.

  2. Python/R: Advanced analysis, statistics, machine learning (Pandas, NumPy, SciPy, scikit-learn, tidyverse).

  3. Excel/Sheets: Quick analysis, prototyping.

  4. BI Tools (Power BI, Tableau): Build interactive dashboards.

  5. Version Control (Git): Track code changes, collaborate. End-to-End Workflow: $$\displaystyle \text{Data Collection} \rightarrow \text{Wrangling} \rightarrow \text{Analysis} \rightarrow \text{Visualization} \rightarrow \text{Reporting} $$.

File Formats

Format Type Structure Pros Cons Best For
CSV/TSV Structured Plain text, delimiter-separated. Human-readable, universal. No schema, no types, inefficient storage. Small-medium data exchange, simple import/export.
JSON Semi-structured Hierarchical (key-value), nested. Flexible schema, web-friendly. Verbose, slower to parse. APIs, config files, NoSQL (MongoDB).
XML Semi-structured Tag-based, hierarchical. Self-describing, validation (XSD). Very verbose, complex. Legacy systems, documents (e.g., SOAP).
Parquet Binary/Columnar Column-oriented storage. Highly compressed, fast column queries, schema embedded. Not human-readable, write-optimized. Big data analytics (Spark, Hive), data lakes.
Avro Binary/Row-based Compact binary, schema stored separately. Fast serialization, schema evolution. Less optimized for column scans than Parquet. Streaming, Kafka, row-based processing.
ORC Binary/Columnar Optimized for Hive (like Parquet). High compression, ACID properties in Hive. Less ecosystem support than Parquet. Hive/Impala workloads.

[!TIP] Rule of Thumb: Use Parquet for analytical queries in big data (column reads). Use CSV/JSON for interchange or small data. Use Avro for row-based streaming or when schema evolves frequently.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in