Unit 1: Fundamentals of Big Data and Hadoop
1.1 Introduction to Big Data
Definition & Evolution:
Big Data refers to datasets that are too large, complex, and fast-changing for traditional data processing tools to handle. Its evolution is driven by the digitization of the world (social media, IoT, sensors) and the need to extract insights from massive, diverse data sources.
Characteristics of Big Data (4 V's):
| V | Definition | Example |
|---|---|---|
| Volume | Scale of data (Terabytes to Zettabytes) | Social media posts, sensor logs |
| Velocity | Speed of data generation & processing needs | Real-time stock trades, GPS tracking |
| Variety | Different formats & types (structured, unstructured, semi-structured) | Text, images, videos, CSV files |
| Veracity | Uncertainty, quality, and trustworthiness of data | Noisy sensor data, incomplete records |
[!TIP] Exam Focus: The 4 V's are a highly frequent 7-mark question. Always define each V with a concrete example.
Applications & Use Cases – Smart Cities:
Big Data analytics enables:
-
Intelligent Traffic Management: Analyzing real-time traffic camera feeds and GPS data to optimize signal timings and reduce congestion.
-
Smart Energy Grids: Balancing load by predicting energy demand from smart meter data.
-
Predictive Waste Management: Optimizing garbage collection routes using fill-level sensor data from bins.
-
Public Safety: Analyzing crime patterns and emergency call data for resource deployment.
1.2 Statistical Foundations for Data Analytics
Regression Analysis:
-
Purpose: To model the relationship between a dependent variable (Y) and one or more independent variables (X). Used for prediction and forecasting.
-
Simple Linear Regression Model:
$$ Y = \beta_0 + \beta_1 X + \epsilon $$
Where:
* $Y$ = Dependent variable
* $X$ = Independent variable
* $$\displaystyle \beta_0 $$ = Intercept
* $$\displaystyle \beta_1 $$ = Slope coefficient
* $\epsilon$ = Error term
- Key Output: Coefficients ($$\displaystyle \beta_0, \beta_1 $$), R-squared (goodness of fit), p-values (significance).
Analysis of Variance (ANOVA):
-
Purpose: To analyze the differences among group means in a sample. Used to determine if at least one group mean is statistically different from others. It partitions total variance into "between-group" and "within-group" variance.
-
One-way ANOVA: Compares means across one categorical independent variable (factor).
-
Key Output: F-statistic, p-value.
Comparative Analysis: Regression vs. ANOVA
| Feature | Regression | ANOVA |
|---|---|---|
| Primary Goal | Model relationship, predict continuous Y | Compare means of groups |
| Dependent Variable (Y) | Continuous | Continuous |
| Independent Variable (X) | Continuous (can be categorical after encoding) | Categorical (factor) |
| Output Focus | Coefficients, prediction equation | F-test for group differences |
| Relationship | Can include categorical X (dummy variables) → ANOVA is a special case of Regression with only categorical X. |
[!TIP] Common Pitfall: Students often confuse their purpose. Remember: Regression predicts/explains a continuous outcome based on X; ANOVA tests if group averages are different.
1.3 Hadoop Ecosystem Overview
Goals & Design Principles:
-
Goals: Store & process massive datasets reliably, scalably, and cost-effectively on commodity hardware.
-
Design Principles:
-
Horizontal Scaling: Add more machines (nodes) to increase capacity.
-
Data Locality: Move computation to where the data resides, not vice-versa.
-
Fault Tolerance: Assume hardware fails; system automatically handles failures.
-
Write Once, Read Many (WORM): Files are immutable once written.
-
Core Architecture Components:
-
HDFS (Hadoop Distributed File System): Distributed storage layer.
-
MapReduce: Distributed computation/processing model.
-
YARN (Yet Another Resource Negotiator): Resource management and job scheduling layer (separates resource management from monitoring).
1.4 Hadoop Distributed File System (HDFS)
Design Goals:
-
Store extremely large files reliably across clusters.
-
Provide high aggregate bandwidth for data access.
-
Tolerate hardware failures gracefully.
Architecture & Components (Master-Slave):
| Component | Role | Key Function |
|---|---|---|
| NameNode (Master) | Master server | Manages file system namespace, metadata (file permissions, block locations). Single Point of Failure (SPOF) in v1. |
| DataNode (Slave) | Worker node | Stores data blocks, serves read/write requests, performs block replication/deletion. |
| Secondary NameNode | Helper to NameNode | Periodically merges fsimage and edit logs (checkpointing). Not a backup. |
Blocks & Replication:
-
Files are split into fixed-size blocks (default 128 MB/256 MB).
-
Each block is replicated across multiple DataNodes (default replication factor = 3).
-
Replacement Policy: First replica on local node (if writer is DataNode), second on a different rack, third on same rack as second but different node. This balances fault tolerance (rack failure) and bandwidth (intra-rack is cheaper).
Fault Tolerance & High Throughput:
-
Fault Tolerance: If a DataNode fails, NameNode detects it (heartbeat mechanism) and automatically re-replicates its blocks to other healthy nodes from existing replicas.
-
High Throughput: Achieved via data locality (MapReduce tasks run on DataNode storing the block) and parallel read/write from multiple DataNodes.
[!DIAGRAM: CANVAS] HDFS Architecture: Draw a cluster with one NameNode (central metadata), multiple DataNodes (each with blocks
blk_1,blk_2replicas). Show Rack 1 and Rack 2. Arrows from client to NameNode (metadata) and directly to DataNodes (data transfer).
1.5 MapReduce Programming Model
Phases:
-
Map Phase:
-
Input:
(key, value)pairs from HDFS (e.g.,(line_offset, line_text)). -
User-defined
map()function processes each pair, emits intermediate(key, value)pairs. -
Example (Word Count):
map(line_offset, "Hello World")→ emit("Hello", 1),("World", 1).
-
-
Shuffle & Sort (Framework-managed):
-
Critical intermediate step. Framework sorts all intermediate keys and groups values for the same key together.
-
Transfers sorted data to the Reduce phase nodes.
-
-
Reduce Phase:
-
Input:
(key, list<values>)from Shuffle (e.g.,("Hello", [1,1,1])). -
User-defined
reduce()function aggregates/summarizes the list for each key. -
Example:
reduce("Hello", [1,1,1])→ emit("Hello", 3). -
Final output written to HDFS.
-
Capacity Scheduler in MapReduce (YARN):
-
Purpose: A pluggable scheduler for YARN that allocates resources (memory, vcores) among multiple organizations/users sharing a cluster.
-
Key Features:
-
Hierarchical Queues: Queues can be parent/child (e.g.,
/engineering,/marketing). -
Capacity Guarantees: Each queue has a minimum capacity (guaranteed share) and can use excess capacity from unused queues (with configurable limits).
-
Resource-Based Scheduling: Schedules based on memory and CPU (vcores), not just slots.
-
Priority Scheduling: Within a queue, applications with higher priority get resources first.
-
Elasticity: Queues can dynamically expand beyond minimum capacity if cluster has free resources.
-
-
Goal: Enable multi-tenancy with fairness and guaranteed SLAs in a shared cluster.
[!TIP] Exam Key: For "Capacity Scheduler," emphasize queues, minimum capacity, elasticity, and multi-tenancy. It's a 5-mark question in past papers.
1.6 Information Management in Big Data Context
Concepts & Challenges:
-
Concept: The processes, systems, and policies for acquiring, storing, securing, organizing, and utilizing large, diverse datasets to derive value.
-
Key Challenges:
-
Volume & Variety: Managing structured, unstructured, and semi-structured data together.
-
Velocity: Handling real-time/streaming data ingestion and processing.
-
Veracity: Ensuring data quality, lineage, and trustworthiness.
-
Storage & Scalability: Cost-effective, elastic storage.
-
Security & Privacy: Access control, encryption, compliance (GDPR).
-
Metadata Management: Tracking data origin, schema, and transformations ("data about data").
-
Integration: Combining data from siloed sources (legacy systems, cloud, IoT).
-
Strategies for Effective Information Management:
-
Adopt a Data Lake/Lakehouse Architecture: Store raw data in its native format (Data Lake) with structured layers for analytics (Lakehouse).
-
Implement Robust Metadata & Catalogs: Use tools (e.g., Apache Atlas, AWS Glue) for data discovery, lineage, and governance.
-
Define Clear Data Governance Policies: Roles, responsibilities, standards, and compliance frameworks.
-
Employ Tiered Storage: Use cheaper storage (e.g., HDFS, S3) for raw data and high-performance storage for hot/active data.
-
Ensure Security by Design: Encryption (at rest, in transit), fine-grained access control (e.g., Apache Ranger, Sentry), and auditing.
-
Use Schema-on-Read: For flexible ingestion of varied data, apply structure during analysis (vs. schema-on-write in traditional RDBMS).
[!TIP] Exam Focus: Link challenges directly to the 4 V's. For strategies, mention Data Lake, Metadata Catalog, and Governance as core pillars.