Skip to content
CE-803 (B) · Data Analytics/Quick Revision Short Notes

Data Analytics (CE-803 (B)) - Unit 1 Short Notes

Unit 1: Fundamentals of Big Data and Hadoop


1.1 Introduction to Big Data

Definition & Evolution:

Big Data refers to datasets that are too large, complex, and fast-changing for traditional data processing tools to handle. Its evolution is driven by the digitization of the world (social media, IoT, sensors) and the need to extract insights from massive, diverse data sources.

Characteristics of Big Data (4 V's):

V Definition Example
Volume Scale of data (Terabytes to Zettabytes) Social media posts, sensor logs
Velocity Speed of data generation & processing needs Real-time stock trades, GPS tracking
Variety Different formats & types (structured, unstructured, semi-structured) Text, images, videos, CSV files
Veracity Uncertainty, quality, and trustworthiness of data Noisy sensor data, incomplete records

[!TIP] Exam Focus: The 4 V's are a highly frequent 7-mark question. Always define each V with a concrete example.

Applications & Use Cases – Smart Cities:

Big Data analytics enables:

  • Intelligent Traffic Management: Analyzing real-time traffic camera feeds and GPS data to optimize signal timings and reduce congestion.

  • Smart Energy Grids: Balancing load by predicting energy demand from smart meter data.

  • Predictive Waste Management: Optimizing garbage collection routes using fill-level sensor data from bins.

  • Public Safety: Analyzing crime patterns and emergency call data for resource deployment.


1.2 Statistical Foundations for Data Analytics

Regression Analysis:

  • Purpose: To model the relationship between a dependent variable (Y) and one or more independent variables (X). Used for prediction and forecasting.

  • Simple Linear Regression Model:

$$ Y = \beta_0 + \beta_1 X + \epsilon $$

Where:

*   $Y$ = Dependent variable

*   $X$ = Independent variable

*   $$\displaystyle \beta_0 $$ = Intercept

*   $$\displaystyle \beta_1 $$ = Slope coefficient

*   $\epsilon$ = Error term
  • Key Output: Coefficients ($$\displaystyle \beta_0, \beta_1 $$), R-squared (goodness of fit), p-values (significance).

Analysis of Variance (ANOVA):

  • Purpose: To analyze the differences among group means in a sample. Used to determine if at least one group mean is statistically different from others. It partitions total variance into "between-group" and "within-group" variance.

  • One-way ANOVA: Compares means across one categorical independent variable (factor).

  • Key Output: F-statistic, p-value.

Comparative Analysis: Regression vs. ANOVA

Feature Regression ANOVA
Primary Goal Model relationship, predict continuous Y Compare means of groups
Dependent Variable (Y) Continuous Continuous
Independent Variable (X) Continuous (can be categorical after encoding) Categorical (factor)
Output Focus Coefficients, prediction equation F-test for group differences
Relationship Can include categorical X (dummy variables) → ANOVA is a special case of Regression with only categorical X.

[!TIP] Common Pitfall: Students often confuse their purpose. Remember: Regression predicts/explains a continuous outcome based on X; ANOVA tests if group averages are different.


1.3 Hadoop Ecosystem Overview

Goals & Design Principles:

  • Goals: Store & process massive datasets reliably, scalably, and cost-effectively on commodity hardware.

  • Design Principles:

    1. Horizontal Scaling: Add more machines (nodes) to increase capacity.

    2. Data Locality: Move computation to where the data resides, not vice-versa.

    3. Fault Tolerance: Assume hardware fails; system automatically handles failures.

    4. Write Once, Read Many (WORM): Files are immutable once written.

Core Architecture Components:

  • HDFS (Hadoop Distributed File System): Distributed storage layer.

  • MapReduce: Distributed computation/processing model.

  • YARN (Yet Another Resource Negotiator): Resource management and job scheduling layer (separates resource management from monitoring).


1.4 Hadoop Distributed File System (HDFS)

Design Goals:

  • Store extremely large files reliably across clusters.

  • Provide high aggregate bandwidth for data access.

  • Tolerate hardware failures gracefully.

Architecture & Components (Master-Slave):

Component Role Key Function
NameNode (Master) Master server Manages file system namespace, metadata (file permissions, block locations). Single Point of Failure (SPOF) in v1.
DataNode (Slave) Worker node Stores data blocks, serves read/write requests, performs block replication/deletion.
Secondary NameNode Helper to NameNode Periodically merges fsimage and edit logs (checkpointing). Not a backup.

Blocks & Replication:

  • Files are split into fixed-size blocks (default 128 MB/256 MB).

  • Each block is replicated across multiple DataNodes (default replication factor = 3).

  • Replacement Policy: First replica on local node (if writer is DataNode), second on a different rack, third on same rack as second but different node. This balances fault tolerance (rack failure) and bandwidth (intra-rack is cheaper).

Fault Tolerance & High Throughput:

  • Fault Tolerance: If a DataNode fails, NameNode detects it (heartbeat mechanism) and automatically re-replicates its blocks to other healthy nodes from existing replicas.

  • High Throughput: Achieved via data locality (MapReduce tasks run on DataNode storing the block) and parallel read/write from multiple DataNodes.

[!DIAGRAM: CANVAS] HDFS Architecture: Draw a cluster with one NameNode (central metadata), multiple DataNodes (each with blocks blk_1, blk_2 replicas). Show Rack 1 and Rack 2. Arrows from client to NameNode (metadata) and directly to DataNodes (data transfer).


1.5 MapReduce Programming Model

Phases:

  1. Map Phase:

    • Input: (key, value) pairs from HDFS (e.g., (line_offset, line_text)).

    • User-defined map() function processes each pair, emits intermediate (key, value) pairs.

    • Example (Word Count): map(line_offset, "Hello World") → emit ("Hello", 1), ("World", 1).

  2. Shuffle & Sort (Framework-managed):

    • Critical intermediate step. Framework sorts all intermediate keys and groups values for the same key together.

    • Transfers sorted data to the Reduce phase nodes.

  3. Reduce Phase:

    • Input: (key, list<values>) from Shuffle (e.g., ("Hello", [1,1,1])).

    • User-defined reduce() function aggregates/summarizes the list for each key.

    • Example: reduce("Hello", [1,1,1]) → emit ("Hello", 3).

    • Final output written to HDFS.

Capacity Scheduler in MapReduce (YARN):

  • Purpose: A pluggable scheduler for YARN that allocates resources (memory, vcores) among multiple organizations/users sharing a cluster.

  • Key Features:

    • Hierarchical Queues: Queues can be parent/child (e.g., /engineering, /marketing).

    • Capacity Guarantees: Each queue has a minimum capacity (guaranteed share) and can use excess capacity from unused queues (with configurable limits).

    • Resource-Based Scheduling: Schedules based on memory and CPU (vcores), not just slots.

    • Priority Scheduling: Within a queue, applications with higher priority get resources first.

    • Elasticity: Queues can dynamically expand beyond minimum capacity if cluster has free resources.

  • Goal: Enable multi-tenancy with fairness and guaranteed SLAs in a shared cluster.

[!TIP] Exam Key: For "Capacity Scheduler," emphasize queues, minimum capacity, elasticity, and multi-tenancy. It's a 5-mark question in past papers.


1.6 Information Management in Big Data Context

Concepts & Challenges:

  • Concept: The processes, systems, and policies for acquiring, storing, securing, organizing, and utilizing large, diverse datasets to derive value.

  • Key Challenges:

    • Volume & Variety: Managing structured, unstructured, and semi-structured data together.

    • Velocity: Handling real-time/streaming data ingestion and processing.

    • Veracity: Ensuring data quality, lineage, and trustworthiness.

    • Storage & Scalability: Cost-effective, elastic storage.

    • Security & Privacy: Access control, encryption, compliance (GDPR).

    • Metadata Management: Tracking data origin, schema, and transformations ("data about data").

    • Integration: Combining data from siloed sources (legacy systems, cloud, IoT).

Strategies for Effective Information Management:

  1. Adopt a Data Lake/Lakehouse Architecture: Store raw data in its native format (Data Lake) with structured layers for analytics (Lakehouse).

  2. Implement Robust Metadata & Catalogs: Use tools (e.g., Apache Atlas, AWS Glue) for data discovery, lineage, and governance.

  3. Define Clear Data Governance Policies: Roles, responsibilities, standards, and compliance frameworks.

  4. Employ Tiered Storage: Use cheaper storage (e.g., HDFS, S3) for raw data and high-performance storage for hot/active data.

  5. Ensure Security by Design: Encryption (at rest, in transit), fine-grained access control (e.g., Apache Ranger, Sentry), and auditing.

  6. Use Schema-on-Read: For flexible ingestion of varied data, apply structure during analysis (vs. schema-on-write in traditional RDBMS).

[!TIP] Exam Focus: Link challenges directly to the 4 V's. For strategies, mention Data Lake, Metadata Catalog, and Governance as core pillars.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in