Skip to content
CE-803 (B) · Data Analytics/Quick Revision Short Notes

Data Analytics (CE-803 (B)) - Unit 3 Short Notes

UNIT 3: Big Data Technologies & Analytics


1.0 Fundamentals of Big Data

1.1 Definition and Need for Big Data

  • Big Data refers to datasets that are too large, complex, and fast-changing for traditional data processing tools to capture, store, manage, and analyze effectively.

  • Need: To uncover hidden patterns, correlations, and insights from vast, diverse data sources (sensors, social media, transactions) for better decision-making, predictive analytics, and innovation.

1.2 Characteristics of Big Data (The 4 V's)

V Description Example
Volume Scale of data (Terabytes, Petabytes, Exabytes) Social media posts, sensor data
Velocity Speed of data generation and processing needs Real-time stock trades, IoT streams
Variety Different formats and types (structured, semi-structured, unstructured) CSV files, JSON logs, images, videos
Veracity Data quality, uncertainty, and trustworthiness Noisy sensor data, incomplete records

[!TIP] Exam Focus: The 4 V's are a very high-frequency question. Be prepared to define each with a relevant example.

1.3 Goals and Drivers of Big Data Adoption

  • Goals: Improve operational efficiency, enhance customer experience, enable new revenue streams, manage risks, and drive innovation.

  • Drivers: Proliferation of connected devices (IoT), cheaper storage/compute, advanced analytics (ML/AI), and competitive pressure.

1.4 Application Domains

  • Smart Cities: Traffic optimization, smart grids, public safety.

  • IoT: Predictive maintenance, asset tracking.

  • Business Analytics: Customer segmentation, recommendation systems, fraud detection.


2.0 Hadoop Ecosystem & Architecture

2.1 Core Components of Hadoop 2.x/3.x Architecture

2.1.1 Hadoop Distributed File System (HDFS)
  • Goals: Store massive files across clusters of commodity hardware reliably and efficiently.

  • Architecture: Master-Slave.

    • NameNode (Master): Manages file system namespace, metadata (file->block mapping), and regulates access.

    • DataNode (Slave): Stores actual data blocks, handles read/write requests, performs block replication.

  • Read Mechanism:

    1. Client contacts NameNode for block locations.

    2. NameNode returns DataNode addresses.

    3. Client reads data from the closest DataNode (rack-aware).

  • Write Mechanism:

    1. Client requests file creation; NameNode checks & grants permission.

    2. Client splits file into blocks (default 128MB/256MB).

    3. Data is pipelined to a series of DataNodes (default replication=3).

    4. DataNode acknowledges receipt and forwards to next in pipeline.

    5. Client notifies NameNode on completion.

2.1.2 Yet Another Resource Negotiator (YARN)
  • Goal: Separate resource management and scheduling from data processing (MapReduce), enabling multiple processing engines (Spark, Tez) on the same cluster.

  • Architecture & Components:

    • ResourceManager (RM): Global master, arbitrates resources among all applications.

    • NodeManager (NM): Per-node agent, manages containers, monitors resource usage.

    • ApplicationMaster (AM): Per-application master (one per job), negotiates resources from RM, works with NM to execute tasks.

2.2 Hadoop Ecosystem Overview

Tool Purpose Key Feature
MapReduce Parallel processing of large datasets Programming model (Map, Shuffle, Reduce)
Hive Data warehouse on Hadoop HiveQL (SQL-like), Metastore
Pig Scripting platform for data flows Pig Latin (high-level scripting)
HBase NoSQL, real-time read/write Columnar store on HDFS, low-latency
ZooKeeper Coordination service Centralized configuration, synchronization
Sqoop RDBMS ↔ Hadoop transfer Structured data import/export
Flume Streaming data ingestion Log/event data collection

3.0 Data Processing Frameworks

3.1 MapReduce in Detail

3.1.1 Programming Paradigm and Workflow
  1. Map Phase: Processes input key-value pairs (<k1, v1>) to produce intermediate key-value pairs (<k2, v2>). Example: Word Count - emit <word, 1>.

  2. Shuffle & Sort: Framework automatically groups all values by their intermediate keys (k2).

  3. Reduce Phase: Aggregates values for each key (<k2, list(v2)>) to produce final output (<k3, v3>). Example: Sum counts for each word.

3.1.2 Job Scheduling
Scheduler Policy Best For
Capacity Scheduler Hierarchical queues with guaranteed capacity, FIFO within queue Multi-tenant clusters with SLAs
Fair Scheduler Allocates equal share of resources to running jobs, can preempt Mixed workloads, fairness priority

3.2 Apache Pig

3.2.1 Pig Latin Data Model & Operators
  • Data Model: Atomic (int, long, float, etc.) or Complex (Tuple, Bag, Map).

  • Key Operators: LOAD, FILTER, GROUP, JOIN, FOREACH...GENERATE, STORE.

  • Example: A = LOAD 'data' AS (name, age); B = FILTER A BY age > 20;

3.2.2 Execution Modes
  • Local Mode: Runs on a single JVM (for development/testing).

  • MapReduce Mode: Submits jobs to Hadoop cluster (default).

3.3 Apache Hive

3.3.1 HiveQL DDL Commands
-- CREATE (Internal Table)

CREATE TABLE employees (id INT, name STRING) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',';

-- CREATE (External Table - schema-only, data outside Hive)

CREATE EXTERNAL TABLE logs (timestamp STRING, message STRING) LOCATION '/user/logs';

-- ALTER (Add column)

ALTER TABLE employees ADD COLUMNS (dept STRING);

-- DROP (Internal table deletes data; External does not)

DROP TABLE employees;

3.3.2 HiveQL DML Commands
  • LOAD DATA [LOCAL] INPATH 'file' [OVERWRITE] INTO TABLE table; (Move/copy data into Hive's warehouse)

  • INSERT INTO TABLE table SELECT ...; (Query-based insert)

  • SELECT ... FROM ... WHERE ... GROUP BY ... HAVING ... ORDER BY ...;

3.3.3 Metastore
  • Structure: Relational database (Derby, MySQL, PostgreSQL) storing metadata: table names, column names & types, partition info, table location, serialization/deserialization (SerDe) info.

  • Importance: Central repository for schema-on-read, enables Hive to query data without moving it. Critical for query planning and optimization.

  • Configuration: hive.metastore.uris in hive-site.xml.


4.0 Storage & Data Management

4.1 HDFS Deep Dive

4.1.1 NameNode and DataNode Architecture
  • NameNode: Single point of failure (SPOF) in v1. Stores FsImage (complete namespace snapshot) and EditLog (all modifications). In v2+ (HA), uses QJM (Quorum Journal Manager) or Shared Storage for HA.

  • DataNode: Sends heartbeats (3 sec default) and block reports to NameNode. Handles checksum verification on read/write.

4.1.2 HDFS Federation & High Availability (HA)
  • Federation: Multiple independent NameNodes (namespaces) share a pool of DataNodes. Solves scalability (metadata limits) and performance (single NN bottleneck).

  • HA (v2+): Active-Standby NameNodes with automatic failover using ZooKeeper and ZKFailoverController.

4.1.3 File Permissions & Security
  • Permissions: Unix-like (r, w, x) for owner, group, others. chmod, chown supported.

  • Security: Integrates with Kerberos for authentication. Service-Level Authorization and HDFS Transparent Encryption (at rest) available.

4.2 NoSQL Databases

4.2.1 NoSQL vs RDBMS
Feature RDBMS NoSQL (e.g., MongoDB)
Schema Fixed (Schema-on-write) Dynamic (Schema-on-read)
Scaling Vertical (Scale-up) Horizontal (Scale-out)
Joins Supported (ACID) Denormalized, no joins (BASE)
Use Case Complex transactions High volume, variety, velocity
4.2.2 MongoDB: Document Model, Indexing, Aggregation
  • Document Model: BSON (binary JSON) documents in collections. Nested, dynamic schema.

  • Indexing Strategies:

    • Single-field, compound, multikey (arrays), geospatial, text, hashed, TTL.

    • db.collection.createIndex({field: 1}) (1=ascending, -1=descending).

    • Rule: Queries without index scan entire collection (COLLSCAN).

  • Aggregation Framework: Pipeline of stages ($match, $group, $sort, $project, $lookup).

    
    db.sales.aggregate([
    
      { $match: { status: "A" } },
    
      { $$\displaystyle group: { _id: " $$item", total: { $sum: "$price" } } }
    
    ])
    
    

5.0 Data Analytics & Text Mining Concepts

5.1 Information Management in Big Data Context

  • Strategies for metadata management (what data exists, where, format), data governance (policies, quality, lineage), and master data management (single source of truth for key entities).

5.2 Text Analytics Fundamentals

5.2.1 Term Frequency (TF)
  • Measure of how frequently a term t appears in a document d.

$$ \text{TF}(t,d) = \frac{\text{frequency of } t \text{ in } d}{\text{total terms in } d} $$

  • Raw count or normalized (0-1).
5.2.2 Inverse Document Frequency (IDF)
  • Measures how common/rare a term is across the entire document corpus D.

$$ \text{IDF}(t, D) = \log \left( \frac{\text{total number of documents in } D}{\text{number of documents containing } t} \right) $$

  • Down-weights common words (e.g., "the", "is").
5.2.3 TF-IDF: Calculation & Significance
  • Formula:

$$ \text{TF-IDF}(t, d, D) = \text{TF}(t,d) \times \text{IDF}(t, D) $$

  • Significance: High TF-IDF score indicates a term is frequent in a specific document but rare across the corpus → good for keyword extraction, document similarity, and search ranking.

6.0 Applications & Case Studies

6.1 Big Data Analytics for Smart Cities Development

  • Traffic Management: Real-time analysis of GPS, camera feeds to optimize signal timing, predict congestion, suggest routes.

  • Energy Grids: Smart meter data for demand forecasting, outage detection, integration of renewables.

  • Public Safety: Predictive policing using crime reports, social media sentiment; emergency response optimization.

  • Waste Management: Sensor-equipped bins for fill-level monitoring → efficient collection routing.

  • Example: Singapore's "Virtual Singapore" 3D model integrates traffic, utility, environmental data for urban planning and emergency simulation.


7.0 Supporting & Operational Tools

7.1 Apache ZooKeeper

  • Purpose: Centralized service for configuration management, naming, synchronization, and group services in distributed systems.

  • Data Model: Hierarchical namespace of Znodes (like files/directories). Each ZNode can hold data and have children.

  • Advantages:

    • Simple API (create, get, set, watch).

    • High performance (in-memory).

    • Reliable (replicated across cluster).

  • Role in Coordination:

    • HDFS HA: Elects Active NameNode.

    • HBase: Manages region server assignment and meta-data.

    • Kafka: Tracks broker/metadata, consumer group coordination.

7.2 Information Management

  • Metadata Management: Tools like Apache Atlas or Hive Metastore catalog data assets (schema, lineage, ownership).

  • Data Governance: Policies for data quality, security, privacy (GDPR), and compliance.

  • Data Lineage: Tracking data flow from source to destination (ETL/ELT pipelines). Crucial for debugging, impact analysis, and trust.

  • Strategy: Implement a centralized metadata repository with automated lineage capture from processing engines (Spark, Hive).

[!TIP] Exam Strategy: For 7-mark application questions (e.g., Smart Cities), structure answer as: Problem → Big Data Source → Analytics Technique → Outcome/Impact. Use specific, real-world examples.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in