Skip to content
CE-803 (B) · Data Analytics/Quick Revision Short Notes

Data Analytics (CE-803 (B)) - Unit 4 Short Notes

UNIT 4: DATA ANALYTICS - SHORT NOTES


I. FUNDAMENTALS OF BIG DATA

Definition and Need for Big Data

Big Data refers to extremely large, complex, and rapidly growing datasets that traditional data processing tools cannot effectively capture, store, manage, and analyze. The need arises from the explosion of digital data from sources like social media, sensors, transactions, and logs, which holds potential for uncovering hidden patterns, correlations, and insights for strategic decision-making.

Characteristics of Big Data (The 4 V's)

V Definition Explanation & Example
Volume The sheer scale of data. Measured in terabytes, petabytes, exabytes. Example: Facebook generates petabytes of user activity data daily.
Velocity The speed at which data is generated, collected, and processed. Real-time or near-real-time streams. Example: Stock market tick data, IoT sensor readings.
Variety The different types and formats of data. Structured (databases), Semi-structured (XML, JSON), Unstructured (text, images, video).
Veracity The quality, accuracy, and trustworthiness of data. Concerns noise, ambiguity, inconsistency, and uncertainty in data sources.

[!TIP] Exam Focus: The 4 V's are a very high-probability question. Be ready to define each with a concrete example.

Applications of Big Data Analytics: Smart Cities

Big Data analytics enables Smart Cities by integrating data from various urban systems (traffic, energy, water, waste, public safety) to improve efficiency, sustainability, and quality of life.

  • Example: Intelligent Traffic Management

    • Data Sources: GPS from vehicles, traffic cameras, road sensors, social media reports.

    • Analysis: Real-time analysis to predict congestion, optimize traffic light timings, and suggest alternative routes.

    • Outcome: Reduced commute times, lower fuel consumption, decreased emissions.


II. HADOOP ECOSYSTEM ARCHITECTURE

Hadoop Architecture Overview

Hadoop is a framework for distributed storage and processing of very large datasets on clusters of commodity hardware. Its core architecture follows a master-slave model.

Key Components & Roles:

  1. HDFS (Hadoop Distributed File System): Distributed storage system.

  2. MapReduce / YARN: Distributed processing engine (MapReduce is the programming model; YARN is the resource manager).

  3. Common: Libraries and utilities shared by other modules.

DiagramCANVAS: A block diagram showing two main layers: Storage Layer (HDFS with NameNode as master and DataNodes as slaves) and Processing Layer (YARN with ResourceManager as master and NodeManagers as slaves). An ApplicationMaster (for a specific job like MapReduce) interacts with the ResourceManager and NodeManagers.

Hadoop Distributed File System (HDFS)

  • Goals & Design Principles:

    • Store massive files (GBs to TBs) across a cluster.

    • Provide fault tolerance through data replication.

    • Enable high throughput access (suitable for batch processing, not low-latency).

    • Follow "Write Once, Read Many" (WORM) model.

  • Architecture:

    • NameNode (Master): Manages the file system namespace (metadata: file names, permissions, block locations). It does not store actual data.

    • DataNode (Slave): Stores data blocks on local disk. Handles read/write requests from clients. Sends heartbeats and block reports to NameNode.

    • Secondary NameNode: Not a backup for NameNode. It periodically checkpoints the NameNode's edit logs and fsimage to prevent file system corruption and reduce NameNode restart time.

  • Features:

    • Fault Tolerance: Default replication factor is 3. If a DataNode fails, data is automatically re-replicated from other replicas.

    • Scalability: Can scale horizontally by adding more DataNodes.

MapReduce Programming Model

A programming paradigm for processing vast amounts of data in parallel on a Hadoop cluster.

  • Core Concepts & Workflow:

    1. Map Phase: Input is split into independent chunks. The map() function processes each key-value pair and emits intermediate key-value pairs.

    2. Shuffle & Sort Phase: The framework automatically groups all intermediate values associated with the same intermediate key and sorts them. This is the "heart" of the process.

    3. Reduce Phase: The reduce() function processes each grouped intermediate key and its list of values to produce the final output key-value pairs.

[!TIP] Exam Tip: Understand the data flow: Input -> Map -> (Shuffle/Sort) -> Reduce -> Output. The Shuffle phase is critical and framework-managed.

Yet Another Resource Negotiator (YARN)

Introduced in Hadoop 2.x to decouple resource management from the MapReduce programming model, enabling multiple processing engines (Spark, Tez) on the same cluster.

  • Architecture:

    • ResourceManager (RM): Global master that arbitrates resources across all applications.

    • NodeManager (NM): Per-node slave that launches and monitors containers (resource allocations for applications).

    • ApplicationMaster (AM): Per-application master (one for each MapReduce job, Spark job). Negotiates resources with RM and works with NM to execute tasks.

  • Capacity Scheduler:

    • A YARN scheduler designed for multi-tenant clusters (e.g., shared by multiple organizations).

    • Features & Working:

      • Uses queues to allocate resources. Each queue gets a guaranteed capacity (percentage of cluster resources).

      • Supports elastic scheduling: unused capacity from one queue can be temporarily borrowed by others (with limits).

      • Provides capacity guarantees, security (ACLs per queue), and priority-based scheduling within a queue.

      • Ensures fairness and predictability in a shared environment.


III. DATA PROCESSING TOOLS: HIVE & PIG

Apache Hive

  • Architecture & Key Features:

    • Provides an SQL-like interface (HiveQL) to query and manage data stored in HDFS.

    • Schema-on-Read: Data is stored as-is; schema is applied at query time.

    • Not for OLTP (low-latency queries). Optimized for OLAP and batch processing.

    • Hive Metastore: Central repository for metadata (table names, column types, partition info, data location). Stores this info in a relational database ( Derby, MySQL). Critical for Hive to function.

  • Hive Query Language (HiveQL) - DDL Commands:

    • CREATE TABLE: Defines a new managed or external table.

      
      CREATE TABLE employees (id INT, name STRING, dept STRING)
      
      ROW FORMAT DELIMITED
      
      FIELDS TERMINATED BY ','
      
      STORED AS TEXTFILE;
      
      
    • ALTER TABLE: Modifies table structure or properties.

      
      ALTER TABLE employees ADD COLUMNS (salary FLOAT);
      
      
    • DROP TABLE: Deletes table metadata and data (for managed tables) or just metadata (for external tables).

      
      DROP TABLE employees;
      
      

Apache Pig

  • Pig Latin: A high-level data flow scripting language for analyzing large datasets.

  • Need & Advantages over Raw MapReduce:

    • Abstraction: Developers write Pig Latin scripts (series of operations like LOAD, FILTER, GROUP, FOREACH, STORE) instead of complex Java MapReduce code.

    • Less Code: A Pig script requiring ~200 lines of MapReduce can be written in ~10 lines of Pig Latin.

    • Automatic Optimization: Pig's execution planner optimizes the logical and physical plan (e.g., combining operations, choosing join algorithms).

    • Flexibility: Handles both structured and semi-structured data easily.


IV. COORDINATION & MANAGEMENT SERVICES

Apache ZooKeeper

  • Definition & Role: A centralized service for maintaining configuration information, naming, providing distributed synchronization, and providing group services. It's the "coordination kernel" for distributed applications (like HBase, Hive, Kafka).

  • Core Concept: A hierarchical namespace of data nodes called znodes, similar to a file system. Data stored in znodes is kept in memory for high speed.

  • Advantages:

    1. Reliable Coordination: Provides strong consistency guarantees. Watches can be set on znodes to notify clients of changes.

    2. Configuration Management: Stores configuration data (e.g., cluster membership, master election info) in a central, accessible location. Changes propagate automatically to all clients.


V. NOSQL DATABASES: MONGODB

MongoDB Indexing

  • Purpose: Improve query performance by allowing the database to quickly locate documents without scanning the entire collection (collection scan). Similar to indexes in relational databases.

  • Types of Indexes:

    • Single Field: Index on a single field.

    • Compound Index: Index on multiple fields (order matters). Supports queries on the prefix of the index.

    • Multikey Index: For fields that hold arrays; indexes each element of the array.

    • Text Index: For text search on string content.

    • Geospatial Index: For location-based queries (2dsphere, 2d).

    • Hashed Index: Indexes the hash of a field's value, used for sharding.

  • Impact: Indexes speed up reads (find, sort) but slow down writes (insert, update) and consume additional disk space and memory.

MongoDB Aggregation Framework

  • Purpose: A powerful tool for data analysis and transformation. It processes data records (documents) and returns computed results. Used for operations like filtering, grouping, sorting, and computing aggregates (sum, avg, max).

  • Aggregation Pipeline: Data passes through a sequence of stages, each transforming the documents as they pass through.

    • Common Stages:

      • $match: Filters documents (like WHERE clause).

      • $$\displaystyle group`: Groups documents by a key and applies accumulator expressions (` $$sum, $avg, $push).

      • $sort: Orders documents.

      • $project: Reshapes documents (includes/excludes fields, adds computed fields).

      • $skip / $limit: For pagination.

      • $unwind: Deconstructs an array field to output a document for each element.


VI. TEXT MINING & ANALYTICAL CONCEPTS

Text Mining Fundamentals: TF, IDF, TF-IDF

  • Term Frequency (TF): Measures how frequently a term appears in a document.

$$ \text{TF}(t,d) = \frac{\text{Number of times term } t \text{ appears in document } d}{\text{Total number of terms in document } d} $$

*   **Intuition:** A term that appears often in a doc is important to that doc.
  • Inverse Document Frequency (IDF): Measures how common a term is across all documents. Downweights terms that appear in many docs (like "the", "is").

$$ \text{IDF}(t,D) = \log \left( \frac{\text{Total number of documents in corpus } D}{\text{Number of documents containing term } t} \right) $$

*   **Intuition:** A term is more significant if it is rare across the corpus.
  • TF-IDF: The product of TF and IDF. It assigns a weight to a term in a document, reflecting its importance within that document relative to the entire corpus.

$$ \text{TF-IDF}(t,d,D) = \text{TF}(t,d) \times \text{IDF}(t,D) \boxed{} $$

*   **Significance:** A high TF-IDF score means the term is **frequent in the document but rare in the corpus**—a good candidate for a keyword or topic.

[!TIP] Common Pitfall: Remember IDF uses a logarithm. This dampens the effect of very common terms. Also, TF is often normalized (e.g., by document length) to avoid bias towards long documents.

Analytical Methods Comparison: Regression vs. ANOVA

Aspect Regression ANOVA (Analysis of Variance)
Primary Purpose To model and predict a continuous dependent variable based on one or more independent variables. To test for significant differences in the means of a continuous dependent variable across two or more groups (categorical independent variable).
Independent Variable(s) Can be continuous or categorical. Often used for continuous predictors. Categorical (e.g., treatment groups).
Dependent Variable Continuous (e.g., sales, temperature, price). Continuous (e.g., test scores, yield, measurement).
Key Output Regression equation (e.g., Y = β₀ + β₁X + ε), coefficients (β), R-squared (goodness of fit), p-values for coefficients. F-statistic and p-value to reject/accept the null hypothesis that all group means are equal. May be followed by post-hoc tests (e.g., Tukey) to see which specific groups differ.
Typical Use Case "How does advertising spend (continuous) affect sales (continuous)?" "Predict house price based on size, location, age." "Is there a difference in crop yield (continuous) between three different fertilizer types (categorical)?" "Do test scores differ across four teaching methods?"

VII. DATA MANAGEMENT CONCEPTS

Information Management in Big Data Context

Managing large-scale data involves overcoming challenges inherent to the 3Vs (Volume, Velocity, Variety).

  • Challenges:

    • Storage & Scalability: Storing petabytes of diverse data cost-effectively.

    • Processing Speed: Analyzing data in near-real-time (Velocity).

    • Data Integration: Combining structured, semi-structured, and unstructured data.

    • Data Quality & Governance: Ensuring accuracy, consistency, and compliance (Veracity).

    • Security & Privacy: Protecting sensitive data across distributed systems.

  • Strategies:

    • Adopt Distributed Frameworks: Use Hadoop (HDFS) for scalable, fault-tolerant storage and Spark for fast in-memory processing.

    • Implement NoSQL Databases: Use MongoDB, Cassandra, HBase for flexible schema and horizontal scaling (Variety).

    • Employ Data Lakes: Store raw data in its native format (e.g., on HDFS or cloud storage like S3) until needed, avoiding premature schema definition.

    • Utilize Metadata Management: Use tools like Hive Metastore or Apache Atlas for data cataloging, lineage, and governance.

    • Leverage Stream Processing: Use Apache Kafka, Flink, Storm for ingesting and processing high-velocity data streams.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in