Skip to content
AD-801 · Big Data/Quick Revision Short Notes

Big Data (AD-801) - Unit 2 Short Notes

How unit 2 is examined

This unit covers what Hadoop is, its components (HDFS, YARN, MapReduce), the ecosystem, and Hadoop against RDBMS; HDFS carries the most marks, then Hive architecture and RDBMS versus Hadoop.

Introduction to Hadoop

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Hadoop is an open-source Apache framework that stores and processes very large data sets across clusters of commodity computers using simple programming models.</mark>

Key points.

  1. Hadoop is written in Java and was created by Doug Cutting, inspired by Google's GFS and MapReduce papers.
  2. It scales out by adding cheap commodity nodes instead of buying one costly server.
  3. It moves computation to the data, so processing happens where the blocks are stored.
  4. It tolerates failure by replicating data, so a dead node does not lose data.
  5. It handles structured, semi-structured and unstructured data with a write-once, read-many model.

Core Hadoop components

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Core Hadoop has four modules: HDFS for storage, YARN for resource management, MapReduce for processing, and Hadoop Common for shared libraries.

Key points.

  1. HDFS stores files as replicated blocks over many machines.
  2. YARN allocates cluster CPU and memory to applications and schedules them.
  3. MapReduce is the batch programming model that processes data in parallel as map and reduce tasks.
  4. Hadoop Common holds the utilities, Java libraries and file-system abstractions that the other modules use.

Hadoop Eco system

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. The Hadoop ecosystem is the set of tools built around the core modules to ingest, store, process, query and manage big data.

Key points.

  1. Hive gives an SQL-like language (HiveQL) over HDFS data, and Pig gives a data-flow scripting language (Pig Latin).
  2. HBase is a column-oriented NoSQL database running on HDFS for random real-time access.
  3. Sqoop moves data between RDBMS and Hadoop, and Flume collects streaming log data into HDFS.
  4. Zookeeper coordinates distributed services, Oozie schedules workflows, and Mahout provides machine learning.
  5. Spark is an in-memory engine that can replace MapReduce for faster processing.

Hive Physical Architecture

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. Hive is a data-warehouse layer on Hadoop that turns HiveQL queries into MapReduce jobs over data in HDFS.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-01" viewBox="0 0 492.8 252" width="492.8" height="252" role="img" aria-label="Hive query flow. UI = CLI/Web UI/HiveServer, Drv = Driver, Cmp = Compiler, MS = Metastore, Exe = Execution engine, YRN = YARN/MapReduce, HD = HDFS"><style>#dsfig-u2-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-01 .t{fill:#16181D;font-weight:500}#dsfig-u2-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-01 .dot{fill:#16181D}#dsfig-u2-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-01 .ah{fill:#454C5A}#dsfig-u2-01 .ah.hi{fill:#2340B8}#dsfig-u2-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-01 .e{stroke:#B1B7C3}html.dark #dsfig-u2-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-01 .t{fill:#E6E8ED}html.dark #dsfig-u2-01 .t.inv{fill:#0F1115}html.dark #dsfig-u2-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-01 .dot{fill:#E6E8ED}html.dark #dsfig-u2-01 .ann{fill:#8FA3FF}html.dark #dsfig-u2-01 .lbl{fill:#858D9C}html.dark #dsfig-u2-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-01 .ah{fill:#B1B7C3}html.dark #dsfig-u2-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,126 L122.2,126" marker-end="url(#ah2)"/><path class="e" d="M157.8,113.8 L230.3,53.4" marker-end="url(#ah2)"/><path class="e" d="M246.4,61 L246.4,191" marker-end="url(#ah2)" marker-start="url(#ah2)"/><path class="e" d="M261,52.2 L333.5,112.6" marker-end="url(#ah2)"/><path class="e" d="M367.1,118.7 L433.4,91.1" marker-end="url(#ah2)"/><path class="e" d="M367.1,133.3 L433.4,160.9" marker-end="url(#ah2)"/><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">UI</text><circle class="n" cx="143.2" cy="126" r="18"/><text class="t" x="143.2" y="126" dy=".35em" text-anchor="middle">Drv</text><circle class="n" cx="246.4" cy="40" r="18"/><text class="t" x="246.4" y="40" dy=".35em" text-anchor="middle">Cmp</text><circle class="n" cx="246.4" cy="212" r="18"/><text class="t" x="246.4" y="212" dy=".35em" text-anchor="middle">MS</text><circle class="n" cx="349.6" cy="126" r="18"/><text class="t" x="349.6" y="126" dy=".35em" text-anchor="middle">Exe</text><circle class="n" cx="452.8" cy="83" r="18"/><text class="t" x="452.8" y="83" dy=".35em" text-anchor="middle">YRN</text><circle class="n" cx="452.8" cy="169" r="18"/><text class="t" x="452.8" y="169" dy=".35em" text-anchor="middle">HD</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Hive query flow. UI = CLI/Web UI/HiveServer, Drv = Driver, Cmp = Compiler, MS = Metastore, Exe = Execution engine, YRN = YARN/MapReduce, HD = HDFS</figcaption></figure>

Key points.

  1. The client (CLI, Web UI, or HiveServer via JDBC/ODBC) submits the query to the Driver.
  2. The Compiler parses it, asks the Metastore (schema and table metadata in an RDBMS) for details, and builds a plan.
  3. The Execution engine runs the plan as MapReduce jobs on YARN, reading and writing HDFS data.
  4. Results return through the Driver to the client.

Asked: [7 marks] (Jun 2025) Describe the physical architecture of Hive, focusing on its components and their interactions within a Hadoop ecosystem.

Hadoop limitations

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Hadoop limitations are the areas where the framework performs poorly or needs extra effort.

Key points.

  1. Hadoop is poor at many small files, because each file's metadata uses NameNode memory.
  2. MapReduce is batch-oriented with disk writes between stages, so it is slow for real-time and iterative work.
  3. The NameNode was a single point of failure in Hadoop 1, needing High Availability in later versions.
  4. Security is weak by default and needs Kerberos, and MapReduce coding in Java is verbose.

RDBMS Versus Hadoop

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>RDBMS stores structured data in tables with a fixed schema and SQL; Hadoop is a distributed framework of HDFS, YARN and MapReduce for huge, varied data.</mark>

RDBMS is a single-server design (query engine over tables, ACID) that scales up; Hadoop is master-slave (NameNode/ResourceManager over DataNodes) that scales out.

Point RDBMS Hadoop
Data Structured only All types
Schema On write On read
Scaling Vertical Horizontal
Hardware Costly servers Commodity
Access Interactive, ACID Batch, write-once

Asked: [7 marks] (Jun 2025) Define RDBMS and Hadoop and explain their architectures.

Hadoop Distributed File system

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>HDFS is the distributed, fault-tolerant file system of Hadoop that splits large files into blocks and stores replicated copies across commodity DataNodes under a master NameNode.</mark>

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-02" viewBox="0 0 510 295" width="510" height="295" role="img" aria-label="HDFS. C = Client, NN = NameNode (metadata), D1-D3 = DataNodes; replication pipeline"><style>#dsfig-u2-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-02 .t{fill:#16181D;font-weight:500}#dsfig-u2-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-02 .dot{fill:#16181D}#dsfig-u2-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-02 .ah{fill:#454C5A}#dsfig-u2-02 .ah.hi{fill:#2340B8}#dsfig-u2-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-02 .e{stroke:#B1B7C3}html.dark #dsfig-u2-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-02 .t{fill:#E6E8ED}html.dark #dsfig-u2-02 .t.inv{fill:#0F1115}html.dark #dsfig-u2-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-02 .dot{fill:#E6E8ED}html.dark #dsfig-u2-02 .ann{fill:#8FA3FF}html.dark #dsfig-u2-02 .lbl{fill:#858D9C}html.dark #dsfig-u2-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-02 .ah{fill:#B1B7C3}html.dark #dsfig-u2-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah3" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh3" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M57,117.5 L193.2,49.4" marker-end="url(#ah3)"/><path class="e" d="M50.5,141.8 L114.4,237.5" marker-end="url(#ah3)"/><path class="e" d="M145,255 L277,255" marker-end="url(#ah3)"/><path class="e" d="M317,255 L449,255" marker-end="url(#ah3)"/><path class="e" d="M204.9,57.6 L133.1,237.4"/><path class="e" d="M219.1,57.6 L290.9,237.4"/><path class="e" d="M226.6,52.2 L455.4,242.8"/><g class="wl"><rect x="105.6" y="74" width="40.8" height="18" rx="9"/><text class="t" x="126" y="83" dy=".35em" text-anchor="middle">meta</text></g><g class="wl"><rect x="62.6" y="181.5" width="40.8" height="18" rx="9"/><text class="t" x="83" y="190.5" dy=".35em" text-anchor="middle">data</text></g><g class="wl"><rect x="191.6" y="246" width="40.8" height="18" rx="9"/><text class="t" x="212" y="255" dy=".35em" text-anchor="middle">repl</text></g><g class="wl"><rect x="363.6" y="246" width="40.8" height="18" rx="9"/><text class="t" x="384" y="255" dy=".35em" text-anchor="middle">repl</text></g><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">NN</text><circle class="n" cx="126" cy="255" r="18"/><text class="t" x="126" y="255" dy=".35em" text-anchor="middle">D1</text><circle class="n" cx="298" cy="255" r="18"/><text class="t" x="298" y="255" dy=".35em" text-anchor="middle">D2</text><circle class="n" cx="470" cy="255" r="18"/><text class="t" x="470" y="255" dy=".35em" text-anchor="middle">D3</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">HDFS. C = Client, NN = NameNode (metadata), D1-D3 = DataNodes; replication pipeline</figcaption></figure>

Key points.

  1. Design goals: store huge files, use commodity hardware, tolerate failure, stream data with high throughput, and follow write-once, read-many.
  2. The NameNode is the master; it holds the namespace and block-to-DataNode map in memory, and it never stores file data.
  3. DataNodes are slaves that store the blocks, serve reads and writes, and send heartbeats and block reports to the NameNode.
  4. A block is the storage unit, 128 MB by default; a large block cuts seek time and metadata.
  5. Each block is replicated, by default 3 times, on different nodes and racks.
  6. Write: the client asks the NameNode, gets DataNodes, and streams the block through a pipeline D1, D2, D3 with acknowledgements.
  7. Read: the client gets block locations from the NameNode and reads directly from the nearest DataNode.
  8. Fault tolerance: missed heartbeats mark a node dead, and the NameNode re-replicates its blocks; the Secondary NameNode checkpoints the edit log.

Commands. hdfs dfs -put file /dir uploads, -get downloads, -ls lists, -cat shows, -rm deletes.

Answer frame. Open with the HDFS definition and goals; draw the diagram above; develop NameNode, DataNode, block, then write, read, replication; close that HDFS gives reliable storage on cheap hardware.

Asked: [14 marks] (Jun 2025) Define HDFS. Describe namenode, datanode and block. Explain HDFS operations in detail.

Processing Data with Hadoop

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Hadoop processes data by shipping MapReduce code to the nodes holding the blocks, so work runs in parallel near the data.

Key points.

  1. The input in HDFS is split, and each split is handled by one map task.
  2. Map tasks run in parallel on many nodes, giving data locality and little network traffic.
  3. Map output is shuffled and sorted by key, then reduce tasks aggregate it.
  4. The final output is written back to HDFS.

Managing Resources and Application with Hadoop YARN

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. YARN (Yet Another Resource Negotiator) is the Hadoop 2 layer that manages cluster resources and schedules applications.

Key points.

  1. The ResourceManager is the master; it allocates resources across all applications.
  2. A NodeManager runs on each node and reports its resource use, launching containers.
  3. A Container is a bundle of CPU and memory in which tasks run.
  4. An ApplicationMaster is created per application, negotiates containers, and monitors its tasks.
  5. YARN separates resource management from processing, so engines like Spark can share the cluster.

MapReduce programming

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. MapReduce is a programming model that processes large data in parallel with a map function and a reduce function on key-value pairs.

Key points.

  1. Map reads input and emits intermediate pairs $(k, v)$, such as (word, 1).
  2. Shuffle and sort groups all values by key.
  3. Reduce combines each key's values into output, such as (word, total).
  4. A Combiner can pre-aggregate map output, and a Partitioner decides which reducer gets each key.

Last-minute revision

  • Hadoop = HDFS + YARN + MapReduce + Common.
  • Default HDFS block size is 128 MB and replication factor is 3.
  • NameNode holds metadata; DataNodes hold blocks.
  • DataNodes send heartbeats, and the Secondary NameNode only checkpoints.
  • HDFS is write-once, read-many.
  • YARN parts: ResourceManager, NodeManager, ApplicationMaster, Container.
  • MapReduce flow: Map, Shuffle and sort, Reduce.
  • Hive parts: Driver, Compiler, Metastore, Execution engine.
  • RDBMS scales vertically, Hadoop horizontally.
  • Hadoop weak points: small files, real-time processing, NameNode failure.

Memory hooks

  • NameNode = librarian catalogue; DataNode = shelves with books.
  • Block 128, copies 3.
  • MSR: Map, Shuffle, Reduce.
  • Hive = SQL on Hadoop.
  • RDBMS = up, Hadoop = out.

Coverage checklist

  • Introduction to Hadoop: definition and key points
  • Core Hadoop components: four modules
  • Hadoop Eco system: tools list
  • Hive Physical Architecture: Jun 2025 7-mark Hive architecture question
  • Hadoop limitations: four limits
  • RDBMS Versus Hadoop: Jun 2025 7-mark define and architecture question
  • Hadoop Distributed File system: Jun 2025 14-mark HDFS question
  • Processing Data with Hadoop: data locality flow
  • Managing Resources and Application with Hadoop YARN: components
  • MapReduce programming: map, shuffle, reduce
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in