How unit 2 is examined
This unit covers what Hadoop is, its components (HDFS, YARN, MapReduce), the ecosystem, and Hadoop against RDBMS; HDFS carries the most marks, then Hive architecture and RDBMS versus Hadoop.
Introduction to Hadoop
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Hadoop is an open-source Apache framework that stores and processes very large data sets across clusters of commodity computers using simple programming models.</mark>
Key points.
- Hadoop is written in Java and was created by Doug Cutting, inspired by Google's GFS and MapReduce papers.
- It scales out by adding cheap commodity nodes instead of buying one costly server.
- It moves computation to the data, so processing happens where the blocks are stored.
- It tolerates failure by replicating data, so a dead node does not lose data.
- It handles structured, semi-structured and unstructured data with a write-once, read-many model.
Core Hadoop components
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Core Hadoop has four modules: HDFS for storage, YARN for resource management, MapReduce for processing, and Hadoop Common for shared libraries.
Key points.
- HDFS stores files as replicated blocks over many machines.
- YARN allocates cluster CPU and memory to applications and schedules them.
- MapReduce is the batch programming model that processes data in parallel as map and reduce tasks.
- Hadoop Common holds the utilities, Java libraries and file-system abstractions that the other modules use.
Hadoop Eco system
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. The Hadoop ecosystem is the set of tools built around the core modules to ingest, store, process, query and manage big data.
Key points.
- Hive gives an SQL-like language (HiveQL) over HDFS data, and Pig gives a data-flow scripting language (Pig Latin).
- HBase is a column-oriented NoSQL database running on HDFS for random real-time access.
- Sqoop moves data between RDBMS and Hadoop, and Flume collects streaming log data into HDFS.
- Zookeeper coordinates distributed services, Oozie schedules workflows, and Mahout provides machine learning.
- Spark is an in-memory engine that can replace MapReduce for faster processing.
Hive Physical Architecture
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. Hive is a data-warehouse layer on Hadoop that turns HiveQL queries into MapReduce jobs over data in HDFS.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-01" viewBox="0 0 492.8 252" width="492.8" height="252" role="img" aria-label="Hive query flow. UI = CLI/Web UI/HiveServer, Drv = Driver, Cmp = Compiler, MS = Metastore, Exe = Execution engine, YRN = YARN/MapReduce, HD = HDFS"><style>#dsfig-u2-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-01 .t{fill:#16181D;font-weight:500}#dsfig-u2-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-01 .dot{fill:#16181D}#dsfig-u2-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-01 .ah{fill:#454C5A}#dsfig-u2-01 .ah.hi{fill:#2340B8}#dsfig-u2-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-01 .e{stroke:#B1B7C3}html.dark #dsfig-u2-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-01 .t{fill:#E6E8ED}html.dark #dsfig-u2-01 .t.inv{fill:#0F1115}html.dark #dsfig-u2-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-01 .dot{fill:#E6E8ED}html.dark #dsfig-u2-01 .ann{fill:#8FA3FF}html.dark #dsfig-u2-01 .lbl{fill:#858D9C}html.dark #dsfig-u2-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-01 .ah{fill:#B1B7C3}html.dark #dsfig-u2-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,126 L122.2,126" marker-end="url(#ah2)"/><path class="e" d="M157.8,113.8 L230.3,53.4" marker-end="url(#ah2)"/><path class="e" d="M246.4,61 L246.4,191" marker-end="url(#ah2)" marker-start="url(#ah2)"/><path class="e" d="M261,52.2 L333.5,112.6" marker-end="url(#ah2)"/><path class="e" d="M367.1,118.7 L433.4,91.1" marker-end="url(#ah2)"/><path class="e" d="M367.1,133.3 L433.4,160.9" marker-end="url(#ah2)"/><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">UI</text><circle class="n" cx="143.2" cy="126" r="18"/><text class="t" x="143.2" y="126" dy=".35em" text-anchor="middle">Drv</text><circle class="n" cx="246.4" cy="40" r="18"/><text class="t" x="246.4" y="40" dy=".35em" text-anchor="middle">Cmp</text><circle class="n" cx="246.4" cy="212" r="18"/><text class="t" x="246.4" y="212" dy=".35em" text-anchor="middle">MS</text><circle class="n" cx="349.6" cy="126" r="18"/><text class="t" x="349.6" y="126" dy=".35em" text-anchor="middle">Exe</text><circle class="n" cx="452.8" cy="83" r="18"/><text class="t" x="452.8" y="83" dy=".35em" text-anchor="middle">YRN</text><circle class="n" cx="452.8" cy="169" r="18"/><text class="t" x="452.8" y="169" dy=".35em" text-anchor="middle">HD</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Hive query flow. UI = CLI/Web UI/HiveServer, Drv = Driver, Cmp = Compiler, MS = Metastore, Exe = Execution engine, YRN = YARN/MapReduce, HD = HDFS</figcaption></figure>
Key points.
- The client (CLI, Web UI, or HiveServer via JDBC/ODBC) submits the query to the Driver.
- The Compiler parses it, asks the Metastore (schema and table metadata in an RDBMS) for details, and builds a plan.
- The Execution engine runs the plan as MapReduce jobs on YARN, reading and writing HDFS data.
- Results return through the Driver to the client.
Asked: [7 marks] (Jun 2025) Describe the physical architecture of Hive, focusing on its components and their interactions within a Hadoop ecosystem.
Hadoop limitations
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Hadoop limitations are the areas where the framework performs poorly or needs extra effort.
Key points.
- Hadoop is poor at many small files, because each file's metadata uses NameNode memory.
- MapReduce is batch-oriented with disk writes between stages, so it is slow for real-time and iterative work.
- The NameNode was a single point of failure in Hadoop 1, needing High Availability in later versions.
- Security is weak by default and needs Kerberos, and MapReduce coding in Java is verbose.
RDBMS Versus Hadoop
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>RDBMS stores structured data in tables with a fixed schema and SQL; Hadoop is a distributed framework of HDFS, YARN and MapReduce for huge, varied data.</mark>
RDBMS is a single-server design (query engine over tables, ACID) that scales up; Hadoop is master-slave (NameNode/ResourceManager over DataNodes) that scales out.
| Point | RDBMS | Hadoop |
|---|---|---|
| Data | Structured only | All types |
| Schema | On write | On read |
| Scaling | Vertical | Horizontal |
| Hardware | Costly servers | Commodity |
| Access | Interactive, ACID | Batch, write-once |
Asked: [7 marks] (Jun 2025) Define RDBMS and Hadoop and explain their architectures.
Hadoop Distributed File system
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>HDFS is the distributed, fault-tolerant file system of Hadoop that splits large files into blocks and stores replicated copies across commodity DataNodes under a master NameNode.</mark>
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-02" viewBox="0 0 510 295" width="510" height="295" role="img" aria-label="HDFS. C = Client, NN = NameNode (metadata), D1-D3 = DataNodes; replication pipeline"><style>#dsfig-u2-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-02 .t{fill:#16181D;font-weight:500}#dsfig-u2-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-02 .dot{fill:#16181D}#dsfig-u2-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-02 .ah{fill:#454C5A}#dsfig-u2-02 .ah.hi{fill:#2340B8}#dsfig-u2-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-02 .e{stroke:#B1B7C3}html.dark #dsfig-u2-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-02 .t{fill:#E6E8ED}html.dark #dsfig-u2-02 .t.inv{fill:#0F1115}html.dark #dsfig-u2-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-02 .dot{fill:#E6E8ED}html.dark #dsfig-u2-02 .ann{fill:#8FA3FF}html.dark #dsfig-u2-02 .lbl{fill:#858D9C}html.dark #dsfig-u2-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-02 .ah{fill:#B1B7C3}html.dark #dsfig-u2-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah3" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh3" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M57,117.5 L193.2,49.4" marker-end="url(#ah3)"/><path class="e" d="M50.5,141.8 L114.4,237.5" marker-end="url(#ah3)"/><path class="e" d="M145,255 L277,255" marker-end="url(#ah3)"/><path class="e" d="M317,255 L449,255" marker-end="url(#ah3)"/><path class="e" d="M204.9,57.6 L133.1,237.4"/><path class="e" d="M219.1,57.6 L290.9,237.4"/><path class="e" d="M226.6,52.2 L455.4,242.8"/><g class="wl"><rect x="105.6" y="74" width="40.8" height="18" rx="9"/><text class="t" x="126" y="83" dy=".35em" text-anchor="middle">meta</text></g><g class="wl"><rect x="62.6" y="181.5" width="40.8" height="18" rx="9"/><text class="t" x="83" y="190.5" dy=".35em" text-anchor="middle">data</text></g><g class="wl"><rect x="191.6" y="246" width="40.8" height="18" rx="9"/><text class="t" x="212" y="255" dy=".35em" text-anchor="middle">repl</text></g><g class="wl"><rect x="363.6" y="246" width="40.8" height="18" rx="9"/><text class="t" x="384" y="255" dy=".35em" text-anchor="middle">repl</text></g><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">NN</text><circle class="n" cx="126" cy="255" r="18"/><text class="t" x="126" y="255" dy=".35em" text-anchor="middle">D1</text><circle class="n" cx="298" cy="255" r="18"/><text class="t" x="298" y="255" dy=".35em" text-anchor="middle">D2</text><circle class="n" cx="470" cy="255" r="18"/><text class="t" x="470" y="255" dy=".35em" text-anchor="middle">D3</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">HDFS. C = Client, NN = NameNode (metadata), D1-D3 = DataNodes; replication pipeline</figcaption></figure>
Key points.
- Design goals: store huge files, use commodity hardware, tolerate failure, stream data with high throughput, and follow write-once, read-many.
- The NameNode is the master; it holds the namespace and block-to-DataNode map in memory, and it never stores file data.
- DataNodes are slaves that store the blocks, serve reads and writes, and send heartbeats and block reports to the NameNode.
- A block is the storage unit, 128 MB by default; a large block cuts seek time and metadata.
- Each block is replicated, by default 3 times, on different nodes and racks.
- Write: the client asks the NameNode, gets DataNodes, and streams the block through a pipeline D1, D2, D3 with acknowledgements.
- Read: the client gets block locations from the NameNode and reads directly from the nearest DataNode.
- Fault tolerance: missed heartbeats mark a node dead, and the NameNode re-replicates its blocks; the Secondary NameNode checkpoints the edit log.
Commands. hdfs dfs -put file /dir uploads, -get downloads, -ls lists, -cat shows, -rm deletes.
Answer frame. Open with the HDFS definition and goals; draw the diagram above; develop NameNode, DataNode, block, then write, read, replication; close that HDFS gives reliable storage on cheap hardware.
Asked: [14 marks] (Jun 2025) Define HDFS. Describe namenode, datanode and block. Explain HDFS operations in detail.
Processing Data with Hadoop
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Hadoop processes data by shipping MapReduce code to the nodes holding the blocks, so work runs in parallel near the data.
Key points.
- The input in HDFS is split, and each split is handled by one map task.
- Map tasks run in parallel on many nodes, giving data locality and little network traffic.
- Map output is shuffled and sorted by key, then reduce tasks aggregate it.
- The final output is written back to HDFS.
Managing Resources and Application with Hadoop YARN
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. YARN (Yet Another Resource Negotiator) is the Hadoop 2 layer that manages cluster resources and schedules applications.
Key points.
- The ResourceManager is the master; it allocates resources across all applications.
- A NodeManager runs on each node and reports its resource use, launching containers.
- A Container is a bundle of CPU and memory in which tasks run.
- An ApplicationMaster is created per application, negotiates containers, and monitors its tasks.
- YARN separates resource management from processing, so engines like Spark can share the cluster.
MapReduce programming
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. MapReduce is a programming model that processes large data in parallel with a map function and a reduce function on key-value pairs.
Key points.
- Map reads input and emits intermediate pairs $(k, v)$, such as (word, 1).
- Shuffle and sort groups all values by key.
- Reduce combines each key's values into output, such as (word, total).
- A Combiner can pre-aggregate map output, and a Partitioner decides which reducer gets each key.
Last-minute revision
- Hadoop = HDFS + YARN + MapReduce + Common.
- Default HDFS block size is 128 MB and replication factor is 3.
- NameNode holds metadata; DataNodes hold blocks.
- DataNodes send heartbeats, and the Secondary NameNode only checkpoints.
- HDFS is write-once, read-many.
- YARN parts: ResourceManager, NodeManager, ApplicationMaster, Container.
- MapReduce flow: Map, Shuffle and sort, Reduce.
- Hive parts: Driver, Compiler, Metastore, Execution engine.
- RDBMS scales vertically, Hadoop horizontally.
- Hadoop weak points: small files, real-time processing, NameNode failure.
Memory hooks
- NameNode = librarian catalogue; DataNode = shelves with books.
- Block 128, copies 3.
- MSR: Map, Shuffle, Reduce.
- Hive = SQL on Hadoop.
- RDBMS = up, Hadoop = out.
Coverage checklist
- Introduction to Hadoop: definition and key points
- Core Hadoop components: four modules
- Hadoop Eco system: tools list
- Hive Physical Architecture: Jun 2025 7-mark Hive architecture question
- Hadoop limitations: four limits
- RDBMS Versus Hadoop: Jun 2025 7-mark define and architecture question
- Hadoop Distributed File system: Jun 2025 14-mark HDFS question
- Processing Data with Hadoop: data locality flow
- Managing Resources and Application with Hadoop YARN: components
- MapReduce programming: map, shuffle, reduce