Skip to content
CS-702 (D) · Big Data/Quick Revision Short Notes

Big Data (CS-702 (D)) - Unit 2 Short Notes

How unit 2 is examined

This unit covers what Hadoop is, its core components and ecosystem, HDFS, YARN and MapReduce; MapReduce, HDFS, the ecosystem and RDBMS versus Hadoop carry the marks.

Introduction to Hadoop

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Hadoop is an open-source Apache framework that stores and processes very large datasets across clusters of commodity computers using simple programming models.</mark>

Key points.

  1. Hadoop is called a Big Data technology because it handles the 3 Vs: volume through distributed storage, velocity through parallel batch processing, and variety through schema-free storage of structured and unstructured data.
  2. HDFS splits files into blocks and stores them on many machines, so storage grows by simply adding nodes (horizontal scaling).
  3. Blocks are replicated (default 3 copies), so a node failure loses no data; this is its fault tolerance.
  4. MapReduce moves the computation to the data and runs it in parallel, and YARN allocates cluster resources to the jobs.
  5. It runs on cheap commodity hardware, so cost per terabyte is low.

Asked: [7 marks] (Nov 2022) Why Hadoop is called a Big data technology? Explain how it supports Big data?

Core Hadoop components

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. Hadoop has four core modules: HDFS for storage, YARN for resource management, MapReduce for processing and Hadoop Common for shared libraries; tools like Pig, Hive and HBase sit on top.

Key points.

  1. HDFS is the storage layer that keeps replicated blocks across DataNodes, managed by a NameNode.
  2. YARN is the resource manager that allocates CPU and memory to applications and schedules jobs.
  3. MapReduce is the processing engine that runs a job as parallel map and reduce tasks.
  4. Hadoop Common is the set of Java libraries and utilities, including the configuration files, that all other modules use.
  5. Pig (scripting), Hive (SQL-like queries) and HBase (NoSQL store) are higher-level tools built over these layers.
  6. Together they give distributed storage and processing of Big Data, for example log analysis, recommendations and web indexing.

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-01" viewBox="0 0 262.5 381" width="262.5" height="381" role="img" aria-label="Hadoop layers: Pig, Hive, HBase on top of MapReduce, YARN and HDFS (Common supports all)"><style>#dsfig-u2-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-01 .t{fill:#16181D;font-weight:500}#dsfig-u2-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-01 .dot{fill:#16181D}#dsfig-u2-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-01 .ah{fill:#454C5A}#dsfig-u2-01 .ah.hi{fill:#2340B8}#dsfig-u2-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-01 .e{stroke:#B1B7C3}html.dark #dsfig-u2-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-01 .t{fill:#E6E8ED}html.dark #dsfig-u2-01 .t.inv{fill:#0F1115}html.dark #dsfig-u2-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-01 .dot{fill:#E6E8ED}html.dark #dsfig-u2-01 .ann{fill:#8FA3FF}html.dark #dsfig-u2-01 .lbl{fill:#858D9C}html.dark #dsfig-u2-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-01 .ah{fill:#B1B7C3}html.dark #dsfig-u2-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M187.3,56.1 L57.6,140.4" marker-end="url(#ah2)"/><path class="e" d="M57.3,159.6 L186.5,217.7" marker-end="url(#ah2)"/><path class="e" d="M212,255.2 L212,313" marker-end="url(#ah2)"/><rect class="n" x="183.5" y="25" width="57" height="30" rx="15"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">Tools</text><circle class="n" cx="40" cy="151.8" r="18"/><text class="t" x="40" y="151.8" dy=".35em" text-anchor="middle">MR</text><rect class="n" x="187" y="214.2" width="50" height="30" rx="15"/><text class="t" x="212" y="229.2" dy=".35em" text-anchor="middle">Yarn</text><rect class="n" x="187" y="326" width="50" height="30" rx="15"/><text class="t" x="212" y="341" dy=".35em" text-anchor="middle">HDFS</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Hadoop layers: Pig, Hive, HBase on top of MapReduce, YARN and HDFS (Common supports all)</figcaption></figure>

Configuration files.

File Purpose
core-site.xml Common settings such as default file system URI (fs.defaultFS)
hdfs-site.xml HDFS settings such as replication factor and block size
mapred-site.xml MapReduce settings such as the framework name (yarn)
yarn-site.xml ResourceManager address and NodeManager settings
hadoop-env.sh Environment variables such as JAVA_HOME

Site-specific values in these files override the built-in defaults (core-default.xml and similar).

Answer frame. Open with the four modules; draw the layered sketch; explain HDFS, YARN, MapReduce, Common, then the tools; close with the purpose (storage plus processing of Big Data). For the configuration question, give the table and the default-versus-site precedence.

Asked: [7 marks] (Dec 2024) List the components of Hadoop, explain its use. Explain in detail building blocks of Hadoop with neat sketch. Asked: [7 marks] (Jun 2025) What are the different types of Hadoop configuration files? Discuss.

Hadoop Eco system

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>The Hadoop ecosystem is the collection of tools built around the core of HDFS, YARN and MapReduce that together ingest, store, process, query and manage Big Data.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-02" viewBox="0 0 510 252" width="510" height="252" role="img" aria-label="Ecosystem: Sqoop (Sq), Flume (Fl), HDFS, YARN, MapReduce (MR), Pig, Hive, HBase"><style>#dsfig-u2-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-02 .t{fill:#16181D;font-weight:500}#dsfig-u2-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-02 .dot{fill:#16181D}#dsfig-u2-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-02 .ah{fill:#454C5A}#dsfig-u2-02 .ah.hi{fill:#2340B8}#dsfig-u2-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-02 .e{stroke:#B1B7C3}html.dark #dsfig-u2-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-02 .t{fill:#E6E8ED}html.dark #dsfig-u2-02 .t.inv{fill:#0F1115}html.dark #dsfig-u2-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-02 .dot{fill:#E6E8ED}html.dark #dsfig-u2-02 .ann{fill:#8FA3FF}html.dark #dsfig-u2-02 .lbl{fill:#858D9C}html.dark #dsfig-u2-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-02 .ah{fill:#B1B7C3}html.dark #dsfig-u2-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah3" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh3" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M53.4,53.4 L197.2,197.2" marker-end="url(#ah3)"/><path class="e" d="M57,134.5 L193.2,202.6" marker-end="url(#ah3)"/><path class="e" d="M229,91.5 L451.2,202.6" marker-end="url(#ah3)"/><path class="e" d="M354.4,96.4 L455.2,197.2" marker-end="url(#ah3)"/><path class="e" d="M453,91.5 L230.8,202.6" marker-end="url(#ah3)"/><path class="e" d="M451,212 L362,212" marker-end="url(#ah3)"/><path class="e" d="M322,212 L233,212" marker-end="url(#ah3)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Sq</text><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">Fl</text><circle class="n" cx="212" cy="212" r="18"/><text class="t" x="212" y="212" dy=".35em" text-anchor="middle">HD</text><circle class="n" cx="341" cy="212" r="18"/><text class="t" x="341" y="212" dy=".35em" text-anchor="middle">YN</text><circle class="n" cx="470" cy="212" r="18"/><text class="t" x="470" y="212" dy=".35em" text-anchor="middle">MR</text><circle class="n" cx="212" cy="83" r="18"/><text class="t" x="212" y="83" dy=".35em" text-anchor="middle">Pg</text><circle class="n" cx="341" cy="83" r="18"/><text class="t" x="341" y="83" dy=".35em" text-anchor="middle">Hv</text><circle class="n" cx="470" cy="83" r="18"/><text class="t" x="470" y="83" dy=".35em" text-anchor="middle">HB</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Ecosystem: Sqoop (Sq), Flume (Fl), HDFS, YARN, MapReduce (MR), Pig, Hive, HBase</figcaption></figure>

Key points.

  1. HDFS is the distributed storage layer and MapReduce is the batch processing engine, with YARN managing resources.
  2. Hive is a data warehouse layer that gives SQL-like queries (HiveQL) that are converted into MapReduce jobs.
  3. Pig is a scripting platform whose Pig Latin scripts express data flows and also compile to MapReduce.
  4. HBase is a column-oriented NoSQL database on HDFS for real-time random read and write.
  5. Sqoop transfers data between relational databases and HDFS, while Flume collects streaming log data into HDFS.
  6. ZooKeeper coordinates distributed services, Oozie schedules workflows, and Mahout provides machine learning.
  7. HIVE was created at Facebook; it stores table metadata in a metastore, lets analysts query huge datasets without writing Java, and suits batch data warehousing, not low-latency transactions.

Answer frame. Open with the definition; draw the ecosystem figure; describe components in order of storage, processing, querying, ingestion, coordination; for HIVE add HiveQL, metastore and warehousing use; close with the point that all tools share HDFS.

Asked: [14 marks] (Dec 2020, Jun 2025) Write short notes on these: i) Hadoop Eco system ii) HIVE. What is Hadoop Ecosystem? Discuss various components of Hadoop Ecosystem.

Hive Physical Architecture

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Hive is a data warehouse tool on Hadoop whose physical architecture is made of a client, driver, compiler, metastore and execution engine over HDFS.

Key points.

  1. The client (CLI, JDBC or web UI) submits HiveQL to the driver.
  2. The compiler parses the query, uses the metastore (a relational database holding table schemas) and builds a plan.
  3. The execution engine runs the plan as MapReduce jobs on Hadoop and returns results.
  4. Table data lives in HDFS, while only the schema lives in the metastore.

Hadoop limitations

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Hadoop is built for large batch jobs, so it is a poor fit for small, interactive or transactional work.

Key points.

  1. It handles many small files badly because the NameNode keeps every file's metadata in memory.
  2. It has high latency, since MapReduce writes to disk between stages, so it is not real-time.
  3. It is unsuitable for iterative or interactive processing and for ACID transactions.
  4. It has a single NameNode as a bottleneck, and security and skilled administration are hard.

RDBMS Versus Hadoop

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Hadoop is a distributed framework for storing and processing huge structured and unstructured data on commodity clusters, whereas an RDBMS stores structured data in tables with a fixed schema on a single server.</mark>

Key points.

  1. Hadoop has HDFS for storage and MapReduce with YARN for processing, so it scales across thousands of nodes.
  2. An RDBMS relies on SQL, tables and ACID transactions, which suit consistent operational data.
Basis RDBMS Hadoop
Data type Structured only Structured, semi-structured, unstructured
Schema Schema on write (fixed before loading) Schema on read (applied when queried)
Scalability Vertical (bigger server) Horizontal (add nodes)
Cost Licensed, costly hardware Open source, commodity hardware
Processing SQL, fast for small data, real-time MapReduce batch, high throughput, high latency
Data size Gigabytes to terabytes Terabytes to petabytes
Access Read and write many times, ACID Write once, read many, no full ACID
  1. Use Hadoop for very large, varied data and batch analytics such as logs and clickstreams; use an RDBMS for transactions, small consistent data and low-latency queries such as banking.

Answer frame. Open with the definition of Hadoop and its HDFS and MapReduce components; draw the seven-row table; close with when to use each (Hadoop for big batch data, RDBMS for transactions).

Asked: [7 marks] (Dec 2020) What is Hadoop? Compare it with RDBMS. Asked: [14 marks] (Nov 2022) Write short notes on any two: i) RDBMS Vs Hadoop ii) Hadoop Ecosystem iii) ETL Processing iv) Variation of NOSAQL

Hadoop Distributed File system

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>HDFS is the distributed, fault-tolerant file system of Hadoop that splits files into large blocks and stores replicated copies across the DataNodes of a cluster, managed by a NameNode.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-03" viewBox="0 0 553 209" width="553" height="209" role="img" aria-label="HDFS: Client (Cl), NameNode (NN) with metadata, Secondary NameNode (SNN), DataNodes D1-D3 holding replicated blocks"><style>#dsfig-u2-03 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-03 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-03 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-03 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-03 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-03 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-03 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-03 .t{fill:#16181D;font-weight:500}#dsfig-u2-03 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-03 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-03 .dot{fill:#16181D}#dsfig-u2-03 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-03 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-03 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-03 .ah{fill:#454C5A}#dsfig-u2-03 .ah.hi{fill:#2340B8}#dsfig-u2-03 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-03 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-03 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-03 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-03 .e{stroke:#B1B7C3}html.dark #dsfig-u2-03 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-03 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-03 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-03 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-03 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-03 .t{fill:#E6E8ED}html.dark #dsfig-u2-03 .t.inv{fill:#0F1115}html.dark #dsfig-u2-03 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-03 .dot{fill:#E6E8ED}html.dark #dsfig-u2-03 .ann{fill:#8FA3FF}html.dark #dsfig-u2-03 .lbl{fill:#858D9C}html.dark #dsfig-u2-03 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-03 .ah{fill:#B1B7C3}html.dark #dsfig-u2-03 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-03 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-03 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-03 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M56.3,159.2 L237,50.8" marker-end="url(#ah4)"/><path class="e" d="M61,169 L234,169" marker-end="url(#ah4)" marker-start="url(#ah4)"/><path class="e" d="M255,59 L255,148" marker-end="url(#ah4)"/><path class="e" d="M268.4,53.4 L369.2,154.2" marker-end="url(#ah4)"/><path class="e" d="M272,48.5 L494.2,159.6" marker-end="url(#ah4)"/><path class="e" d="M276,169 L363,169" marker-end="url(#ah4)" marker-start="url(#ah4)"/><path class="e" d="M405,169 L492,169" marker-end="url(#ah4)" marker-start="url(#ah4)"/><path class="e" d="M451,40 L276,40" marker-end="url(#ah4)"/><g class="wl"><rect x="113.2" y="95.5" width="68.7" height="18" rx="9"/><text class="t" x="147.5" y="104.5" dy=".35em" text-anchor="middle">metadata</text></g><g class="wl"><rect x="127.1" y="160" width="40.8" height="18" rx="9"/><text class="t" x="147.5" y="169" dy=".35em" text-anchor="middle">data</text></g><g class="wl"><rect x="281.6" y="160" width="75.9" height="18" rx="9"/><text class="t" x="319.5" y="169" dy=".35em" text-anchor="middle">replicate</text></g><g class="wl"><rect x="321.4" y="31" width="82.2" height="18" rx="9"/><text class="t" x="362.5" y="40" dy=".35em" text-anchor="middle">checkpoint</text></g><circle class="n" cx="40" cy="169" r="18"/><text class="t" x="40" y="169" dy=".35em" text-anchor="middle">Cl</text><circle class="n" cx="255" cy="40" r="18"/><text class="t" x="255" y="40" dy=".35em" text-anchor="middle">NN</text><circle class="n" cx="255" cy="169" r="18"/><text class="t" x="255" y="169" dy=".35em" text-anchor="middle">D1</text><circle class="n" cx="384" cy="169" r="18"/><text class="t" x="384" y="169" dy=".35em" text-anchor="middle">D2</text><circle class="n" cx="513" cy="169" r="18"/><text class="t" x="513" y="169" dy=".35em" text-anchor="middle">D3</text><circle class="n" cx="470" cy="40" r="18"/><text class="t" x="470" y="40" dy=".35em" text-anchor="middle">SNN</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">HDFS: Client (Cl), NameNode (NN) with metadata, Secondary NameNode (SNN), DataNodes D1-D3 holding replicated blocks</figcaption></figure>

Key points.

  1. The NameNode is the master; it stores metadata (file names, block locations, permissions) in memory and never stores file data.
  2. DataNodes are slaves that store the actual blocks, serve read and write requests and send heartbeats to the NameNode.
  3. A file is split into blocks of 128 MB by default, and each block is replicated 3 times on different nodes and racks.
  4. The Secondary NameNode periodically merges the edit log with the fsimage (checkpointing); it is not a hot standby.
  5. Read: the client asks the NameNode for block locations and then reads the blocks directly from the nearest DataNodes.
  6. Write: the NameNode picks DataNodes, the client streams a block to the first DataNode, and the block is forwarded in a pipeline to the replicas, which acknowledge back.
  7. If a DataNode stops sending heartbeats, its blocks are re-replicated from other copies, which gives fault tolerance.
  8. Features are scalability, high throughput, write-once read-many access and low-cost hardware.

Commands.

hdfs dfs -mkdir /user/data            # create a directory
hdfs dfs -put file.txt /user/data     # copy local file into HDFS
hdfs dfs -ls /user/data               # list contents
hdfs dfs -cat /user/data/file.txt     # show file
hdfs dfs -get /user/data/file.txt .   # copy from HDFS to local

Answer frame. Open with the definition; draw the architecture figure with client, NameNode, DataNodes and replicas; develop points 1-4, then read and write flow, then fault tolerance and features; for the command question give syntax and one example for two commands; close with why HDFS suits Big Data.

Asked: [7 marks] (Nov 2022, Nov 2023) Explain in detail about HDFS. Describe the structure of HDFS in a Hadoop ecosystem using a diagram. Asked: [7 marks] (Dec 2024, Jun 2025) Define HDFS. Discuss the HDFS Architecture and HDFS Commands in brief. Draw HDFS architecture. Explain any two commands of HDFS with syntax and an example.

Processing Data with Hadoop

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. Hadoop processes data by MapReduce: the input is split, mapped in parallel on the nodes that hold the data, shuffled and sorted, then reduced to the final output.

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-04" viewBox="0 0 596 252" width="596" height="252" role="img" aria-label="MapReduce: input splits, mappers (M), shuffle and sort (Sh), reducers (R), output"><style>#dsfig-u2-04 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-04 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-04 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-04 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-04 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-04 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-04 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-04 .t{fill:#16181D;font-weight:500}#dsfig-u2-04 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-04 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-04 .dot{fill:#16181D}#dsfig-u2-04 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-04 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-04 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-04 .ah{fill:#454C5A}#dsfig-u2-04 .ah.hi{fill:#2340B8}#dsfig-u2-04 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-04 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-04 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-04 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-04 .e{stroke:#B1B7C3}html.dark #dsfig-u2-04 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-04 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-04 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-04 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-04 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-04 .t{fill:#E6E8ED}html.dark #dsfig-u2-04 .t.inv{fill:#0F1115}html.dark #dsfig-u2-04 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-04 .dot{fill:#E6E8ED}html.dark #dsfig-u2-04 .ann{fill:#8FA3FF}html.dark #dsfig-u2-04 .lbl{fill:#858D9C}html.dark #dsfig-u2-04 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-04 .ah{fill:#B1B7C3}html.dark #dsfig-u2-04 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-04 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-04 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-04 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.8,115.5 L151.5,51.6" marker-end="url(#ah5)"/><path class="e" d="M55.8,136.5 L151.5,200.4" marker-end="url(#ah5)"/><path class="e" d="M184.8,50.5 L280.5,114.4" marker-end="url(#ah5)"/><path class="e" d="M184.8,201.5 L280.5,137.6" marker-end="url(#ah5)"/><path class="e" d="M313.8,115.5 L409.5,51.6" marker-end="url(#ah5)"/><path class="e" d="M313.8,136.5 L409.5,200.4" marker-end="url(#ah5)"/><path class="e" d="M442.8,50.5 L538.5,114.4" marker-end="url(#ah5)"/><path class="e" d="M442.8,201.5 L538.5,137.6" marker-end="url(#ah5)"/><g class="wl"><rect x="81" y="74" width="47.1" height="18" rx="9"/><text class="t" x="104.5" y="83" dy=".35em" text-anchor="middle">split</text></g><g class="wl"><rect x="81" y="160" width="47.1" height="18" rx="9"/><text class="t" x="104.5" y="169" dy=".35em" text-anchor="middle">split</text></g><g class="wl"><rect x="220.3" y="74" width="26.4" height="18" rx="9"/><text class="t" x="233.5" y="83" dy=".35em" text-anchor="middle">kv</text></g><g class="wl"><rect x="220.3" y="160" width="26.4" height="18" rx="9"/><text class="t" x="233.5" y="169" dy=".35em" text-anchor="middle">kv</text></g><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">In</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">M1</text><circle class="n" cx="169" cy="212" r="18"/><text class="t" x="169" y="212" dy=".35em" text-anchor="middle">M2</text><circle class="n" cx="298" cy="126" r="18"/><text class="t" x="298" y="126" dy=".35em" text-anchor="middle">Sh</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">R1</text><circle class="n" cx="427" cy="212" r="18"/><text class="t" x="427" y="212" dy=".35em" text-anchor="middle">R2</text><circle class="n" cx="556" cy="126" r="18"/><text class="t" x="556" y="126" dy=".35em" text-anchor="middle">Out</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">MapReduce: input splits, mappers (M), shuffle and sort (Sh), reducers (R), output</figcaption></figure>

Key points.

  1. Input splits are processed by separate map tasks that emit intermediate (key, value) pairs.
  2. Shuffle and sort groups all values of the same key and sends them to one reducer.
  3. Reducers aggregate each key's values and write the output to HDFS.
  4. Example: for "car car river", map emits (car,1),(car,1),(river,1) and reduce gives (car,2),(river,1).
  5. Big data processing works on huge, varied data with distributed storage, automatic fault tolerance and data locality, whereas traditional distributed processing moves data to compute nodes, expects structured data and handles failures manually.
Basis Big data processing (Hadoop) Traditional distributed processing
Data Petabytes, any type Limited size, mostly structured
Data movement Code moves to data Data moves to code
Fault tolerance Automatic re-execution Programmer handles it
Scalability Add commodity nodes Costly, limited

Answer frame. Open with the MapReduce definition; draw the phase figure; explain split, map, shuffle and sort, reduce and output with the word-count example; for the comparison question give the table; close with the advantages of parallelism and locality.

Asked: [7 marks] (Nov 2023) Explain how big data processing differs from distributed processing. Asked: [7 marks] (Jun 2025) What is MapReduce? Explain working of various phases of MapReduce with appropriate example and diagram.

Managing Resources and Application with Hadoop YARN

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. YARN (Yet Another Resource Negotiator) is Hadoop's cluster resource manager that separates resource management from job processing.

Key points.

  1. The ResourceManager is the master that allocates cluster resources and schedules applications.
  2. A NodeManager runs on each node, launches containers and reports usage to the ResourceManager.
  3. An ApplicationMaster is started per application to negotiate containers and track its tasks.
  4. MapReduce runs on YARN as one application, so YARN also lets other engines such as Spark share the cluster.

MapReduce programming

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>MapReduce is a programming model for processing large datasets in parallel across a cluster, where a Map function turns input into (key, value) pairs and a Reduce function aggregates all values of each key.</mark>

Key points.

  1. Job flow is: input split, map, (combine), shuffle and sort, reduce, output; all data moves as (key, value) pairs.
  2. The Mapper reads one input split record by record and emits intermediate (key, value) pairs, for example (word, 1).
  3. The Reducer receives each key with its grouped values after shuffle and sort and aggregates them, for example a sum.
  4. The Combiner is a mini-reducer run on each mapper's local output; it cuts the data sent across the network but must not change the result.
  5. It is used for large-scale parallel batch processing such as word count, log analysis, indexing and sorting of unstructured data.
  6. Why: it scales by adding nodes, tolerates failures by re-running failed tasks, and uses data locality by running maps where the block is stored.
  7. YARN runs MapReduce: the ResourceManager allocates containers, NodeManagers run them and an ApplicationMaster tracks the job.

Example (word count). Input lines: "deer bear river", "car car river", "deer car bear".

Stage Result
Map (deer,1)(bear,1)(river,1) (car,1)(car,1)(river,1) (deer,1)(car,1)(bear,1)
Shuffle and sort bear:[1,1] car:[1,1,1] deer:[1,1] river:[1,1]
Reduce bear 2, car 3, deer 2, river 2

Types and formats. Input formats decide how a split becomes records: TextInputFormat (key is byte offset, value is the line; the default), KeyValueTextInputFormat (line split at a tab into key and value), SequenceFileInputFormat (binary key-value files, fast and compact) and NLineInputFormat (fixed N lines per mapper). Output formats mirror them: TextOutputFormat (key tab value per line), SequenceFileOutputFormat (binary, to chain jobs) and NullOutputFormat (no output).

Answer frame. Open with the definition; draw the data-flow figure (as in Processing Data with Hadoop) with word count; develop points 1-4, then use and why; for Mapper, Reducer and Combiner give three short paragraphs; for MapReduce and YARN add point 7; close with scalability and fault tolerance.

Asked: [7 marks] (Dec 2020, Jun 2025) What is Map reduce programming? Where is it used and why? Asked: [7 marks] (Nov 2022) Briefly discuss about MapReduce and YARN. Asked: [7 marks] (Nov 2023) Discuss the following: i) Mapper ii) Reducer iii) Combiner Asked: [7 marks] (Dec 2024) Explain concept of Map Reduce using an example. Asked: [7 marks] (Dec 2024) Discuss on the different types and formats of Map-reduce with an example each one.

Last-minute revision

  • Hadoop = HDFS (storage) + YARN (resources) + MapReduce (processing) + Common (libraries).
  • HDFS block is 128 MB by default and replication factor is 3.
  • NameNode holds metadata; DataNodes hold blocks; Secondary NameNode only checkpoints.
  • HDFS is write-once, read-many; reads go directly to DataNodes.
  • MapReduce flow: split, map, combine, shuffle and sort, reduce.
  • Combiner is a local mini-reducer used to cut network traffic.
  • RDBMS: schema on write, vertical scaling; Hadoop: schema on read, horizontal scaling.
  • YARN: ResourceManager, NodeManager, ApplicationMaster, containers.
  • Hive = HiveQL over HDFS with a metastore; Pig = Pig Latin scripts; HBase = NoSQL; Sqoop = RDBMS import; Flume = logs.
  • Config files: core-site, hdfs-site, mapred-site, yarn-site, hadoop-env.sh.
  • Hadoop limits: small files, high latency, no ACID.

Memory hooks

  • "HYMC": HDFS, YARN, MapReduce, Common.
  • NameNode = librarian with the catalogue; DataNodes = shelves with the books.
  • "SMSR": Split, Map, Shuffle, Reduce.
  • Sqoop for SQL, Flume for logs.
  • Schema on write = RDBMS, schema on read = Hadoop.

Coverage checklist

  • Introduction to Hadoop: Nov 2022 why Big Data technology.
  • Core Hadoop components: Dec 2024 components and building blocks; Jun 2025 configuration files.
  • Hadoop Eco system: Dec 2020 and Jun 2025 ecosystem and HIVE; Nov 2022 short note.
  • Hive Physical Architecture: no past question.
  • Hadoop limitations: no past question.
  • RDBMS Versus Hadoop: Dec 2020 compare; Nov 2022 short note.
  • Hadoop Distributed File system: Nov 2022, Nov 2023, Dec 2024, Jun 2025.
  • Processing Data with Hadoop: Nov 2023 differs from distributed; Jun 2025 MapReduce phases.
  • Managing Resources and Application with Hadoop YARN: no past question (Nov 2022 MapReduce and YARN is under MapReduce).
  • MapReduce programming: Dec 2020, Jun 2025, Nov 2022, Nov 2023, Dec 2024 (two).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in