Skip to content
CS-802 (B) · Cloud Computing/Quick Revision Short Notes

Cloud Computing (CS-802 (B)) - Unit 3 Short Notes

How unit 3 is examined

This unit covers how the cloud stores and processes big data; GFS/HDFS, Big Table, Map-Reduce and batch processing carry most of the marks.

Relational databases

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Cloud data management is the storage, access, consistency and lifecycle control of data spread over many cloud servers, and a relational database is its structured part: data kept in tables of rows and columns, queried with SQL and protected by ACID.</mark>

Key points.

  1. A relational database stores structured data as tables linked by keys, and SQL gives declarative queries, joins and aggregation.
  2. ACID (Atomicity, Consistency, Isolation, Durability) guarantees that transactions such as payments are all-or-nothing and never leave the data inconsistent.
  3. In the cloud it is offered as a managed service (Amazon RDS, Azure SQL, Google Cloud SQL) with automatic backup, patching and replication.
  4. Its limits are that scaling out across many servers is hard, joins and strict consistency slow down at cloud scale, and the fixed schema suits poorly unstructured data, so NoSQL stores (Big Table, Dynamo) complement it.

Asked: [7 marks] (May 2026) Define cloud data management and discuss the role of relational databases in cloud computing systems.

Cloud file systems: GFS and HDFS

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>GFS (Google File System) and HDFS (Hadoop Distributed File System) are scalable distributed file systems that split very large files into big chunks, replicate them over cheap commodity servers, and tolerate failures automatically.</mark>

Diagram. GFS architecture (Client asks the Master for metadata, then reads and writes data directly on chunkservers). <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-01" viewBox="0 0 424 312.2" width="424" height="312.2" role="img" aria-label="GFS: Cl = client, M = single master, C1-C3 = chunkservers (64 MB chunks, 3 replicas)"><style>#dsfig-u3-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-01 .t{fill:#16181D;font-weight:500}#dsfig-u3-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-01 .dot{fill:#16181D}#dsfig-u3-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-01 .ah{fill:#454C5A}#dsfig-u3-01 .ah.hi{fill:#2340B8}#dsfig-u3-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-01 .e{stroke:#B1B7C3}html.dark #dsfig-u3-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-01 .t{fill:#E6E8ED}html.dark #dsfig-u3-01 .t.inv{fill:#0F1115}html.dark #dsfig-u3-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-01 .dot{fill:#E6E8ED}html.dark #dsfig-u3-01 .ann{fill:#8FA3FF}html.dark #dsfig-u3-01 .lbl{fill:#858D9C}html.dark #dsfig-u3-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-01 .ah{fill:#B1B7C3}html.dark #dsfig-u3-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M57,117.5 L193.2,49.4" marker-end="url(#ah8)"/><path class="e" d="M60.7,122.4 L363.3,69.4" marker-end="url(#ah8)" marker-start="url(#ah8)"/><path class="e" d="M60.8,128.6 L363.2,166.4" marker-end="url(#ah8)" marker-start="url(#ah8)"/><path class="e" d="M59.3,134.2 L364.7,264" marker-end="url(#ah8)" marker-start="url(#ah8)"/><path class="e" d="M230.8,42.8 L363.2,62.7" marker-end="url(#ah8)"/><path class="e" d="M227.2,51.4 L367.2,156.4" marker-end="url(#ah8)"/><path class="e" d="M223.3,55.3 L371.5,255.3" marker-end="url(#ah8)"/><g class="wl"><rect x="105.6" y="74" width="40.8" height="18" rx="9"/><text class="t" x="126" y="83" dy=".35em" text-anchor="middle">meta</text></g><g class="wl"><rect x="191.6" y="86.9" width="40.8" height="18" rx="9"/><text class="t" x="212" y="95.9" dy=".35em" text-anchor="middle">data</text></g><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">Cl</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">M</text><circle class="n" cx="384" cy="65.8" r="18"/><text class="t" x="384" y="65.8" dy=".35em" text-anchor="middle">C1</text><circle class="n" cx="384" cy="169" r="18"/><text class="t" x="384" y="169" dy=".35em" text-anchor="middle">C2</text><circle class="n" cx="384" cy="272.2" r="18"/><text class="t" x="384" y="272.2" dy=".35em" text-anchor="middle">C3</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">GFS: Cl = client, M = single master, C1-C3 = chunkservers (64 MB chunks, 3 replicas)</figcaption></figure>

Key points (GFS).

  1. A GFS cluster has one master, many chunkservers and many clients; files are divided into fixed 64 MB chunks, each with a unique 64-bit chunk handle.
  2. The single master keeps all metadata in memory: the namespace, file-to-chunk mapping and chunk locations, and it grants chunk leases to a primary replica.
  3. Chunkservers store chunks as Linux files and each chunk is replicated, by default on three chunkservers.
  4. Clients ask the master only for metadata and then read or write data directly with chunkservers, so the master is not a bottleneck.
  5. Fault tolerance comes from replication, heartbeat messages, re-replication of lost chunks, an operation log with checkpoints and a shadow master.
  6. GFS is optimised for huge files, append-heavy writes and high throughput, not low latency.

Key points (HDFS).

  1. HDFS is the open-source Hadoop equivalent with a master NameNode holding the namespace and block map, and many DataNodes storing blocks.
  2. Files are split into blocks of 128 MB (older default 64 MB), replicated three times by default and placed across racks (rack awareness).
  3. It follows write-once-read-many: a file is written once and appended, never edited in place, which keeps replicas consistent.
  4. DataNodes send heartbeats and block reports to the NameNode; a dead node's blocks are re-replicated, and a Secondary or Standby NameNode protects the metadata.
  5. Data locality moves computation to the node holding the block, giving high throughput; storage scales by adding nodes.

Answer frame. Open with the definition; draw the GFS diagram (master, chunkservers, clients) and a similar NameNode/DataNode sketch; develop the points in order architecture, chunk/block size, replication, read/write path, fault tolerance, features (WORM, locality, scalability); close with a line that both make cloud storage of huge data cheap, reliable and scalable. For "features of HDFS", list the HDFS points as headings, each with its sentence.

Asked: [7 marks] (May 2022, May 2026) Write short note: i) GFS ii) HDFS; explain GFS and HDFS and their significance in cloud data storage. Asked: [7 marks] (May 2023) Enlist and explain the features of HDFS. Asked: [? marks] (May 2024) What is Bigtable? Describe the architecture of Google File System.

Features and comparisons among GFS, HDFS etc

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>GFS and HDFS share a master-slave design and replication, differing mainly in origin, chunk/block size and write model.</mark>

Point GFS HDFS
Owner Google, proprietary Apache Hadoop, open source
Master Master NameNode
Slaves Chunkservers DataNodes
Unit size 64 MB chunk 128 MB block (256 MB possible)
Writes Multiple writers, record append Single writer, write-once, append only
Replication 3 by default 3 by default, rack aware
Metadata log Operation log, checkpoint EditLog and FsImage
Fault tolerance Heartbeats, re-replication, shadow master Heartbeats, block reports, Standby NameNode
Use Google search and services Hadoop/MapReduce clusters

Data store vs simple database.

Point Data store Simple (relational) database
Data Any type: files, key-value, documents Structured rows and columns
Schema Flexible or none Fixed schema
Scaling Scales out horizontally Mostly scales up
Query API or simple lookups SQL with joins
Consistency Often eventual Strong ACID

Answer frame. Open by defining both systems in one line; draw the two architectures side by side; give the table, then close with one line that HDFS is the open-source counterpart of GFS.

Asked: [7 marks] (Dec 2024) i) Compare GFS and HDFS. ii) Differentiate data store and simple database. Asked: [? marks] (May 2024) Describe relational database. Compare GFS and HDFS.

Big Table

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Bigtable is Google's distributed, column-oriented NoSQL storage system for structured data, whose data model is a sparse, distributed, persistent multi-dimensional sorted map indexed by (row key, column key, timestamp).</mark>

Key points.

  1. Data is stored in tables whose rows are sorted by row key; columns are grouped into column families, and each cell keeps versions distinguished by timestamps.
  2. A table is split into row ranges called tablets (about 100-200 MB), which are the units of distribution and load balancing.
  3. A single master assigns tablets to tablet servers, balances load and detects server failure; tablet servers handle reads and writes.
  4. Data files (SSTables) and logs are stored in GFS, which supplies replication, so Bigtable is built on GFS.
  5. Chubby, a distributed lock service, elects the master, tracks live tablet servers and stores the root tablet location.
  6. It is called a distributed storage system because data is partitioned into tablets over thousands of servers, replicated through GFS, and scales and recovers from failures automatically.
  7. Bigtable and HBase matter in the cloud because they store petabytes of sparse data with low latency, scale by adding servers and suit web indexing, maps and analytics, unlike an RDBMS that lacks joins-free horizontal scaling.

Answer frame. Open with the definition; draw a box diagram of Client, Master, Tablet servers, Chubby and GFS; develop points 1-5, then scalability and fault tolerance; for the importance question add a two-line comparison with RDBMS; close with the value for big-data cloud applications.

Asked: [7 marks] (Dec 2024) Why does Big Table known as distributed storage system? Asked: [7 marks] (Jun 2025) Discuss the importance of data storage solutions like Big Table and HBase in cloud computing.

H Base and Dynamo

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. HBase is the open-source Bigtable clone that runs on HDFS, and Dynamo is Amazon's highly available key-value store.

Key points.

  1. HBase stores tables of column families in regions served by RegionServers, coordinated by an HMaster and ZooKeeper, with data kept in HDFS.
  2. Dynamo is a key-value store that partitions data with consistent hashing and replicates each item on several nodes.
  3. Dynamo favours availability over consistency: it is eventually consistent, uses vector clocks to reconcile versions, and has no master.

Parallel computing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Parallel computing performs many computations at the same time by dividing a problem into parts solved simultaneously on multiple processing units.</mark>

Key points.

  1. Parallelism means tasks truly execute at the same instant on different cores, while concurrency only means tasks are interleaved and progress together.
  2. On a single machine, parallelism comes from multiple threads and processes, multicore CPUs, and SIMD instructions that apply one operation to many data items.
  3. Its limits are the number of cores, shared memory and bus contention, synchronisation overhead, and Amdahl's law: the serial part bounds the speedup.
  4. Once one machine is not enough, the cloud extends the same idea across many machines, which is what Map-Reduce does.

Asked: [7 marks] (Dec 2024) Summarize the parallelism for single machine computation.

The Map-Reduce model

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>MapReduce is a programming model and framework for processing very large data sets in parallel on a cluster: a map function turns input records into intermediate key-value pairs, and a reduce function merges all values of each key.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-02" viewBox="0 0 596 252" width="596" height="252" role="img" aria-label="In = input splits, M = map tasks, Sh = shuffle and sort, R = reduce tasks, Out = output files"><style>#dsfig-u3-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-02 .t{fill:#16181D;font-weight:500}#dsfig-u3-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-02 .dot{fill:#16181D}#dsfig-u3-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-02 .ah{fill:#454C5A}#dsfig-u3-02 .ah.hi{fill:#2340B8}#dsfig-u3-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-02 .e{stroke:#B1B7C3}html.dark #dsfig-u3-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-02 .t{fill:#E6E8ED}html.dark #dsfig-u3-02 .t.inv{fill:#0F1115}html.dark #dsfig-u3-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-02 .dot{fill:#E6E8ED}html.dark #dsfig-u3-02 .ann{fill:#8FA3FF}html.dark #dsfig-u3-02 .lbl{fill:#858D9C}html.dark #dsfig-u3-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-02 .ah{fill:#B1B7C3}html.dark #dsfig-u3-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.8,115.5 L151.5,51.6" marker-end="url(#ah9)"/><path class="e" d="M59,126 L148,126" marker-end="url(#ah9)"/><path class="e" d="M55.8,136.5 L151.5,200.4" marker-end="url(#ah9)"/><path class="e" d="M184.8,50.5 L280.5,114.4" marker-end="url(#ah9)"/><path class="e" d="M188,126 L277,126" marker-end="url(#ah9)"/><path class="e" d="M184.8,201.5 L280.5,137.6" marker-end="url(#ah9)"/><path class="e" d="M316,120 L407.1,89.6" marker-end="url(#ah9)"/><path class="e" d="M316,132 L407.1,162.4" marker-end="url(#ah9)"/><path class="e" d="M445,89 L536.1,119.4" marker-end="url(#ah9)"/><path class="e" d="M445,163 L536.1,132.6" marker-end="url(#ah9)"/><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">In</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">M1</text><circle class="n" cx="169" cy="126" r="18"/><text class="t" x="169" y="126" dy=".35em" text-anchor="middle">M2</text><circle class="n" cx="169" cy="212" r="18"/><text class="t" x="169" y="212" dy=".35em" text-anchor="middle">M3</text><circle class="n" cx="298" cy="126" r="18"/><text class="t" x="298" y="126" dy=".35em" text-anchor="middle">Sh</text><circle class="n" cx="427" cy="83" r="18"/><text class="t" x="427" y="83" dy=".35em" text-anchor="middle">R1</text><circle class="n" cx="427" cy="169" r="18"/><text class="t" x="427" y="169" dy=".35em" text-anchor="middle">R2</text><circle class="n" cx="556" cy="126" r="18"/><text class="t" x="556" y="126" dy=".35em" text-anchor="middle">Out</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">In = input splits, M = map tasks, Sh = shuffle and sort, R = reduce tasks, Out = output files</figcaption></figure>

Key points.

  1. Input is divided into splits and each split is processed by a map task, giving intermediate pairs $(k, v)$.
  2. Shuffle collects the pairs, partitions them by key (hash of key mod number of reducers), sends them over the network to the reducers and sorts them so all values of one key arrive together.
  3. Each reduce task applies the reduce function to $(k, [v_1, v_2, ...])$ and writes the result to the file system.
  4. A master (JobTracker) schedules tasks near the data, and a failed or slow task is simply re-executed on another node, which gives fault tolerance.
  5. An optional combiner does local reduction after map, cutting network traffic.
  6. Applications are inverted indexing, log analysis, web-link graphs, machine learning and sorting.
  7. Unlike a traditional (single-machine or MPI) model, the programmer writes only map and reduce, and the framework handles partitioning, scheduling, communication and failures, so it scales on cheap commodity nodes.

Answer frame. Open with the definition and the three phases Map, Shuffle, Reduce; draw the diagram; walk through the word-count example (see Example/Application); list applications; close with the difference from other models (automatic parallelism and fault tolerance). For "explain shuffling" give point 2 in detail.

Asked: [7 marks] (May 2022, Dec 2024, May 2024) Discuss and explain the MapReduce model and its applications; how does MapReduce work, explain shuffling; how does it differ from other models.

Parallel efficiency of Map-Reduce

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Parallel efficiency measures how well the $p$ processors are used: $E = S/p = T_1/(p\,T_p)$, where speedup $S = T_1/T_p$.

Key points.

  1. $T_1$ is the run time on one machine and $T_p$ the time on $p$ machines; ideal efficiency is 1.
  2. Example: $T_1 = 100$ s, $p = 4$, $T_p = 30$ s gives $S = 3.33$ and E = 0.83.
  3. Map-Reduce loses efficiency to shuffle traffic, stragglers, skewed keys and start-up cost, and it stays efficient when data is large and the reduce work is small.

Relational operations

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Relational operations are the SQL-style operations (selection, projection, union, intersection, difference, join, grouping) that can be expressed as Map-Reduce jobs.

Key points.

  1. Selection and projection need only map: the mapper keeps rows that satisfy the condition, or emits chosen columns.
  2. Union, intersection and difference emit each tuple as the key so the reducer sees identical tuples together and outputs by count.
  3. Join maps each row to its join key tagged with its table name, and the reducer pairs rows from the two tables with the same key.
  4. Group-by and aggregation map rows to the group key, and the reducer applies sum, count or average.

Enterprise batch processing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Enterprise batch processing is the execution of large volumes of similar jobs as a group, with no user interaction, at scheduled times, such as payroll, billing and report generation.</mark>

Key points.

  1. Jobs are collected, queued and run in bulk, usually off-peak (nightly), and the results are delivered later rather than instantly.
  2. Its characteristics are large data volume, scheduling, no user interaction, throughput over response time, and restart from failure points.
  3. In the cloud, batch jobs run on elastic pools of virtual machines that are rented for the run and released afterwards, so the enterprise pays only for what it uses.
  4. Map-Reduce (Hadoop) suits batch work because it splits a job across many nodes (parallelism), re-runs failed tasks (fault tolerance) and scales with cheap machines (cost efficiency).
  5. Typical use cases are payroll, billing, log analysis, ETL for data warehouses, end-of-day banking and web indexing.
  6. Advantages are efficient resource use, lower cost and automation; the drawback is high latency, so it is not for real-time needs.

Answer frame. Open with the definition; draw a small flow Input data, Scheduler, Map-Reduce cluster, Output store; develop points 1-4 in order; then use cases; close with cost and scalability benefits. For the Map-Reduce benefit question, expand point 4.

Asked: [7 marks] (May 2023, Jun 2025) Write a brief note on enterprise batch processing; how does it benefit from the Map-Reduce model?

Example/Application of Map-Reduce

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Word count is the standard MapReduce example: map emits (word, 1) for every word and reduce sums the counts for each word.</mark>

Steps.

Step 1: Split the input text into blocks, one per map task.
Step 2: Map emits (word, 1) for each word.
Step 3: Shuffle groups pairs by word: (cloud, [1,1]).
Step 4: Reduce sums the list: (cloud, 2).

Example. Input "cloud data cloud" gives map output (cloud,1), (data,1), (cloud,1); after shuffle cloud:[1,1], data:[1]; reduce output cloud 2, data 1.

Key points.

  1. The data flow is input splits, map, shuffle and sort, reduce, output, as in the diagram of the model.
  2. A failed map or reduce task is re-run on another node, so the job still finishes correctly.

Asked: [7 marks] (May 2023) Take a suitable example and the concept of Map Reduce.

Last-minute revision

  • GFS: 1 master, chunkservers, 64 MB chunks, 3 replicas; clients get metadata from the master, data from chunkservers.
  • HDFS: NameNode + DataNodes, 128 MB blocks, replication 3, write-once-read-many, data locality.
  • GFS vs HDFS: proprietary vs open source; chunk vs block; record append vs append-only.
  • Bigtable: sparse, distributed, sorted map (row, column, timestamp); tablets, master, Chubby, GFS.
  • HBase = open-source Bigtable on HDFS; Dynamo = Amazon key-value, eventually consistent.
  • MapReduce phases: Map, Shuffle (partition, sort, transfer), Reduce.
  • Word count: map (word,1), reduce sum.
  • Efficiency $E = S/p$, speedup $S = T_1/T_p$.
  • Batch processing: bulk, scheduled, non-interactive, off-peak.
  • RDBMS = ACID + SQL; NoSQL = scale-out, flexible schema.

Memory hooks

  • GFS: "One Master, Many Chunks": Master, Chunkservers, Clients.
  • HDFS features: "B-R-F-W-L": Blocks, Replication, Fault tolerance, Write-once, Locality.
  • MapReduce: "Split, Map, Shuffle, Reduce".
  • Bigtable: "Tablets on servers, Chubby locks, GFS stores".

Coverage checklist

  • Relational databases: Q9 (May 2026), relational database part of May 2024.
  • Cloud file systems: GFS and HDFS: Q3, Q4, Q11.
  • Features and comparisons among GFS, HDFS etc: Q7, Q12.
  • Big Table: Q1, Q2.
  • H Base and Dynamo: taught, not asked.
  • Parallel computing: Q8.
  • The Map-Reduce model: Q10.
  • Parallel efficiency of Map-Reduce: taught, not asked.
  • Relational operations: taught, not asked.
  • Enterprise batch processing: Q5.
  • Example/Application of Map-Reduce: Q6.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in