How unit 2 is examined
This unit defines Big Data and the 4 V's, then covers analytics, Hadoop, the open-source ecosystem and the named analytics types; the marks sit in the 4 V's, Predictive Analytics, Hadoop, applications and the Big Data definition.
Big Data and its Importance
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Big Data is a collection of datasets so large, fast-growing and varied that traditional databases and tools cannot store, process or analyse them within acceptable time.</mark>
Key points.
- Big Data comes from social media, sensors, mobile phones, machine logs, transactions and the web, so it is mostly unstructured or semi-structured.
- Its characteristics are Volume, Velocity, Variety, Veracity and Value (see the 4 V's section), and each one breaks a traditional tool.
- Benefit: analysis of large data improves decision making because decisions rest on evidence instead of guesses.
- Benefit: cheap commodity clusters and open-source tools lower storage and processing cost, and automation improves operational efficiency.
- Benefit: customer behaviour analysis gives customer insight, personalised offers, fraud detection and new products.
- Challenge: storing and moving huge volume, handling high-velocity streams and mixing many formats.
- Challenge: poor veracity (noisy, incomplete data), privacy and security of personal data, and a shortage of skilled analysts.
- These challenges create the need for analytics frameworks such as Hadoop and Spark.
5 P's of Big Data. The 5 P's are Purpose (the business question the data must answer), People (data scientists, analysts and owners), Processes (collect, clean, store and govern the data), Platforms (Hadoop, Spark, cloud and NoSQL tools) and Performance (the measurable value or outcome delivered). Together they cover the data-management lifecycle and make analytics succeed.
Answer frame. Open with the definition; list benefits (points 3-5) with one example each; then challenges (6-7); close with point 8. For "characteristics" write the 5 V's from the next section. For the 5 P's give one line per P.
Asked: [7 marks] (Dec 2020) What are the benefits of Big Data? Discuss challenges under Big Data. Asked: [7 marks] (Jun 2020) What is Big Data? Explain characteristics of Big Data. Asked: [7 marks] (Jun 2020) Explain 5 P's of Big data in brief.
Four V's of Big Data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>The 4 V's of Big Data are Volume, Velocity, Variety and Veracity (Value is often added as the fifth V), and they describe what makes data "big" and hard to handle.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-01" viewBox="0 0 408 130" width="408" height="130" role="img" aria-label="The 4 V's of Big Data (Value is the fifth V)"><style>#dsfig-u2-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-01 .t{fill:#16181D;font-weight:500}#dsfig-u2-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-01 .dot{fill:#16181D}#dsfig-u2-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-01 .ah{fill:#454C5A}#dsfig-u2-01 .ah.hi{fill:#2340B8}#dsfig-u2-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-01 .e{stroke:#B1B7C3}html.dark #dsfig-u2-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-01 .t{fill:#E6E8ED}html.dark #dsfig-u2-01 .t.inv{fill:#0F1115}html.dark #dsfig-u2-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-01 .dot{fill:#E6E8ED}html.dark #dsfig-u2-01 .ann{fill:#8FA3FF}html.dark #dsfig-u2-01 .lbl{fill:#858D9C}html.dark #dsfig-u2-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-01 .ah{fill:#B1B7C3}html.dark #dsfig-u2-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="188" y1="37" x2="47.5" y2="101"/><line class="e" x1="188" y1="37" x2="138.5" y2="101"/><line class="e" x1="188" y1="37" x2="233.5" y2="101"/><line class="e" x1="188" y1="37" x2="328.5" y2="101"/><rect class="n" x="146.5" y="22" width="83" height="30" rx="8"/><text class="t" x="188" y="37" dy=".35em" text-anchor="middle">Big Data</text><rect class="n" x="14" y="86" width="67" height="30" rx="8"/><text class="t" x="47.5" y="101" dy=".35em" text-anchor="middle">Volume</text><rect class="n" x="97" y="86" width="83" height="30" rx="8"/><text class="t" x="138.5" y="101" dy=".35em" text-anchor="middle">Velocity</text><rect class="n" x="196" y="86" width="75" height="30" rx="8"/><text class="t" x="233.5" y="101" dy=".35em" text-anchor="middle">Variety</text><rect class="n" x="287" y="86" width="83" height="30" rx="8"/><text class="t" x="328.5" y="101" dy=".35em" text-anchor="middle">Veracity</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">The 4 V's of Big Data (Value is the fifth V)</figcaption></figure>
Key points.
- Volume is the sheer size of data, measured in terabytes, petabytes and zettabytes; Facebook, for example, stores photos and messages in petabytes.
- Velocity is the speed at which data is generated and must be processed, often in real time, such as stock ticks, sensor streams and tweets.
- Variety is the many forms of data: structured (tables), semi-structured (JSON, XML, logs) and unstructured (text, images, audio, video).
- Veracity is the trustworthiness and quality of data, because noise, bias, duplicates and missing values make results unreliable; social media posts are a typical uncertain source.
- Value (the fifth V) is the useful insight that can be turned into profit or decisions; data with no value is only cost.
- Each V creates a challenge: volume needs distributed storage, velocity needs streaming engines, variety needs flexible schemas and veracity needs cleaning.
Answer frame. Open with the definition; write one heading per V with meaning, characteristic and example (points 1-5); close with point 6 and the link to Big Data challenges. For Q1 (any three of five) write this section, Information management and one of Pig Latin, Sharding or Spark from the open-source section.
Pitfall: Do not write Veracity as "velocity again"; veracity is data quality, and Value is the fifth V.
Asked: [7 marks] (Dec 2020) Explain the 4V's of Big data. Asked: [14 marks] (Jun 2020) Write short note on any three: i) Pig latin ii) Information management iii) Sharding process iv) SPARK v) 4 V's of Big data.
Drivers for Big Data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Drivers of Big Data are the trends that increase how much data is generated and how it is acquired.</mark>
Key points.
- Generation grows from IoT devices, social media, smartphones and machine-generated logs, which produce data continuously.
- Acquisition is moving to real-time ingestion through sensors, streaming pipelines and cloud platforms instead of periodic batch loads.
- Cheap storage, cloud computing and open-source frameworks make it affordable to keep and process everything.
- The result is growth in volume, velocity and variety together, which forces new tools.
Asked: [7 marks] (Nov 2023) Discuss the trends in big data generation and acquisition.
Introduction to Big Data Analytics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Big Data analytics is the process of examining large, varied datasets to find patterns, correlations and insights that support decisions.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-02" viewBox="0 0 596 166" width="596" height="166" role="img" aria-label="Big data analytics life cycle: Discovery, Ingestion, Processing, Analysis, Visualization, Deployment (repeats as an iterative loop)"><style>#dsfig-u2-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-02 .t{fill:#16181D;font-weight:500}#dsfig-u2-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-02 .dot{fill:#16181D}#dsfig-u2-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-02 .ah{fill:#454C5A}#dsfig-u2-02 .ah.hi{fill:#2340B8}#dsfig-u2-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-02 .e{stroke:#B1B7C3}html.dark #dsfig-u2-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-02 .t{fill:#E6E8ED}html.dark #dsfig-u2-02 .t.inv{fill:#0F1115}html.dark #dsfig-u2-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-02 .dot{fill:#E6E8ED}html.dark #dsfig-u2-02 .ann{fill:#8FA3FF}html.dark #dsfig-u2-02 .lbl{fill:#858D9C}html.dark #dsfig-u2-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-02 .ah{fill:#B1B7C3}html.dark #dsfig-u2-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh2" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M54.6,52.2 L127.1,112.6" marker-end="url(#ah2)"/><path class="e" d="M157.8,113.8 L230.3,53.4" marker-end="url(#ah2)"/><path class="e" d="M261,52.2 L333.5,112.6" marker-end="url(#ah2)"/><path class="e" d="M364.2,113.8 L436.7,53.4" marker-end="url(#ah2)"/><path class="e" d="M467.4,52.2 L539.9,112.6" marker-end="url(#ah2)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Dis</text><circle class="n" cx="143.2" cy="126" r="18"/><text class="t" x="143.2" y="126" dy=".35em" text-anchor="middle">Ing</text><circle class="n" cx="246.4" cy="40" r="18"/><text class="t" x="246.4" y="40" dy=".35em" text-anchor="middle">Prc</text><circle class="n" cx="349.6" cy="126" r="18"/><text class="t" x="349.6" y="126" dy=".35em" text-anchor="middle">Ana</text><circle class="n" cx="452.8" cy="40" r="18"/><text class="t" x="452.8" y="40" dy=".35em" text-anchor="middle">Viz</text><circle class="n" cx="556" cy="126" r="18"/><text class="t" x="556" y="126" dy=".35em" text-anchor="middle">Dep</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Big data analytics life cycle: Discovery, Ingestion, Processing, Analysis, Visualization, Deployment (repeats as an iterative loop)</figcaption></figure>
Key points.
- Discovery defines the business problem, goals and data sources.
- Ingestion collects and loads data from databases, logs, sensors and streams into storage such as HDFS.
- Processing cleans, transforms and integrates the data into an analysable form.
- Analysis applies statistics, mining and machine learning to find patterns.
- Visualization presents results as charts and dashboards, and deployment puts the model into production; the cycle then repeats as results are refined.
Asked: [7 marks] (Nov 2023) What are the various stages in big data analytics life cycle? Illustrate with a figure, explaining each of them.
Big Data Analytics applications
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Big Data analytics applications use large-scale data from many sources to predict, monitor and optimise real-world systems in health, cities, weather, retail, finance and social media.</mark>
Key points.
- Smart cities: a smart city uses IoT sensors and analytics to run traffic, energy, water, waste and governance services efficiently.
- Smart-city data sources are traffic cameras and GPS, smart meters, environmental sensors and citizen-service records.
- Smart traffic example: live sensor and GPS data are analysed to predict congestion, retime signals and suggest routes, which cuts travel time and fuel; challenges are privacy, integration cost and data quality.
- Weather forecasting: data comes from ground sensors, satellites, radar and historical records, and the pipeline is collection, storage (HDFS), then predictive modelling.
- Weather applications are storm and cyclone prediction, crop planning in agriculture and disaster management; Big Data handles the huge volume and velocity of readings.
- Social media analytics is the analysis of posts, likes, shares and comments to understand opinion and behaviour; sources are Twitter, Facebook and Instagram, and metrics are sentiment, reach and engagement.
- Social media example: brand monitoring, where sentiment analysis of tweets shows how customers react to a product launch; election prediction is another; tools include Hootsuite, Google Analytics and Hadoop or Spark.
Answer frame. For smart city: define smart city and analytics, list sources, develop one application (traffic), close with benefits and challenges. For weather: sources, pipeline, applications, benefits. For social media: definition and need, sources and metrics, brand-monitoring example, tools and outcome.
Asked: [7 marks] (Dec 2020) How Big Data analytics can be useful in development of smart cities? (Discuss one application). Asked: [7 marks] (Dec 2020) Discuss the applications of big data analytics in weather forecasting. Asked: [7 marks] (Nov 2023) With an example, explain the term social media analytics.
Hadoop's Parallel World
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Hadoop is an open-source Apache framework that stores huge data across a cluster of commodity machines with HDFS and processes it in parallel with MapReduce, managed by YARN.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u2-03" viewBox="0 0 553 338" width="553" height="338" role="img" aria-label="Hadoop architecture. Cli = client, NN = NameNode, DN = DataNode (HDFS storage layer); RM = ResourceManager, NM = NodeManager (YARN); MapReduce tasks run in containers on the NodeManagers beside the data"><style>#dsfig-u2-03 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u2-03 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u2-03 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u2-03 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u2-03 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u2-03 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u2-03 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u2-03 .t{fill:#16181D;font-weight:500}#dsfig-u2-03 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u2-03 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u2-03 .dot{fill:#16181D}#dsfig-u2-03 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u2-03 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u2-03 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u2-03 .ah{fill:#454C5A}#dsfig-u2-03 .ah.hi{fill:#2340B8}#dsfig-u2-03 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u2-03 .wl .t{font-size:12px;font-weight:700}#dsfig-u2-03 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u2-03 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u2-03 .e{stroke:#B1B7C3}html.dark #dsfig-u2-03 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u2-03 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u2-03 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u2-03 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u2-03 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u2-03 .t{fill:#E6E8ED}html.dark #dsfig-u2-03 .t.inv{fill:#0F1115}html.dark #dsfig-u2-03 .kd{stroke:#E6E8ED}html.dark #dsfig-u2-03 .dot{fill:#E6E8ED}html.dark #dsfig-u2-03 .ann{fill:#8FA3FF}html.dark #dsfig-u2-03 .lbl{fill:#858D9C}html.dark #dsfig-u2-03 .ptr{fill:#8FA3FF}html.dark #dsfig-u2-03 .ah{fill:#B1B7C3}html.dark #dsfig-u2-03 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u2-03 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u2-03 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u2-03 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah3" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh3" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.2,157.6 L195.2,52.6" marker-end="url(#ah3)"/><path class="e" d="M55.2,180.4 L195.2,285.4" marker-end="url(#ah3)"/><path class="e" d="M231,40 L365,40"/><path class="e" d="M231,40 L494,40"/><path class="e" d="M231,298 L365,298"/><path class="e" d="M231,298 L494,298"/><path class="e" d="M384,59 L384,279"/><path class="e" d="M513,59 L513,279"/><circle class="n" cx="40" cy="169" r="18"/><text class="t" x="40" y="169" dy=".35em" text-anchor="middle">Cli</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">NN</text><circle class="n" cx="212" cy="298" r="18"/><text class="t" x="212" y="298" dy=".35em" text-anchor="middle">RM</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">DN1</text><circle class="n" cx="513" cy="40" r="18"/><text class="t" x="513" y="40" dy=".35em" text-anchor="middle">DN2</text><circle class="n" cx="384" cy="298" r="18"/><text class="t" x="384" y="298" dy=".35em" text-anchor="middle">NM1</text><circle class="n" cx="513" cy="298" r="18"/><text class="t" x="513" y="298" dy=".35em" text-anchor="middle">NM2</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Hadoop architecture. Cli = client, NN = NameNode, DN = DataNode (HDFS storage layer); RM = ResourceManager, NM = NodeManager (YARN); MapReduce tasks run in containers on the NodeManagers beside the data</figcaption></figure>
Key points.
- Hadoop has two main layers: HDFS for storage and MapReduce (with YARN) for processing.
- HDFS has one NameNode that holds metadata (file names, block locations) and many DataNodes that store the blocks, which are replicated (default three copies).
- YARN manages resources: the ResourceManager allocates cluster resources and each NodeManager runs containers on its node.
- MapReduce processes data in two phases: Map turns input into key-value pairs in parallel, and Reduce aggregates them by key.
- Hadoop is chosen because it scales out by adding cheap commodity nodes, tolerates failure by replication, is open source and moves computation to the data.
- Distributed computing challenges are node and network failure, data placement and consistency, and communication overhead; parallel computing challenges are splitting work, synchronisation and load balancing.
- Hadoop answers these: HDFS replication handles failure, MapReduce handles splitting and scheduling automatically, and YARN balances the load.
Answer frame. For architecture: define Hadoop and its layers, draw the diagram, explain HDFS, YARN, MapReduce, close with scalability and fault tolerance. For "why Hadoop": list point 5, contrast distributed and parallel challenges (6), relate to Hadoop (7), close with use cases such as log analysis and recommendation.
Asked: [7 marks] (Jun 2020, Nov 2023) Explain Hadoop architecture and its components with proper diagram. What is Hadoop? Describe the role of Hadoop in big data analysis and explain its core components. Asked: [7 marks] (Nov 2023) Why to choose Hadoop for processing Big Data in detail and explain the concept of distributed and parallel computing challenges?
Data discovery
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data discovery is the exploratory process of finding, visualising and understanding patterns and outliers in data before formal modelling.</mark>
Key points.
- It is the first stage of the analytics life cycle and uses profiling and visual tools.
- Analysts explore data quality, relationships and trends to form hypotheses.
- Tools such as Tableau and Power BI let business users do it without programming.
Open source technology for Big Data Analytics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Open-source Big Data technologies are freely available, community-maintained tools, built around the Hadoop ecosystem, for storing, processing, coordinating and analysing large data.</mark>
Key points.
- ZooKeeper is a centralised coordination service for naming, configuration, synchronisation and leader election among distributed nodes.
- ZooKeeper benefits are reliability (replicated servers), scalability, simplicity of its small API and ordered, fast updates; it is used in Hadoop, HBase and Kafka.
- HDFS stores data and YARN schedules resources; clients interact by writing data into HDFS and then submitting jobs through YARN.
- Hive gives SQL-like queries and Pig gives the Pig Latin data-flow scripting language, both compiled into batch jobs; HBase is a NoSQL column store for random real-time access.
- Spark is a fast in-memory engine for batch, streaming and machine learning; Kafka carries real-time streams; Sqoop and Flume load data in.
- Sharding is splitting a large database horizontally into shards stored on different servers so load and storage are shared.
- Processing technologies are MapReduce and Hive (batch) and Spark, Storm and Flink (real-time or streaming).
Answer frame. For ZooKeeper: define, features, benefits, usage. For the ecosystem: draw HDFS and YARN at the base with Hive, Pig, HBase and Spark above, explain the workflow and end with the batch versus real-time list.
Asked: [7 marks] (Jun 2020) What is Zoo keeper? List the benefits of it. Asked: [7 marks] (Nov 2023) Explain in detail the interacting process with Hadoop Ecosystem. List out various big data processing technologies.
cloud and Big Data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Cloud Big Data means running storage and analytics on rented cloud infrastructure instead of owned clusters.</mark>
Key points.
- It gives elastic scaling and pay-as-you-go cost, so small firms can analyse large data.
- Examples are Amazon EMR, Google BigQuery and Azure HDInsight.
- Concerns are data security, latency and vendor lock-in.
Predictive Analytics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Predictive analytics uses historical data, statistical models and machine learning to forecast the likelihood of future outcomes.</mark>
Key points.
- Its purpose is to answer "what is likely to happen", going beyond descriptive analytics, which only says what happened.
- The steps are define the goal, collect and clean data, build a model, validate it on test data and deploy it.
- Key techniques are regression, decision trees, time-series forecasting, neural networks and classification.
- Applications are credit scoring and fraud detection in banking, churn prediction in telecom, demand forecasting in retail and disease-risk prediction in health.
- Example: a bank trains a model on past loan repayments to predict which new applicants may default.
- Accuracy depends on data quality and volume, so Big Data platforms improve prediction.
Four-part answer for Q2. Write predictive analytics here, and inter-/trans-firewall analytics, information management and crowd sourcing analytics from their own sections below: each as a definition, purpose or technique, and one example. All four are ways of extracting value from data and relate directly to data analytics.
Answer frame. Open with the definition; list techniques; give the bank example; close with the link to Big Data. For the 14-mark question give about 3-4 lines to each of the four terms.
Pitfall: Predictive analytics forecasts probabilities, not certainties; say "likelihood".
Asked: [14 marks] (Nov 2023) Explain the following: a) Predictive analytics b) Inter-and Trans-firewall analytics c) Information management d) Crowd sourcing analytics
Mobile Business Intelligence and Big Data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Mobile BI delivers business intelligence reports and dashboards to smartphones and tablets so decisions can be made anywhere.</mark>
Key points.
- It gives real-time access to KPIs and alerts for managers on the move.
- Big Data supplies fresh data to the dashboards, and apps display them.
- Issues are small screens, security of data on devices and connectivity.
Crowd Sourcing Analytics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Crowd sourcing analytics collects data, labels or analysis from a large group of people, usually online, and analyses it to gain insight.</mark>
Key points.
- Contributors supply opinions, ratings, tagging or data; examples are Wikipedia, Waze traffic reports and product reviews.
- It is cheap and scalable and captures diverse views.
- Its weakness is uneven quality, so results must be validated and bias controlled.
- Example: a map app uses drivers' reports to predict jams.
Inter- and Trans-Firewall Analytics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Inter- and trans-firewall analytics analyses security and traffic logs collected from firewalls to detect threats, where inter-firewall analytics works on data inside one organisation across its own firewalls and trans-firewall analytics works across the firewalls of different organisations.</mark>
Key points.
- Firewalls generate huge logs, and analysing them reveals attacks that one device alone cannot see.
- Inter-firewall analytics correlates logs across the firewalls and network zones within a single organisation.
- Trans-firewall analytics shares and correlates data across organisational boundaries, for example partners or industry groups, while respecting privacy.
- Use case in threat detection: linking the same suspicious IP address seen at several sites reveals a coordinated attack early.
Asked: [7 marks] (Dec 2020) What do you mean by inter and trans fire wall analytics.
Information Management
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Information management is the process of collecting, storing, governing and delivering data so it is accurate, secure and usable.</mark>
Key points.
- It covers the data life cycle: capture, storage, quality, security, sharing and archiving.
- Data governance sets the rules, owners and policies for data.
- For Big Data it needs scalable storage, metadata and privacy controls.
- Good information management raises trust in analytics results.
Last-minute revision
- Big Data: data too large, fast or varied for traditional tools.
- 4 V's: Volume, Velocity, Variety, Veracity; Value is the fifth.
- 5 P's: Purpose, People, Processes, Platforms, Performance.
- Analytics life cycle: Discovery, Ingestion, Processing, Analysis, Visualization, Deployment.
- Hadoop = HDFS (storage) + MapReduce (processing) + YARN (resources).
- HDFS: one NameNode (metadata), many DataNodes (blocks, default three replicas).
- ZooKeeper: coordination service for naming, synchronisation and leader election.
- Spark is in-memory; Hive is SQL-like; Pig uses Pig Latin; HBase is NoSQL.
- Sharding: horizontal split of a database across servers.
- Predictive analytics forecasts future likelihood using history and models.
- Inter-firewall is within one organisation; trans-firewall is across organisations.
- Applications asked: smart cities, weather forecasting, social media analytics.
Memory hooks
- 4 V's: "Very Vast Vehicles Vroom" for Volume, Velocity, Variety, Veracity.
- 5 P's: "Purpose People Process Platform Performance".
- Hadoop: NameNode is the "boss with the map", DataNodes are the "workers with the boxes".
- Inter = inside, Trans = through and across firewalls.
- Life cycle: "Discover, Ingest, Process, Analyse, Visualise, Deploy" (DIPAVD).
Coverage checklist
- Big Data and its Importance: Dec 2020 benefits and challenges; Jun 2020 what is Big Data; Jun 2020 5 P's.
- Four V's of Big Data: Dec 2020 4V's; Jun 2020 short note.
- Drivers for Big Data: Nov 2023 trends in generation and acquisition.
- Introduction to Big Data Analytics: Nov 2023 life cycle.
- Big Data Analytics applications: smart cities, weather forecasting, social media analytics.
- Hadoop's Parallel World: architecture with diagram; why choose Hadoop.
- Data discovery: not asked recently.
- Open source technology for Big Data Analytics: ZooKeeper; Hadoop ecosystem.
- cloud and Big Data: not asked recently.
- Predictive Analytics: Nov 2023 four-part question.
- Mobile Business Intelligence and Big Data: not asked recently.
- Crowd Sourcing Analytics: Nov 2023 four-part question.
- Inter- and Trans-Firewall Analytics: Dec 2020 definition.
- Information Management: Nov 2023 four-part question; Jun 2020 short note.