How unit 1 is examined
This unit covers what Big Data is, its Vs, types, challenges, technologies and desired system properties; challenges with Big Data carries the marks, then the significance and veracity questions.
Introduction to Big data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Big Data is a collection of data so large, fast-growing and varied that traditional database tools cannot store, process or analyse it within acceptable time.</mark>
Key points.
- It is measured in terabytes, petabytes and beyond, and it grows continuously from social media, sensors, transactions and logs.
- It needs distributed storage and parallel processing, for example Hadoop, rather than a single server.
- Healthcare uses it for patient records and disease prediction, and finance uses it for fraud detection and risk scoring.
- Retail uses it for recommendations and demand forecasting, and manufacturing uses it for predictive maintenance from sensor data.
- Government uses it for smart cities, traffic and welfare targeting, so it improves decisions and drives innovation.
Asked: [7 marks] (Jun 2025) Explain the significance of Big Data in various domains and industries.
Big data characteristics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. The characteristics are the Vs of Big Data: Volume, Velocity, Variety, Veracity and Value.
Key points.
- Volume is the huge size of data, from terabytes to zettabytes.
- Velocity is the speed at which data is generated and must be processed, for example streaming sensor data.
- Variety is the mix of structured, semi-structured and unstructured data such as tables, JSON, text and video.
- <mark>Veracity is the accuracy and trustworthiness of data; noise, bias, duplicates and missing values make analysis unreliable.</mark>
- Value is the useful insight extracted from data. Poor veracity leads to wrong analysis and bad decisions, so it is reduced by cleaning, validation, deduplication and source checks before analysis.
Asked: [7 marks] (Jun 2025) Explain the concept of data veracity in the context of Big Data and discuss its implications for data analysis and decision-making.
Types of big data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Big Data is classified by structure into structured, semi-structured and unstructured data.</mark>
Key points.
- Structured data has a fixed schema in rows and columns, for example an RDBMS table.
- Semi-structured data has tags but no rigid schema, for example XML, JSON and log files.
- Unstructured data has no predefined model, for example images, video, audio and free text.
Traditional versus Big data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Traditional data is small, structured and handled by a centralised RDBMS, while Big Data is huge, varied and handled by distributed systems.</mark>
| Basis | Traditional data | Big Data |
|---|---|---|
| Size | GB to TB | TB to PB and above |
| Type | Structured | Structured, semi-structured, unstructured |
| Architecture | Centralised | Distributed |
| Scaling | Vertical | Horizontal |
Evolution of Big data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Big Data evolved as data volume outgrew databases, moving from file systems and RDBMS to Hadoop-style distributed processing.</mark>
Key points.
- Early data was small and lived in files and relational databases.
- The web and social media in the 2000s made data explode in volume, velocity and variety.
- Google's GFS and MapReduce papers led to Hadoop, and later NoSQL, Spark and cloud platforms arrived.
challenges with Big Data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>The challenges of Big Data are the difficulties in capturing, storing, processing, securing and extracting value from data that is huge, fast and varied.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-01" viewBox="0 0 424 338" width="424" height="338" role="img" aria-label="Challenges of Big Data: Cap capture and integration, Sto storage, Pro processing speed, Sec security and privacy, Qua data quality, Ski skills and cost"><style>#dsfig-u1-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-01 .t{fill:#16181D;font-weight:500}#dsfig-u1-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-01 .dot{fill:#16181D}#dsfig-u1-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-01 .ah{fill:#454C5A}#dsfig-u1-01 .ah.hi{fill:#2340B8}#dsfig-u1-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-01 .e{stroke:#B1B7C3}html.dark #dsfig-u1-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-01 .t{fill:#E6E8ED}html.dark #dsfig-u1-01 .t.inv{fill:#0F1115}html.dark #dsfig-u1-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-01 .dot{fill:#E6E8ED}html.dark #dsfig-u1-01 .ann{fill:#8FA3FF}html.dark #dsfig-u1-01 .lbl{fill:#858D9C}html.dark #dsfig-u1-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-01 .ah{fill:#B1B7C3}html.dark #dsfig-u1-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M196.8,157.6 L56.8,52.6" marker-end="url(#ah1)"/><path class="e" d="M212,150 L212,61" marker-end="url(#ah1)"/><path class="e" d="M227.2,157.6 L367.2,52.6" marker-end="url(#ah1)"/><path class="e" d="M196.8,180.4 L56.8,285.4" marker-end="url(#ah1)"/><path class="e" d="M212,188 L212,277" marker-end="url(#ah1)"/><path class="e" d="M227.2,180.4 L367.2,285.4" marker-end="url(#ah1)"/><circle class="n" cx="212" cy="169" r="18"/><text class="t" x="212" y="169" dy=".35em" text-anchor="middle">BD</text><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Cap</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">Sto</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">Pro</text><circle class="n" cx="40" cy="298" r="18"/><text class="t" x="40" y="298" dy=".35em" text-anchor="middle">Sec</text><circle class="n" cx="212" cy="298" r="18"/><text class="t" x="212" y="298" dy=".35em" text-anchor="middle">Qua</text><circle class="n" cx="384" cy="298" r="18"/><text class="t" x="384" y="298" dy=".35em" text-anchor="middle">Ski</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Challenges of Big Data: Cap capture and integration, Sto storage, Pro processing speed, Sec security and privacy, Qua data quality, Ski skills and cost</figcaption></figure>
Key points.
- Storage is a challenge because data grows exponentially, so scalable and affordable distributed storage is required.
- Capturing and integrating data from many sources and formats is hard because of variety.
- Processing speed is a challenge because real-time velocity needs parallel and streaming frameworks.
- Data quality suffers because data is noisy, incomplete and inconsistent, which is the veracity problem.
- Security and privacy are at risk because large sensitive datasets attract attacks and must follow privacy laws.
- Analysis is difficult because unstructured data needs special tools and techniques.
- There is a shortage of skilled data scientists and engineers.
- Infrastructure, tools and maintenance are costly for organisations.
Answer frame. Open with the definition; draw the challenges diagram with the full names in boxes; develop points 1-8 in order, one or two sentences each with an example such as social media storage; close with the remark that Hadoop, NoSQL and cloud tools address these challenges.
Asked: [14 marks] (Jun 2025) Write a short note (any three): i) Hadoop distributed file ii) Challenges with big data iii) Properties of Big data systems iv) Data types of pig
Technologies available for Big Data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Big Data technologies are the tools used to store, process, analyse and visualise large datasets.</mark>
Key points.
- Storage and processing use Hadoop HDFS and MapReduce, and Spark for fast in-memory processing.
- NoSQL databases such as MongoDB and Cassandra store varied data.
- Hive and Pig query and script over Hadoop, Kafka handles streaming, and Tableau and Power BI visualise.
Infrastructure for Big data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Big Data infrastructure is the hardware, network and software platform that stores and processes large data at scale.</mark>
Key points.
- It uses clusters of commodity servers, so capacity grows by adding nodes (horizontal scaling).
- Distributed storage such as HDFS keeps replicated copies for fault tolerance.
- High-bandwidth networks, cloud platforms and resource managers such as YARN complete the stack.
Use of Data Analytics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data analytics is the process of examining data to find patterns and insights that support decisions.</mark>
Key points.
- Descriptive analytics explains what happened and predictive analytics forecasts what will happen.
- Prescriptive analytics recommends what to do.
- Uses include fraud detection, customer recommendations, healthcare diagnosis and demand forecasting.
Desired properties of Big Data system
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A Big Data system should be robust, scalable, fault tolerant and able to handle large, varied data with low latency.</mark>
Key points.
- Robustness and fault tolerance mean the system keeps working despite hardware failures, using replication.
- Scalability means performance grows by adding machines, and low latency means fast reads and updates.
- Extensibility, generalisation, ad hoc queries, minimal maintenance and debuggability are also desired.
Last-minute revision
- Big Data is data too large, fast and varied for traditional tools.
- The Vs are Volume, Velocity, Variety, Veracity and Value.
- Veracity means accuracy and trustworthiness of data.
- Types by structure are structured, semi-structured and unstructured.
- Traditional data is centralised and vertically scaled; Big Data is distributed and horizontally scaled.
- Hadoop grew from Google's GFS and MapReduce papers.
- Main challenges are storage, capture, processing, quality, security, skills and cost.
- Technologies include HDFS, MapReduce, Spark, NoSQL, Hive, Pig and Kafka.
- Analytics types are descriptive, predictive and prescriptive.
- Desired system properties are robustness, fault tolerance, scalability and low latency.
Memory hooks
- VVVVV: Volume, Velocity, Variety, Veracity, Value.
- Structured has Schema, semi has Tags, un has Nothing.
- Scale out (add machines), not up.
- Challenges: SPQ-SSC, storage, processing, quality, security, skills, cost.
Coverage checklist
- Introduction to Big data: significance in domains (Q3).
- Big data characteristics: veracity (Q2).
- Types of big data: structured, semi-structured, unstructured.
- Traditional versus Big data: comparison table.
- Evolution of Big data: file systems to Hadoop.
- challenges with Big Data: challenges short note (Q1).
- Technologies available for Big Data: Hadoop, Spark, NoSQL.
- Infrastructure for Big data: clusters and HDFS.
- Use of Data Analytics: types and uses.
- Desired properties of Big Data system: robustness, scalability (Q1 part iii).