How unit 1 is examined
This unit covers what big data is, its Vs, challenges, technologies, infrastructure and analytics; characteristics, challenges, technologies and analytics carry the marks.
Introduction to Big data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Big data is data so large, fast-growing and varied that traditional database tools cannot capture, store, manage or analyse it within acceptable time.</mark>
Key points.
- Big data comes from social media, sensors, transactions, logs, mobile phones and video.
- It is usually described by the Vs model: Volume, Velocity, Variety, Veracity and Value.
- It needs distributed storage and parallel processing on clusters instead of a single server.
- Its purpose is to turn raw data into insight for decisions.
Big data characteristics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>The characteristics of big data are its Vs: Volume, Velocity, Variety, Veracity, Value and Variability, which together separate big data from ordinary data.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-01" viewBox="0 0 510 424" width="510" height="424" role="img" aria-label="Vs of big data. Vol=Volume, Vel=Velocity, Var=Variety, Ver=Veracity, Val=Value, Vty=Variability"><style>#dsfig-u1-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-01 .t{fill:#16181D;font-weight:500}#dsfig-u1-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-01 .dot{fill:#16181D}#dsfig-u1-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-01 .ah{fill:#454C5A}#dsfig-u1-01 .ah.hi{fill:#2340B8}#dsfig-u1-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-01 .e{stroke:#B1B7C3}html.dark #dsfig-u1-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-01 .t{fill:#E6E8ED}html.dark #dsfig-u1-01 .t.inv{fill:#0F1115}html.dark #dsfig-u1-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-01 .dot{fill:#E6E8ED}html.dark #dsfig-u1-01 .ann{fill:#8FA3FF}html.dark #dsfig-u1-01 .lbl{fill:#858D9C}html.dark #dsfig-u1-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-01 .ah{fill:#B1B7C3}html.dark #dsfig-u1-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M255,193 L255,59"/><path class="e" d="M272.6,204.9 L452.4,133.1"/><path class="e" d="M272.6,219.1 L452.4,290.9"/><path class="e" d="M255,231 L255,365"/><path class="e" d="M237.4,219.1 L57.6,290.9"/><path class="e" d="M237.4,204.9 L57.6,133.1"/><circle class="n" cx="255" cy="212" r="18"/><text class="t" x="255" y="212" dy=".35em" text-anchor="middle">BD</text><circle class="n" cx="255" cy="40" r="18"/><text class="t" x="255" y="40" dy=".35em" text-anchor="middle">Vol</text><circle class="n" cx="470" cy="126" r="18"/><text class="t" x="470" y="126" dy=".35em" text-anchor="middle">Vel</text><circle class="n" cx="470" cy="298" r="18"/><text class="t" x="470" y="298" dy=".35em" text-anchor="middle">Var</text><circle class="n" cx="255" cy="384" r="18"/><text class="t" x="255" y="384" dy=".35em" text-anchor="middle">Ver</text><circle class="n" cx="40" cy="298" r="18"/><text class="t" x="40" y="298" dy=".35em" text-anchor="middle">Val</text><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">Vty</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Vs of big data. Vol=Volume, Vel=Velocity, Var=Variety, Ver=Veracity, Val=Value, Vty=Variability</figcaption></figure>
Key points.
- Volume is the sheer size of data, measured in terabytes, petabytes and zettabytes; Facebook, for example, stores petabytes of photos and posts.
- Velocity is the speed at which data is generated and must be processed, often in real time; stock-market ticks and sensor streams are examples.
- Variety is the many formats of data: structured tables, semi-structured JSON or XML, and unstructured text, images, audio and video.
- Veracity is the trustworthiness and quality of data, since noise, bias and missing values make it uncertain; unverified tweets are an example.
- Value is the useful business insight that can be extracted; data that yields no decision is worthless however large it is, as with purchase history used for recommendations.
- Variability is the changing meaning and flow of data over time, such as the word "sick" meaning different things in different posts, or seasonal traffic peaks.
- Handling these Vs needs new technologies, which is why Hadoop, Spark and NoSQL exist.
Answer frame. Open with the definition of big data and the Vs model; draw the six-V diagram; develop points 1-6 with one example each; close with the need for big data technologies. For Q2 (Nov 2022) list the Vs briefly, then add the challenges topic below.
Pitfall: Listing the Vs without an example for each loses marks; the examiner expects an example per V.
Asked: [7 marks] (Dec 2020, Nov 2023, Dec 2024) Write and explain main characteristics of big data; describe any five characteristics; what are the V's of Big Data, explain with examples. Asked: [7 marks] (Nov 2022) List out the characteristics of Big data and challenges in handling big data.
Types of big data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Big data is classified by structure into structured, semi-structured and unstructured data.</mark>
Key points.
- Structured data has a fixed schema in rows and columns, such as an SQL bank table.
- Semi-structured data has tags or keys but no rigid schema, such as JSON, XML and CSV logs.
- Unstructured data has no predefined model, such as emails, images, video and social posts, and forms most (about 80%) of big data.
- Each type needs different storage: RDBMS for structured, NoSQL for semi-structured, HDFS or object stores for unstructured.
Traditional versus Big data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Traditional data is small, structured, RDBMS-based enterprise data; big data is huge, fast, varied data beyond the reach of traditional tools.</mark>
| Basis | Traditional data | Big data |
|---|---|---|
| Volume | GB to TB | TB to PB and beyond |
| Velocity | Batch, slow growth | Real-time, streaming |
| Variety | Structured only | Structured, semi, unstructured |
| Storage | Central RDBMS | Distributed HDFS, NoSQL |
| Processing | Single server, SQL | Parallel MapReduce, Spark |
| Scalability | Vertical (bigger machine) | Horizontal (more nodes) |
| Example | Payroll, bank ledger | Tweets, sensor logs, video |
Asked: [7 marks] (Dec 2020) What do you mean by Big Data? How is it different from traditional data?
Evolution of Big data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Big data evolved as data volume outgrew databases, moving from RDBMS to distributed systems.</mark>
Key points.
- Big data 1.0 was RDBMS and data warehouses handling structured enterprise data.
- Big data 2.0 came with the web and social media, producing unstructured data; Google's GFS and MapReduce papers led to Hadoop in 2006.
- Big data 3.0 is mobile, IoT and real-time data, with Spark, cloud and machine learning.
challenges with Big Data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>The challenges of big data are the technical, security and human difficulties in capturing, storing, processing and using huge, varied data.</mark>
Key points.
- Storage: petabytes of data need scalable, low-cost distributed storage, which a single server cannot give.
- Processing: analysing data quickly, especially streams, needs parallel computing.
- Heterogeneity and scalability: data arrives in many formats and the system must grow by adding nodes without redesign.
- Capture, curation, search and visualization: collecting, cleaning, finding and displaying billions of records is hard.
- Data quality: noisy, duplicate and inconsistent data lowers trust in results.
- Privacy and security: personal data must be protected from breaches and misuse.
- Skill gap and cost: data scientists are scarce and clusters are expensive.
Answer frame. Open with the definition; list the seven challenges grouped as technical (1-4), data (5), security (6) and people and cost (7); close with the impact on processing and the need for tools like Hadoop.
Asked: [7 marks] (Dec 2020, Nov 2023) What are the challenges while handling Big data; enlist and explain the various challenges with big data.
Technologies available for Big Data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Big data technologies are the tools that store, process, ingest and manage large data on distributed clusters.</mark>
Key points.
- HDFS is the Hadoop distributed file system that stores large files in replicated blocks across nodes.
- NoSQL databases such as MongoDB, Cassandra and HBase store semi-structured data with horizontal scaling.
- MapReduce is the batch processing model that splits work into map and reduce tasks run in parallel.
- Spark is a fast in-memory processing engine for batch, streaming and machine learning.
- Hive and Pig give SQL-like and scripting interfaces over Hadoop data.
- Sqoop imports and exports relational data, Flume collects log streams, and YARN manages cluster resources.
Answer frame. Open by defining the big data technology stack; group points as storage (1, 2), processing (3-5) and ingestion and management (6); close with how together they store and process large data.
Asked: [7 marks] (Nov 2022, Nov 2023) Discuss various technologies available for Big Data; explain the technologies available for big data.
Infrastructure for Big data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Big data infrastructure is the hardware and software platform that stores and processes large data reliably and at scale.</mark>
Key points.
- Hardware is a cluster of commodity servers with large storage and high-speed networking.
- Software includes Hadoop, NoSQL databases, and virtualization or cloud platforms.
- It scales horizontally by adding nodes, and gives fault tolerance by replicating data.
- Cloud services offer elastic capacity on pay-as-you-use terms.
Asked: [7 marks] (Dec 2024) Explain the required infrastructure for Big Data.
Use of Data Analytics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Data analysis is inspecting, cleaning and modelling data to find useful information; data analytics is the wider use of tools and techniques on data to support decisions.</mark>
Key points.
- Descriptive analytics answers what happened, using reports and dashboards, such as monthly sales totals.
- Diagnostic analytics answers why it happened, by drilling into causes, such as why sales fell in one region.
- Predictive analytics answers what will happen, using statistics and machine learning on past data, such as forecasting demand or credit risk.
- Prescriptive analytics answers what should be done, by recommending actions, such as the best delivery route.
- Predictive workflow: collect data, clean it, build a model (regression, classification, time series), validate it, then deploy and predict.
- Role today: better decisions, business insight, efficiency and cost saving.
- Applications: fraud detection in industry and banking, disease prediction in healthcare, recommendations in e-commerce.
Answer frame. Open by defining analysis and analytics; develop the four types with an example each; for Q7 expand points 5 and 3 with techniques and uses; for Q6 close with points 6-7.
Asked: [7 marks] (Dec 2020) What is Data Analysis? Explain the role of data analytics in present scenario. Asked: [7 marks] (Nov 2022) Explain various types of Big Data analytics and explain about Predictive analytics.
Desired properties of Big Data system
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A good big data system is robust, scalable and fault tolerant while supporting low-latency reads and updates.</mark>
Key points.
- Scalability means handling growth by adding machines.
- Fault tolerance means surviving node failures through replication.
- Low latency, extensibility, ad hoc queries and minimal maintenance are also wanted.
- Debuggability and generalization to many workloads complete the list.
Last-minute revision
- Big data is data too large, fast or varied for traditional tools.
- Vs: Volume, Velocity, Variety, Veracity, Value, Variability.
- Types: structured, semi-structured, unstructured; unstructured is the majority.
- Traditional data scales vertically on an RDBMS; big data scales horizontally on clusters.
- Hadoop began in 2006 after Google's GFS and MapReduce papers.
- Challenges: storage, processing, quality, privacy, security, skills, cost.
- Storage tech: HDFS, NoSQL; processing: MapReduce, Spark, Hive, Pig; ingestion: Sqoop, Flume; resources: YARN.
- Analytics types: descriptive, diagnostic, predictive, prescriptive.
- Infrastructure: commodity cluster, fast network, replication, cloud.
Memory hooks
- Vs: "Very Very Very Valuable Vast Variety" for the Vs.
- Analytics ladder: What happened, Why, What will, What to do.
- Traditional is Vertical, big data is Horizontal.
- Storage HDFS, Speed Spark, Move Sqoop, Logs Flume.
Coverage checklist
- Introduction to Big data: definition and Vs.
- Big data characteristics: Q1, Q2.
- Types of big data: structured, semi-structured, unstructured.
- Traditional versus Big data: Q5.
- Evolution of Big data: three phases.
- challenges with Big Data: Q8, and Q2 challenges part.
- Technologies available for Big Data: Q4.
- Infrastructure for Big data: Q3.
- Use of Data Analytics: Q6, Q7.
- Desired properties of Big Data system: properties list.