Skip to content
CS-702 (D) · Big Data/Quick Revision Short Notes

Big Data (CS-702 (D)) - Unit 1 Short Notes

How unit 1 is examined

This unit covers what big data is, its Vs, challenges, technologies, infrastructure and analytics; characteristics, challenges, technologies and analytics carry the marks.

Introduction to Big data

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Big data is data so large, fast-growing and varied that traditional database tools cannot capture, store, manage or analyse it within acceptable time.</mark>

Key points.

  1. Big data comes from social media, sensors, transactions, logs, mobile phones and video.
  2. It is usually described by the Vs model: Volume, Velocity, Variety, Veracity and Value.
  3. It needs distributed storage and parallel processing on clusters instead of a single server.
  4. Its purpose is to turn raw data into insight for decisions.

Big data characteristics

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>The characteristics of big data are its Vs: Volume, Velocity, Variety, Veracity, Value and Variability, which together separate big data from ordinary data.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-01" viewBox="0 0 510 424" width="510" height="424" role="img" aria-label="Vs of big data. Vol=Volume, Vel=Velocity, Var=Variety, Ver=Veracity, Val=Value, Vty=Variability"><style>#dsfig-u1-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-01 .t{fill:#16181D;font-weight:500}#dsfig-u1-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-01 .dot{fill:#16181D}#dsfig-u1-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-01 .ah{fill:#454C5A}#dsfig-u1-01 .ah.hi{fill:#2340B8}#dsfig-u1-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-01 .e{stroke:#B1B7C3}html.dark #dsfig-u1-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-01 .t{fill:#E6E8ED}html.dark #dsfig-u1-01 .t.inv{fill:#0F1115}html.dark #dsfig-u1-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-01 .dot{fill:#E6E8ED}html.dark #dsfig-u1-01 .ann{fill:#8FA3FF}html.dark #dsfig-u1-01 .lbl{fill:#858D9C}html.dark #dsfig-u1-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-01 .ah{fill:#B1B7C3}html.dark #dsfig-u1-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M255,193 L255,59"/><path class="e" d="M272.6,204.9 L452.4,133.1"/><path class="e" d="M272.6,219.1 L452.4,290.9"/><path class="e" d="M255,231 L255,365"/><path class="e" d="M237.4,219.1 L57.6,290.9"/><path class="e" d="M237.4,204.9 L57.6,133.1"/><circle class="n" cx="255" cy="212" r="18"/><text class="t" x="255" y="212" dy=".35em" text-anchor="middle">BD</text><circle class="n" cx="255" cy="40" r="18"/><text class="t" x="255" y="40" dy=".35em" text-anchor="middle">Vol</text><circle class="n" cx="470" cy="126" r="18"/><text class="t" x="470" y="126" dy=".35em" text-anchor="middle">Vel</text><circle class="n" cx="470" cy="298" r="18"/><text class="t" x="470" y="298" dy=".35em" text-anchor="middle">Var</text><circle class="n" cx="255" cy="384" r="18"/><text class="t" x="255" y="384" dy=".35em" text-anchor="middle">Ver</text><circle class="n" cx="40" cy="298" r="18"/><text class="t" x="40" y="298" dy=".35em" text-anchor="middle">Val</text><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">Vty</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Vs of big data. Vol=Volume, Vel=Velocity, Var=Variety, Ver=Veracity, Val=Value, Vty=Variability</figcaption></figure>

Key points.

  1. Volume is the sheer size of data, measured in terabytes, petabytes and zettabytes; Facebook, for example, stores petabytes of photos and posts.
  2. Velocity is the speed at which data is generated and must be processed, often in real time; stock-market ticks and sensor streams are examples.
  3. Variety is the many formats of data: structured tables, semi-structured JSON or XML, and unstructured text, images, audio and video.
  4. Veracity is the trustworthiness and quality of data, since noise, bias and missing values make it uncertain; unverified tweets are an example.
  5. Value is the useful business insight that can be extracted; data that yields no decision is worthless however large it is, as with purchase history used for recommendations.
  6. Variability is the changing meaning and flow of data over time, such as the word "sick" meaning different things in different posts, or seasonal traffic peaks.
  7. Handling these Vs needs new technologies, which is why Hadoop, Spark and NoSQL exist.

Answer frame. Open with the definition of big data and the Vs model; draw the six-V diagram; develop points 1-6 with one example each; close with the need for big data technologies. For Q2 (Nov 2022) list the Vs briefly, then add the challenges topic below.

Pitfall: Listing the Vs without an example for each loses marks; the examiner expects an example per V.

Asked: [7 marks] (Dec 2020, Nov 2023, Dec 2024) Write and explain main characteristics of big data; describe any five characteristics; what are the V's of Big Data, explain with examples. Asked: [7 marks] (Nov 2022) List out the characteristics of Big data and challenges in handling big data.

Types of big data

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Big data is classified by structure into structured, semi-structured and unstructured data.</mark>

Key points.

  1. Structured data has a fixed schema in rows and columns, such as an SQL bank table.
  2. Semi-structured data has tags or keys but no rigid schema, such as JSON, XML and CSV logs.
  3. Unstructured data has no predefined model, such as emails, images, video and social posts, and forms most (about 80%) of big data.
  4. Each type needs different storage: RDBMS for structured, NoSQL for semi-structured, HDFS or object stores for unstructured.

Traditional versus Big data

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Traditional data is small, structured, RDBMS-based enterprise data; big data is huge, fast, varied data beyond the reach of traditional tools.</mark>

Basis Traditional data Big data
Volume GB to TB TB to PB and beyond
Velocity Batch, slow growth Real-time, streaming
Variety Structured only Structured, semi, unstructured
Storage Central RDBMS Distributed HDFS, NoSQL
Processing Single server, SQL Parallel MapReduce, Spark
Scalability Vertical (bigger machine) Horizontal (more nodes)
Example Payroll, bank ledger Tweets, sensor logs, video

Asked: [7 marks] (Dec 2020) What do you mean by Big Data? How is it different from traditional data?

Evolution of Big data

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Big data evolved as data volume outgrew databases, moving from RDBMS to distributed systems.</mark>

Key points.

  1. Big data 1.0 was RDBMS and data warehouses handling structured enterprise data.
  2. Big data 2.0 came with the web and social media, producing unstructured data; Google's GFS and MapReduce papers led to Hadoop in 2006.
  3. Big data 3.0 is mobile, IoT and real-time data, with Spark, cloud and machine learning.

challenges with Big Data

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>The challenges of big data are the technical, security and human difficulties in capturing, storing, processing and using huge, varied data.</mark>

Key points.

  1. Storage: petabytes of data need scalable, low-cost distributed storage, which a single server cannot give.
  2. Processing: analysing data quickly, especially streams, needs parallel computing.
  3. Heterogeneity and scalability: data arrives in many formats and the system must grow by adding nodes without redesign.
  4. Capture, curation, search and visualization: collecting, cleaning, finding and displaying billions of records is hard.
  5. Data quality: noisy, duplicate and inconsistent data lowers trust in results.
  6. Privacy and security: personal data must be protected from breaches and misuse.
  7. Skill gap and cost: data scientists are scarce and clusters are expensive.

Answer frame. Open with the definition; list the seven challenges grouped as technical (1-4), data (5), security (6) and people and cost (7); close with the impact on processing and the need for tools like Hadoop.

Asked: [7 marks] (Dec 2020, Nov 2023) What are the challenges while handling Big data; enlist and explain the various challenges with big data.

Technologies available for Big Data

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Big data technologies are the tools that store, process, ingest and manage large data on distributed clusters.</mark>

Key points.

  1. HDFS is the Hadoop distributed file system that stores large files in replicated blocks across nodes.
  2. NoSQL databases such as MongoDB, Cassandra and HBase store semi-structured data with horizontal scaling.
  3. MapReduce is the batch processing model that splits work into map and reduce tasks run in parallel.
  4. Spark is a fast in-memory processing engine for batch, streaming and machine learning.
  5. Hive and Pig give SQL-like and scripting interfaces over Hadoop data.
  6. Sqoop imports and exports relational data, Flume collects log streams, and YARN manages cluster resources.

Answer frame. Open by defining the big data technology stack; group points as storage (1, 2), processing (3-5) and ingestion and management (6); close with how together they store and process large data.

Asked: [7 marks] (Nov 2022, Nov 2023) Discuss various technologies available for Big Data; explain the technologies available for big data.

Infrastructure for Big data

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Big data infrastructure is the hardware and software platform that stores and processes large data reliably and at scale.</mark>

Key points.

  1. Hardware is a cluster of commodity servers with large storage and high-speed networking.
  2. Software includes Hadoop, NoSQL databases, and virtualization or cloud platforms.
  3. It scales horizontally by adding nodes, and gives fault tolerance by replicating data.
  4. Cloud services offer elastic capacity on pay-as-you-use terms.

Asked: [7 marks] (Dec 2024) Explain the required infrastructure for Big Data.

Use of Data Analytics

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Data analysis is inspecting, cleaning and modelling data to find useful information; data analytics is the wider use of tools and techniques on data to support decisions.</mark>

Key points.

  1. Descriptive analytics answers what happened, using reports and dashboards, such as monthly sales totals.
  2. Diagnostic analytics answers why it happened, by drilling into causes, such as why sales fell in one region.
  3. Predictive analytics answers what will happen, using statistics and machine learning on past data, such as forecasting demand or credit risk.
  4. Prescriptive analytics answers what should be done, by recommending actions, such as the best delivery route.
  5. Predictive workflow: collect data, clean it, build a model (regression, classification, time series), validate it, then deploy and predict.
  6. Role today: better decisions, business insight, efficiency and cost saving.
  7. Applications: fraud detection in industry and banking, disease prediction in healthcare, recommendations in e-commerce.

Answer frame. Open by defining analysis and analytics; develop the four types with an example each; for Q7 expand points 5 and 3 with techniques and uses; for Q6 close with points 6-7.

Asked: [7 marks] (Dec 2020) What is Data Analysis? Explain the role of data analytics in present scenario. Asked: [7 marks] (Nov 2022) Explain various types of Big Data analytics and explain about Predictive analytics.

Desired properties of Big Data system

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A good big data system is robust, scalable and fault tolerant while supporting low-latency reads and updates.</mark>

Key points.

  1. Scalability means handling growth by adding machines.
  2. Fault tolerance means surviving node failures through replication.
  3. Low latency, extensibility, ad hoc queries and minimal maintenance are also wanted.
  4. Debuggability and generalization to many workloads complete the list.

Last-minute revision

  • Big data is data too large, fast or varied for traditional tools.
  • Vs: Volume, Velocity, Variety, Veracity, Value, Variability.
  • Types: structured, semi-structured, unstructured; unstructured is the majority.
  • Traditional data scales vertically on an RDBMS; big data scales horizontally on clusters.
  • Hadoop began in 2006 after Google's GFS and MapReduce papers.
  • Challenges: storage, processing, quality, privacy, security, skills, cost.
  • Storage tech: HDFS, NoSQL; processing: MapReduce, Spark, Hive, Pig; ingestion: Sqoop, Flume; resources: YARN.
  • Analytics types: descriptive, diagnostic, predictive, prescriptive.
  • Infrastructure: commodity cluster, fast network, replication, cloud.

Memory hooks

  • Vs: "Very Very Very Valuable Vast Variety" for the Vs.
  • Analytics ladder: What happened, Why, What will, What to do.
  • Traditional is Vertical, big data is Horizontal.
  • Storage HDFS, Speed Spark, Move Sqoop, Logs Flume.

Coverage checklist

  • Introduction to Big data: definition and Vs.
  • Big data characteristics: Q1, Q2.
  • Types of big data: structured, semi-structured, unstructured.
  • Traditional versus Big data: Q5.
  • Evolution of Big data: three phases.
  • challenges with Big Data: Q8, and Q2 challenges part.
  • Technologies available for Big Data: Q4.
  • Infrastructure for Big data: Q3.
  • Use of Data Analytics: Q6, Q7.
  • Desired properties of Big Data system: properties list.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in