How unit 4 is examined
This unit covers what NoSQL is and how it differs from RDBMS, why industry adopts it, the four NoSQL types, and MongoDB; the marks sit in NoSQL vs relational, types of NoSQL, and MongoDB (features, CRUD, sharding, MongoDB vs Hadoop).
Introduction to NoSQL
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>NoSQL ("Not Only SQL") databases are non-relational, distributed databases that store data without a fixed table schema and scale horizontally across commodity servers, so they can handle large, varied and fast-growing data.</mark>
Key points.
- NoSQL is schema-less: each record can have different fields, so the structure can change without altering a table definition.
- It scales horizontally (scale-out) by adding cheap servers, whereas an RDBMS mostly scales vertically by buying a bigger machine.
- Data is distributed and replicated over many nodes, which gives high availability and fault tolerance.
- It follows the BASE model (Basically Available, Soft state, Eventually consistent) and trades strict consistency for availability, as the CAP theorem allows.
- It stores data in flexible models (key-value, document, column-family, graph) instead of joined tables, so reads and writes are fast at large volume.
- It is needed because RDBMS struggles with the volume, variety and velocity of Big Data such as social feeds, logs and sensor streams.
- Advantages are performance, flexibility, easy scaling and low cost; limitations are weak joins, no standard query language and weaker transactions.
- Examples are MongoDB (document), Redis (key-value), Cassandra and HBase (column-family) and Neo4j (graph).
Comparison.
| Basis | NoSQL | Relational (RDBMS) |
|---|---|---|
| Data model | Key-value, document, column, graph | Tables of rows and columns |
| Schema | Dynamic, schema-less | Fixed, predefined schema |
| Scalability | Horizontal (scale-out) | Vertical (scale-up) |
| Consistency | BASE, eventual consistency | ACID, strong consistency |
| Query language | No standard, database-specific APIs | Standard SQL |
| Joins | Rare or absent; data is denormalised | Strong support for joins |
| Best suited to | Big, unstructured, fast-changing data | Structured data, banking, transactions |
Answer frame. Open with the definition and the need over RDBMS; for a short note list points 1-5, the four types with examples and advantages; for "differentiate" draw the table above and end with use-case suitability; close with "NoSQL suits Big Data, RDBMS suits transactional data".
Pitfall: Do not say NoSQL "has no SQL" or "has no consistency"; it means "Not Only SQL" and consistency is eventual.
Asked: [7 marks] (Nov 2023, Dec 2024) Write a short note on NoSQL databases. What is a NoSQL database? Discuss key characteristics and advantages of NoSQL database. Asked: [7 marks] (Nov 2023) Differentiate between NoSQL and relational database.
NoSQL Business Drivers
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. Business drivers are the market and business pressures that push organisations to adopt NoSQL and Big Data analytics: growing data volume, velocity and variety, and the need for faster decisions at lower cost.
Key points.
- Volume: companies collect terabytes of clicks, logs and transactions, and NoSQL stores them by adding servers.
- Velocity: real-time needs such as live recommendations and fraud alerts demand fast writes and reads.
- Variety: text, images, JSON and sensor data do not fit rigid tables, and NoSQL stores them together.
- Competition and personalisation: firms analyse customer behaviour to personalise offers and win customers.
- Operational efficiency and cost reduction: commodity clusters and open-source software are cheaper than large licensed RDBMS servers.
- New business models and better decisions: data-driven products and analytics improve decisions, and this links directly to Big Data adoption.
- Industry applications: e-commerce catalogues and carts (MongoDB), social media feeds and messaging (Cassandra), IoT sensor streams, banking fraud detection, and analytics on HBase.
Answer frame. For drivers, open with a definition, develop points 1-6 and close by linking them to Big Data adoption; for applications, open with the industrial need, give one example per sector (points 7), and close with benefits: scalability, flexibility and low cost.
Asked: [7 marks] (Nov 2023, Dec 2024) Describe applications of NoSQL databases in industry. Write a short note on use of NoSQL database in industry. Asked: [7 marks] (Nov 2022) Explain in detail about market and business drivers for Big Data analytics.
NoSQL Data architectural patterns
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>NoSQL architectural patterns are the four data models in which NoSQL databases organise data: key-value, document, column-family and graph.</mark>
Diagram.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-01" viewBox="0 0 455 130" width="455" height="130" role="img" aria-label="Four types of NoSQL databases"><style>#dsfig-u4-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-01 .t{fill:#16181D;font-weight:500}#dsfig-u4-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-01 .dot{fill:#16181D}#dsfig-u4-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-01 .ah{fill:#454C5A}#dsfig-u4-01 .ah.hi{fill:#2340B8}#dsfig-u4-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-01 .e{stroke:#B1B7C3}html.dark #dsfig-u4-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-01 .t{fill:#E6E8ED}html.dark #dsfig-u4-01 .t.inv{fill:#0F1115}html.dark #dsfig-u4-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-01 .dot{fill:#E6E8ED}html.dark #dsfig-u4-01 .ann{fill:#8FA3FF}html.dark #dsfig-u4-01 .lbl{fill:#858D9C}html.dark #dsfig-u4-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-01 .ah{fill:#B1B7C3}html.dark #dsfig-u4-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><line class="e" x1="223.5" y1="37" x2="59.5" y2="101"/><line class="e" x1="223.5" y1="37" x2="162.5" y2="101"/><line class="e" x1="223.5" y1="37" x2="281" y2="101"/><line class="e" x1="223.5" y1="37" x2="387.5" y2="101"/><rect class="n" x="170.5" y="22" width="106" height="30" rx="8"/><text class="t" x="223.5" y="37" dy=".35em" text-anchor="middle">NoSQL types</text><rect class="n" x="14" y="86" width="91" height="30" rx="8"/><text class="t" x="59.5" y="101" dy=".35em" text-anchor="middle">Key-value</text><rect class="n" x="121" y="86" width="83" height="30" rx="8"/><text class="t" x="162.5" y="101" dy=".35em" text-anchor="middle">Document</text><rect class="n" x="220" y="86" width="122" height="30" rx="8"/><text class="t" x="281" y="101" dy=".35em" text-anchor="middle">Column-family</text><rect class="n" x="358" y="86" width="59" height="30" rx="8"/><text class="t" x="387.5" y="101" dy=".35em" text-anchor="middle">Graph</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Four types of NoSQL databases</figcaption></figure>
Key points.
- Key-value stores keep data as a unique key mapped to an opaque value, so lookups are very fast; use them for caching, sessions and shopping carts (Redis, DynamoDB).
- Document stores keep self-describing JSON or BSON documents grouped in collections and can query on fields inside them; use them for catalogues, content and user profiles (MongoDB, CouchDB).
- Column-family stores keep data in rows with dynamic columns grouped into column families, and are built for huge write-heavy, distributed data such as logs and time series (Cassandra, HBase).
- Graph stores keep nodes and edges with properties, so relationship traversal is fast; use them for social networks, recommendations and fraud rings (Neo4j).
- Key-value is the simplest model and graph the richest in relationships, so the choice depends on the access pattern of the application.
- Each type sacrifices joins and strict schema to gain scale, and all can be distributed with replication.
| Type | Data unit | Example | Typical use |
|---|---|---|---|
| Key-value | Key and value | Redis | Cache, sessions |
| Document | JSON document | MongoDB | Catalogue, profiles |
| Column-family | Row with column families | Cassandra | Logs, time series |
| Graph | Node and edge | Neo4j | Social links |
Answer frame. For "types", open with the definition, draw the tree, describe points 1-4 in that order with an example each, and close with the table of use cases. For the 14-mark question, first define NoSQL (topic 1), draw the NoSQL vs relational table from "Introduction to NoSQL", then give the types briefly.
Asked: [7 marks] (Dec 2024) Discuss various types of NoSQL databases with example. Asked: [14 marks] (Jun 2025) What is a NoSQL database? List the differences between NoSQL and relational databases. Explain in brief various types of NoSQL databases in practice.
Variations of NoSQL architectural patterns for Big Data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. These variations are the ways NoSQL patterns are applied to Big Data problems, using scale-out distribution, replication and flexible schemas to handle volume, variety and velocity.
Key points.
- CAP theorem: a distributed store can guarantee only two of Consistency, Availability and Partition tolerance, and NoSQL systems usually choose availability with eventual consistency (AP), while some such as MongoDB and HBase favour consistency (CP).
- Volume: data is sharded across many nodes, so capacity grows by adding servers rather than by upgrading one machine.
- Velocity: key-value and column-family stores accept very fast writes, and replication keeps reads quick.
- Variety: document and graph stores accept unstructured and semi-structured data without a fixed schema.
- Variations by type are key-value (Redis), document (MongoDB), column-family (Cassandra) and graph (Neo4j).
- Big Data problems solved: real-time recommendations, social feeds, log and clickstream storage, sensor data, and fraud detection.
Answer frame. For "introduction and variations", define NoSQL and its need over RDBMS, state CAP briefly, then list the four variations with one example each; for "useful for Big Data", develop volume, velocity and variety (points 2-4), then examples, and close with scalability and schema flexibility.
Asked: [7 marks] (Dec 2020) Give an introduction to NoSQL. Write variations of NoSQL. Asked: [7 marks] (Nov 2022) How is NoSQL useful for Big Data problems?
Introduction to MongoDB
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>MongoDB is an open-source, document-oriented NoSQL database that stores data as flexible JSON-like BSON documents grouped into collections, and supports indexing, replication and sharding.</mark>
Key points.
- A document is a set of field-value pairs stored in BSON (binary JSON), a collection is a group of documents like a table, and a database holds collections.
- The schema is dynamic, so documents in one collection may have different fields.
- Indexing on any field, including secondary indexes, speeds up queries.
- Replication uses a replica set with one primary and several secondaries, giving automatic failover and high availability.
- Sharding splits data across servers for horizontal scaling.
- The aggregation framework processes data in pipeline stages (
$match,$group,$sort) for analytics. - It has a rich query language, is fast for reads and writes, and is used for content management, catalogues, mobile apps, IoT and real-time analytics.
Example. CRUD operations in the mongo shell:
db.students.insertOne({name:"Ravi", age:20}) // Create
db.students.find({age:{$gt:18}}) // Read
db.students.updateOne({name:"Ravi"}, {$set:{age:21}}) // Update
db.students.deleteOne({name:"Ravi"}) // Delete
Sharding. Sharding is the horizontal partitioning of a collection across several servers, needed when data or load exceeds one machine.
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-02" viewBox="0 0 338 252" width="338" height="252" role="img" aria-label="App = client application, Mgs = mongos query router, Cfg = config servers (metadata), S1 and S2 = shards"><style>#dsfig-u4-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-02 .t{fill:#16181D;font-weight:500}#dsfig-u4-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-02 .dot{fill:#16181D}#dsfig-u4-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-02 .ah{fill:#454C5A}#dsfig-u4-02 .ah.hi{fill:#2340B8}#dsfig-u4-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-02 .e{stroke:#B1B7C3}html.dark #dsfig-u4-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-02 .t{fill:#E6E8ED}html.dark #dsfig-u4-02 .t.inv{fill:#0F1115}html.dark #dsfig-u4-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-02 .dot{fill:#E6E8ED}html.dark #dsfig-u4-02 .ann{fill:#8FA3FF}html.dark #dsfig-u4-02 .lbl{fill:#858D9C}html.dark #dsfig-u4-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-02 .ah{fill:#B1B7C3}html.dark #dsfig-u4-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,126 L148,126" marker-end="url(#ah8)"/><path class="e" d="M184.8,115.5 L280.5,51.6" marker-end="url(#ah8)"/><path class="e" d="M188,126 L277,126" marker-end="url(#ah8)"/><path class="e" d="M184.8,136.5 L280.5,200.4" marker-end="url(#ah8)"/><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">App</text><circle class="n" cx="169" cy="126" r="18"/><text class="t" x="169" y="126" dy=".35em" text-anchor="middle">Mgs</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">Cfg</text><circle class="n" cx="298" cy="126" r="18"/><text class="t" x="298" y="126" dy=".35em" text-anchor="middle">S1</text><circle class="n" cx="298" cy="212" r="18"/><text class="t" x="298" y="212" dy=".35em" text-anchor="middle">S2</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">App = client application, Mgs = mongos query router, Cfg = config servers (metadata), S1 and S2 = shards</figcaption></figure>
- Components: shards hold the data, mongos routes queries, config servers store the metadata of which chunk lives where, and the shard key decides distribution.
- Steps: start config servers and shards, start mongos, run
sh.addShard()for each shard, runsh.enableSharding("db"), thensh.shardCollection("db.coll", {key:1}). - MongoDB divides data into chunks by shard key range or hash and balances them across shards automatically.
- Advantages are scalability, higher throughput and no single-machine limit.
Comparison.
| Basis | MongoDB | Hadoop |
|---|---|---|
| Nature | Document-oriented NoSQL database | Distributed framework (HDFS and MapReduce) |
| Data model | JSON/BSON documents | Files in HDFS |
| Access | Real-time queries by field | Batch processing |
| Architecture | Replica sets and shards | Master-slave (NameNode and DataNodes) |
| Use case | Operational apps, catalogues | Large-scale offline analytics |
| Scalability | Scales by sharding | Scales by adding DataNodes |
Answer frame. For "what is MongoDB", define it, list features 1-7, then give the CRUD example with output; for sharding, define it, draw the graph, give components, steps and advantages; for MongoDB vs Hadoop, draw the table and close with "MongoDB for real-time data, Hadoop for batch analytics".
Asked: [14 marks] (Jun 2025) What is MongoDB? Explain in brief key features of MongoDB. Show basic CRUD operations in MongoDB with proper example. Asked: [7 marks] (Dec 2020) What is MongoDB? Explain it in detail. Asked: [7 marks] (Nov 2022) Explain the process of sharding in MongoDB. Asked: [7 marks] (Jun 2025) Write difference between MongoDB and Hadoop.
Last-minute revision
- NoSQL means "Not Only SQL": schema-less, distributed, horizontally scalable, BASE.
- RDBMS is ACID, fixed schema, SQL and vertical scaling; NoSQL is the opposite on each.
- Four types: key-value (Redis), document (MongoDB), column-family (Cassandra, HBase), graph (Neo4j).
- CAP theorem: choose two of Consistency, Availability and Partition tolerance.
- Business drivers: volume, velocity, variety, competition, cost and personalisation.
- MongoDB stores BSON documents in collections and has a dynamic schema.
- Replica set: one primary and secondaries with automatic failover.
- Sharding parts: shards, mongos, config servers and shard key.
- CRUD: insertOne, find, updateOne with
$set, deleteOne. - MongoDB is for real-time operational data; Hadoop is for batch analytics.
Memory hooks
- "KDCG" for key-value, document, column, graph.
- CAP: pick any two of three.
- Sharding: mongos is the traffic police, config servers are the map, shards are the warehouses.
- CRUD: Create-Read-Update-Delete maps to insert-find-update-delete.
Coverage checklist
- Introduction to NoSQL: NoSQL short note (Nov 2023, Dec 2024) and NoSQL vs relational (Nov 2023).
- NoSQL Business Drivers: industry applications (Nov 2023, Dec 2024) and market and business drivers (Nov 2022).
- NoSQL Data architectural patterns: types of NoSQL (Dec 2024) and NoSQL definition, differences and types (Jun 2025).
- Variations of NOSQL architectural patterns using NoSQL to Manage Big Data: introduction and variations (Dec 2020) and NoSQL for Big Data (Nov 2022).
- Introduction to MangoDB: MongoDB features and CRUD (Jun 2025), MongoDB in detail (Dec 2020), sharding (Nov 2022) and MongoDB vs Hadoop (Jun 2025).