Skip to content
AD-801 · Big Data/Quick Revision Short Notes

Big Data (AD-801) - Unit 5 Short Notes

How unit 5 is examined

This unit covers social network mining (uses, fraud detection, difference from traditional mining), the social network as a graph (with Hive QL and YARN, asked with it), network types, community discovery and recommender systems. Most marks sit in the graph topic (14 marks) and the mining applications (two 7-mark questions).

Introduction and applications of social network mining

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Social network mining is the extraction of useful patterns, communities, influential users and trends from the relationships and interactions among people in online networks, by treating the network as a graph.</mark>

Key points.

  1. Traditional data mining works on independent records in flat tables, such as sales rows, and assumes each row is unrelated to the others.
  2. Social network mining works on linked data: the links between users carry as much meaning as the users' own attributes.
  3. Structure differs: traditional data is tabular, while social data is a graph of nodes (users) and edges (friend, follow, like, share).
  4. Scale differs: social networks hold billions of nodes and edges, so distributed tools such as Hadoop, Spark and graph engines are needed.
  5. Dynamics differ: a social graph changes every second as users join, post and connect, so analysis is often streaming and real-time, whereas traditional data is mostly static or batch.
  6. Techniques differ: traditional mining uses classification, clustering and association rules; social mining uses community detection, link analysis, centrality (PageRank) and link prediction.
  7. Data types differ: traditional mining handles numbers and categories; social mining handles text, images, timestamps and connections.
  8. Applications include friend and product recommendation, targeted advertising, influencer detection, trend and sentiment analysis, and fraud detection.
Basis Traditional data mining Social network mining
Data Independent records, tables Interlinked nodes and edges
Structure Flat, relational Graph
Scale Moderate Very large (billions of links)
Dynamics Mostly static, batch Constantly changing, real-time
Techniques Classification, clustering, association Community detection, link analysis, PageRank
Example Market-basket analysis Finding influencers on Twitter

Fraud and cyber-attack detection.

  1. Fraudsters behave differently in the graph: fake accounts form dense, tightly connected clusters with few links to genuine users, and anomaly detection flags such patterns.
  2. Community detection isolates whole rings of bot or scam accounts that act together.
  3. Link analysis follows suspicious connections, for example many accounts sharing one device, IP address or phone number.
  4. Pattern detection spots phishing: one account sending the same malicious link to many contacts in a short time.
  5. Real-time monitoring of the stream lets the platform block accounts, warn users and stop the spread of malware before it reaches many people.
  6. Examples are fake-account detection, phishing-link spread, credit-card fraud rings and spam campaigns.

Answer frame. For the difference question, open with the definition of both terms, draw the table above, then close with one example of each. For the fraud question, open with "Social networks are graphs, so fraud shows up as abnormal structure", develop points 1-5 in order, and close with fake-account and phishing examples.

Asked: [7 marks] (Jun 2025) What is social network mining, and how does it differ from traditional data mining? Asked: [7 marks] (Jun 2025) Describe the role of social network mining in detecting and preventing online fraud and cyber-attacks.

Social Networks as a Graph

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>A social network as a graph is a model in which each person or entity is a node (vertex) and each relationship or interaction, such as friendship, follow or message, is an edge, so that the network can be analysed with graph theory.</mark>

Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-01" viewBox="0 0 467 252" width="467" height="252" role="img" aria-label="Social graph. Nodes A to E are users; undirected edges are friendships, arrows are follows."><style>#dsfig-u5-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-01 .t{fill:#16181D;font-weight:500}#dsfig-u5-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-01 .dot{fill:#16181D}#dsfig-u5-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-01 .ah{fill:#454C5A}#dsfig-u5-01 .ah.hi{fill:#2340B8}#dsfig-u5-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-01 .e{stroke:#B1B7C3}html.dark #dsfig-u5-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-01 .t{fill:#E6E8ED}html.dark #dsfig-u5-01 .t.inv{fill:#0F1115}html.dark #dsfig-u5-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-01 .dot{fill:#E6E8ED}html.dark #dsfig-u5-01 .ann{fill:#8FA3FF}html.dark #dsfig-u5-01 .lbl{fill:#858D9C}html.dark #dsfig-u5-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-01 .ah{fill:#B1B7C3}html.dark #dsfig-u5-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh5" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.8,115.5 L153.2,50.5"/><path class="e" d="M55.8,136.5 L153.2,201.5"/><path class="e" d="M169,59 L169,193"/><path class="e" d="M184.8,50.5 L280.5,114.4" marker-end="url(#ah5)"/><path class="e" d="M184.8,201.5 L280.5,137.6" marker-end="url(#ah5)"/><path class="e" d="M317,126 L408,126"/><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">A</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">B</text><circle class="n" cx="169" cy="212" r="18"/><text class="t" x="169" y="212" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="298" cy="126" r="18"/><text class="t" x="298" y="126" dy=".35em" text-anchor="middle">D</text><circle class="n" cx="427" cy="126" r="18"/><text class="t" x="427" y="126" dy=".35em" text-anchor="middle">E</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Social graph. Nodes A to E are users; undirected edges are friendships, arrows are follows.</figcaption></figure>

Key points.

  1. Nodes represent users, pages or groups, and edges represent the relations between them.
  2. Edges may be undirected (Facebook friendship, mutual) or directed (Twitter follow, one-way).
  3. Edges may carry weights, such as the number of messages exchanged, to show tie strength.
  4. The degree of a node is its number of connections, and high-degree nodes are hubs or influencers.
  5. Dense parts of the graph form communities, and short paths between users explain how information spreads.
  6. Stored as an adjacency list or matrix, the graph is too large for one machine, so it is processed on Big Data platforms.
  7. Uses include recommendation, influence analysis, community finding and fraud detection.

ii) Hive Query Language (HiveQL). Hive is a data warehouse on Hadoop, and HiveQL is its SQL-like language with SELECT, WHERE, GROUP BY and JOIN. A query is converted into MapReduce (or Tez/Spark) jobs, so analysts query huge data stored in HDFS without writing Java, for example counting followers per user from an edges table.

iii) YARN (Yet Another Resource Negotiator). YARN is the resource-management layer of Hadoop 2. The global ResourceManager allocates cluster resources to applications, and a NodeManager on each node launches containers and reports usage. A per-application ApplicationMaster negotiates resources, which lets MapReduce, Spark and graph jobs share one cluster.

Answer frame. Write three headed parts (i, ii, iii) of about 4-5 marks each. For (i) open with the definition, draw the graph, then points 1-5. For (ii) and (iii) give the definition, the component names and one use in Big Data processing, and close by tying them together: the graph is stored in HDFS, queried with Hive and run on YARN.

Pitfall: Do not leave YARN without naming ResourceManager and NodeManager, or Hive without saying it is SQL-like; these are the marks examiners look for.

Asked: [14 marks] (Jun 2025) Explain the following term: i) Social Network as Graph ii) Hive query language iii) YARN

Types of social Networks

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Social networks are classed by their purpose and by the kind of ties they hold.

Key points.

  1. Social-connection networks such as Facebook and LinkedIn link people who know each other, personally or professionally.
  2. Media-sharing networks such as YouTube and Instagram are built around sharing photos and videos.
  3. Micro-blogging and discussion networks such as Twitter (X) and Reddit spread short posts and opinions through follow links.
  4. Messaging networks such as WhatsApp and Telegram connect small private groups.
  5. Structurally, networks are undirected (friendship) or directed (follow), and unweighted or weighted.

Clustering of social Graphs and direct discovery of communities

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Clustering a social graph groups nodes into communities that have many links inside the group and few links to other groups; direct discovery finds these communities from the graph structure alone.

Key points.

  1. Members of a community are more densely connected to each other than to outsiders, for example a college class or a fan group.
  2. Betweenness-based (Girvan-Newman) methods repeatedly remove the edges with highest betweenness, the ones that bridge groups, until communities separate.
  3. Modularity-based methods, such as Louvain, group nodes so that the fraction of internal edges exceeds what chance would give.
  4. Clique-based methods look for complete sub-graphs where everyone knows everyone.
  5. Communities are used for targeted advertising, recommendation and detecting fraud rings.

Introduction to recommender system

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>A recommender system is an information-filtering system that predicts the items a user is likely to prefer, such as movies, products or friends, and suggests them.</mark>

Key points.

  1. Collaborative filtering recommends items liked by users similar to you, for example "people who bought this also bought that"; it needs no item details but fails for new users and new items (cold start).
  2. Content-based filtering recommends items whose features resemble those you liked before, such as more action films; it needs no other users but gives narrow, repetitive suggestions.
  3. Hybrid systems combine both approaches to remove each one's weakness, as Netflix does.
  4. Knowledge-based systems use explicit user requirements and domain rules, for example choosing a laptop by budget and RAM, and suit rarely bought items.
  5. Applications are Amazon, Netflix, YouTube and friend suggestions on Facebook; common limitations are cold start, sparse data and privacy.

Answer frame. Open with the definition, take the four types in the order above with one example each, and close with applications and limitations.

Asked: [7 marks] (Jun 2025) Discuss the types of recommendation approaches commonly used in recommender systems.

Last-minute revision

  1. Social network mining extracts patterns and communities from linked user data treated as a graph.
  2. Traditional mining works on independent flat records; social mining works on linked, huge, fast-changing graphs.
  3. Fraud shows as dense clusters of fake accounts and repeated links spread in bursts.
  4. Techniques for fraud: anomaly detection, community detection, link analysis, real-time monitoring.
  5. Nodes are users; edges are relationships; edges can be directed or weighted.
  6. HiveQL is SQL-like and runs as MapReduce jobs over HDFS.
  7. YARN has a ResourceManager (cluster) and NodeManagers (each node).
  8. Girvan-Newman removes high-betweenness edges to find communities.
  9. Recommender types: collaborative, content-based, hybrid, knowledge-based.
  10. Cold start is the main weakness of collaborative filtering.

Memory hooks

  • Graph = "Nodes are Neighbours, Edges are Ties".
  • YARN: Resource Manager is the Manager, Node Managers are the Workers.
  • Recommenders: "CCHK" - Collaborative, Content, Hybrid, Knowledge.
  • Social mining differs in three S: Structure, Scale, Speed of change.

Coverage checklist

  • Introduction Applications of social Network mining: difference from traditional mining (Jun 2025); fraud and cyber-attack detection (Jun 2025).
  • Social Networks as a Graph: Social Network as Graph, Hive query language, YARN (Jun 2025).
  • Types of social Networks: no past question.
  • Clustering of social Graphs Direct Discovery of communities in a social graph: no past question.
  • Introduction to recommender system: recommendation approaches (Jun 2025).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in