Skip to content
CS-702 (D) · Big Data/Quick Revision Short Notes

Big Data (CS-702 (D)) - Unit 5 Short Notes

How unit 5 is examined

This unit covers what social network mining is and where it is used, how a social network is modelled as a graph, the types of networks, clustering to find communities, and recommender systems; social network mining and recommender systems carry the marks.

Social Network Mining: Introduction and Applications

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Social network mining is the process of extracting useful patterns, communities, influential people and relationships from social network data, which is stored as a graph of actors (nodes) joined by relations (edges).</mark>

Key points.

  1. Graph data means the data is a set of vertices (users, pages, authors) and edges (friend, follow, like, message, co-authorship), so the relationships matter as much as the individual records.
  2. Community detection groups nodes that are densely connected inside the group and sparsely connected to the rest, for example a college class or a fan club.
  3. Link prediction estimates which pair of nodes will connect next, which is how "people you may know" suggestions are produced.
  4. Influence analysis finds the most influential nodes using measures such as degree (number of connections) and centrality, so that a message reaches the most people.
  5. Frequent subgraph mining and link mining find repeated structures and classify or rank links and nodes.
  6. Applications include viral marketing, friend and product recommendation, fraud and spam detection, terrorist or criminal network analysis, epidemic tracking and opinion or sentiment analysis.
  7. Challenges are huge size (billions of nodes), constant change of the network, noisy or fake accounts, sparse data and privacy of users, so distributed tools such as Hadoop and Spark are used.

Most frequently used method. Community detection (clustering of the graph) is the method used most often on social graphs. Steps.

Step 1: Collect data and build the graph (users as nodes, relations as edges).
Step 2: Clean it: remove duplicate, fake and isolated nodes.
Step 3: Measure closeness: common neighbours, edge density or edge betweenness.
Step 4: Group the nodes so that links inside groups are dense and links between groups are few.
Step 5: Interpret each community and act on it (target ads, recommend friends).

Example. A telecom company builds a call graph, finds a community of customers who all call each other, sees that one member has left the network, and targets offers at the rest of that community to prevent churn.

Answer frame. Open with the definition of social network mining and graph data; draw a small graph with two communities joined by one edge; then develop points 1-5 (tasks) followed by 6 (applications); close with point 7 (challenges) and one line on why community detection is the most used method.

Asked: [14 marks] (Dec 2020, Nov 2023, Dec 2024, Jun 2025) What do you mean by social network mining? Write its applications. Also asked as: explain applications of social network mining and recommender system; explain social network as a graph, social network mining and recommender system. Asked: [7 marks] (Nov 2022) Which mining method is most frequently used for social network graph? Explain.

Social Networks as a Graph

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. ==A social network is modelled as a graph $G=(V,E)$ in which each vertex is an actor (person, page, organisation) and each edge is a relation between two actors.==

Key points.

  1. In an undirected graph an edge is mutual, such as a Facebook friendship, while in a directed graph an edge has a direction, such as "A follows B" on Twitter.
  2. In a weighted graph each edge carries a number, such as the count of messages exchanged or the strength of a tie.
  3. The degree of a node is the number of edges on it, and a node with high degree is a hub or influencer.
  4. A path, a connected component and clusters of dense edges are the graph structures that mining analyses.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u5-01" viewBox="0 0 467 252" width="467" height="252" role="img" aria-label="Actors A-E; A, B, C form a tight group; D to E is a directed follow edge"><style>#dsfig-u5-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u5-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u5-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u5-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u5-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u5-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u5-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u5-01 .t{fill:#16181D;font-weight:500}#dsfig-u5-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u5-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u5-01 .dot{fill:#16181D}#dsfig-u5-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u5-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u5-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u5-01 .ah{fill:#454C5A}#dsfig-u5-01 .ah.hi{fill:#2340B8}#dsfig-u5-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u5-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u5-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u5-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u5-01 .e{stroke:#B1B7C3}html.dark #dsfig-u5-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u5-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u5-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u5-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u5-01 .t{fill:#E6E8ED}html.dark #dsfig-u5-01 .t.inv{fill:#0F1115}html.dark #dsfig-u5-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u5-01 .dot{fill:#E6E8ED}html.dark #dsfig-u5-01 .ann{fill:#8FA3FF}html.dark #dsfig-u5-01 .lbl{fill:#858D9C}html.dark #dsfig-u5-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u5-01 .ah{fill:#B1B7C3}html.dark #dsfig-u5-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u5-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u5-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u5-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh9" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.8,115.5 L153.2,50.5"/><path class="e" d="M55.8,136.5 L153.2,201.5"/><path class="e" d="M169,59 L169,193"/><path class="e" d="M184.8,50.5 L282.2,115.5"/><path class="e" d="M184.8,201.5 L282.2,136.5"/><path class="e" d="M317,126 L406,126" marker-end="url(#ah9)"/><g class="wl"><rect x="331.8" y="117" width="61.5" height="18" rx="9"/><text class="t" x="362.5" y="126" dy=".35em" text-anchor="middle">follows</text></g><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">A</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">B</text><circle class="n" cx="169" cy="212" r="18"/><text class="t" x="169" y="212" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="298" cy="126" r="18"/><text class="t" x="298" y="126" dy=".35em" text-anchor="middle">D</text><circle class="n" cx="427" cy="126" r="18"/><text class="t" x="427" y="126" dy=".35em" text-anchor="middle">E</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Actors A-E; A, B, C form a tight group; D to E is a directed follow edge</figcaption></figure>

Asked: [7 marks] (Dec 2020) Explain the term "Social networks as a Graph". Write types of social networks. Asked: [14 marks] (Jun 2025) Short note: Social Network as a graph (any three of four).

Types of Social Networks

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. Social networks are classified by the kind of relation their edges represent.

Key points.

  1. Friendship or social-connection networks link people who know each other, for example Facebook and LinkedIn.
  2. Collaboration networks link people who work together, for example authors who co-write papers.
  3. Information networks link content, such as web pages joined by hyperlinks or posts joined by shares, and follower networks such as Twitter are directed.
  4. Communication networks link people by calls, emails or messages.

Clustering of Social Graphs and Direct Discovery of Communities

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Clustering in social network analysis is the grouping of nodes into clusters (communities) so that nodes inside a cluster are strongly connected to each other and weakly connected to nodes outside.</mark>

Key points.

  1. The purpose is to identify communities, such as friend circles or interest groups, so that behaviour of a group can be studied and targeted.
  2. Ordinary distance-based clustering does not fit well because a graph has edges rather than coordinates, so closeness is measured by shared neighbours or by edge betweenness.
  3. In the Girvan-Newman method (direct discovery by edge betweenness) the edge lying on the most shortest paths is removed repeatedly, and the connected pieces left behind are the communities.
  4. Other approaches are hierarchical clustering, modularity maximisation and finding cliques or dense subgraphs.
  5. Its significance lies in targeted marketing, detecting fake-account rings and understanding how information spreads.

Asked: [7 marks] (Nov 2022) What do you understand by clustering in Social network analysis?

Introduction to Recommender Systems

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>A recommender system is an information-filtering system that predicts the rating or preference a user would give to an item and suggests the items the user is most likely to like.</mark>

Key points.

  1. It solves information overload by showing each user a short personalised list instead of millions of items.
  2. Collaborative filtering recommends items liked by users with similar tastes, using only the user-item rating matrix, and can be user-based or item-based.
  3. Content-based filtering recommends items whose features (genre, author, keywords) resemble items the user liked earlier.
  4. Hybrid systems combine both to overcome the weakness of each, and Netflix and Amazon use them.
  5. Data comes from explicit feedback (ratings, reviews) and implicit feedback (clicks, watch time, purchases).
  6. Problems are the cold start of new users or items, sparse rating matrices, scalability and filter bubbles, and Big Data tools such as Hadoop and Spark handle the scale.
  7. Applications are Amazon products, Netflix movies, YouTube videos, Spotify songs and friend suggestions on social networks.

Example. Ravi rated A=5 and B=4. Sita rated A=5, B=4 and C=5. Ravi and Sita agree on A and B, so collaborative filtering recommends C to Ravi.

Basis Collaborative Content-based
Uses Ratings of similar users Features of items
Needs item details No Yes
New item Cold start Works
Diversity High Low

Answer frame. Open with the definition; draw the user-item rating table or a user-item graph; then develop points 1-4 (types) and 5-6 (data and problems); close with applications (point 7) and the Ravi-Sita example.

Asked: [14 marks] (Dec 2020) Short notes on any two: (i) Recommender system (ii) ETL processing (iii) Hadoop YARN. Asked: [14 marks] (Jun 2025) Short notes (any three): (i) Recommender system (ii) Social Network as a graph (iii) Traditional versus Big data (iv) Data types of Pig.

Last-minute revision

  • Social network mining extracts communities, influencers and links from a graph of actors and relations.
  • Graph $G=(V,E)$: vertices are actors and edges are relations.
  • Undirected means mutual friendship; directed means follow; weighted means strength of tie.
  • Three tasks: community detection, link prediction, influence analysis.
  • The most used method is community detection (clustering).
  • Girvan-Newman removes the edge with the highest betweenness to split communities.
  • Types: friendship, collaboration, information and communication networks.
  • Applications: marketing, recommendation, fraud detection, churn prediction.
  • Recommender types: collaborative, content-based and hybrid.
  • Cold start and sparsity are the main problems of recommenders.

Memory hooks

  • Nodes are Nice people, Edges are their Every relation.
  • Community = crowded inside, quiet outside.
  • Three tasks spell CLI: Community, Link, Influence.
  • Collaborative = "people like you"; content-based = "items like this".

Coverage checklist

  • Mining social Network Graphs: Introduction Applications of social Network mining: Dec 2020, Nov 2023, Dec 2024, Jun 2025 (social network mining and applications); Nov 2022 (most used method).
  • Social Networks as a Graph: Dec 2020 (social networks as a graph, types); Jun 2025 (short note).
  • Types of social Networks: not asked; covered with Dec 2020.
  • Clustering of social Graphs Direct Discovery of communities in a social graph: Nov 2022.
  • Introduction to recommender system: Dec 2020 and Jun 2025 short notes; recommender part of the 14-mark explain question.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in