How unit 4 is examined
This unit covers link-based ranking (hubs, authorities, PageRank, HITS), web relevance scoring, Hadoop MapReduce, evaluation and personalization, recommenders, invisible web with snippets and summaries, and QA with cross-lingual retrieval. No topic was asked in recent papers, so each is short but complete.
Link Analysis - hubs and authorities
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Link analysis ranks web pages using the hyperlink structure of the web, treating a link from page A to page B as a vote for B.</mark>
Key points.
- An authority is a page that many good hubs point to, because it holds trusted content on a topic.
- A hub is a page that points to many good authorities, such as a directory or resource list.
- The two scores reinforce each other: good hubs raise authorities, and good authorities raise hubs.
- In-links carry more weight than out-links, and links from important pages count more.
Page Rank and HITS algorithms
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>PageRank gives each page a query-independent importance equal to the chance that a random surfer is on it; HITS gives each page query-dependent hub and authority scores.</mark>
Formula. With damping $d$ (about 0.85), $N$ pages and $L(q)$ out-links of $q$: $$PR(p)=\frac{1-d}{N}+d\sum_{q\to p}\frac{PR(q)}{L(q)}$$ HITS: $a(p)=\sum_{q\to p}h(q)$ and $h(p)=\sum_{p\to q}a(q)$, normalised each round.
Key points.
- PageRank is computed offline by power iteration until the scores stop changing.
- The damping term models a surfer who jumps to a random page, so dead ends and rank sinks do not trap the score.
- HITS runs at query time on a small base set of pages fetched for the query and their neighbours.
- HITS is topic-sensitive but slow and easy to spam; PageRank is fast but ignores the query.
Relevance Scoring and ranking for Web - Similarity
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Web ranking combines a content similarity score, such as TF-IDF cosine similarity, with link-based scores like PageRank and other quality signals.</mark>
Key points.
- Content score: $\cos(q,d)=\frac{q\cdot d}{|q||d|}$ measures how close the query vector is to the document vector.
- Anchor text of in-links is added to the target page, since it describes the page in others' words.
- Extra signals include term position in title or headings, URL, freshness, click data and spam filters.
- The final score is a weighted mix of these features, often learned from data.
Hadoop & Map Reduce
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>MapReduce is a programming model that processes huge data on a cluster in two phases, map and reduce; Hadoop is its open-source implementation with the HDFS distributed file system.</mark>
Key points.
- Map reads input splits in parallel and emits key-value pairs; the framework then shuffles and groups them by key.
- Reduce combines all values of one key into the result, for example summing counts.
- For IR it builds inverted indexes: map emits (term, docID) and reduce forms each posting list.
- HDFS replicates blocks across nodes, so failed machines do not lose data or stop the job.
Evaluation - Personalized search
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Evaluation measures retrieval quality with precision and recall; personalized search re-ranks results using the individual user's profile and history.</mark>
Formula. $$P=\frac{|\text{relevant}\cap\text{retrieved}|}{|\text{retrieved}|},\quad R=\frac{|\text{relevant}\cap\text{retrieved}|}{|\text{relevant}|},\quad F=\frac{2PR}{P+R}$$
Key points.
- Precision is the fraction of retrieved documents that are relevant; recall is the fraction of relevant documents that were retrieved.
- Ranked lists are also judged by precision at k and mean average precision.
- Personalization builds a profile from past queries, clicks, location and bookmarks.
- It improves relevance for ambiguous queries, but raises privacy concerns and can trap users in a filter bubble.
Collaborative filtering and content-based recommendation of documents And products
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Collaborative filtering recommends items liked by users with similar tastes, while content-based recommendation suggests items whose features resemble those the user liked before.</mark>
Key points.
- Collaborative filtering uses the user-item rating matrix and finds neighbours by cosine or Pearson correlation.
- It needs no item content, but suffers from cold start and sparse ratings.
- Content-based methods compare item feature vectors, such as TF-IDF of text, with the user profile.
- They handle new items well, but recommend only similar items and lack surprise; hybrids combine both.
handling invisible Web - Snippet generation, Summarization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The invisible (deep) web is content that crawlers cannot reach, and a snippet is the short query-biased extract shown under each result.</mark>
Key points.
- Invisible content sits behind search forms, logins and scripts, so it is reached by submitting queries to forms or by using database feeds.
- A snippet is built by picking sentences containing the query terms, scored by term density and position, and highlighting them.
- Summarization is extractive, selecting key sentences, or abstractive, rewriting the text in new words.
- Sentences are ranked by term frequency, position and similarity to the title.
Question Answering, Cross-Lingual Retrieval
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Question answering returns a direct answer to a natural-language question, and cross-lingual retrieval finds documents in a language different from the query.</mark>
Key points.
- A QA pipeline has question analysis (finding the answer type), passage retrieval and answer extraction.
- Answers are ranked by how well the passage matches the question and the expected type.
- CLIR translates the query, the documents, or both into a common language or representation.
- Translation ambiguity and missing terms are the main problems, reduced by bilingual dictionaries and machine translation.
Last-minute revision
- Link analysis treats a hyperlink as a vote for the target page.
- Authority: pointed to by good hubs; hub: points to good authorities.
- PageRank: $PR(p)=\frac{1-d}{N}+d\sum PR(q)/L(q)$ with $d\approx0.85$.
- PageRank is query-independent and offline; HITS is query-dependent and online.
- Web ranking mixes cosine similarity, anchor text, PageRank and other signals.
- MapReduce has map, shuffle and reduce phases; HDFS stores replicated blocks.
- Precision = relevant retrieved / retrieved; recall = relevant retrieved / relevant.
- Collaborative filtering uses similar users; content-based uses item features.
- The invisible web lies behind forms and logins; snippets are query-biased.
- QA finds answer passages; CLIR translates the query or documents.
Memory hooks
- Hub = signpost pointing out; Authority = destination pointed at.
- PageRank = random surfer with a 15% teleport.
- Map splits, shuffle sorts, reduce sums.
- Collaborative = "people like you"; content-based = "things like this".
Coverage checklist
- Link Analysis -hubs and authorities: covered (no past questions).
- Page Rank and HITS algorithms: covered (no past questions).
- Relevance Scoring and ranking for Web - Similarity: covered (no past questions).
- Hadoop & Map Reduce: covered (no past questions).
- Evaluation - Personalized search: covered (no past questions).
- Collaborative filtering and content-based recommendation of documents And products: covered (no past questions).
- handling invisible Web - Snippet generation, Summarization: covered (no past questions).
- Question Answering, Cross-Lingual Retrieval: covered (no past questions).