Skip to content
AD-603 (C) · Information Retrieval/Quick Revision Short Notes

Information Retrieval (AD-603 (C)) - Unit 4 Short Notes

How unit 4 is examined

This unit covers link-based ranking (hubs, authorities, PageRank, HITS), web relevance scoring, Hadoop MapReduce, evaluation and personalization, recommenders, invisible web with snippets and summaries, and QA with cross-lingual retrieval. No topic was asked in recent papers, so each is short but complete.

Link Analysis - hubs and authorities

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Link analysis ranks web pages using the hyperlink structure of the web, treating a link from page A to page B as a vote for B.</mark>

Key points.

  1. An authority is a page that many good hubs point to, because it holds trusted content on a topic.
  2. A hub is a page that points to many good authorities, such as a directory or resource list.
  3. The two scores reinforce each other: good hubs raise authorities, and good authorities raise hubs.
  4. In-links carry more weight than out-links, and links from important pages count more.

Page Rank and HITS algorithms

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>PageRank gives each page a query-independent importance equal to the chance that a random surfer is on it; HITS gives each page query-dependent hub and authority scores.</mark>

Formula. With damping $d$ (about 0.85), $N$ pages and $L(q)$ out-links of $q$: $$PR(p)=\frac{1-d}{N}+d\sum_{q\to p}\frac{PR(q)}{L(q)}$$ HITS: $a(p)=\sum_{q\to p}h(q)$ and $h(p)=\sum_{p\to q}a(q)$, normalised each round.

Key points.

  1. PageRank is computed offline by power iteration until the scores stop changing.
  2. The damping term models a surfer who jumps to a random page, so dead ends and rank sinks do not trap the score.
  3. HITS runs at query time on a small base set of pages fetched for the query and their neighbours.
  4. HITS is topic-sensitive but slow and easy to spam; PageRank is fast but ignores the query.

Relevance Scoring and ranking for Web - Similarity

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Web ranking combines a content similarity score, such as TF-IDF cosine similarity, with link-based scores like PageRank and other quality signals.</mark>

Key points.

  1. Content score: $\cos(q,d)=\frac{q\cdot d}{|q||d|}$ measures how close the query vector is to the document vector.
  2. Anchor text of in-links is added to the target page, since it describes the page in others' words.
  3. Extra signals include term position in title or headings, URL, freshness, click data and spam filters.
  4. The final score is a weighted mix of these features, often learned from data.

Hadoop & Map Reduce

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>MapReduce is a programming model that processes huge data on a cluster in two phases, map and reduce; Hadoop is its open-source implementation with the HDFS distributed file system.</mark>

Key points.

  1. Map reads input splits in parallel and emits key-value pairs; the framework then shuffles and groups them by key.
  2. Reduce combines all values of one key into the result, for example summing counts.
  3. For IR it builds inverted indexes: map emits (term, docID) and reduce forms each posting list.
  4. HDFS replicates blocks across nodes, so failed machines do not lose data or stop the job.

Evaluation - Personalized search

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Evaluation measures retrieval quality with precision and recall; personalized search re-ranks results using the individual user's profile and history.</mark>

Formula. $$P=\frac{|\text{relevant}\cap\text{retrieved}|}{|\text{retrieved}|},\quad R=\frac{|\text{relevant}\cap\text{retrieved}|}{|\text{relevant}|},\quad F=\frac{2PR}{P+R}$$

Key points.

  1. Precision is the fraction of retrieved documents that are relevant; recall is the fraction of relevant documents that were retrieved.
  2. Ranked lists are also judged by precision at k and mean average precision.
  3. Personalization builds a profile from past queries, clicks, location and bookmarks.
  4. It improves relevance for ambiguous queries, but raises privacy concerns and can trap users in a filter bubble.

Collaborative filtering and content-based recommendation of documents And products

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Collaborative filtering recommends items liked by users with similar tastes, while content-based recommendation suggests items whose features resemble those the user liked before.</mark>

Key points.

  1. Collaborative filtering uses the user-item rating matrix and finds neighbours by cosine or Pearson correlation.
  2. It needs no item content, but suffers from cold start and sparse ratings.
  3. Content-based methods compare item feature vectors, such as TF-IDF of text, with the user profile.
  4. They handle new items well, but recommend only similar items and lack surprise; hybrids combine both.

handling invisible Web - Snippet generation, Summarization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>The invisible (deep) web is content that crawlers cannot reach, and a snippet is the short query-biased extract shown under each result.</mark>

Key points.

  1. Invisible content sits behind search forms, logins and scripts, so it is reached by submitting queries to forms or by using database feeds.
  2. A snippet is built by picking sentences containing the query terms, scored by term density and position, and highlighting them.
  3. Summarization is extractive, selecting key sentences, or abstractive, rewriting the text in new words.
  4. Sentences are ranked by term frequency, position and similarity to the title.

Question Answering, Cross-Lingual Retrieval

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Question answering returns a direct answer to a natural-language question, and cross-lingual retrieval finds documents in a language different from the query.</mark>

Key points.

  1. A QA pipeline has question analysis (finding the answer type), passage retrieval and answer extraction.
  2. Answers are ranked by how well the passage matches the question and the expected type.
  3. CLIR translates the query, the documents, or both into a common language or representation.
  4. Translation ambiguity and missing terms are the main problems, reduced by bilingual dictionaries and machine translation.

Last-minute revision

  • Link analysis treats a hyperlink as a vote for the target page.
  • Authority: pointed to by good hubs; hub: points to good authorities.
  • PageRank: $PR(p)=\frac{1-d}{N}+d\sum PR(q)/L(q)$ with $d\approx0.85$.
  • PageRank is query-independent and offline; HITS is query-dependent and online.
  • Web ranking mixes cosine similarity, anchor text, PageRank and other signals.
  • MapReduce has map, shuffle and reduce phases; HDFS stores replicated blocks.
  • Precision = relevant retrieved / retrieved; recall = relevant retrieved / relevant.
  • Collaborative filtering uses similar users; content-based uses item features.
  • The invisible web lies behind forms and logins; snippets are query-biased.
  • QA finds answer passages; CLIR translates the query or documents.

Memory hooks

  • Hub = signpost pointing out; Authority = destination pointed at.
  • PageRank = random surfer with a 15% teleport.
  • Map splits, shuffle sorts, reduce sums.
  • Collaborative = "people like you"; content-based = "things like this".

Coverage checklist

  • Link Analysis -hubs and authorities: covered (no past questions).
  • Page Rank and HITS algorithms: covered (no past questions).
  • Relevance Scoring and ranking for Web - Similarity: covered (no past questions).
  • Hadoop & Map Reduce: covered (no past questions).
  • Evaluation - Personalized search: covered (no past questions).
  • Collaborative filtering and content-based recommendation of documents And products: covered (no past questions).
  • handling invisible Web - Snippet generation, Summarization: covered (no past questions).
  • Question Answering, Cross-Lingual Retrieval: covered (no past questions).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in