Skip to content
AD-603 (C) · Information Retrieval/Quick Revision Short Notes

Information Retrieval (AD-603 (C)) - Unit 1 Short Notes

How unit 1 is examined

This unit covers what information retrieval is, its history and components, open source search frameworks, the web's impact, AI's role and IR versus web search; no topic has been asked recently, so each is kept short but complete.

Introduction, History of IR, Components of IR, Issues

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Information retrieval (IR) is finding material of an unstructured nature, usually text documents, that satisfies an information need from within large collections.</mark>

Key points.

  1. History: manual library catalogues led to Luhn's automatic indexing (1950s), the SMART system and the Cranfield tests (1960s), then web search engines in the 1990s.
  2. Components: document collection, indexer, query interface, retrieval model that ranks documents, and evaluation by precision and recall.
  3. Issues: vague queries, ambiguity of language, relevance being subjective, scale of data and evaluation cost.

Open source Search engine Frameworks

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>An open source search framework is freely available software that indexes documents and answers ranked queries, so an application need not build retrieval from scratch.</mark>

Key points.

  1. Apache Lucene is the core Java library that builds inverted indexes and ranks with TF-IDF or BM25.
  2. Solr and Elasticsearch are servers built on Lucene that add REST APIs, faceting and distributed search.
  3. Other examples are Terrier, Indri and Xapian, often used in research.

The Impact of the web on IR

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>The web turned IR from searching small curated collections into searching a huge, open, changing and linked collection used by everyone.</mark>

Key points.

  1. Scale grew enormously, so indexing and crawling had to be distributed.
  2. Content is heterogeneous, unedited, multilingual and often duplicated or spam.
  3. Hyperlinks give a new relevance signal, used by PageRank and HITS.
  4. Users are non-experts with short queries, so ranking quality and speed matter most.

The role of artificial intelligence (AI) in IR

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>AI in IR means using machine learning and natural language processing to understand queries and documents and rank results better.</mark>

Key points.

  1. Machine-learned ranking (learning to rank) combines many signals into one score.
  2. NLP helps with query understanding, spelling correction, synonyms and entity recognition.
  3. Neural models and embeddings match meaning rather than exact words.
  4. Recommendation, question answering and conversational assistants build on the same techniques.

IR Versus Web Search, Components of a search engine, characterizing the web

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Web search is IR applied to the web, where the collection is huge, dynamic, distributed and linked, so a crawler is needed and link analysis helps ranking.</mark>

Key points.

  1. Classic IR has a small, static, controlled collection; web search has billions of changing pages under no central control.
  2. Search engine components are the crawler, indexer, query processor and ranker, with the user interface on top.
  3. The web is characterized by a power-law link structure, large duplication, spam and rapid change.

Last-minute revision

  • IR finds relevant unstructured documents for an information need.
  • Milestones: Luhn, SMART and Cranfield, then web engines.
  • Components: collection, indexer, query interface, ranking model, evaluation.
  • Lucene is the core library; Solr and Elasticsearch build on it.
  • The web brings scale, links, spam and non-expert users.
  • Links give PageRank and HITS.
  • AI adds learning to rank, NLP and neural embeddings.
  • Search engine parts: crawler, indexer, query processor, ranker.

Memory hooks

  • CIQRE: Collection, Indexer, Query, Ranking, Evaluation.
  • Lucene is the engine; Solr and Elasticsearch are the car.
  • Crawl, Index, Query, Rank: the search engine order.

Coverage checklist

  • Introduction - History of IR- Components of IR - Issues: no past questions.
  • Open source Search engine Frameworks: no past questions.
  • The Impact of the web on IR: no past questions.
  • The role of artificial intelligence (AI) in IR: no past questions.
  • IR Versus Web Search - Components of a search engine, characterizing the web: no past questions.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in