How unit 1 is examined
This unit covers what information retrieval is, its history and components, open source search frameworks, the web's impact, AI's role and IR versus web search; no topic has been asked recently, so each is kept short but complete.
Introduction, History of IR, Components of IR, Issues
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Information retrieval (IR) is finding material of an unstructured nature, usually text documents, that satisfies an information need from within large collections.</mark>
Key points.
- History: manual library catalogues led to Luhn's automatic indexing (1950s), the SMART system and the Cranfield tests (1960s), then web search engines in the 1990s.
- Components: document collection, indexer, query interface, retrieval model that ranks documents, and evaluation by precision and recall.
- Issues: vague queries, ambiguity of language, relevance being subjective, scale of data and evaluation cost.
Open source Search engine Frameworks
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>An open source search framework is freely available software that indexes documents and answers ranked queries, so an application need not build retrieval from scratch.</mark>
Key points.
- Apache Lucene is the core Java library that builds inverted indexes and ranks with TF-IDF or BM25.
- Solr and Elasticsearch are servers built on Lucene that add REST APIs, faceting and distributed search.
- Other examples are Terrier, Indri and Xapian, often used in research.
The Impact of the web on IR
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The web turned IR from searching small curated collections into searching a huge, open, changing and linked collection used by everyone.</mark>
Key points.
- Scale grew enormously, so indexing and crawling had to be distributed.
- Content is heterogeneous, unedited, multilingual and often duplicated or spam.
- Hyperlinks give a new relevance signal, used by PageRank and HITS.
- Users are non-experts with short queries, so ranking quality and speed matter most.
The role of artificial intelligence (AI) in IR
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>AI in IR means using machine learning and natural language processing to understand queries and documents and rank results better.</mark>
Key points.
- Machine-learned ranking (learning to rank) combines many signals into one score.
- NLP helps with query understanding, spelling correction, synonyms and entity recognition.
- Neural models and embeddings match meaning rather than exact words.
- Recommendation, question answering and conversational assistants build on the same techniques.
IR Versus Web Search, Components of a search engine, characterizing the web
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Web search is IR applied to the web, where the collection is huge, dynamic, distributed and linked, so a crawler is needed and link analysis helps ranking.</mark>
Key points.
- Classic IR has a small, static, controlled collection; web search has billions of changing pages under no central control.
- Search engine components are the crawler, indexer, query processor and ranker, with the user interface on top.
- The web is characterized by a power-law link structure, large duplication, spam and rapid change.
Last-minute revision
- IR finds relevant unstructured documents for an information need.
- Milestones: Luhn, SMART and Cranfield, then web engines.
- Components: collection, indexer, query interface, ranking model, evaluation.
- Lucene is the core library; Solr and Elasticsearch build on it.
- The web brings scale, links, spam and non-expert users.
- Links give PageRank and HITS.
- AI adds learning to rank, NLP and neural embeddings.
- Search engine parts: crawler, indexer, query processor, ranker.
Memory hooks
- CIQRE: Collection, Indexer, Query, Ranking, Evaluation.
- Lucene is the engine; Solr and Elasticsearch are the car.
- Crawl, Index, Query, Rank: the search engine order.
Coverage checklist
- Introduction - History of IR- Components of IR - Issues: no past questions.
- Open source Search engine Frameworks: no past questions.
- The Impact of the web on IR: no past questions.
- The role of artificial intelligence (AI) in IR: no past questions.
- IR Versus Web Search - Components of a search engine, characterizing the web: no past questions.