How unit 1 is examined
This unit covers what NLP is, where it is used, why it is hard, the levels of language analysis, the three classic retrieval models (Boolean, vector, probabilistic) and the classical NLP approaches (rule-based, statistical, IR, rule-based MT, graphical models). No question was asked in the supplied papers, so every topic is short but complete.
Natural Language Processing (NLP): Definition and scope
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Natural Language Processing is the branch of artificial intelligence and linguistics that enables computers to understand, interpret and generate human language, spoken or written.</mark>
Key points.
- NLP has two halves: Natural Language Understanding (NLU) maps text to meaning, and Natural Language Generation (NLG) produces text from meaning.
- Its scope covers text and speech, from tokenization and parsing up to translation, question answering and dialogue.
- It draws on linguistics, computer science, machine learning and logic.
- Its aim is to let people talk to machines in ordinary language instead of formal commands.
Applications in various domains
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>NLP applications are the practical systems that use language technology to solve problems in a domain.</mark>
Key points.
- Web and search: search engines, spell checking, autocomplete and question answering.
- Communication: machine translation, chatbots, voice assistants and speech recognition.
- Business and social media: sentiment analysis, opinion mining, spam filtering and text summarization.
- Healthcare, legal and education: clinical note extraction, contract analysis, automated essay grading and grammar tools.
Challenges and limitations
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The central challenge of NLP is ambiguity: one piece of text can have several valid interpretations, and machines lack the world knowledge people use to choose.</mark>
Key points.
- Lexical ambiguity: one word has many senses ("bank" is a river edge or a lender).
- Syntactic ambiguity: one sentence has many parses ("I saw the man with a telescope").
- Semantic and pragmatic ambiguity: meaning depends on context, sarcasm, idiom, and reference such as "it" or "he".
- Other limits: many languages and dialects, spelling variation, missing labelled data for low-resource languages, and bias learned from training data.
NLP tasks in syntax, semantics, and pragmatics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>NLP tasks are grouped by level of language analysis: syntax studies structure, semantics studies literal meaning, and pragmatics studies meaning in context.</mark>
Key points.
- Syntax tasks: tokenization, POS tagging, stemming, and parsing a sentence into a parse tree.
- Semantics tasks: word sense disambiguation, named entity recognition, semantic role labelling and meaning representation.
- Pragmatics tasks: coreference resolution, discourse analysis, intent and speech-act recognition, and sarcasm detection.
- Each level feeds the next, so errors in syntax carry into semantics and pragmatics.
Boolean Model
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The Boolean model is an information retrieval model in which a document is a set of index terms and a query is a Boolean expression using AND, OR and NOT; a document is either relevant or not.</mark>
Key points.
- A term is present (1) or absent (0) in a document, and the query is evaluated with set logic, so the answer is exact match.
- Example: the query "nlp AND parsing AND NOT speech" returns only documents containing nlp and parsing but not speech.
- Advantages: simple, fast, and gives the user precise control.
- Drawbacks: no ranking, no partial match, no term weights, and hard queries for ordinary users.
Vector model
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The vector space model represents each document and query as a vector of weighted terms and ranks documents by their similarity to the query.</mark>
Formula. Term weight is TF-IDF, and similarity is the cosine of the angle between the vectors:
$$w_{t,d}=tf_{t,d}\times\log\frac{N}{df_t},\qquad \cos(q,d)=\frac{q\cdot d}{|q|\,|d|}$$
Key points.
- $tf$ is how often term $t$ occurs in document $d$, $N$ is the number of documents and $df_t$ is how many contain $t$.
- Rare terms get a high IDF, so they count more than common words.
- Documents are ranked by cosine score, giving partial matching.
- It ignores word order and assumes terms are independent.
Probabilistic Model
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The probabilistic model ranks documents by the estimated probability that a document is relevant to the query, using the probability ranking principle.</mark>
Formula. Documents are ordered by $P(R\mid d,q)$; BM25 is a popular practical form:
$$score(d,q)=\sum_{t\in q} IDF_t\cdot\frac{tf\,(k_1+1)}{tf+k_1\left(1-b+b\frac{|d|}{avgdl}\right)}$$
Key points.
- The model needs estimates of term occurrence in relevant and non-relevant documents, refined by user feedback.
- BM25 saturates term frequency and normalizes by document length.
- It ranks documents and needs no arbitrary term weights.
- It still assumes term independence.
Comparison of classical NLP models: Rule-based model
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A rule-based model processes language with hand-written linguistic rules, such as grammar rules, dictionaries and pattern matching.</mark>
Key points.
- Experts write the rules, for example "a determiner followed by a noun forms a noun phrase".
- It is transparent and easy to debug, and needs no training data.
- It is costly to build, brittle on unseen input, and hard to scale to all exceptions of a language.
- Compared with statistical models: rules are precise but rigid, statistics are robust but need data.
Statistical model
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A statistical model learns language patterns from large text corpora and uses probabilities to choose the most likely analysis.</mark>
Key points.
- It counts events in a corpus, for example the n-gram estimate $P(w_n\mid w_{n-1})=\dfrac{C(w_{n-1}w_n)}{C(w_{n-1})}$.
- Examples are n-gram language models, Naive Bayes classifiers and Hidden Markov Models for tagging.
- It handles ambiguity by picking the highest-probability choice and degrades gracefully on noisy input.
- It needs large data and smoothing for unseen events, and its decisions are less interpretable.
Information retrieval model
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>An information retrieval model is a formal framework that describes how documents and queries are represented and how relevance between them is computed.</mark>
Key points.
- Its parts are the document representation, the query representation and a matching or ranking function.
- The classic families are the Boolean, vector space and probabilistic models.
- The pipeline is: collect documents, index terms, process the query, rank, and return results.
- Quality is measured by precision (share of retrieved documents that are relevant) and recall (share of relevant documents that are retrieved).
Rule-based machine translation model
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Rule-based machine translation (RBMT) translates text using bilingual dictionaries and hand-written grammar rules for the source and target languages.</mark>
Key points.
- The three approaches are direct (word by word), transfer (parse the source, apply transfer rules, generate the target) and interlingua (source to a language-neutral meaning to target).
- Stages are analysis, transfer and generation.
- Output is grammatically controlled and predictable, and it works without parallel corpora.
- It needs huge manual effort, handles idioms and ambiguity poorly, and is largely replaced by statistical and neural MT.
Probabilistic Graphical model
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A probabilistic graphical model (PGM) represents random variables as nodes and their dependencies as edges, so a joint probability is factorized into simple local terms.</mark>
<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u1-01" viewBox="0 0 252 209" width="252" height="209" role="img" aria-label="Bayesian network, A and B are parents of C, so P(A,B,C)=P(A)P(B)P(C|A,B)"><style>#dsfig-u1-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u1-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u1-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u1-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u1-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u1-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u1-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u1-01 .t{fill:#16181D;font-weight:500}#dsfig-u1-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u1-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u1-01 .dot{fill:#16181D}#dsfig-u1-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u1-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u1-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u1-01 .ah{fill:#454C5A}#dsfig-u1-01 .ah.hi{fill:#2340B8}#dsfig-u1-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u1-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u1-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u1-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u1-01 .e{stroke:#B1B7C3}html.dark #dsfig-u1-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u1-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u1-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u1-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u1-01 .t{fill:#E6E8ED}html.dark #dsfig-u1-01 .t.inv{fill:#0F1115}html.dark #dsfig-u1-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u1-01 .dot{fill:#E6E8ED}html.dark #dsfig-u1-01 .ann{fill:#8FA3FF}html.dark #dsfig-u1-01 .lbl{fill:#858D9C}html.dark #dsfig-u1-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u1-01 .ah{fill:#B1B7C3}html.dark #dsfig-u1-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u1-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u1-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u1-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M50.5,55.8 L114.4,151.5" marker-end="url(#ah1)"/><path class="e" d="M201.5,55.8 L137.6,151.5" marker-end="url(#ah1)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">A</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">B</text><circle class="n" cx="126" cy="169" r="18"/><text class="t" x="126" y="169" dy=".35em" text-anchor="middle">C</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Bayesian network, A and B are parents of C, so P(A,B,C)=P(A)P(B)P(C|A,B)</figcaption></figure>
Key points.
- Directed PGMs are Bayesian networks, with factorization $P(X_1..X_n)=\prod_i P(X_i\mid Parents(X_i))$.
- Undirected PGMs are Markov random fields.
- In NLP, Hidden Markov Models (tagging) and Conditional Random Fields (NER) are PGMs.
- They model uncertainty and dependencies compactly, but inference can be expensive.
Last-minute revision
- NLP = NLU (understanding) plus NLG (generation).
- Ambiguity types: lexical, syntactic, semantic, pragmatic.
- Levels: syntax (structure), semantics (meaning), pragmatics (context).
- Boolean model: exact match with AND, OR, NOT; no ranking.
- Vector model: TF-IDF weights and cosine similarity.
- $w=tf\times\log(N/df)$.
- Probabilistic model ranks by probability of relevance; BM25 is the practical form.
- Rule-based: precise but brittle; statistical: robust but data-hungry.
- IR quality: precision and recall.
- RBMT approaches: direct, transfer, interlingua.
- PGM: nodes are variables, edges are dependencies; Bayesian network is directed.
Memory hooks
- Boolean = "yes or no": in or out, never ranked.
- TF-IDF = "frequent here, rare elsewhere" is important.
- Vaquois idea for MT: direct at the bottom, transfer in the middle, interlingua at the top.
- Rules = teacher's grammar book; statistics = reading millions of pages.
- Ambiguity has four faces: word, sentence, meaning, context.
Coverage checklist
- Natural Language Processing(NLP): Definition and scope: definition, NLU and NLG, scope (no past question).
- Applications in various domains: search, translation, chatbots, sentiment, healthcare (no past question).
- Challenges and limitations: ambiguity types, low-resource languages, bias (no past question).
- NLP tasks in syntax, semantics, and pragmatics: task list per level (no past question).
- Boolean Model: exact match, pros and cons (no past question).
- Vector model: TF-IDF and cosine (no past question).
- Probabilistic Model: probability ranking, BM25 (no past question).
- Comparison of classical NLP models: Rule-based model: rules versus statistics (no past question).
- Statistical model: corpus counts, n-gram estimate (no past question).
- Information retrieval model: components, precision and recall (no past question).
- Rule-based machine translation model: direct, transfer, interlingua (no past question).
- Probabilistic Graphical model: Bayesian network and factorization (no past question).