I. Introduction to Natural Language Processing
Definition and Scope:
Natural Language Processing (NLP) is a subfield of Artificial Intelligence (AI) and Computational Linguistics focused on enabling computers to understand, interpret, manipulate, and generate human language in a valuable way. Its scope bridges linguistics and computer science, covering tasks from low-level text processing to high-level semantic understanding.
Need and Importance:
-
Volume of Text Data: Exponential growth of digital text (social media, documents, web).
-
Human-Computer Interaction: Enables intuitive interfaces (chatbots, voice assistants).
-
Automation: Streamlines tasks like document classification, translation, and information retrieval.
-
Insight Extraction: Analyzes unstructured text for business intelligence, research, and sentiment.
Broad Classes of NLP Applications:
-
Information Extraction (IE): Automatically extracting structured information (e.g., entities, relationships) from unstructured text.
-
Machine Translation (MT): Automatic translation of text/speech between languages (e.g., Google Translate).
-
Sentiment Analysis: Determining the emotional tone or opinion expressed in text (positive/negative/neutral).
-
Question Answering (QA) Systems: Providing precise answers to natural language questions (e.g., IBM Watson, Siri).
Commercial Uses (Examples):
-
Customer Support Chatbots: Automate responses, reduce operational costs.
-
Sentiment Analysis for Brand Monitoring: Track public opinion on social media.
-
Grammar & Style Checkers: Tools like Grammarly, Microsoft Editor.
[!TIP]
Exam Focus: Be prepared to define NLP, list its need (3–4 points), name four application classes with brief examples, and cite three commercial uses.
II. Text Preprocessing in NLP
Steps in Text Preprocessing:
-
Tokenization: Splitting text into basic units (words, subwords, or sentences).
-
Word Segmentation: Dividing a continuous string of characters into words (critical for languages without spaces like Chinese).
-
Stemming: Crudely chopping word endings to get a root form (e.g., "running" → "run").
-
Lemmatization: Using vocabulary/morphology to get the dictionary form (e.g., "better" → "good").
-
Stop Word Removal: Eliminating common, low-meaning words (e.g., "the", "is", "and").
-
Case Folding: Converting all text to lowercase (or uppercase) for uniformity.
Detailed Discussion:
Preprocessing cleans and standardizes raw text, reducing noise and dimensionality. Stemming is faster but less accurate; lemmatization is slower but linguistically informed. Stop words removal depends on task (often kept in sentiment analysis). Tokenization varies by language: whitespace-based for English, but requires specialized algorithms for Chinese/Japanese.
Difference: Word Segmentation vs. Tokenization
| Aspect | Tokenization | Word Segmentation |
|---|---|---|
| Definition | Breaking text into tokens (words, symbols). | Specifically splitting character streams into words. |
| Languages | Straightforward for space-delimited (English). | Essential for languages without spaces (Chinese, Thai). |
| Complexity | Lower (often regex-based). | Higher (requires statistical/ML models). |
| Example | "Hello, world!" → ["Hello", ",", "world", "!"] |
"我爱自然语言处理" → ["我", "爱", "自然语言处理"] |
[!TIP]
Common Pitfall: Students often confuse these. Remember: Tokenization is the broader task; Word Segmentation is a specific, challenging subtask for non-space languages.
III. Corpora in NLP
Introduction to Corpora:
A corpus (plural: corpora) is a large, structured collection of digital text used for linguistic analysis and NLP model training.
Types of Corpora:
-
Monolingual Corpus: Text in a single language (e.g., British National Corpus).
-
Parallel Corpus: Same text in multiple languages, aligned sentence-by-sentence (e.g., EUROPARL).
-
Annotated Corpus: Text enriched with linguistic labels (POS tags, parse trees, named entities). Examples: Penn Treebank (POS+parse), CoNLL-2003 (NER).
Significance of Corpora Analysis:
-
For Model Training: Provides empirical data to learn statistical patterns (e.g., word frequencies, transition probabilities in HMMs).
-
For Linguistic Research: Enables study of language usage, variation, and change (e.g., corpus linguistics).
-
Benchmarking: Standard corpora allow fair comparison of NLP systems (e.g., GLUE benchmark).
[!TIP]
Exam Focus: Know definitions of corpus types and at least two significance points for each (training & research).
IV. Morphology
Basics of Morphology:
-
Morpheme: Smallest meaningful unit (e.g., "un-", "break", "-able").
-
Inflection: Modifying a word to express grammar (tense, number) without changing core meaning (e.g., "walk" → "walked").
-
Derivation: Creating a new word with changed meaning/part-of-speech (e.g., "happy" → "happiness").
Morphology of Indian Languages:
-
Agglutinative: Words formed by stringing morphemes (e.g., Telugu, Tamil).
-
Inflectional: Rich morphology for case, gender, tense (e.g., Sanskrit, Marathi).
-
Script: Primarily abugida (e.g., Devanagari) where consonants have inherent vowels.
-
Challenges: Complex sandhi (sound changes), compound words, high OOV (out-of-vocabulary) rates.
Finite State Automata (FSA) in Morphology:
-
Relationship: FSA models the finite set of morphological rules in a language. Each state represents a morpheme position; transitions represent affixation or inflection rules.
-
Finite State Morphological Analyzers: Use FSA/transducers to generate all valid word forms from a root or analyze a surface form into root+features. Efficient for agglutinative languages.
[!TIP]
Key Point: FSA provides a compact, computationally efficient representation of morphological knowledge, crucial for languages with rich morphology.
V. Part-of-Speech (POS) Tagging
Definition and Example:
Assigning a grammatical category (noun, verb, adjective, etc.) to each word in a sentence.
Example: "The/DET quick/ADJ fox/NOUN jumps/VERB over/ADP the/DET lazy/ADJ dog/NOUN ."
Models for POS Tagging:
-
Maximum Entropy Model (MaxEnt):
-
Principle: Choose the probability distribution with maximum entropy (least biased) subject to constraints from training data (feature expectations).
-
Application: Features can include word itself, prefixes/suffixes, surrounding words, capitalization. Model estimates $P(tag | features)$.
-
Formula: Maximize $$\displaystyle H(p) = -\sum_{x,t} p(x,t) \log p(x,t) $$ subject to $$\displaystyle \sum_{x,t} p(x,t) f_i(x,t) = E_{\text{train}}(f_i) $$.
-
-
Transformation-Based Learning (TBL):
-
Algorithm:
-
Start with a baseline tagger (e.g., assign most frequent tag).
-
Error-driven: Learn transformation rules (if condition then change tag) that fix tagging errors in training data.
-
Apply rules in order of highest accuracy gain.
-
-
Advantage: Captures contextual rules, interpretable.
-
Example Rule:
If previous tag is 'DET' and current word ends with 'ly' → change current tag to 'ADV'.
-
[!TIP]
Comparison: MaxEnt is probabilistic and feature-based; TBL is rule-based and error-driven. Both handle context better than simple HMMs.
VI. Parsing
What is Parsing?
Syntactic analysis of a sentence to determine its grammatical structure, typically represented as a Parse Tree.
Types of Parsers:
| Synthetic Parser (Rule-Based) | Statistical Parser (Data-Driven) |
|---|---|
| Uses hand-crafted grammar rules (e.g., CFG). | Learns probabilities from annotated corpora (e.g., Treebank). |
| Precision high for covered constructs, but brittle (fails on unseen patterns). | Robust to variation, but requires large training data. |
| Example: Chart parsers with grammar rules. | Example: Probabilistic Context-Free Grammar (PCFG) parsers, Neural parsers. |
[!TIP]
Modern Trend: Statistical/Neural parsers dominate due to scalability and performance on real-world text.
VII. Semantic Analysis
Need for Semantic Analysis:
To move beyond syntax to meaning—understand word senses, relationships, and discourse coherence. Essential for tasks like QA, MT, and information retrieval.
Bootstrapping Methods in Semantic Analysis:
-
Self-Training: Use a model trained on limited labeled data to pseudo-label unlabeled data, then retrain.
-
Co-Training: Train two models on different feature views; each labels data for the other, iteratively.
Word Sense Disambiguation (WSD):
-
Goal: Determine the correct sense of an ambiguous word in context.
-
Approaches:
-
Supervised: Train on sense-annotated data (e.g., using SVM, MaxEnt).
-
Unsupervised: Cluster word contexts to induce senses (e.g., LDA).
-
Knowledge-Based: Use lexical resources (WordNet) to measure relatedness (e.g., Lesk algorithm).
-
[!TIP]
WSD Challenge: "Bank" (river vs. financial) – supervised needs labeled data; knowledge-based relies on resource coverage.
VIII. Discourse Analysis
Anaphora Resolution:
-
Task: Link pronouns (anaphora) to their antecedents (e.g., "John arrived. He was tired." → "He" → "John").
-
Techniques:
-
Rule-Based: Use syntactic patterns (e.g., proximity, gender agreement).
-
Machine Learning: Features: distance, number agreement, semantic compatibility.
-
-
Challenges: Ambiguity, non-local references, bridging anaphora.
Named Entity Resolution (NER) vs. Anaphora Resolution:
| Named Entity Recognition (NER) | Anaphora Resolution |
|---|---|
| Identifies and classifies entities (PERSON, LOCATION, etc.) in text. | Links pronouns/nouns to previously mentioned entities. |
| Local to a sentence/phrase. | Cross-sentential (discourse-level). |
Output: [ORG Apple] [LOC Cupertino] |
Output: "It" → [ORG Apple] |
[!TIP]
Key Difference: NER finds entities; Anaphora Resolution connects references to entities.
IX. Phonology and Phonetics in NLP
Phonological Rules:
-
Transform underlying phonemic representations into surface phonetic forms (e.g., assimilation, deletion).
-
Significance: Crucial for speech recognition (acoustic model mapping) and text-to-speech (pronunciation generation). Also aids in handling spelling variations.
Minimum Edit Distance (Levenshtein Distance):
-
Definition: Minimum number of single-character operations (insertion, deletion, substitution) to transform one string into another.
-
Formula: Computed via dynamic programming.
Let $D[i,j]$ = distance between first $i$ chars of string A and first $j$ chars of B.
$$D[i,j] = \min \begin{cases} D[i-1,j] + 1 \text{ (deletion)} \\ D[i,j-1] + 1 \text{ (insertion)} \\ D[i-1,j-1] + \text{cost} \text{ (substitution)} \end{cases}$$
where $$\displaystyle \text{cost}=0 $$ if $$\displaystyle A[i]=B[j] $$, else $1$.
- Applications: Spell checking, speech recognition (aligning recognized vs. reference text), DNA sequence alignment.
Bayesian Method of Pronunciation:
-
Models pronunciation as a probabilistic mapping from graphemes (letters) to phonemes.
-
Uses Bayes' rule: $$\displaystyle P(\text{phonemes}|\text{spelling}) \propto P(\text{spelling}|\text{phonemes}) P(\text{phonemes}) $$.
-
Application: Pronunciation modeling for OOV words in speech synthesis/recognition, especially for irregular spellings (English).
[!TIP]
Edit Distance: Remember recurrence and base case $$\displaystyle D[0,j]=j $$, $$\displaystyle D[i,0]=i $$. Bayesian pronunciation uses generative model of spelling given phonemes.
X. Models and Algorithms in NLP
Overview of Key Models:
| Category | Models | Use Cases |
|---|---|---|
| Probabilistic | HMM, Bayesian Networks | POS tagging, language modeling |
| Machine Learning | Maximum Entropy, CRF, SVM | Sequence labeling, classification |
| Finite State | FSA, Finite State Transducers (FST) | Morphology, tokenization, chunking |
Algorithmic Approaches:
-
Rule-Based: Hand-crafted linguistic rules. High precision, low recall, maintenance heavy.
-
Statistical: Learn from data (n-grams, PCFGs). Data-hungry but robust.
-
Machine Learning: Use features + algorithms (MaxEnt, CRF). Balance between rule and pure stats.
Evaluation Metrics for NLP Models:
-
Accuracy: $$\displaystyle \frac{\text{Correct}}{\text{Total}} $$ (for classification).
-
Precision, Recall, F1-Score: For tasks like NER, parsing.
$$\text{Precision} = \frac{TP}{TP+FP},\quad \text{Recall} = \frac{TP}{TP+FN},\quad F1 = 2 \cdot \frac{P \cdot R}{P+R}$$
-
BLEU: For machine translation (n-gram overlap).
-
Perplexity: For language models (lower is better).
-
Parsing: Exact Match, Labeled Attachment Score (LAS).
[!TIP]
Metric Choice: Use F1 for imbalanced tasks (NER), BLEU for MT, perplexity for LM. Always consider task-specific metrics.