UNIT 4: Natural Language Processing (NLP) – Core Concepts, Techniques, and Applications
1. Introduction to Natural Language Processing
Definition: NLP is a field of AI and linguistics focused on enabling computers to understand, interpret, manipulate, and generate human language in a valuable way.
Need & Importance:
-
Challenges: Ambiguity (lexical, syntactic, semantic), variability (spelling, grammar), context-dependence, and non-standard language (slang, typos).
-
Opportunities: Automate text/speech processing, extract insights from unstructured data, enable human-computer interaction.
Broad Classes of Applications:
-
Information Retrieval (IR): Search engines, document ranking.
-
Machine Translation (MT): Translating text/speech between languages (e.g., Google Translate).
-
Sentiment Analysis: Determining opinion polarity (positive/negative) from text.
-
Question Answering (QA): Systems like IBM Watson that answer natural language questions.
Commercial Uses:
-
Chatbots & virtual assistants (Siri, Alexa).
-
Social media monitoring & brand sentiment.
-
Grammar/style checkers (Grammarly).
-
Automated customer support.
[!TIP] Exam Focus: Always link challenges (ambiguity, variability) to the need for NLP. List 4 applications with brief examples.
2. Text Pre-processing and Normalization
Role: First step in NLP pipeline; converts raw text into a clean, normalized format suitable for analysis. Reduces noise and standardizes input.
Tokenization:
-
Splitting text into tokens (words, sentences, subwords).
-
Challenges: Handling punctuation, contractions ("don't" → "do", "n't"), emoticons, and language-specific issues (e.g., Chinese/Japanese lack spaces).
Word Segmentation vs. Tokenization:
| Aspect | Tokenization | Word Segmentation |
|---|---|---|
| Definition | General splitting into tokens | Specifically splitting continuous text into words (for languages without explicit word boundaries). |
| Languages | Applicable to all (space-delimited or not) | Crucial for Chinese, Japanese, Thai. |
| Example | "I love NLP!" → ["I", "love", "NLP", "!"] | Chinese: "我爱自然语言处理" → ["我", "爱", "自然语言处理"] |
Additional Pre-processing Steps:
-
Case Folding: Convert to lowercase (except for proper nouns sometimes).
-
Stop-word Removal: Remove frequent, low-meaning words (the, is, and).
-
Stemming: Crude heuristic chopping (e.g., "running" → "run").
-
Lemmatization: Vocabulary-based normalization using POS (e.g., "better" → "good").
[!TIP] Common Pitfall: Stemming can produce non-words ("argue" → "argu"). Lemmatization is slower but more accurate.
3. Morphological Analysis
Fundamentals:
-
Morpheme: Smallest meaningful unit (e.g., "un-", "break", "-able").
-
Inflection: Grammatical variants (walk, walks, walked).
-
Derivation: Creating new words (happy → unhappiness).
-
Compounding: Combining words (bookstore, smartphone).
Morphology of Indian Languages:
-
Often agglutinative (Tamil, Telugu) or inflectional (Sanskrit, Hindi).
-
Rich case markings, verb conjugations, compound words.
-
NLP Challenges: High morphological complexity, large number of inflected forms, resource scarcity.
Finite State Automata (FSA) & Transducers (FST):
-
FSA: Accepts/rejects strings (recognizer). States represent morphological conditions.
-
FST: Maps input string to output string (transducer). Used for morphological analysis/generation.
-
Relationship: FSTs model morphology by encoding rules (e.g., pluralization:
(stem) + "s" → plural). They efficiently handle regular morphological processes.
[!TIP] Exam Key: FST = FSA with output. Example: English plural FST: input "cat" → output "cats".
4. Corpora and Corpus Linguistics
What is a Corpus? A structured, machine-readable collection of authentic texts (spoken/written) used for linguistic analysis and NLP training.
Types:
-
Monolingual: Single language (e.g., British National Corpus).
-
Parallel: Translations aligned (e.g., Europarl).
-
Annotated: With linguistic labels (POS, parse trees, named entities).
-
Spoken: Transcribed speech (e.g., Switchboard).
Significance of Corpora Analysis:
-
Model Training: Provide data for statistical models (n-grams, HMMs).
-
Evaluation: Benchmark system performance (e.g., parsing accuracy on Penn Treebank).
-
Linguistic Research: Study language usage, frequency, variation.
Corpus Creation & Annotation:
-
Collection: Sampling (balanced by genre, domain).
-
Annotation: Adding linguistic tags (manual or automatic with quality control).
-
Quality Assessment: Inter-annotator agreement (Kappa), consistency, coverage.
Use in NLP Systems: Corpora are the foundation for supervised learning (taggers, parsers, MT systems).
5. Part-of-Speech (POS) Tagging
Purpose: Assign each word a syntactic category (noun, verb, adjective) based on context.
Tag Sets:
-
Penn Treebank: 45 tags (e.g.,
NNfor singular noun,VBZfor 3rd person verb). -
Indian Languages: Often based on EAGLES or custom tags (e.g.,
NNPfor proper noun in Hindi).
Tagging Approaches:
| Approach | Method | Pros & Cons |
|---|---|---|
| Rule-based | Hand-crafted linguistic rules | High precision, low recall, labor-intensive. |
| Stochastic | n-gram models, HMMs (probability-based) | Data-driven, but sparse data issues. |
| Maximum Entropy | Features + max entropy principle | Flexible features, robust, needs training. |
| Transformation-Based (TBL) | Error-driven rule learning | Fast, interpretable rules, order matters. |
Maximum Entropy Model:
-
Features: Binary functions capturing context (e.g.,
isPreviousWordThe,wordEndsWith“ing”,nextTagIsNoun). -
Training: Maximize entropy subject to constraints from training data (using GIS or L-BFGS).
-
Advantages: Can incorporate diverse features without independence assumptions.
Transformation-Based Tagging (TBL):
-
Start with baseline (e.g., assign most frequent tag).
-
Learn transformation rules (if condition → change tag) from errors.
-
Apply rules in order until no improvement.
- Example Rule:
Change tag from NN to VB if previous word is "to".
[!TIP] Exam Distinction: MaxEnt uses probabilistic model with features; TBL uses symbolic rules learned from errors.
6. Syntactic Analysis (Parsing)
What is Parsing? Analyzing sentence structure to derive a syntactic representation.
Goals:
-
Constituency (Phrase Structure): Hierarchical grouping into phrases (NP, VP). Output: parse tree.
-
Dependency: Head-dependent relationships. Output: dependency graph.
Synthetic (Rule-based) vs. Statistical Parsers:
| Aspect | Synthetic Parsers | Statistical Parsers |
|---|---|---|
| Method | Hand-written grammar rules (e.g., CFG) | Probabilistic models trained on corpora |
| Strengths | Linguistically interpretable | Robust to noise, higher accuracy |
| Weaknesses | Brittle, cannot handle ambiguity well | Requires large annotated data |
Parsing Algorithms:
-
CYK: For CFG in Chomsky Normal Form; $$\displaystyle O(n^3) $$ time.
-
Earley: For any CFG; handles left-recursion efficiently.
-
Shift-Reduce: Linear time; uses stack (shift) and reduce actions.
Evaluation Metrics:
-
Precision: % of produced constituents that are correct.
-
Recall: % of gold constituents that are produced.
-
PARSEVAL: F1-score (harmonic mean of precision/recall) on labeled/bracketed trees.
7. Semantic Analysis
Need: Syntax alone insufficient for meaning; e.g., "Time flies like an arrow" vs. "Fruit flies like a banana".
Bootstrapping Methods:
-
Start with seed examples (e.g., known synonyms).
-
Use patterns to extract new examples (e.g., "X and Y" → X,Y may be similar).
-
Iteratively expand lexicon/ontology (e.g., Snowball algorithm).
Word Sense Disambiguation (WSD):
-
Task: Determine correct sense of ambiguous word in context.
-
Approaches:
-
Knowledge-based: Use dictionaries/ontologies (e.g., Lesk algorithm – overlap in dictionary definitions).
-
Supervised: Train classifier on sense-tagged data (features: context words, POS).
-
Unsupervised: Cluster contexts (e.g., WSD using word embeddings).
-
Semantic Role Labeling (SRL): Identify thematic roles (Agent, Patient, Instrument) in a predicate-argument structure.
8. Pragmatic and Discourse Processing
Anaphora Resolution:
-
Definition: Linking anaphor (pronoun, definite NP) to its antecedent.
-
Types: Pronominal ("he", "she"), definite noun phrase ("the company").
-
Methods:
-
Centering Theory: Track discourse center (most salient entity).
-
ML-based: Features: gender/number agreement, distance, syntactic role.
-
Named Entity Recognition (NER):
-
Definition: Identify and classify named entities (PERSON, ORGANIZATION, LOCATION, DATE, etc.).
-
Techniques:
-
Rule-based (gazetteers, patterns).
-
ML-based (CRFs, BiLSTM-CRF) using contextual features.
-
Differences:
| Aspect | Anaphora Resolution | Named Entity Resolution |
|---|---|---|
| Goal | Link reference to antecedent | Identify entity mentions & types |
| Context Needed | Discourse-level (multi-sentence) | Often local (sentence/paragraph) |
| Challenges | Coreference across sentences, ambiguity | Ambiguous names ("Apple" as fruit/company) |
9. Phonological Processing and Edit Distance
Phonological Rules:
-
Describe sound changes in a language (e.g., plural "s" pronounced /s/ after voiceless, /z/ after voiced).
-
Significance: Crucial for pronunciation modeling in speech recognition, text-to-speech.
Minimum Edit Distance (Levenshtein Distance):
-
Definition: Minimum number of operations (insert, delete, substitute) to transform string A into B.
-
Dynamic Programming Recurrence:
$$ d[i,j] = \min \left\{ \begin{array}{ll} d[i-1,j] + 1 & \text{(deletion)} \\ d[i,j-1] + 1 & \text{(insertion)} \\ d[i-1,j-1] + \text{cost} & \text{(substitution; cost=0 if match)} \end{array} \right. $$
where $d[i,j]$ = distance between first $i$ chars of A and first $j$ chars of B.
Applications:
-
Spelling correction (suggestions within edit distance).
-
Speech recognition (aligning phoneme sequences).
-
String matching / DNA sequence alignment.
[!TIP] Exam Calculation: Always set up matrix with strings on axes; fill row/column-wise.
10. Probabilistic and Bayesian Methods in NLP
Bayesian Inference:
- Bayes' Theorem:
$$ P(A|B) = \frac{P(B|A) P(A)}{P(B)} $$
-
$P(A)$: Prior probability.
-
$P(A|B)$: Posterior probability (updated belief after evidence B).
-
$P(B|A)$: Likelihood.
Bayesian Method for Pronunciation Modeling:
-
Model pronunciation variants probabilistically (e.g., "tomato" as /təˈmɑːtoʊ/ vs /təˈmeɪtoʊ/).
-
Use Bayesian networks to combine acoustic, linguistic, and lexical probabilities in speech recognition.
Other Bayesian Applications:
- Naïve Bayes for Text Classification: Assumes feature independence.
$$ P(\text{class}|\text{doc}) \propto P(\text{class}) \prod_{i} P(\text{word}_i|\text{class}) $$
- Bayesian Networks: Graphical models representing conditional dependencies (e.g., for WSD using context features).
11. Overview of NLP Models and Algorithms
| Model Type | Examples & Use Cases | Key Idea |
|---|---|---|
| Finite State | FSA (morphology checking), FST (pluralization) | State transitions with symbols. |
| Probabilistic | n-grams (language modeling), HMMs (POS tagging) | Sequence probabilities, hidden states. |
| Maximum Entropy | POS tagging, NER | Maximize entropy subject to feature constraints. |
| Transformation-Based (TBL) | POS tagging, chunking | Learn error-correcting rules. |
| Statistical Parsing | PCFGs, dependency parsers | Probabilistic context-free grammars. |
| WSD Models | Knowledge-based (Lesk), supervised (SVM) | Contextual similarity or classification. |
Evaluation:
-
Intrinsic: Measure component performance (e.g., tagging accuracy).
-
Extrinsic: Measure impact on end task (e.g., MT quality improvement).
12. Advanced Topics and Current Trends (Contextual from Questions)
Integration of Machine Learning in NLP:
- Shift from rule-based to data-driven approaches (e.g., neural networks replacing hand-crafted rules).
Ensemble Methods & Model Combination:
- Bagging/Boosting: Combine multiple models (e.g., parsers) to improve robustness and accuracy.
Dimensionality Reduction for NLP Features:
-
PCA: Linear projection to reduce feature space (e.g., for word vectors).
-
LLE: Nonlinear reduction, preserves local manifold structure (useful for semantic spaces).
Ethical Considerations & Bias:
-
Bias in Training Data: Models perpetuate societal biases (gender, race).
-
Fairness: Ensure equitable performance across demographics.
-
Transparency: Interpretability of black-box models (e.g., transformers).
[!TIP] Exam Relevance: Advanced topics may be asked in "Discuss models and algorithms" – mention neural trends (word2vec, transformers) briefly if context allows.