Skip to content
IT-802 (B) · Natural Language Processing/Quick Revision Short Notes

Natural Language Processing (IT-802 (B)) - Unit 4 Short Notes

UNIT 4: Natural Language Processing (NLP) – Core Concepts, Techniques, and Applications


1. Introduction to Natural Language Processing

Definition: NLP is a field of AI and linguistics focused on enabling computers to understand, interpret, manipulate, and generate human language in a valuable way.

Need & Importance:

  • Challenges: Ambiguity (lexical, syntactic, semantic), variability (spelling, grammar), context-dependence, and non-standard language (slang, typos).

  • Opportunities: Automate text/speech processing, extract insights from unstructured data, enable human-computer interaction.

Broad Classes of Applications:

  • Information Retrieval (IR): Search engines, document ranking.

  • Machine Translation (MT): Translating text/speech between languages (e.g., Google Translate).

  • Sentiment Analysis: Determining opinion polarity (positive/negative) from text.

  • Question Answering (QA): Systems like IBM Watson that answer natural language questions.

Commercial Uses:

  • Chatbots & virtual assistants (Siri, Alexa).

  • Social media monitoring & brand sentiment.

  • Grammar/style checkers (Grammarly).

  • Automated customer support.

[!TIP] Exam Focus: Always link challenges (ambiguity, variability) to the need for NLP. List 4 applications with brief examples.


2. Text Pre-processing and Normalization

Role: First step in NLP pipeline; converts raw text into a clean, normalized format suitable for analysis. Reduces noise and standardizes input.

Tokenization:

  • Splitting text into tokens (words, sentences, subwords).

  • Challenges: Handling punctuation, contractions ("don't" → "do", "n't"), emoticons, and language-specific issues (e.g., Chinese/Japanese lack spaces).

Word Segmentation vs. Tokenization:

Aspect Tokenization Word Segmentation
Definition General splitting into tokens Specifically splitting continuous text into words (for languages without explicit word boundaries).
Languages Applicable to all (space-delimited or not) Crucial for Chinese, Japanese, Thai.
Example "I love NLP!" → ["I", "love", "NLP", "!"] Chinese: "我爱自然语言处理" → ["我", "爱", "自然语言处理"]

Additional Pre-processing Steps:

  • Case Folding: Convert to lowercase (except for proper nouns sometimes).

  • Stop-word Removal: Remove frequent, low-meaning words (the, is, and).

  • Stemming: Crude heuristic chopping (e.g., "running" → "run").

  • Lemmatization: Vocabulary-based normalization using POS (e.g., "better" → "good").

[!TIP] Common Pitfall: Stemming can produce non-words ("argue" → "argu"). Lemmatization is slower but more accurate.


3. Morphological Analysis

Fundamentals:

  • Morpheme: Smallest meaningful unit (e.g., "un-", "break", "-able").

  • Inflection: Grammatical variants (walk, walks, walked).

  • Derivation: Creating new words (happy → unhappiness).

  • Compounding: Combining words (bookstore, smartphone).

Morphology of Indian Languages:

  • Often agglutinative (Tamil, Telugu) or inflectional (Sanskrit, Hindi).

  • Rich case markings, verb conjugations, compound words.

  • NLP Challenges: High morphological complexity, large number of inflected forms, resource scarcity.

Finite State Automata (FSA) & Transducers (FST):

  • FSA: Accepts/rejects strings (recognizer). States represent morphological conditions.

  • FST: Maps input string to output string (transducer). Used for morphological analysis/generation.

  • Relationship: FSTs model morphology by encoding rules (e.g., pluralization: (stem) + "s" → plural). They efficiently handle regular morphological processes.

[!TIP] Exam Key: FST = FSA with output. Example: English plural FST: input "cat" → output "cats".


4. Corpora and Corpus Linguistics

What is a Corpus? A structured, machine-readable collection of authentic texts (spoken/written) used for linguistic analysis and NLP training.

Types:

  • Monolingual: Single language (e.g., British National Corpus).

  • Parallel: Translations aligned (e.g., Europarl).

  • Annotated: With linguistic labels (POS, parse trees, named entities).

  • Spoken: Transcribed speech (e.g., Switchboard).

Significance of Corpora Analysis:

  • Model Training: Provide data for statistical models (n-grams, HMMs).

  • Evaluation: Benchmark system performance (e.g., parsing accuracy on Penn Treebank).

  • Linguistic Research: Study language usage, frequency, variation.

Corpus Creation & Annotation:

  1. Collection: Sampling (balanced by genre, domain).

  2. Annotation: Adding linguistic tags (manual or automatic with quality control).

  3. Quality Assessment: Inter-annotator agreement (Kappa), consistency, coverage.

Use in NLP Systems: Corpora are the foundation for supervised learning (taggers, parsers, MT systems).


5. Part-of-Speech (POS) Tagging

Purpose: Assign each word a syntactic category (noun, verb, adjective) based on context.

Tag Sets:

  • Penn Treebank: 45 tags (e.g., NN for singular noun, VBZ for 3rd person verb).

  • Indian Languages: Often based on EAGLES or custom tags (e.g., NNP for proper noun in Hindi).

Tagging Approaches:

Approach Method Pros & Cons
Rule-based Hand-crafted linguistic rules High precision, low recall, labor-intensive.
Stochastic n-gram models, HMMs (probability-based) Data-driven, but sparse data issues.
Maximum Entropy Features + max entropy principle Flexible features, robust, needs training.
Transformation-Based (TBL) Error-driven rule learning Fast, interpretable rules, order matters.

Maximum Entropy Model:

  • Features: Binary functions capturing context (e.g., isPreviousWordThe, wordEndsWith“ing”, nextTagIsNoun).

  • Training: Maximize entropy subject to constraints from training data (using GIS or L-BFGS).

  • Advantages: Can incorporate diverse features without independence assumptions.

Transformation-Based Tagging (TBL):

  1. Start with baseline (e.g., assign most frequent tag).

  2. Learn transformation rules (if condition → change tag) from errors.

  3. Apply rules in order until no improvement.

  • Example Rule: Change tag from NN to VB if previous word is "to".

[!TIP] Exam Distinction: MaxEnt uses probabilistic model with features; TBL uses symbolic rules learned from errors.


6. Syntactic Analysis (Parsing)

What is Parsing? Analyzing sentence structure to derive a syntactic representation.

Goals:

  • Constituency (Phrase Structure): Hierarchical grouping into phrases (NP, VP). Output: parse tree.

  • Dependency: Head-dependent relationships. Output: dependency graph.

Synthetic (Rule-based) vs. Statistical Parsers:

Aspect Synthetic Parsers Statistical Parsers
Method Hand-written grammar rules (e.g., CFG) Probabilistic models trained on corpora
Strengths Linguistically interpretable Robust to noise, higher accuracy
Weaknesses Brittle, cannot handle ambiguity well Requires large annotated data

Parsing Algorithms:

  • CYK: For CFG in Chomsky Normal Form; $$\displaystyle O(n^3) $$ time.

  • Earley: For any CFG; handles left-recursion efficiently.

  • Shift-Reduce: Linear time; uses stack (shift) and reduce actions.

Evaluation Metrics:

  • Precision: % of produced constituents that are correct.

  • Recall: % of gold constituents that are produced.

  • PARSEVAL: F1-score (harmonic mean of precision/recall) on labeled/bracketed trees.


7. Semantic Analysis

Need: Syntax alone insufficient for meaning; e.g., "Time flies like an arrow" vs. "Fruit flies like a banana".

Bootstrapping Methods:

  • Start with seed examples (e.g., known synonyms).

  • Use patterns to extract new examples (e.g., "X and Y" → X,Y may be similar).

  • Iteratively expand lexicon/ontology (e.g., Snowball algorithm).

Word Sense Disambiguation (WSD):

  • Task: Determine correct sense of ambiguous word in context.

  • Approaches:

    • Knowledge-based: Use dictionaries/ontologies (e.g., Lesk algorithm – overlap in dictionary definitions).

    • Supervised: Train classifier on sense-tagged data (features: context words, POS).

    • Unsupervised: Cluster contexts (e.g., WSD using word embeddings).

Semantic Role Labeling (SRL): Identify thematic roles (Agent, Patient, Instrument) in a predicate-argument structure.


8. Pragmatic and Discourse Processing

Anaphora Resolution:

  • Definition: Linking anaphor (pronoun, definite NP) to its antecedent.

  • Types: Pronominal ("he", "she"), definite noun phrase ("the company").

  • Methods:

    • Centering Theory: Track discourse center (most salient entity).

    • ML-based: Features: gender/number agreement, distance, syntactic role.

Named Entity Recognition (NER):

  • Definition: Identify and classify named entities (PERSON, ORGANIZATION, LOCATION, DATE, etc.).

  • Techniques:

    • Rule-based (gazetteers, patterns).

    • ML-based (CRFs, BiLSTM-CRF) using contextual features.

Differences:

Aspect Anaphora Resolution Named Entity Resolution
Goal Link reference to antecedent Identify entity mentions & types
Context Needed Discourse-level (multi-sentence) Often local (sentence/paragraph)
Challenges Coreference across sentences, ambiguity Ambiguous names ("Apple" as fruit/company)

9. Phonological Processing and Edit Distance

Phonological Rules:

  • Describe sound changes in a language (e.g., plural "s" pronounced /s/ after voiceless, /z/ after voiced).

  • Significance: Crucial for pronunciation modeling in speech recognition, text-to-speech.

Minimum Edit Distance (Levenshtein Distance):

  • Definition: Minimum number of operations (insert, delete, substitute) to transform string A into B.

  • Dynamic Programming Recurrence:

$$ d[i,j] = \min \left\{ \begin{array}{ll} d[i-1,j] + 1 & \text{(deletion)} \\ d[i,j-1] + 1 & \text{(insertion)} \\ d[i-1,j-1] + \text{cost} & \text{(substitution; cost=0 if match)} \end{array} \right. $$

where $d[i,j]$ = distance between first $i$ chars of A and first $j$ chars of B.

Applications:

  • Spelling correction (suggestions within edit distance).

  • Speech recognition (aligning phoneme sequences).

  • String matching / DNA sequence alignment.

[!TIP] Exam Calculation: Always set up matrix with strings on axes; fill row/column-wise.


10. Probabilistic and Bayesian Methods in NLP

Bayesian Inference:

  • Bayes' Theorem:

$$ P(A|B) = \frac{P(B|A) P(A)}{P(B)} $$

  • $P(A)$: Prior probability.

  • $P(A|B)$: Posterior probability (updated belief after evidence B).

  • $P(B|A)$: Likelihood.

Bayesian Method for Pronunciation Modeling:

  • Model pronunciation variants probabilistically (e.g., "tomato" as /təˈmɑːtoʊ/ vs /təˈmeɪtoʊ/).

  • Use Bayesian networks to combine acoustic, linguistic, and lexical probabilities in speech recognition.

Other Bayesian Applications:

  • Naïve Bayes for Text Classification: Assumes feature independence.

$$ P(\text{class}|\text{doc}) \propto P(\text{class}) \prod_{i} P(\text{word}_i|\text{class}) $$

  • Bayesian Networks: Graphical models representing conditional dependencies (e.g., for WSD using context features).

11. Overview of NLP Models and Algorithms

Model Type Examples & Use Cases Key Idea
Finite State FSA (morphology checking), FST (pluralization) State transitions with symbols.
Probabilistic n-grams (language modeling), HMMs (POS tagging) Sequence probabilities, hidden states.
Maximum Entropy POS tagging, NER Maximize entropy subject to feature constraints.
Transformation-Based (TBL) POS tagging, chunking Learn error-correcting rules.
Statistical Parsing PCFGs, dependency parsers Probabilistic context-free grammars.
WSD Models Knowledge-based (Lesk), supervised (SVM) Contextual similarity or classification.

Evaluation:

  • Intrinsic: Measure component performance (e.g., tagging accuracy).

  • Extrinsic: Measure impact on end task (e.g., MT quality improvement).


12. Advanced Topics and Current Trends (Contextual from Questions)

Integration of Machine Learning in NLP:

  • Shift from rule-based to data-driven approaches (e.g., neural networks replacing hand-crafted rules).

Ensemble Methods & Model Combination:

  • Bagging/Boosting: Combine multiple models (e.g., parsers) to improve robustness and accuracy.

Dimensionality Reduction for NLP Features:

  • PCA: Linear projection to reduce feature space (e.g., for word vectors).

  • LLE: Nonlinear reduction, preserves local manifold structure (useful for semantic spaces).

Ethical Considerations & Bias:

  • Bias in Training Data: Models perpetuate societal biases (gender, race).

  • Fairness: Ensure equitable performance across demographics.

  • Transparency: Interpretability of black-box models (e.g., transformers).

[!TIP] Exam Relevance: Advanced topics may be asked in "Discuss models and algorithms" – mention neural trends (word2vec, transformers) briefly if context allows.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in