Skip to content
IT-802 (B) · Natural Language Processing/Quick Revision Short Notes

Natural Language Processing (IT-802 (B)) - Unit 2 Short Notes

UNIT 2: Natural Language Processing Core Concepts


I. Introduction to NLP

Definition and Purpose

Natural Language Processing (NLP) is a subfield of artificial intelligence and computational linguistics that enables computers to understand, interpret, manipulate, and generate human language. Its interdisciplinary nature combines computer science, linguistics, and AI.

Need for NLP

  • Human-Computer Interaction (HCI): Bridge the gap between natural human communication and rigid computer commands.

  • Data Explosion: Process and extract insights from massive volumes of unstructured text data (social media, emails, documents).

  • Automation: Automate language-intensive tasks (e.g., customer support, document summarization, translation).

Broad Classes of Applications (4 Examples)

  1. Information Retrieval (IR) & Extraction (IE): Search engines, extracting structured data from text.

  2. Machine Translation (MT): Automatic translation between languages (e.g., Google Translate).

  3. Sentiment Analysis / Opinion Mining: Determining sentiment (positive/negative/neutral) from text.

  4. Question Answering (QA) & Dialogue Systems: Systems like Siri, Alexa, or customer service chatbots.

Commercial Uses (3 Examples)

  • Customer Support Chatbots: Automate responses to common queries.

  • Grammar & Style Checkers: Tools like Grammarly, Microsoft Editor.

  • Voice Assistants: Siri, Google Assistant, Alexa for voice-based interaction.

  • Social Media Monitoring Tools: Brand sentiment tracking, trend analysis.

[!TIP] Exam Focus: Be prepared to list and briefly explain 4 application classes and 3 commercial uses with real-world examples.


II. Text Pre-processing

Importance and Goals

To reduce noise, standardize text, and convert raw, unstructured text into a format suitable for downstream NLP tasks. It improves accuracy and efficiency of models.

Key Steps:

  1. Tokenization: Splitting text into basic units (tokens: words, numbers, symbols).

    • Basic Approach: Split on whitespace and punctuation.

    • Challenges: Handling contractions ("don't" → "do", "n't"), hyphenated words, emoticons (😊), and URLs.

  2. Word Segmentation: Splitting a string of characters into words for languages without explicit word boundaries (e.g., Chinese, Japanese, Thai, Indian languages like Telugu, Tamil).

    • Contrast with Tokenization: Tokenization assumes spaces exist; segmentation is needed where they don't.
  3. Stemming: Crude heuristic process of chopping off word endings to get a root/stem (e.g., "running" → "run"). Rule-based (e.g., Porter Stemmer). May produce non-words.

  4. Lemmatization: Sophisticated process using vocabulary and morphological analysis to return the dictionary form (lemma) (e.g., "better" → "good"). More accurate but slower.

  5. Stop Word Removal: Filtering out common, low-meaning words (e.g., "the", "is", "in").

  6. Case Normalization: Converting all text to lowercase (or preserving case for proper nouns).

  7. Handling Numbers & Special Characters: Removing, replacing, or keeping based on task.

[!TIP] Common Pitfall: Confusing Stemming (fast, rule-based, may be inaccurate) with Lemmatization (slower, vocabulary-based, accurate). Know Porter Stemmer as a key example.


III. Corpora in NLP

What are Corpora?

A corpus (plural: corpora) is a structured, large, and machine-readable collection of texts used for linguistic analysis and NLP model training.

Types of Corpora:

Type Description Example
Monolingual Text in a single language. English Wikipedia dump.
Parallel Same text translated into multiple languages. Europarl Corpus (EU proceedings).
Annotated Text with added linguistic labels. POS-tagged corpus, Parse treebank, semantically tagged.

Significance of Corpora Analysis

  • Training Data: Essential for statistical and machine learning models (e.g., training a POS tagger on an annotated corpus).

  • Linguistic Research: Study language frequencies, patterns, and distributions.

  • System Evaluation & Benchmarking: Provide standard test sets (e.g., Penn Treebank) to compare different NLP systems.

  • Resource for Low-Resource Languages: Creating corpora for Indian languages is a critical step to develop their NLP tools.


IV. Morphology and Phonology

A. Morphology

Definition: Study of word structure and formation.

  • Morphemes: Smallest meaningful units (e.g., root "teach", prefix "un-", suffix "-er").

  • Processes:

    • Inflection: Modifying a word for grammatical function (e.g., "cat" → "cats" [number]).

    • Derivation: Creating a new word/lexeme (e.g., "teach" → "teacher").

    • Compounding: Combining words (e.g., "notebook").

Morphology of Indian Languages

  • Agglutinative Nature (e.g., Telugu, Tamil, Kannada): Words are formed by stringing together morphemes (root + suffixes) in a linear, predictable fashion. A single complex word can encode what English expresses in a full sentence.

  • Rich Inflectional Morphology: Extensive use of suffixes to mark case, gender, number, tense, aspect, mood.

  • Challenges for NLP:

    1. Large Number of Word Forms: A single root can generate thousands of surface forms.

    2. Sandhi/Sandhi Rules: Phonological changes at morpheme boundaries (e.g., "राम" + "अय" → "रामाय"). Complicates segmentation.

Finite State Automata (FSA) in Morphology

  • Relationship: FSAs are a natural and efficient formal model for morphological generators (produce valid word forms) and recognizers (analyze surface forms into morphemes).

  • Construction:

    • States represent morphemes (root, suffixes).

    • Transitions represent the concatenation of morphemes, often labeled with the output string.

    • Sandhi rules can be encoded as transitions.

  • Applications: Morphological analysis, spell checking, morphological generation for MT.

[!TIP] Key Link: FSA provides a computational solution to the morphological complexity (agglutination + sandhi) of Indian languages.

B. Phonology

Phonological Rules

  • Definition: Rules that describe systematic sound pattern changes in a language's phonological system.

  • Examples:

    • Assimilation: A sound becomes more like a neighboring sound (e.g., "in-" + "possible" → "impossible").

    • Elision: Sound deletion (e.g., "and" → "an'" in "rock 'n' roll").

    • Epenthesis: Sound insertion (e.g., "hamster" often pronounced "hampster").

  • Significance: Crucial for speech recognition (modeling pronunciation variants) and speech synthesis (generating natural-sounding speech).

Minimum Edit Distance (Levenshtein Distance)

  • Definition: The minimum number of single-character edit operations (insert, delete, substitute) required to change one string into another.

  • Calculation: Solved using dynamic programming.

    • Let D[i, j] be the distance between first i chars of string A and first j chars of string B.

    • Recurrence:

$$D[i,j] = \min \begin{cases} D[i-1, j] + 1 & \text{(deletion)} \\ D[i, j-1] + 1 & \text{(insertion)} \\ D[i-1, j-1] + \text{cost} & \text{(substitution)} \end{cases}$$

*   `cost = 0` if chars match, `1` otherwise.
  • Applications: Spelling correction, approximate string matching, speech recognition (hypothesis scoring).

[!TIP] Formula to Remember: The dynamic programming recurrence for Minimum Edit Distance is a classic exam problem. Know the three operations and base cases (D[0,j]=j, D[i,0]=i).


V. Part-of-Speech (POS) Tagging

Definition and Examples

Assigning a syntactic category tag (Noun, Verb, Adjective, etc.) to each word in a sentence.

  • Example: "They /PRP can /MD /VB fish /NN ."

  • Ambiguity: "book" can be a Noun ("read a book") or Verb ("book a flight"). Context resolves this.

Models and Algorithms

1. Maximum Entropy Model

  • Principle: Given constraints (features), choose the distribution with maximum entropy (most uniform, least assumptions).

  • Features: Contextual features like previous word, next word, word suffixes, capitalization.

  • Training: Learn weights for features from annotated training data to maximize likelihood.

  • Decoding: For a given word, compute probability for each tag using features and weights, choose tag with max probability.

  • Advantage: Flexible feature design, good accuracy.

  • Disadvantage: Training can be computationally intensive.

2. Transformation-Based Learning (TBL) / Brill Tagger

  • Algorithm: Error-driven, rule induction.

    1. Start with a simple baseline tagger (e.g., assign most frequent tag).

    2. Apply transformation rules (e.g., "Change tag from VBD to NN if previous tag is DT").

    3. Learn rules that most reduce errors on training data compared to current state.

    4. Apply learned rules in order.

  • Template Representation: Rules are instantiated from templates like: "Change tag A to B if:

    • previous tag = X

    • next word = Y

    • current word ends with suffix Z"

  • Advantages: Interpretable rules, fast tagging, high accuracy.

  • Disadvantages: Rule ordering is critical, can overfit.

[!TIP] Comparison: Maximum Entropy is a probabilistic discriminative model; TBL is an error-driven rule learner. Both are prominent for POS tagging.


VI. Parsing

Definition and Goals

Analyzing a sentence to determine its syntactic structure according to a grammar. Goals: Identify phrases (constituents) and grammatical relations (dependencies).

Types of Parsers

Feature Synthetic (Rule-Based) Parsers Statistical Parsers
Basis Hand-crafted grammars (CFG, Feature-Based, Lexicalized). Data-driven, probabilistic models trained on treebanks.
Model Deterministic (e.g., chart parsers with grammar rules). Probabilistic (e.g., PCFG - Probabilistic CFG).
Advantages High precision on covered sentences, linguistically insightful. Robustness & coverage, handles ambiguity via probabilities.
Limitations Poor coverage (grammar cannot handle all real-world sentences), expensive to build/maintain. Data sparsity (rare constructions), domain sensitivity (performance drops on new domains).

Key Comparison:

  • Synthetic: Grammar → Parser. Top-down or bottom-up application of rules.

  • Statistical: Treebank → Model → Parser. Learns probabilities for rule applications (e.g., P(NP -> DT NN)).

[!TIP] Remember: Statistical parsers rely on treebanks (annotated corpora with parse trees) for training. PCFG is the foundational statistical model.


VII. Semantic Analysis and Disambiguation

A. Semantic Analysis

Need for Semantic Analysis

To move beyond syntactic structure and capture meaning. Essential for tasks requiring true understanding: Question Answering, Textual Entailment, Machine Translation, Information Extraction.

Bootstrapping Methods

  • Definition: Self-improving, iterative learning methods that start from a small set of seed examples or rules, and iteratively expand the knowledge base.

  • Approaches:

    1. Iterative Expansion: Use current extractor to find new instances → add high-confidence ones to training set → retrain.

    2. Co-training: Use two independent "views" of the data (e.g., words and parse patterns) to label unlabeled data for each other.

  • Example: Word Sense Disambiguation (WSD). Start with few examples of a sense → train classifier → apply to new contexts → add confident predictions → retrain.

B. Word Sense Disambiguation (WSD)
  • Definition: Identifying the correct sense/meaning of a word in a given context.

  • Approaches:

    • Knowledge-Based: Use lexical resources (WordNet). Lesk Algorithm: Compare dictionary definitions (glosses) of possible senses with the context sentence; choose sense with max overlap.

    • Supervised: Treat as classification. Train on sense-annotated corpora using context features (surrounding words, POS).

    • Unsupervised: Word Sense Induction - cluster word occurrences based on context similarity to discover senses.

C. Anaphora and Named Entity Resolution

Anaphora Resolution

  • Goal: Link a pronoun (he, she, it, they) or definite noun phrase ("the company") to its antecedent (the entity it refers to) in the text.

  • Techniques:

    • Centering Theory: Track the "center" of attention in discourse.

    • Hobbs Algorithm: A syntactic, search-based algorithm over parse trees.

    • ML-Based: Use features like distance, gender/number agreement, syntactic role.

Named Entity Recognition/Resolution (NER/NER)

  • Recognition (NER): Identifying and classifying named entities in text into predefined categories (PERSON, ORGANIZATION, LOCATION, DATE, etc.).

  • Resolution (NEL): Linking a recognized entity mention to a real-world entity in a knowledge base (e.g., "Apple" → Apple_Inc. not Apple_(fruit)).

  • Challenges: Ambiguity, nested entities, missing entities.

Comparison: Anaphora vs Named Entity Resolution

Aspect Anaphora Resolution Named Entity Resolution
Target Pronouns (he, it) & definite NPs ("the president"). Named Entity Mentions ("Barack Obama", "Google").
Core Task Coreference: Determine if multiple mentions refer to the same discourse entity. Linking: Map a mention to a unique KB entry (disambiguation + linking).
Key Techniques Centering, syntactic patterns, agreement features. NER first, then entity linking using KBs (Wikipedia, DBpedia), context similarity.

VIII. Models and Algorithms Overview

Categorization of NLP Models

  1. Rule-Based: Finite-state transducers, hand-crafted grammars.

  2. Statistical: n-grams, HMMs (for tagging, ASR), PCFGs (parsing).

  3. Machine Learning: Maximum Entropy, Conditional Random Fields (CRF), SVM.

  4. Deep Learning: RNNs/LSTMs, Transformers (BERT, GPT).

Bayesian Method of Pronunciation

  • A probabilistic framework for modeling pronunciation variants in speech recognition.

  • Goal: Compute P(word | acoustic_evidence) using Bayes' rule:

$$P(w|a) \propto P(a|w) \cdot P(w)$$

  • P(a|w): Acoustic model probability of sound given word.

  • P(w): Language model probability of the word sequence.

  • Application: Handles grapheme-to-phoneme (G2P) conversion and pronunciation variations (e.g., "tomato" /təˈmɑːtoʊ/ vs /təˈmeɪtoʊ/).

Other Key Algorithms (from other sections)

  • Finite State Automata (FSA): Morphology, finite-state parsing.

  • Maximum Entropy: POS tagging, segmentation.

  • Transformation-Based Learning (TBL): POS tagging, chunking.

  • Minimum Edit Distance: Spelling correction, string matching.


IX. Synthesis and Applications

Integration of Components

Typical NLP pipeline:

Raw Text → Tokenization → Morphological Analysis → POS Tagging → Parsing → Semantic Analysis → Application Output.

Challenges in Indian Languages

  1. Morphological Richness: Agglutination, sandhi → huge vocabulary, complex analysis.

  2. Resource Scarcity: Lack of large, high-quality annotated corpora and lexical resources compared to English.

  3. Code-Mixing: Frequent mixing of Indian languages with English in text/speech.

  4. Script Diversity: Multiple scripts (Devanagari, Tamil, Telugu, etc.) require separate processing tools.

Evaluation Metrics

  • Tagging/NER: Accuracy, Precision, Recall, F1-Score.

  • Parsing:

    • UAS (Unlabeled Attachment Score): % of words with correct head.

    • LAS (Labeled Attachment Score): % of words with correct head and dependency label.

  • Machine Translation: BLEU (n-gram precision), METEOR (synonym-aware).

Future Directions

  • Cross-lingual Transfer: Leveraging resources from high-resource languages for low-resource ones.

  • Low-resource NLP: Developing methods needing less annotated data (semi-supervised, unsupervised).

  • Multimodal Language: Integrating text with vision, speech.

  • Explainable & Fair NLP.

[!TIP] Synthesis Point: Always connect components. E.g., Morphological complexity in Indian languages affects POS tagging (more tags) and parsing (more complex structures). Corpora scarcity impacts all statistical/ML models.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in