Skip to content
IT-802 (B) · Natural Language Processing/Quick Revision Short Notes

Natural Language Processing (IT-802 (B)) - Unit 5 Short Notes

I. Introduction to Natural Language Processing

Definition and Scope:

Natural Language Processing (NLP) is a subfield of Artificial Intelligence (AI) and Computational Linguistics focused on enabling computers to understand, interpret, manipulate, and generate human language in a valuable way. Its scope bridges linguistics and computer science, covering tasks from low-level text processing to high-level semantic understanding.

Need and Importance:

  • Volume of Text Data: Exponential growth of digital text (social media, documents, web).

  • Human-Computer Interaction: Enables intuitive interfaces (chatbots, voice assistants).

  • Automation: Streamlines tasks like document classification, translation, and information retrieval.

  • Insight Extraction: Analyzes unstructured text for business intelligence, research, and sentiment.

Broad Classes of NLP Applications:

  1. Information Extraction (IE): Automatically extracting structured information (e.g., entities, relationships) from unstructured text.

  2. Machine Translation (MT): Automatic translation of text/speech between languages (e.g., Google Translate).

  3. Sentiment Analysis: Determining the emotional tone or opinion expressed in text (positive/negative/neutral).

  4. Question Answering (QA) Systems: Providing precise answers to natural language questions (e.g., IBM Watson, Siri).

Commercial Uses (Examples):

  • Customer Support Chatbots: Automate responses, reduce operational costs.

  • Sentiment Analysis for Brand Monitoring: Track public opinion on social media.

  • Grammar & Style Checkers: Tools like Grammarly, Microsoft Editor.

[!TIP]

Exam Focus: Be prepared to define NLP, list its need (3–4 points), name four application classes with brief examples, and cite three commercial uses.


II. Text Preprocessing in NLP

Steps in Text Preprocessing:

  1. Tokenization: Splitting text into basic units (words, subwords, or sentences).

  2. Word Segmentation: Dividing a continuous string of characters into words (critical for languages without spaces like Chinese).

  3. Stemming: Crudely chopping word endings to get a root form (e.g., "running" → "run").

  4. Lemmatization: Using vocabulary/morphology to get the dictionary form (e.g., "better" → "good").

  5. Stop Word Removal: Eliminating common, low-meaning words (e.g., "the", "is", "and").

  6. Case Folding: Converting all text to lowercase (or uppercase) for uniformity.

Detailed Discussion:

Preprocessing cleans and standardizes raw text, reducing noise and dimensionality. Stemming is faster but less accurate; lemmatization is slower but linguistically informed. Stop words removal depends on task (often kept in sentiment analysis). Tokenization varies by language: whitespace-based for English, but requires specialized algorithms for Chinese/Japanese.

Difference: Word Segmentation vs. Tokenization

Aspect Tokenization Word Segmentation
Definition Breaking text into tokens (words, symbols). Specifically splitting character streams into words.
Languages Straightforward for space-delimited (English). Essential for languages without spaces (Chinese, Thai).
Complexity Lower (often regex-based). Higher (requires statistical/ML models).
Example "Hello, world!" → ["Hello", ",", "world", "!"] "我爱自然语言处理" → ["我", "爱", "自然语言处理"]

[!TIP]

Common Pitfall: Students often confuse these. Remember: Tokenization is the broader task; Word Segmentation is a specific, challenging subtask for non-space languages.


III. Corpora in NLP

Introduction to Corpora:

A corpus (plural: corpora) is a large, structured collection of digital text used for linguistic analysis and NLP model training.

Types of Corpora:

  • Monolingual Corpus: Text in a single language (e.g., British National Corpus).

  • Parallel Corpus: Same text in multiple languages, aligned sentence-by-sentence (e.g., EUROPARL).

  • Annotated Corpus: Text enriched with linguistic labels (POS tags, parse trees, named entities). Examples: Penn Treebank (POS+parse), CoNLL-2003 (NER).

Significance of Corpora Analysis:

  • For Model Training: Provides empirical data to learn statistical patterns (e.g., word frequencies, transition probabilities in HMMs).

  • For Linguistic Research: Enables study of language usage, variation, and change (e.g., corpus linguistics).

  • Benchmarking: Standard corpora allow fair comparison of NLP systems (e.g., GLUE benchmark).

[!TIP]

Exam Focus: Know definitions of corpus types and at least two significance points for each (training & research).


IV. Morphology

Basics of Morphology:

  • Morpheme: Smallest meaningful unit (e.g., "un-", "break", "-able").

  • Inflection: Modifying a word to express grammar (tense, number) without changing core meaning (e.g., "walk" → "walked").

  • Derivation: Creating a new word with changed meaning/part-of-speech (e.g., "happy" → "happiness").

Morphology of Indian Languages:

  • Agglutinative: Words formed by stringing morphemes (e.g., Telugu, Tamil).

  • Inflectional: Rich morphology for case, gender, tense (e.g., Sanskrit, Marathi).

  • Script: Primarily abugida (e.g., Devanagari) where consonants have inherent vowels.

  • Challenges: Complex sandhi (sound changes), compound words, high OOV (out-of-vocabulary) rates.

Finite State Automata (FSA) in Morphology:

  • Relationship: FSA models the finite set of morphological rules in a language. Each state represents a morpheme position; transitions represent affixation or inflection rules.

  • Finite State Morphological Analyzers: Use FSA/transducers to generate all valid word forms from a root or analyze a surface form into root+features. Efficient for agglutinative languages.

[!TIP]

Key Point: FSA provides a compact, computationally efficient representation of morphological knowledge, crucial for languages with rich morphology.


V. Part-of-Speech (POS) Tagging

Definition and Example:

Assigning a grammatical category (noun, verb, adjective, etc.) to each word in a sentence.

Example: "The/DET quick/ADJ fox/NOUN jumps/VERB over/ADP the/DET lazy/ADJ dog/NOUN ."

Models for POS Tagging:

  1. Maximum Entropy Model (MaxEnt):

    • Principle: Choose the probability distribution with maximum entropy (least biased) subject to constraints from training data (feature expectations).

    • Application: Features can include word itself, prefixes/suffixes, surrounding words, capitalization. Model estimates $P(tag | features)$.

    • Formula: Maximize $$\displaystyle H(p) = -\sum_{x,t} p(x,t) \log p(x,t) $$ subject to $$\displaystyle \sum_{x,t} p(x,t) f_i(x,t) = E_{\text{train}}(f_i) $$.

  2. Transformation-Based Learning (TBL):

    • Algorithm:

      1. Start with a baseline tagger (e.g., assign most frequent tag).

      2. Error-driven: Learn transformation rules (if condition then change tag) that fix tagging errors in training data.

      3. Apply rules in order of highest accuracy gain.

    • Advantage: Captures contextual rules, interpretable.

    • Example Rule: If previous tag is 'DET' and current word ends with 'ly' → change current tag to 'ADV'.

[!TIP]

Comparison: MaxEnt is probabilistic and feature-based; TBL is rule-based and error-driven. Both handle context better than simple HMMs.


VI. Parsing

What is Parsing?

Syntactic analysis of a sentence to determine its grammatical structure, typically represented as a Parse Tree.

Types of Parsers:

Synthetic Parser (Rule-Based) Statistical Parser (Data-Driven)
Uses hand-crafted grammar rules (e.g., CFG). Learns probabilities from annotated corpora (e.g., Treebank).
Precision high for covered constructs, but brittle (fails on unseen patterns). Robust to variation, but requires large training data.
Example: Chart parsers with grammar rules. Example: Probabilistic Context-Free Grammar (PCFG) parsers, Neural parsers.

[!TIP]

Modern Trend: Statistical/Neural parsers dominate due to scalability and performance on real-world text.


VII. Semantic Analysis

Need for Semantic Analysis:

To move beyond syntax to meaning—understand word senses, relationships, and discourse coherence. Essential for tasks like QA, MT, and information retrieval.

Bootstrapping Methods in Semantic Analysis:

  • Self-Training: Use a model trained on limited labeled data to pseudo-label unlabeled data, then retrain.

  • Co-Training: Train two models on different feature views; each labels data for the other, iteratively.

Word Sense Disambiguation (WSD):

  • Goal: Determine the correct sense of an ambiguous word in context.

  • Approaches:

    • Supervised: Train on sense-annotated data (e.g., using SVM, MaxEnt).

    • Unsupervised: Cluster word contexts to induce senses (e.g., LDA).

    • Knowledge-Based: Use lexical resources (WordNet) to measure relatedness (e.g., Lesk algorithm).

[!TIP]

WSD Challenge: "Bank" (river vs. financial) – supervised needs labeled data; knowledge-based relies on resource coverage.


VIII. Discourse Analysis

Anaphora Resolution:

  • Task: Link pronouns (anaphora) to their antecedents (e.g., "John arrived. He was tired." → "He" → "John").

  • Techniques:

    • Rule-Based: Use syntactic patterns (e.g., proximity, gender agreement).

    • Machine Learning: Features: distance, number agreement, semantic compatibility.

  • Challenges: Ambiguity, non-local references, bridging anaphora.

Named Entity Resolution (NER) vs. Anaphora Resolution:

Named Entity Recognition (NER) Anaphora Resolution
Identifies and classifies entities (PERSON, LOCATION, etc.) in text. Links pronouns/nouns to previously mentioned entities.
Local to a sentence/phrase. Cross-sentential (discourse-level).
Output: [ORG Apple] [LOC Cupertino] Output: "It" → [ORG Apple]

[!TIP]

Key Difference: NER finds entities; Anaphora Resolution connects references to entities.


IX. Phonology and Phonetics in NLP

Phonological Rules:

  • Transform underlying phonemic representations into surface phonetic forms (e.g., assimilation, deletion).

  • Significance: Crucial for speech recognition (acoustic model mapping) and text-to-speech (pronunciation generation). Also aids in handling spelling variations.

Minimum Edit Distance (Levenshtein Distance):

  • Definition: Minimum number of single-character operations (insertion, deletion, substitution) to transform one string into another.

  • Formula: Computed via dynamic programming.

    Let $D[i,j]$ = distance between first $i$ chars of string A and first $j$ chars of B.

$$D[i,j] = \min \begin{cases} D[i-1,j] + 1 \text{ (deletion)} \\ D[i,j-1] + 1 \text{ (insertion)} \\ D[i-1,j-1] + \text{cost} \text{ (substitution)} \end{cases}$$

where $$\displaystyle \text{cost}=0 $$ if $$\displaystyle A[i]=B[j] $$, else $1$.

  • Applications: Spell checking, speech recognition (aligning recognized vs. reference text), DNA sequence alignment.

Bayesian Method of Pronunciation:

  • Models pronunciation as a probabilistic mapping from graphemes (letters) to phonemes.

  • Uses Bayes' rule: $$\displaystyle P(\text{phonemes}|\text{spelling}) \propto P(\text{spelling}|\text{phonemes}) P(\text{phonemes}) $$.

  • Application: Pronunciation modeling for OOV words in speech synthesis/recognition, especially for irregular spellings (English).

[!TIP]

Edit Distance: Remember recurrence and base case $$\displaystyle D[0,j]=j $$, $$\displaystyle D[i,0]=i $$. Bayesian pronunciation uses generative model of spelling given phonemes.


X. Models and Algorithms in NLP

Overview of Key Models:

Category Models Use Cases
Probabilistic HMM, Bayesian Networks POS tagging, language modeling
Machine Learning Maximum Entropy, CRF, SVM Sequence labeling, classification
Finite State FSA, Finite State Transducers (FST) Morphology, tokenization, chunking

Algorithmic Approaches:

  • Rule-Based: Hand-crafted linguistic rules. High precision, low recall, maintenance heavy.

  • Statistical: Learn from data (n-grams, PCFGs). Data-hungry but robust.

  • Machine Learning: Use features + algorithms (MaxEnt, CRF). Balance between rule and pure stats.

Evaluation Metrics for NLP Models:

  • Accuracy: $$\displaystyle \frac{\text{Correct}}{\text{Total}} $$ (for classification).

  • Precision, Recall, F1-Score: For tasks like NER, parsing.

$$\text{Precision} = \frac{TP}{TP+FP},\quad \text{Recall} = \frac{TP}{TP+FN},\quad F1 = 2 \cdot \frac{P \cdot R}{P+R}$$

  • BLEU: For machine translation (n-gram overlap).

  • Perplexity: For language models (lower is better).

  • Parsing: Exact Match, Labeled Attachment Score (LAS).

[!TIP]

Metric Choice: Use F1 for imbalanced tasks (NER), BLEU for MT, perplexity for LM. Always consider task-specific metrics.


Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in