Skip to content
AL-504 (C) · Computational Intelligence/Quick Revision Short Notes

Computational Intelligence (AL-504 (C)) - Unit 5 Short Notes

UNIT 5: Computational Intelligence Applications

I. Natural Language Processing (NLP)

A. Introduction to NLP

Natural Language Processing (NLP) is a subfield of AI focused on enabling computers to understand, interpret, and generate human language.

Core Challenges:

  • Ambiguity: Words/phrases have multiple meanings (lexical, syntactic, pragmatic).

  • Variability: Infinite ways to express the same idea.

  • Implicit Knowledge: Requires vast world knowledge not explicitly stated.

  • Non-standard Text: Handling slang, typos, and informal language.

  • Context Dependence: Meaning heavily relies on surrounding text and discourse.

[!TIP] Exam Focus: Be prepared to list and elaborate on these 4-5 core challenges with examples (e.g., "I saw the man with the telescope" – syntactic ambiguity).

B. Text Preprocessing and Fundamental Techniques

Tokenization

Splitting text into meaningful units (words, subwords, sentences).

  • Word Tokenization: "Don't hesitate." → ["Don't", "hesitate", "."]

  • Sentence Tokenization: Requires handling abbreviations (e.g., "Dr.").

Regular Expressions (Regex)

Patterns for matching and extracting text.

  • Common Patterns:

    • \w+ : Word characters

    • \d : Digit

    • \s : Whitespace

    • ^...$ : Start/end of string

    • [...] : Character set

    • * : Zero or more repetitions

  • Applications: Information extraction (emails, phone numbers), text cleaning, simple tokenization.

Spell Checking and Correction

  • Challenges: Non-word errors ("recieve") vs. real-word errors ("their" vs. "there"), context dependence.

  • Methods:

    1. Edit Distance: Find candidate corrections within a threshold.

    2. Language Model Scoring: Choose candidate with highest N-gram probability in context.

    3. Dictionary Lookup: For non-word errors.

Edit Distance (Minimum Edit Distance)

Minimum number of operations to transform string A into string B.

  • Operations: Insertion, Deletion, Substitution (cost=1 each). Transposition (cost=1, in Damerau-Levenshtein).

  • Algorithm (Dynamic Programming):

    Let D[i,j] be distance between first i chars of A and first j chars of B.

$$D[i,j] = \min \begin{cases} D[i-1,j] + 1 \quad \text{(deletion)} \\ D[i,j-1] + 1 \quad \text{(insertion)} \\ D[i-1,j-1] + \text{cost} \quad \text{(substitution)} \end{cases}$$

where `cost = 0` if `A[i] == B[j]`, else `1`.
  • Backtracking from D[m,n] gives the alignment/sequence of edits.

Dictionaries and Thesauri

  • Dictionary: Provides definitions, part-of-speech, pronunciation. Used for morphological analysis and spell checking.

  • Thesaurus (e.g., WordNet): Provides synonyms, antonyms, hypernyms (is-a), hyponyms. Crucial for Word Sense Disambiguation (WSD) and semantic similarity.

Finite-State Automata (FSA)

Mathematical models (finite states, transitions) for pattern recognition.

  • Deterministic (DFA): One transition per symbol.

  • Non-deterministic (NFA): Multiple possible transitions.

  • Applications: Lexical analysis (tokenization, morphological parsing), simple pattern matching, speech recognition (phoneme modeling).

C. Language Modeling

N-gram Models

Probabilistic models predicting next word based on previous n-1 words.

  • Unigram: $$\displaystyle P(w_i) $$ independent of context.

  • Bigram: $$\displaystyle P(w_i | w_{i-1}) = \frac{\text{count}(w_{i-1}, w_i)}{\text{count}(w_{i-1})} $$

  • Trigram: $$\displaystyle P(w_i | w_{i-2}, w_{i-1}) $$

  • Chain Rule: $$\displaystyle P(w_1^n) = \prod_{i=1}^n P(w_i | w_1^{i-1}) \approx \prod_{i=1}^n P(w_i | w_{i-n+1}^{i-1}) $$

Smoothing Techniques

Necessary because N-gram counts are sparse; unseen N-grams get probability 0, which is problematic.

  • Laplace (Add-One) Smoothing:

$$P_{\text{Laplace}}(w_i | w_{i-1}) = \frac{C(w_{i-1}, w_i) + 1}{C(w_{i-1}) + V}$$

where `V` is vocabulary size.
  • Other Methods: Good-Turing, Kneser-Ney, Backoff, Interpolation.

  • Goal: Distribute some probability mass from seen to unseen events.

Perplexity

Intrinsic evaluation metric for language models. Measures how "surprised" the model is by test data.

  • Definition: Perplexity = $$\displaystyle 2^{-\frac{1}{N} \sum_{i=1}^{N} \log_2 P(w_i | \text{context})} $$

  • Interpretation: Lower perplexity = better model (less perplexed by test data).

  • Relation to Cross-Entropy: Perplexity = $$\displaystyle 2^{\text{Cross-Entropy}} $$.

Grammar-based Language Models (Grammarians LM)

Use formal grammars (e.g., Context-Free Grammars) to define legal sentence structures. Assign probabilities to grammar rules (see PCFG). More structured than N-grams but harder to estimate.

D. Part-of-Speech (POS) Tagging

Assigning a grammatical tag (Noun, Verb, Adj, etc.) to each word.

Rule-based Tagging

  • Principle: Apply hand-crafted lexical and contextual rules sequentially.

  • Example Rules: "If word ends in '-ing', tag as Verb (gerund)." "If previous tag is DT (determiner), current word likely JJ (adjective) or NN (noun)."

  • Pros: Transparent, no training data needed. Cons: Labor-intensive, brittle, low accuracy (~90%).

Transformation-based Tagging (Brill's Algorithm)

  • Principle: Start with a simple baseline (e.g., most frequent tag). Learn a sequence of transformations (if-then rules) from annotated corpus to correct errors.

  • Transformation: (Trigger_tag, New_tag, Condition).

    • E.g., (NN, VB, Previous_tag = TO) → Change Noun to Verb if preceded by 'to'.
  • Contribution: Achieves high accuracy (~97%) with a small, interpretable rule set. Bridges rule-based and statistical approaches.

Hidden Markov Models (HMM) for POS Tagging

  • Model: States = POS tags, Observations = words.

  • Key Components:

    • Transition Probabilities: $$\displaystyle A_{ij} = P(tag_j | tag_i) $$ (from tag i to j)

    • Emission Probabilities: $$\displaystyle B_{j}(w) = P(word | tag_j) $$

    • Initial Probabilities: $$\displaystyle \pi_i = P(first\_tag = i) $$

  • Decoding (Viterbi Algorithm): Finds most likely sequence of tags $$\displaystyle T^* $$ for observed word sequence $W$:

$$T^* = \arg\max_T P(T|W) = \arg\max_T P(W|T) P(T)$$

Solved efficiently with dynamic programming.

[!TIP] Exam Focus: Contrast Rule-based, Transformation-based, and HMM tagging. Know Viterbi's role in HMM decoding.

E. Syntactic Analysis

Context-Free Grammars (CFGs)

  • Components: Set of non-terminals (syntactic categories like S, NP, VP), terminals (words), production rules (e.g., S -> NP VP), start symbol (S).

  • Capturing Structure: Defines hierarchical phrase structure (constituency). A sentence is a parse tree where each node is a non-terminal expanding via rules.

Probabilistic CFGs (PCFG)

Extends CFG by assigning probability $P(\text{rhs} | \text{lhs})$ to each rule.

  • Probability of a parse tree: Product of probabilities of all rules used.

  • Goal: Find most probable parse tree (not just any valid tree). Helps resolve ambiguity.

Syntactic Parsing

Process of analyzing a sentence to produce its syntactic structure (parse tree).

  • Constituency Parsing: Builds tree based on CFG/phrase structure (nodes are phrasal categories).

  • Dependency Parsing: Represents grammatical relations as directed edges between words (head-dependent relationships).

Parsing Algorithms

  • CYK Algorithm (Cocke-Younger-Kasami):

    • Input: Sentence in Chomsky Normal Form (CNF: A -> B C or A -> word).

    • Process: Dynamic programming. P[i,j,A] = probability that words i to j form constituent A. Fill triangular table.

    • Time Complexity: $$\displaystyle O(n^3) $$ for sentence length n.

  • Probabilistic CYK: Uses PCFG probabilities in the CYK table to find the most probable parse.

Treebanks

  • Construction: Manually or semi-automatically annotate large corpora with syntactic parse trees (constituency or dependency).

  • Role:

    1. Development: Provide training data for statistical parsers (PCFGs, modern neural parsers).

    2. Evaluation: Standard test sets (e.g., Penn Treebank) allow fair comparison of parser accuracy (metrics: Precision, Recall, F1-score on labeled constituents/dependencies).

Dependency Grammar

  • Principles: Sentence structure is a set of dependency relations (directed links) between words.

  • Key Concepts: Every word (except root) has exactly one head. Relations labeled (e.g., nsubj, dobj, amod).

  • Advantages: More directly models predicate-argument structure; no need for phrasal nodes; better for free-word-order languages.

Ambiguity in Parse Trees

  • Structural Ambiguity: A sentence has multiple valid parse trees.

    • Example: "I saw the man with the telescope."

      • [S [NP I] [VP saw [NP [Det the] [N man] [PP with [Det the] [N telescope]]]]] (I used a telescope)

      • [S [NP I] [VP saw [NP [Det the] [N man]] [PP with [Det the] [N telescope]]]] (The man had a telescope)

  • Resolution: Requires semantic knowledge, world knowledge, or probabilistic models (PCFGs) to choose the most likely parse.

F. Semantic Analysis

Word Sense Disambiguation (WSD)

Task: Determine the intended meaning (sense) of a word in context.

  • Supervised Methods: Train a classifier (e.g., using features: surrounding words, POS, syntax) on sense-annotated data (e.g., SemCor).

  • Dictionary-Based (Knowledge-Based): Use Lesk algorithm. Choose sense whose dictionary definition has most word overlap with the target sentence.

  • Thesaurus-Based (Knowledge-Based): Use WordNet. Select sense whose synset is most related (via path length, information content) to senses of neighboring words in the same text.

Compositional Semantics

  • Principle: The meaning of a complex expression (e.g., a sentence) is a function of the meanings of its parts and the rules used to combine them.

  • Example: Meaning("red car") = combine(meaning("red"), meaning("car")) via modification relation.

  • Contribution: Provides a systematic, rule-based way to build sentence meaning from word meanings and syntactic structure.

First-Order Logic (FOL) in NLP

Formal language for representing meaning with objects, properties, relations, and quantification.

  • Key Elements:

    • Constants: john, mary

    • Predicates: Loves(john, mary), Green(ball)

    • Variables: x, y

    • Quantifiers: ∀x (Student(x) → Smart(x)) (All students are smart), ∃y (Loves(john, y)) (John loves someone).

  • Use: Precise representation for question answering, information extraction, and reasoning.

Quantifiers

Logical operators specifying quantity/scope.

  • Universal (∀): "For all" / "Every".

  • Existential (∃): "There exists" / "Some".

  • Scope Ambiguity: "Every student read a book."

    • ∀x (Student(x) → ∃y (Book(y) ∧ Read(x,y))) (Each student read some book, possibly different).

    • ∃y (Book(y) ∧ ∀x (Student(x) → Read(x,y))) (There is one specific book that all students read).

Word Sense

The specific meaning a word has in a given context. A polysemous word (e.g., "bank") has multiple senses (financial institution, river edge). WSD aims to select the correct sense from a sense inventory (like WordNet's synsets).

G. Applications of NLP

Speech Recognition

  • Process: Audio → Feature Extraction → Acoustic Model (phonemes) → Language Model → Word Sequence.

  • NLP Enhancement: N-gram or neural LMs constrain possible word sequences, dramatically improving accuracy by resolving acoustic ambiguities (e.g., "recognize speech" vs. "wreck a nice beach").

Machine Translation (Transfer Model)

Classical pipeline approach:

  1. Analysis: Source sentence → syntactic/semantic representation.

  2. Transfer: Transfer representation from source to target language (may involve rule-based or statistical mapping).

  3. Generation: Target language representation → fluent target sentence.

  • Modern Note: Largely superseded by End-to-End Neural MT (Seq2Seq).

Commercial Applications (Improving UX)

  • Search Engines: Query understanding, document ranking.

  • Chatbots & Virtual Assistants: Intent recognition, dialogue management.

  • Sentiment Analysis: Monitor brand/review sentiment.

  • Content Recommendation: Topic modeling, content-based filtering.

  • Grammar/Spell Checkers: As in word processors.

Word Processors (Smarter Features)

  • Advanced grammar/style checking (beyond spell-check).

  • Smart autocomplete/suggestion (using contextual LMs).

  • Text summarization.

  • Style/tone adjustment suggestions.

II. AI in Healthcare

A. Overview and General Applications

  • Disease Detection & Diagnosis: Analyzing medical images (X-ray, MRI), genomic data, EHRs for early signs (e.g., diabetic retinopathy, cancer).

  • Patient Triage in ER: AI systems analyze vital signs, symptoms, history to prioritize patients by severity, reducing wait times for critical cases.

  • Heart Disease Risk Prediction: Models (logistic regression, ML) use features (age, cholesterol, BP, smoking) to predict 10-year risk (e.g., Framingham score).

B. Medical Imaging and Diagnostics

  • AI in Segmentation: Automatically delineating anatomical structures (organs, tumors) in scans.

  • Techniques: Primarily Convolutional Neural Networks (CNNs) like U-Net, SegNet.

  • Process: Train on images with manual segmentations (ground truth). Model learns pixel-wise classification.

  • Impact: Faster, more consistent segmentation for diagnosis, surgery planning, radiotherapy.

C. Prognostics and Survival Analysis

Linear Prognostic Models

  • Concept: Use linear regression (or Cox proportional hazards) to predict a continuous prognostic score or survival time based on patient features (age, biomarkers, stage).

  • Example: Predicted_Survival_Months = β₀ + β₁*Age + β₂*Tumor_Size + ...

Survival Models vs. Time Survival Models

  • Survival Model (Cox PH): Models hazard function (instantaneous risk of event). Outputs hazard ratio for features. Does not directly predict survival time.

$$h(t|X) = h_0(t) \exp(\beta^T X)$$

  • Time Survival Model: Directly models survival time as outcome (e.g., using accelerated failure time models). Outputs predicted survival time distribution.

Survival Trees

  • Concept: Decision trees adapted for censored survival data.

  • Splitting Criterion: Maximize difference in survival distributions between child nodes (e.g., log-rank test, log-rank score).

  • Example: Split on "Age > 65" if the two resulting groups have significantly different Kaplan-Meier survival curves.

  • Output: Risk groups, not precise time.

Nelson-Aalen Estimator

  • Purpose: Non-parametric estimate of the cumulative hazard function $H(t)$.

  • Calculation: Sum of observed hazard increments over time.

$$\hat{H}(t) = \sum_{t_i \leq t} \frac{d_i}{n_i}$$

where $$\displaystyle d_i $$ = events at time $$\displaystyle t_i $$, $$\displaystyle n_i $$ = patients at risk just before $$\displaystyle t_i $$.
  • Use: Complement to Kaplan-Meier (estimates survival function $S(t)$). $S(t) \approx \exp(-\hat{H}(t))$.

Conditional Average Treatment Effect (CATE)

  • Definition: The average difference in outcome between treatment and control for a specific subpopulation defined by covariates $$\displaystyle X=x $$.

$$\tau(x) = E[Y(1) - Y(0) | X=x]$$

where $Y(1)$, $Y(0)$ are potential outcomes.
  • Why Use It? In healthcare, treatment effect is rarely uniform. CATE identifies which patient groups benefit most from a drug/procedure (personalized medicine). Estimated using methods like causal forests, T-learner.

[!TIP] Exam Focus: Distinguish Survival Model (hazard) vs. Time Survival Model (time). Know Nelson-Aalen formula. Explain CATE's role in personalization.

D. Health Monitoring and Wearable Technology

Wearable Health Technology & AI

  • Devices: Smartwatches, ECG patches, glucose monitors, fitness trackers.

  • AI Connection: Sensors generate continuous time-series data (heart rate, activity, SpO2). AI models (RNNs, Transformers) process this stream to:

    • Detect anomalies (atrial fibrillation).

    • Predict events (hypoglycemia).

    • Classify activity/sleep stages.

    • Provide personalized feedback.

Remote Patient Monitoring (RPM)

  • How AI Enables It: Wearables/IoT devices transmit patient data to cloud. AI algorithms:

    1. Process raw sensor data (noise removal, feature extraction).

    2. Analyze for trends, deterioration, or acute events.

    3. Alert clinicians/patients automatically.

    4. Reduce false alarms via sophisticated pattern recognition.

  • Benefits: Chronic disease management (diabetes, CHF), post-discharge monitoring, elderly care.

Electronic Health Record (EHR) Systems & AI

  • AI for Efficiency:

    • Clinical Documentation: NLP to auto-populate fields from clinician notes.

    • Coding & Billing: Auto-assign ICD/CPT codes.

    • Information Retrieval: Smart search across patient histories.

    • Alert Fatigue Reduction: ML to prioritize meaningful alerts (e.g., sepsis risk).

    • Predictive Analytics: Flag patients at risk of readmission, deterioration.

E. Hospital Resource Management

  • Optimization with AI:

    • Staff Scheduling: ML forecasts patient influx; optimization algorithms assign shifts.

    • Bed Management: Predict discharge times/occupancy to optimize flow.

    • Inventory & Supply Chain: Predict demand for medicines, equipment.

    • Operating Room Scheduling: Maximize utilization via scheduling algorithms.

    • Energy Management: Optimize HVAC/lighting based on occupancy predictions.

F. Evaluation Metrics and Methodological Challenges

Evaluation Metrics for Model Efficiency (Healthcare Context)

  • Clinical Utility Focus: Beyond accuracy.

    • Sensitivity/Recall (True Positive Rate): Critical for disease detection (miss fewer cases).

    • Specificity (True Negative Rate): Avoid false alarms.

    • AUC-ROC: Overall discrimination ability.

    • Positive Predictive Value (Precision): "If flagged, how likely is disease?"

    • Calibration: Do predicted probabilities match actual frequencies? (Reliability diagrams).

    • Decision Curve Analysis: Net benefit across threshold probabilities.

Overfitting & Techniques to Fix

  • Problem: Model learns noise in training data, fails on new patients.

  • Techniques:

    1. Regularization: L1/L2 penalties on model weights.

    2. Cross-Validation: Robust performance estimation (k-fold).

    3. Hold-out Validation: Separate test set never seen during training/tuning.

    4. Feature Selection: Remove irrelevant/noisy features.

    5. Early Stopping: For iterative models (neural nets).

    6. Ensemble Methods: Bagging (e.g., Random Forest) reduces variance.

    7. Simplify Model: Reduce complexity (e.g., fewer layers/trees).

G. NLP-driven Healthcare Applications

Virtual Health Assistants

  • AI Components: NLP (intent recognition, entity extraction from patient speech/text), dialogue management, knowledge base.

  • Functions: Symptom checking, medication reminders, answering FAQs, mental health support (chatbots), appointment scheduling.

  • Challenge: Requires medical accuracy, safety, and handling of sensitive information.

Medication Adherence

  • AI Approaches:

    • NLP: Analyze patient notes/reports for non-adherence mentions.

    • Predictive Modeling: Identify patients at high risk of non-adherence (using EHR features: demographics, comorbidities, social determinants).

    • Intervention Personalization: Tailor reminder messages/education based on predicted barriers (cost, forgetfulness, side effects).

Sentiment Analysis in Healthcare

  • Applications:

    • Patient Feedback: Analyze reviews/surveys to gauge satisfaction, identify service issues.

    • Social Media Monitoring: Track public sentiment about diseases, vaccines, policies.

    • Clinical Notes: Detect patient sentiment/concern expressed in notes (though challenging due to clinical jargon).

  • Methods: Lexicon-based (medical sentiment lexicons) or ML classifiers (trained on annotated healthcare-specific text).


\boxed{\text{Study all listed topics from the outline, focusing on definitions, key formulas (Edit Distance, N-gram, Perplexity, HMM, CYK, Nelson-Aalen, CATE), algorithms (Viterbi, CYK), and applications. Prioritize questions from past papers.}}

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in