UNIT 5: Computational Intelligence Applications
I. Natural Language Processing (NLP)
A. Introduction to NLP
Natural Language Processing (NLP) is a subfield of AI focused on enabling computers to understand, interpret, and generate human language.
Core Challenges:
-
Ambiguity: Words/phrases have multiple meanings (lexical, syntactic, pragmatic).
-
Variability: Infinite ways to express the same idea.
-
Implicit Knowledge: Requires vast world knowledge not explicitly stated.
-
Non-standard Text: Handling slang, typos, and informal language.
-
Context Dependence: Meaning heavily relies on surrounding text and discourse.
[!TIP] Exam Focus: Be prepared to list and elaborate on these 4-5 core challenges with examples (e.g., "I saw the man with the telescope" – syntactic ambiguity).
B. Text Preprocessing and Fundamental Techniques
Tokenization
Splitting text into meaningful units (words, subwords, sentences).
-
Word Tokenization:
"Don't hesitate."→["Don't", "hesitate", "."] -
Sentence Tokenization: Requires handling abbreviations (e.g.,
"Dr.").
Regular Expressions (Regex)
Patterns for matching and extracting text.
-
Common Patterns:
-
\w+: Word characters -
\d: Digit -
\s: Whitespace -
^...$: Start/end of string -
[...]: Character set -
*: Zero or more repetitions
-
-
Applications: Information extraction (emails, phone numbers), text cleaning, simple tokenization.
Spell Checking and Correction
-
Challenges: Non-word errors (
"recieve") vs. real-word errors ("their"vs."there"), context dependence. -
Methods:
-
Edit Distance: Find candidate corrections within a threshold.
-
Language Model Scoring: Choose candidate with highest N-gram probability in context.
-
Dictionary Lookup: For non-word errors.
-
Edit Distance (Minimum Edit Distance)
Minimum number of operations to transform string A into string B.
-
Operations: Insertion, Deletion, Substitution (cost=1 each). Transposition (cost=1, in Damerau-Levenshtein).
-
Algorithm (Dynamic Programming):
Let
D[i,j]be distance between firstichars of A and firstjchars of B.
$$D[i,j] = \min \begin{cases} D[i-1,j] + 1 \quad \text{(deletion)} \\ D[i,j-1] + 1 \quad \text{(insertion)} \\ D[i-1,j-1] + \text{cost} \quad \text{(substitution)} \end{cases}$$
where `cost = 0` if `A[i] == B[j]`, else `1`.
- Backtracking from
D[m,n]gives the alignment/sequence of edits.
Dictionaries and Thesauri
-
Dictionary: Provides definitions, part-of-speech, pronunciation. Used for morphological analysis and spell checking.
-
Thesaurus (e.g., WordNet): Provides synonyms, antonyms, hypernyms (is-a), hyponyms. Crucial for Word Sense Disambiguation (WSD) and semantic similarity.
Finite-State Automata (FSA)
Mathematical models (finite states, transitions) for pattern recognition.
-
Deterministic (DFA): One transition per symbol.
-
Non-deterministic (NFA): Multiple possible transitions.
-
Applications: Lexical analysis (tokenization, morphological parsing), simple pattern matching, speech recognition (phoneme modeling).
C. Language Modeling
N-gram Models
Probabilistic models predicting next word based on previous n-1 words.
-
Unigram: $$\displaystyle P(w_i) $$ independent of context.
-
Bigram: $$\displaystyle P(w_i | w_{i-1}) = \frac{\text{count}(w_{i-1}, w_i)}{\text{count}(w_{i-1})} $$
-
Trigram: $$\displaystyle P(w_i | w_{i-2}, w_{i-1}) $$
-
Chain Rule: $$\displaystyle P(w_1^n) = \prod_{i=1}^n P(w_i | w_1^{i-1}) \approx \prod_{i=1}^n P(w_i | w_{i-n+1}^{i-1}) $$
Smoothing Techniques
Necessary because N-gram counts are sparse; unseen N-grams get probability 0, which is problematic.
- Laplace (Add-One) Smoothing:
$$P_{\text{Laplace}}(w_i | w_{i-1}) = \frac{C(w_{i-1}, w_i) + 1}{C(w_{i-1}) + V}$$
where `V` is vocabulary size.
-
Other Methods: Good-Turing, Kneser-Ney, Backoff, Interpolation.
-
Goal: Distribute some probability mass from seen to unseen events.
Perplexity
Intrinsic evaluation metric for language models. Measures how "surprised" the model is by test data.
-
Definition: Perplexity = $$\displaystyle 2^{-\frac{1}{N} \sum_{i=1}^{N} \log_2 P(w_i | \text{context})} $$
-
Interpretation: Lower perplexity = better model (less perplexed by test data).
-
Relation to Cross-Entropy: Perplexity = $$\displaystyle 2^{\text{Cross-Entropy}} $$.
Grammar-based Language Models (Grammarians LM)
Use formal grammars (e.g., Context-Free Grammars) to define legal sentence structures. Assign probabilities to grammar rules (see PCFG). More structured than N-grams but harder to estimate.
D. Part-of-Speech (POS) Tagging
Assigning a grammatical tag (Noun, Verb, Adj, etc.) to each word.
Rule-based Tagging
-
Principle: Apply hand-crafted lexical and contextual rules sequentially.
-
Example Rules: "If word ends in '-ing', tag as Verb (gerund)." "If previous tag is DT (determiner), current word likely JJ (adjective) or NN (noun)."
-
Pros: Transparent, no training data needed. Cons: Labor-intensive, brittle, low accuracy (~90%).
Transformation-based Tagging (Brill's Algorithm)
-
Principle: Start with a simple baseline (e.g., most frequent tag). Learn a sequence of transformations (if-then rules) from annotated corpus to correct errors.
-
Transformation:
(Trigger_tag, New_tag, Condition).- E.g.,
(NN, VB, Previous_tag = TO)→ Change Noun to Verb if preceded by 'to'.
- E.g.,
-
Contribution: Achieves high accuracy (~97%) with a small, interpretable rule set. Bridges rule-based and statistical approaches.
Hidden Markov Models (HMM) for POS Tagging
-
Model: States = POS tags, Observations = words.
-
Key Components:
-
Transition Probabilities: $$\displaystyle A_{ij} = P(tag_j | tag_i) $$ (from tag
itoj) -
Emission Probabilities: $$\displaystyle B_{j}(w) = P(word | tag_j) $$
-
Initial Probabilities: $$\displaystyle \pi_i = P(first\_tag = i) $$
-
-
Decoding (Viterbi Algorithm): Finds most likely sequence of tags $$\displaystyle T^* $$ for observed word sequence $W$:
$$T^* = \arg\max_T P(T|W) = \arg\max_T P(W|T) P(T)$$
Solved efficiently with dynamic programming.
[!TIP] Exam Focus: Contrast Rule-based, Transformation-based, and HMM tagging. Know Viterbi's role in HMM decoding.
E. Syntactic Analysis
Context-Free Grammars (CFGs)
-
Components: Set of non-terminals (syntactic categories like S, NP, VP), terminals (words), production rules (e.g.,
S -> NP VP), start symbol (S). -
Capturing Structure: Defines hierarchical phrase structure (constituency). A sentence is a parse tree where each node is a non-terminal expanding via rules.
Probabilistic CFGs (PCFG)
Extends CFG by assigning probability $P(\text{rhs} | \text{lhs})$ to each rule.
-
Probability of a parse tree: Product of probabilities of all rules used.
-
Goal: Find most probable parse tree (not just any valid tree). Helps resolve ambiguity.
Syntactic Parsing
Process of analyzing a sentence to produce its syntactic structure (parse tree).
-
Constituency Parsing: Builds tree based on CFG/phrase structure (nodes are phrasal categories).
-
Dependency Parsing: Represents grammatical relations as directed edges between words (head-dependent relationships).
Parsing Algorithms
-
CYK Algorithm (Cocke-Younger-Kasami):
-
Input: Sentence in Chomsky Normal Form (CNF:
A -> B CorA -> word). -
Process: Dynamic programming.
P[i,j,A]= probability that wordsitojform constituentA. Fill triangular table. -
Time Complexity: $$\displaystyle O(n^3) $$ for sentence length
n.
-
-
Probabilistic CYK: Uses PCFG probabilities in the CYK table to find the most probable parse.
Treebanks
-
Construction: Manually or semi-automatically annotate large corpora with syntactic parse trees (constituency or dependency).
-
Role:
-
Development: Provide training data for statistical parsers (PCFGs, modern neural parsers).
-
Evaluation: Standard test sets (e.g., Penn Treebank) allow fair comparison of parser accuracy (metrics: Precision, Recall, F1-score on labeled constituents/dependencies).
-
Dependency Grammar
-
Principles: Sentence structure is a set of dependency relations (directed links) between words.
-
Key Concepts: Every word (except root) has exactly one head. Relations labeled (e.g.,
nsubj,dobj,amod). -
Advantages: More directly models predicate-argument structure; no need for phrasal nodes; better for free-word-order languages.
Ambiguity in Parse Trees
-
Structural Ambiguity: A sentence has multiple valid parse trees.
-
Example:
"I saw the man with the telescope."-
[S [NP I] [VP saw [NP [Det the] [N man] [PP with [Det the] [N telescope]]]]](I used a telescope) -
[S [NP I] [VP saw [NP [Det the] [N man]] [PP with [Det the] [N telescope]]]](The man had a telescope)
-
-
-
Resolution: Requires semantic knowledge, world knowledge, or probabilistic models (PCFGs) to choose the most likely parse.
F. Semantic Analysis
Word Sense Disambiguation (WSD)
Task: Determine the intended meaning (sense) of a word in context.
-
Supervised Methods: Train a classifier (e.g., using features: surrounding words, POS, syntax) on sense-annotated data (e.g., SemCor).
-
Dictionary-Based (Knowledge-Based): Use Lesk algorithm. Choose sense whose dictionary definition has most word overlap with the target sentence.
-
Thesaurus-Based (Knowledge-Based): Use WordNet. Select sense whose synset is most related (via path length, information content) to senses of neighboring words in the same text.
Compositional Semantics
-
Principle: The meaning of a complex expression (e.g., a sentence) is a function of the meanings of its parts and the rules used to combine them.
-
Example: Meaning(
"red car") = combine(meaning("red"), meaning("car")) via modification relation. -
Contribution: Provides a systematic, rule-based way to build sentence meaning from word meanings and syntactic structure.
First-Order Logic (FOL) in NLP
Formal language for representing meaning with objects, properties, relations, and quantification.
-
Key Elements:
-
Constants:
john,mary -
Predicates:
Loves(john, mary),Green(ball) -
Variables:
x,y -
Quantifiers:
∀x (Student(x) → Smart(x))(All students are smart),∃y (Loves(john, y))(John loves someone).
-
-
Use: Precise representation for question answering, information extraction, and reasoning.
Quantifiers
Logical operators specifying quantity/scope.
-
Universal (
∀): "For all" / "Every". -
Existential (
∃): "There exists" / "Some". -
Scope Ambiguity:
"Every student read a book."-
∀x (Student(x) → ∃y (Book(y) ∧ Read(x,y)))(Each student read some book, possibly different). -
∃y (Book(y) ∧ ∀x (Student(x) → Read(x,y)))(There is one specific book that all students read).
-
Word Sense
The specific meaning a word has in a given context. A polysemous word (e.g., "bank") has multiple senses (financial institution, river edge). WSD aims to select the correct sense from a sense inventory (like WordNet's synsets).
G. Applications of NLP
Speech Recognition
-
Process: Audio → Feature Extraction → Acoustic Model (phonemes) → Language Model → Word Sequence.
-
NLP Enhancement: N-gram or neural LMs constrain possible word sequences, dramatically improving accuracy by resolving acoustic ambiguities (e.g., "recognize speech" vs. "wreck a nice beach").
Machine Translation (Transfer Model)
Classical pipeline approach:
-
Analysis: Source sentence → syntactic/semantic representation.
-
Transfer: Transfer representation from source to target language (may involve rule-based or statistical mapping).
-
Generation: Target language representation → fluent target sentence.
- Modern Note: Largely superseded by End-to-End Neural MT (Seq2Seq).
Commercial Applications (Improving UX)
-
Search Engines: Query understanding, document ranking.
-
Chatbots & Virtual Assistants: Intent recognition, dialogue management.
-
Sentiment Analysis: Monitor brand/review sentiment.
-
Content Recommendation: Topic modeling, content-based filtering.
-
Grammar/Spell Checkers: As in word processors.
Word Processors (Smarter Features)
-
Advanced grammar/style checking (beyond spell-check).
-
Smart autocomplete/suggestion (using contextual LMs).
-
Text summarization.
-
Style/tone adjustment suggestions.
II. AI in Healthcare
A. Overview and General Applications
-
Disease Detection & Diagnosis: Analyzing medical images (X-ray, MRI), genomic data, EHRs for early signs (e.g., diabetic retinopathy, cancer).
-
Patient Triage in ER: AI systems analyze vital signs, symptoms, history to prioritize patients by severity, reducing wait times for critical cases.
-
Heart Disease Risk Prediction: Models (logistic regression, ML) use features (age, cholesterol, BP, smoking) to predict 10-year risk (e.g., Framingham score).
B. Medical Imaging and Diagnostics
-
AI in Segmentation: Automatically delineating anatomical structures (organs, tumors) in scans.
-
Techniques: Primarily Convolutional Neural Networks (CNNs) like U-Net, SegNet.
-
Process: Train on images with manual segmentations (ground truth). Model learns pixel-wise classification.
-
Impact: Faster, more consistent segmentation for diagnosis, surgery planning, radiotherapy.
C. Prognostics and Survival Analysis
Linear Prognostic Models
-
Concept: Use linear regression (or Cox proportional hazards) to predict a continuous prognostic score or survival time based on patient features (age, biomarkers, stage).
-
Example:
Predicted_Survival_Months = β₀ + β₁*Age + β₂*Tumor_Size + ...
Survival Models vs. Time Survival Models
- Survival Model (Cox PH): Models hazard function (instantaneous risk of event). Outputs hazard ratio for features. Does not directly predict survival time.
$$h(t|X) = h_0(t) \exp(\beta^T X)$$
- Time Survival Model: Directly models survival time as outcome (e.g., using accelerated failure time models). Outputs predicted survival time distribution.
Survival Trees
-
Concept: Decision trees adapted for censored survival data.
-
Splitting Criterion: Maximize difference in survival distributions between child nodes (e.g., log-rank test, log-rank score).
-
Example: Split on
"Age > 65"if the two resulting groups have significantly different Kaplan-Meier survival curves. -
Output: Risk groups, not precise time.
Nelson-Aalen Estimator
-
Purpose: Non-parametric estimate of the cumulative hazard function $H(t)$.
-
Calculation: Sum of observed hazard increments over time.
$$\hat{H}(t) = \sum_{t_i \leq t} \frac{d_i}{n_i}$$
where $$\displaystyle d_i $$ = events at time $$\displaystyle t_i $$, $$\displaystyle n_i $$ = patients at risk just before $$\displaystyle t_i $$.
- Use: Complement to Kaplan-Meier (estimates survival function $S(t)$). $S(t) \approx \exp(-\hat{H}(t))$.
Conditional Average Treatment Effect (CATE)
- Definition: The average difference in outcome between treatment and control for a specific subpopulation defined by covariates $$\displaystyle X=x $$.
$$\tau(x) = E[Y(1) - Y(0) | X=x]$$
where $Y(1)$, $Y(0)$ are potential outcomes.
- Why Use It? In healthcare, treatment effect is rarely uniform. CATE identifies which patient groups benefit most from a drug/procedure (personalized medicine). Estimated using methods like causal forests, T-learner.
[!TIP] Exam Focus: Distinguish Survival Model (hazard) vs. Time Survival Model (time). Know Nelson-Aalen formula. Explain CATE's role in personalization.
D. Health Monitoring and Wearable Technology
Wearable Health Technology & AI
-
Devices: Smartwatches, ECG patches, glucose monitors, fitness trackers.
-
AI Connection: Sensors generate continuous time-series data (heart rate, activity, SpO2). AI models (RNNs, Transformers) process this stream to:
-
Detect anomalies (atrial fibrillation).
-
Predict events (hypoglycemia).
-
Classify activity/sleep stages.
-
Provide personalized feedback.
-
Remote Patient Monitoring (RPM)
-
How AI Enables It: Wearables/IoT devices transmit patient data to cloud. AI algorithms:
-
Process raw sensor data (noise removal, feature extraction).
-
Analyze for trends, deterioration, or acute events.
-
Alert clinicians/patients automatically.
-
Reduce false alarms via sophisticated pattern recognition.
-
-
Benefits: Chronic disease management (diabetes, CHF), post-discharge monitoring, elderly care.
Electronic Health Record (EHR) Systems & AI
-
AI for Efficiency:
-
Clinical Documentation: NLP to auto-populate fields from clinician notes.
-
Coding & Billing: Auto-assign ICD/CPT codes.
-
Information Retrieval: Smart search across patient histories.
-
Alert Fatigue Reduction: ML to prioritize meaningful alerts (e.g., sepsis risk).
-
Predictive Analytics: Flag patients at risk of readmission, deterioration.
-
E. Hospital Resource Management
-
Optimization with AI:
-
Staff Scheduling: ML forecasts patient influx; optimization algorithms assign shifts.
-
Bed Management: Predict discharge times/occupancy to optimize flow.
-
Inventory & Supply Chain: Predict demand for medicines, equipment.
-
Operating Room Scheduling: Maximize utilization via scheduling algorithms.
-
Energy Management: Optimize HVAC/lighting based on occupancy predictions.
-
F. Evaluation Metrics and Methodological Challenges
Evaluation Metrics for Model Efficiency (Healthcare Context)
-
Clinical Utility Focus: Beyond accuracy.
-
Sensitivity/Recall (True Positive Rate): Critical for disease detection (miss fewer cases).
-
Specificity (True Negative Rate): Avoid false alarms.
-
AUC-ROC: Overall discrimination ability.
-
Positive Predictive Value (Precision): "If flagged, how likely is disease?"
-
Calibration: Do predicted probabilities match actual frequencies? (Reliability diagrams).
-
Decision Curve Analysis: Net benefit across threshold probabilities.
-
Overfitting & Techniques to Fix
-
Problem: Model learns noise in training data, fails on new patients.
-
Techniques:
-
Regularization: L1/L2 penalties on model weights.
-
Cross-Validation: Robust performance estimation (k-fold).
-
Hold-out Validation: Separate test set never seen during training/tuning.
-
Feature Selection: Remove irrelevant/noisy features.
-
Early Stopping: For iterative models (neural nets).
-
Ensemble Methods: Bagging (e.g., Random Forest) reduces variance.
-
Simplify Model: Reduce complexity (e.g., fewer layers/trees).
-
G. NLP-driven Healthcare Applications
Virtual Health Assistants
-
AI Components: NLP (intent recognition, entity extraction from patient speech/text), dialogue management, knowledge base.
-
Functions: Symptom checking, medication reminders, answering FAQs, mental health support (chatbots), appointment scheduling.
-
Challenge: Requires medical accuracy, safety, and handling of sensitive information.
Medication Adherence
-
AI Approaches:
-
NLP: Analyze patient notes/reports for non-adherence mentions.
-
Predictive Modeling: Identify patients at high risk of non-adherence (using EHR features: demographics, comorbidities, social determinants).
-
Intervention Personalization: Tailor reminder messages/education based on predicted barriers (cost, forgetfulness, side effects).
-
Sentiment Analysis in Healthcare
-
Applications:
-
Patient Feedback: Analyze reviews/surveys to gauge satisfaction, identify service issues.
-
Social Media Monitoring: Track public sentiment about diseases, vaccines, policies.
-
Clinical Notes: Detect patient sentiment/concern expressed in notes (though challenging due to clinical jargon).
-
-
Methods: Lexicon-based (medical sentiment lexicons) or ML classifiers (trained on annotated healthcare-specific text).
\boxed{\text{Study all listed topics from the outline, focusing on definitions, key formulas (Edit Distance, N-gram, Perplexity, HMM, CYK, Nelson-Aalen, CATE), algorithms (Viterbi, CYK), and applications. Prioritize questions from past papers.}}