UNIT 4: Applied Computational Intelligence (AI in Healthcare & Natural Language Processing)
I. AI IN HEALTHCARE
A. Overview & Key Application Domains
AI in healthcare leverages machine learning, natural language processing, and computer vision to improve patient outcomes, reduce costs, and enhance system efficiency.
| Application Domain | Core AI Function | Key Benefit |
|---|---|---|
| Remote Patient Monitoring (RPM) | Analyzes real-time sensor data (heart rate, glucose) from IoT devices. | Early detection of anomalies, reduced hospital readmissions. |
| Wearable Health Tech | Integrates with smartwatches/fitness bands for continuous vitals tracking & predictive alerts. | Proactive health management, personalized insights. |
| Hospital Resource Mgmt. | Optimizes staff scheduling, bed allocation, inventory (e.g., using reinforcement learning). | Cost reduction, improved operational flow. |
| Patient Triage (ER) | NLP & ML models prioritize cases based on severity from clinical notes & vitals. | Faster critical care, reduced wait times. |
| EHR System Efficiency | NLP extracts structured data from unstructured clinical notes; automates coding/billing. | Reduced clinician burden, better data usability. |
| Disease Detection & Diagnosis | ML classifiers (e.g., CNNs on images, ensemble models on lab data) for early detection (cancer, diabetic retinopathy). | Higher accuracy, earlier intervention. |
| Virtual Health Assistants | AI chatbots (NLP) for symptom checking, appointment scheduling, patient education. | 24/7 support, reduced administrative load. |
| Medication Adherence | Predictive models identify non-adherent patients; smart pill bottles with reminders. | Improved treatment outcomes, reduced complications. |
| Sentiment Analysis | NLP analyzes patient feedback/surveys to gauge satisfaction and identify service issues. | Enhanced patient experience, service improvement. |
Exam Tip: Be ready to illustrate RPM with a concrete example: "A patient with CHF uses a Bluetooth scale & blood pressure cuff. AI model detects rapid weight gain (+2kg in 24h) + rising BP → alerts cardiologist for early diuretic adjustment, preventing hospitalization."
B. Predictive Modeling & Prognostic Techniques
1. Linear Prognostic Models
-
Definition: Statistical models (e.g., linear regression, Cox proportional hazards) that predict a continuous health outcome (e.g., survival time, risk score) based on multiple patient features (age, biomarkers, comorbidities).
-
Formula (Cox PH Model):
$$h(t|X) = h_0(t) \exp(\beta_1 X_1 + \beta_2 X_2 + ... + \beta_p X_p)$$
* $h(t|X)$: Hazard function at time $t$ for patient with features $X$.
* $$\displaystyle h_0(t) $$: Baseline hazard.
* $\beta$: Coefficients indicating feature impact on hazard.
- Use: Provides interpretable risk factors; foundation for more complex models.
2. Survival Analysis Models
-
Core Concept: Analyzes time-to-event data, handling censoring (patients lost to follow-up or event not occurred by study end).
-
Survival Model vs. Time Survival Model:
-
Survival Model: Predicts probability of surviving beyond a given time $t$, $$\displaystyle S(t) = P(T > t) $$.
-
Time Survival Model: Predicts the actual survival time $T$ itself (a regression task).
-
-
Nelson-Aalen Estimator:
-
Purpose: Non-parametric estimator of the cumulative hazard function $H(t)$.
-
Formula:
-
$$\hat{H}(t) = \sum_{i: t_i \leq t} \frac{d_i}{n_i}$$
* $$\displaystyle t_i $$: Distinct observed event times.
* $$\displaystyle d_i $$: Number of events at $$\displaystyle t_i $$.
* $$\displaystyle n_i $$: Number of patients "at risk" just before $$\displaystyle t_i $$.
* **Relation to Survival:** $$\displaystyle \hat{S}(t) = \exp(-\hat{H}(t)) $$.
-
Survival Trees:
-
Principle: Decision tree adapted for survival data. Splits are chosen to maximize difference in survival distributions between child nodes (using log-rank test or log-rank score).
-
Example: A tree splits first on "Age > 65". Left node (younger) has median survival 10 years; right node (older) has median survival 4 years. Further splits on "Stage" refine predictions.
-
Advantage: Handles non-linear interactions, provides intuitive rules.
-
3. Conditional Average Treatment Effect (CATE)
-
Definition: Estimates the individual-level effect of a treatment (e.g., drug A vs. B) for a patient with specific covariates $$\displaystyle X=x $$. It's the conditional expectation: $$\displaystyle \tau(x) = E[Y(1) - Y(0) | X=x] $$.
-
Why Use It? Moves beyond average treatment effect (ATE). Enables personalized medicine—predicting which patients will benefit most from a specific intervention.
-
Methods: Meta-learners (e.g., T-learner, X-learner) using ML base models (random forests, neural nets) to estimate $\tau(x)$.
4. Heart Disease Risk Prediction
-
Common Models: Framingham Risk Score (traditional logistic regression), ML models (Random Forest, XGBoost, Neural Networks) using features like age, cholesterol, BP, ECG results.
-
AI Contribution: Handles complex interactions, uses imaging data (CT angiography) for more precise plaque analysis, integrates longitudinal EHR data for dynamic risk updating.
C. Medical Imaging & Specialized Tasks
AI in Medical Image Segmentation
-
Task: Pixel/voxel-level classification to delineate anatomical structures (tumors, organs) or pathologies in scans (MRI, CT, X-ray).
-
Key Architecture: U-Net
-
Structure: Encoder (contracting path) captures context; Decoder (expanding path) enables precise localization. Skip connections between corresponding encoder-decoder layers preserve high-resolution features.
-
Loss Function: Often Dice Loss (or Dice Coefficient) for imbalanced data:
-
$$\text{Dice} = \frac{2|X \cap Y|}{|X| + |Y|}$$
where $X$=prediction, $Y$=ground truth.
- Impact: Automates tedious manual contouring, ensures consistency, aids in surgical planning and radiotherapy targeting.
D. Model Evaluation & Robustness
1. Evaluation Metrics for Model Efficiency (Healthcare Context)
| Metric | Definition/Formula | Best For |
|---|---|---|
| Accuracy | $(TP+TN)/(TP+TN+FP+FN)$ | Balanced classes (rare in healthcare). |
| Precision | $TP/(TP+FP)$ | Minimizing false positives (e.g., cancer screening: avoid unnecessary biopsies). |
| Recall (Sensitivity) | $TP/(TP+FN)$ | Minimizing false negatives (e.g., critical disease detection: miss no cases). |
| F1-Score | $2 \times (Precision \times Recall)/(Precision + Recall)$ | Balancing P & R. |
| Specificity | $TN/(TN+FP)$ | True negative rate (important for ruling out disease). |
| AUC-ROC | Area under ROC curve (TPR vs. FPR). | Overall discriminative power across thresholds. |
| Calibration | Agreement between predicted probability & actual frequency (e.g., reliability diagram). | Trustworthy probability estimates for risk communication. |
2. Techniques to Address Overfitting
-
Regularization: Add penalty to loss function.
-
L1 (Lasso): $$\displaystyle \lambda \sum |\beta_j| $$ → induces sparsity.
-
L2 (Ridge): $$\displaystyle \lambda \sum \beta_j^2 $$ → shrinks coefficients.
-
-
Dropout (NNs): Randomly "drop" (set to zero) a fraction of neurons during training.
-
Early Stopping: Halt training when validation performance degrades.
-
Data Augmentation: Increase training data diversity (e.g., rotate/flip medical images).
-
Cross-Validation: Robust performance estimation (k-fold CV).
-
Ensemble Methods: Bagging (e.g., Random Forest) reduces variance.
Common Pitfall: Using Accuracy on imbalanced healthcare data (e.g., 95% healthy, 5% diseased). A model predicting "all healthy" gets 95% accuracy but is useless. Always use Precision, Recall, F1, or AUC-ROC.
II. NATURAL LANGUAGE PROCESSING (NLP)
A. Text Preprocessing & Foundational Concepts
1. Tokenization
-
Definition: Splitting text into basic units (tokens: words, subwords, characters).
-
Challenges: Ambiguous boundaries (e.g., "U.S.A.", "don't"), language-specific rules (Chinese no spaces), handling punctuation/numbers.
-
Tools: Rule-based (split on whitespace/punctuation), statistical (e.g., Byte-Pair Encoding for subwords).
2. Regular Expressions (Regex)
-
Purpose: Pattern matching & text manipulation using symbolic rules.
-
Key Patterns & Examples:
| Pattern | Matches | Example Use | | :--- | :--- | :--- | |
\d+| One or more digits | Extract ages: "Age: 45 years" | |\w+| Word characters (letters, digits, _) | Find all words. | |\s+| Whitespace | Normalize spaces. | |^/$$\displaystyle ` | Start/end of string | Validate format: `^\d{5} $$(5-digit zip). | |[abc]| Any of a, b, c | Match "cat", "bat", "rat". | |(ing\|ed\|s)$| Ends with ing, ed, or s | Find verb forms. | -
Applications: Information extraction (emails, phones), text cleaning, tokenization rules.
3. Spelling Error Detection & Correction Challenges
-
Detection Challenges: Noisy text (social media), proper nouns, out-of-vocabulary words, homophones ("their" vs. "there").
-
Correction Challenges:
-
Non-word errors: "recieve" → "receive" (dictionary lookup).
-
Real-word errors: "I want to sea the movie" (context needed).
-
Candidate Generation: Finding plausible corrections within edit distance (insert, delete, substitute, transpose).
-
Disambiguation: Choosing correct candidate from many (e.g., "to", "too", "two").
-
-
Algorithms: Minimum Edit Distance (Levenshtein Distance) – dynamic programming to find least-cost transformation sequence (insert/delete/substitute cost=1).
4. Dictionary & Thesaurus Applications
-
Dictionary (e.g., WordNet):
-
WSD: Provides sense definitions & examples.
-
Text Expansion: Synonyms for query expansion in IR.
-
POS Disambiguation: Lists possible POS tags per word.
-
-
Thesaurus:
-
Feature Generation: Replace words with synonyms to reduce sparsity.
-
Semantic Similarity: Measure relatedness between words/concepts.
-
5. Finite-State Automata (FSA)
-
Definition: Abstract machine (states, transitions) for recognizing regular languages.
-
Components: Set of states $Q$, alphabet $\Sigma$, transition function $$\displaystyle \delta: Q \times \Sigma \rightarrow Q $$, start state $$\displaystyle q_0 $$, accept states $F$.
-
NLP Use: Efficient pattern matching (grep), morphological analysis (recognizing verb conjugations), tokenization rules.
-
Example: FSA to recognize words ending in "ing":
(q0) --a-z--> (q0) | --i--> (q1) --n--> (q2) --g--> (q_accept)Any sequence of letters leading to
q_acceptis accepted.
B. Language Modeling
1. N-gram Models
-
Definition: Probabilistic model that predicts next word based on previous $n-1$ words. Assumes Markov property (next word depends only on last $n-1$ words).
-
Bigram Probability Calculation:
$$P(w_i | w_{i-1}) = \frac{C(w_{i-1}, w_i)}{C(w_{i-1})}$$
* $$\displaystyle C(w_{i-1}, w_i) $$: Count of bigram.
* $$\displaystyle C(w_{i-1}) $$: Count of unigram $$\displaystyle w_{i-1} $$.
* **Example:** $$\displaystyle P(\text{"quick"}|\text{"the"}) = \frac{C(\text{"the quick"})}{C(\text{"the"})} $$.
-
Perplexity (Evaluation Metric):
-
Definition: Measures how well a probability model predicts a sample. Lower perplexity = better model.
-
Formula for test set $W$:
-
$$\text{PP}(W) = P(w_1 w_2 ... w_N)^{-1/N} = \sqrt[N]{\frac{1}{P(W)}}$$
* **Interpretation:** Geometric average of inverse probability per word. Equivalent to exponential of cross-entropy.
-
Smoothing (Necessity & Example):
-
Problem: N-grams with zero count in training → $$\displaystyle P=0 $$ for entire sequence → model fails.
-
Solution: Reallocate some probability mass from seen to unseen N-grams.
-
Laplace (Add-One) Smoothing:
-
$$P_{\text{Laplace}}(w_i | w_{i-1}) = \frac{C(w_{i-1}, w_i) + 1}{C(w_{i-1}) + V}$$
* $V$: Vocabulary size.
* Adds 1 to all counts, denominator adjusted by $V$.
* **Example:** If $$\displaystyle C(\text{"the quick"})=0 $$, $$\displaystyle C(\text{"the"})=100 $$, $$\displaystyle V=10000 $$, then $$\displaystyle P_{\text{Laplace}}(\text{"quick"}|\text{"the"}) = \frac{0+1}{100+10000} \approx 0.000099 $$.
2. Grammarians Language Model
-
Components:
-
Lexicon: Dictionary mapping words to their possible Parts-of-Speech (POS).
-
Grammar Rules: Context-Free Grammar (CFG) rules defining how POS tags/phrases combine (e.g., $$\displaystyle S \rightarrow NP\ VP $$).
-
-
Function: Parses a sentence by:
-
Tokenizing & assigning all possible POS tags from lexicon.
-
Applying CFG rules to combine tags/phrases into a valid parse tree.
-
Limitation: Highly ambiguous; generates many parses without probabilities.
-
3. Hidden Markov Models (HMMs) – Principles
-
Principle: Probabilistic model for sequence labeling (e.g., POS tagging). Assumes:
-
Markov assumption: Current hidden state depends only on previous state.
-
Output independence: Current observation depends only on current hidden state.
-
-
Components:
-
States: Hidden POS tags (Noun, Verb, Adj...).
-
Observations: Words in the sentence.
-
Parameters:
-
Transition probabilities $$\displaystyle A = \{a_{ij}\} = P(q_j \text{ at } t | q_i \text{ at } t-1) $$.
-
Emission probabilities $$\displaystyle B = \{b_j(o)\} = P(\text{word } o | \text{ state } q_j) $$.
-
Initial state distribution $\pi$.
-
-
-
Use in NLP: Given a sentence (observations), find the most likely sequence of hidden states (POS tags) using Viterbi algorithm (dynamic programming).
-
Discussion: Simple, efficient, but limited by first-order Markov assumption and inability to model long-range dependencies.
C. Syntactic Processing
1. Part-of-Speech (POS) Tagging
-
Goal: Assign a POS tag (Noun, Verb, Adj...) to each word in a sentence.
-
Rule-Based Tagging Principles:
-
Use hand-crafted lexical and contextual rules.
-
Example Rules:
-
If word ends in "-ing" → tag as Verb (unless in exception list).
-
If previous tag is DT (Determiner) → current word likely Noun/Adj.
-
-
Limitation: Low accuracy (~77%), brittle, rules hard to maintain.
-
-
Transformation-Based Tagging (Brill Tagger):
-
Principle: Start with a baseline tagger (e.g., assign most frequent tag). Apply ordered transformation rules to correct errors.
-
Rule Format:
(Trigger Condition) → (Change Tag).- Example:
(PrevTag = "VB" & CurrWord = "ing") → Change CurrTag from "NN" to "VBG".
- Example:
-
Learning: Rules are learned from a tagged corpus by comparing output to gold standard.
-
Advantage: More accurate than pure rule-based (~92-93%), rules are human-readable.
-
2. Context-Free Grammars (CFGs)
-
Components:
-
Terminals ($T$): Words/tokens (e.g., "cat", "runs").
-
Non-terminals ($N$): Syntactic categories (e.g., S, NP, VP).
-
Productions ($R$): Rules of the form $$\displaystyle A \rightarrow \beta $$ where $A \in N$, $$\displaystyle \beta \in (N \cup T)^* $$.
-
Start Symbol ($S$): Root non-terminal (usually S for Sentence).
-
-
Capturing Syntactic Structure: Generates phrase structure trees. Rules define how phrases combine.
-
Example: $$\displaystyle S \rightarrow NP\ VP $$, $$\displaystyle NP \rightarrow Det\ N $$, $$\displaystyle VP \rightarrow V\ NP $$.
-
Sentence "The cat runs" → Tree:
[S [NP [Det The] [N cat]] [VP [V runs]]].
-
3. Syntactic Parsing
-
Process: Given a sentence and a grammar (CFG), find the parse tree(s) that generate the sentence.
-
Algorithms:
-
Top-Down (e.g., Recursive Descent): Start from $S$, expand using productions. Can get stuck on left-recursive rules.
-
Bottom-Up (e.g., Shift-Reduce): Start with words, combine using productions in reverse.
-
CYK Algorithm (Cocke-Younger-Kasami): Dynamic programming, $$\displaystyle O(n^3) $$ for CFG in Chomsky Normal Form (CNF: $$\displaystyle A \rightarrow BC $$ or $$\displaystyle A \rightarrow a $$). Fills a triangular table.
-
-
Ambiguity in Parse Trees:
-
Definition: A sentence has more than one valid parse tree under the grammar.
-
Example: "I saw the man with the telescope."
-
Reading 1 (Instrument): I used a telescope to see the man. →
[S [NP I] [VP saw [NP [Det the] [N man] [PP with [Det the] [N telescope]]]]] -
Reading 2 (Modifier): I saw the man who had a telescope. →
[S [NP I] [VP [VP saw [NP [Det the] [N man]]] [PP with [Det the] [N telescope]]]]
-
-
Problem: Parser must choose or rank parses. Probabilistic CFGs (PCFGs) help.
-
4. Dependency Grammar
-
Idea: Represents syntactic structure as directed labeled dependencies between words (not phrases).
-
Root: Main verb of the sentence.
-
Arc: From head (governor) to dependent (modifier), labeled with relation (nsubj, dobj, amod, etc.).
-
Example (Universal Dependencies): "She eats apples."
eats (root) ├── nsubj → She └── dobj → apples -
Advantage: More directly models predicate-argument structure; useful for many languages; efficient parsing algorithms.
5. Treebanks
-
Construction Process:
-
Annotation: Human linguists manually parse sentences (phrase structure or dependencies) from a corpus.
-
Adjudication: Disagreements resolved by senior annotators.
-
Format: Standardized (e.g., Penn Treebank for CFG, Universal Dependencies for deps).
-
-
Role:
-
Development: Provides gold-standard training data for statistical parsers (PCFGs, neural parsers).
-
Assessment: Standard test sets for evaluating parser accuracy (e.g., ParsEval metrics: labeled precision/recall/F1 for constituent spans; UAS/LAS for dependencies).
-
6. Probabilistic Parsing
-
Probabilistic Context-Free Grammar (PCFG):
-
Extension of CFG: Each production $$\displaystyle A \rightarrow \beta $$ has a probability $$\displaystyle P(A \rightarrow \beta) $$.
-
Constraints: $$\displaystyle \sum_{\beta} P(A \rightarrow \beta) = 1 $$ for each $A$.
-
Parse Probability: $$\displaystyle P(\text{tree}) = \prod_{\text{rules in tree}} P(\text{rule}) $$.
-
Goal: Find most probable parse (Viterbi parse).
-
-
Probabilistic CYK Algorithm:
-
Modification to CYK: Instead of boolean table cell $[i,j,A]$ meaning "A can span words i..j", store probability $$\displaystyle P(A \overset{w_i...w_j}{\Rightarrow}) $$.
-
Recalculation: For split point $k$:
-
$$P(A \rightarrow BC) \times P(B \overset{w_i...w_k}{\Rightarrow}) \times P(C \overset{w_{k+1}...w_j}{\Rightarrow})$$
* Take **max** over all $B,C,k$ for $[i,j,A]$.
* **Backpointers** store argmax choices to recover best parse.
D. Semantic Processing & Meaning
1. Word Sense Disambiguation (WSD)
-
Definition: Task of determining which sense (meaning) of a polysemous word is used in a given context.
- Example: "bank" → financial institution vs. river side.
-
Methods:
-
Supervised Methods:
-
Approach: Treat as classification task. Train on sense-tagged corpus (e.g., SemCor).
-
Features: Surrounding words (bag-of-words), POS tags, dependency relations, collocations.
-
Models: SVM, Naive Bayes, Neural Networks.
-
-
Dictionary-Based (Knowledge-Based):
-
Approach: Use sense definitions from machine-readable dictionaries (e.g., WordNet).
-
Algorithm: Lesk Algorithm – Select sense whose definition has most word overlap with target sentence context.
-
Limitation: Definitions are short, overlap sparse.
-
-
Thesaurus-Based:
-
Approach: Use semantic relations (synonymy, hypernymy) from thesaurus (WordNet).
-
Algorithm: Similarity-based – For each candidate sense, compute similarity to context words' senses (using measures like Path Similarity, Wu-Palmer). Choose sense with max average similarity.
-
-
2. Compositional Semantics
-
Principle: Meaning of a complex expression is a function of the meanings of its parts and the rules used to combine them.
-
Example: "red ball"
-
Meanings:
red= λx. red(x),ball= λx. ball(x). -
Composition:
red(ball)= λx. (red(x) ∧ ball(x)).
-
-
Contribution: Provides formal framework (often using lambda calculus) to build logical forms from parse trees, enabling logical inference and question answering.
3. First-Order Logic (FOL) in NLP
-
Why FOL? More expressive than propositional logic. Can represent objects, properties, relations, quantification.
-
Key Components:
-
Constants:
john,mary,fido. -
Predicates:
Likes(john, mary),Dog(fido). -
Variables:
x,y. -
Quantifiers:
-
Universal: $\forall x$ (for all x).
-
Existential: $\exists x$ (there exists x).
-
-
Connectives: $$\displaystyle \land, \lor, \rightarrow, \neg $$.
-
-
NLP Use: Represent meaning of sentences for natural language inference (recognizing textual entailment), semantic role labeling, and knowledge base population.
E. Advanced NLP Applications & Systems
1. Speech Recognition & NLP Enhancement
-
How it Works (Pipeline):
-
Acoustic Model: Converts audio signal to phoneme probabilities (HMMs/DNNs).
-
Pronunciation Lexicon: Maps phonemes to words.
-
Language Model (N-gram/Neural): Scores word sequences for plausibility.
-
Decoder: Finds most likely word sequence given acoustic & LM scores.
-
-
NLP Enhancement:
-
Better LMs: Neural LMs (RNNs, Transformers) capture longer context → lower word error rate.
-
Punctuation & Capitalization: NLP post-processing.
-
Semantic Understanding: ASR output fed to NLP for intent recognition (voice assistants).
-
2. Machine Translation: Transfer Model (Phases)
-
Classic Transfer-Based Approach:
-
Analysis: Source sentence → syntactic/semantic representation (parse tree, logical form).
-
Transfer: Map source representation to target language representation (using bilingual dictionaries & transfer rules for structural differences).
-
Generation: Target representation → fluent target sentence (using target language grammar).
-
-
Limitation: Requires deep linguistic analysis for both languages; error-prone; largely superseded by statistical (SMT) and neural (NMT) methods.
3. Making Word Processors Smarter (NLP Integration)
-
Grammar & Style Check: Beyond spell-check. Detects passive voice overuse, complex sentences, jargon, plagiarism (using text similarity).
-
Contextual Suggestions: "Find synonyms" based on sentence context (WSD).
-
Auto-summarization: Extract key sentences from long documents.
-
Smart Templates: Auto-fill forms by extracting entities from emails/letters (NER).
-
Readability Scoring: Estimate grade level (Flesch-Kincaid).
4. Commercial Applications of NLP (Improving UX)
-
Search Engines: Query understanding, semantic search, featured snippets.
-
Chatbots & Virtual Assistants: Intent recognition, dialogue management (Siri, Alexa, customer service bots).
-
Sentiment Analysis: Brand monitoring on social media, product review summarization.
-
Recommendation Systems: Analyze reviews, user-generated content for better item profiling.
-
Content Moderation: Automatically detect hate speech, spam, fake news.
-
HR Tech: Resume parsing, matching JD to CVs, screening questions.
-
Legal Tech: Contract review, e-discovery (document clustering & relevance).
Exam Tip: For "commercial applications," structure answer by industry (e.g., Retail: sentiment analysis & chatbots; Healthcare: clinical note extraction; Finance: fraud detection in news). Always link to user experience: faster search, 24/7 support, personalized content.