Skip to content
AL-504 (C) · Computational Intelligence/Quick Revision Short Notes

Computational Intelligence (AL-504 (C)) - Unit 4 Short Notes

UNIT 4: Applied Computational Intelligence (AI in Healthcare & Natural Language Processing)


I. AI IN HEALTHCARE

A. Overview & Key Application Domains

AI in healthcare leverages machine learning, natural language processing, and computer vision to improve patient outcomes, reduce costs, and enhance system efficiency.

Application Domain Core AI Function Key Benefit
Remote Patient Monitoring (RPM) Analyzes real-time sensor data (heart rate, glucose) from IoT devices. Early detection of anomalies, reduced hospital readmissions.
Wearable Health Tech Integrates with smartwatches/fitness bands for continuous vitals tracking & predictive alerts. Proactive health management, personalized insights.
Hospital Resource Mgmt. Optimizes staff scheduling, bed allocation, inventory (e.g., using reinforcement learning). Cost reduction, improved operational flow.
Patient Triage (ER) NLP & ML models prioritize cases based on severity from clinical notes & vitals. Faster critical care, reduced wait times.
EHR System Efficiency NLP extracts structured data from unstructured clinical notes; automates coding/billing. Reduced clinician burden, better data usability.
Disease Detection & Diagnosis ML classifiers (e.g., CNNs on images, ensemble models on lab data) for early detection (cancer, diabetic retinopathy). Higher accuracy, earlier intervention.
Virtual Health Assistants AI chatbots (NLP) for symptom checking, appointment scheduling, patient education. 24/7 support, reduced administrative load.
Medication Adherence Predictive models identify non-adherent patients; smart pill bottles with reminders. Improved treatment outcomes, reduced complications.
Sentiment Analysis NLP analyzes patient feedback/surveys to gauge satisfaction and identify service issues. Enhanced patient experience, service improvement.

Exam Tip: Be ready to illustrate RPM with a concrete example: "A patient with CHF uses a Bluetooth scale & blood pressure cuff. AI model detects rapid weight gain (+2kg in 24h) + rising BP → alerts cardiologist for early diuretic adjustment, preventing hospitalization."


B. Predictive Modeling & Prognostic Techniques

1. Linear Prognostic Models

  • Definition: Statistical models (e.g., linear regression, Cox proportional hazards) that predict a continuous health outcome (e.g., survival time, risk score) based on multiple patient features (age, biomarkers, comorbidities).

  • Formula (Cox PH Model):

$$h(t|X) = h_0(t) \exp(\beta_1 X_1 + \beta_2 X_2 + ... + \beta_p X_p)$$

*   $h(t|X)$: Hazard function at time $t$ for patient with features $X$.

*   $$\displaystyle h_0(t) $$: Baseline hazard.

*   $\beta$: Coefficients indicating feature impact on hazard.
  • Use: Provides interpretable risk factors; foundation for more complex models.

2. Survival Analysis Models

  • Core Concept: Analyzes time-to-event data, handling censoring (patients lost to follow-up or event not occurred by study end).

  • Survival Model vs. Time Survival Model:

    • Survival Model: Predicts probability of surviving beyond a given time $t$, $$\displaystyle S(t) = P(T > t) $$.

    • Time Survival Model: Predicts the actual survival time $T$ itself (a regression task).

  • Nelson-Aalen Estimator:

    • Purpose: Non-parametric estimator of the cumulative hazard function $H(t)$.

    • Formula:

$$\hat{H}(t) = \sum_{i: t_i \leq t} \frac{d_i}{n_i}$$

    *   $$\displaystyle t_i $$: Distinct observed event times.

    *   $$\displaystyle d_i $$: Number of events at $$\displaystyle t_i $$.

    *   $$\displaystyle n_i $$: Number of patients "at risk" just before $$\displaystyle t_i $$.

*   **Relation to Survival:** $$\displaystyle \hat{S}(t) = \exp(-\hat{H}(t)) $$.
  • Survival Trees:

    • Principle: Decision tree adapted for survival data. Splits are chosen to maximize difference in survival distributions between child nodes (using log-rank test or log-rank score).

    • Example: A tree splits first on "Age > 65". Left node (younger) has median survival 10 years; right node (older) has median survival 4 years. Further splits on "Stage" refine predictions.

    • Advantage: Handles non-linear interactions, provides intuitive rules.

3. Conditional Average Treatment Effect (CATE)

  • Definition: Estimates the individual-level effect of a treatment (e.g., drug A vs. B) for a patient with specific covariates $$\displaystyle X=x $$. It's the conditional expectation: $$\displaystyle \tau(x) = E[Y(1) - Y(0) | X=x] $$.

  • Why Use It? Moves beyond average treatment effect (ATE). Enables personalized medicine—predicting which patients will benefit most from a specific intervention.

  • Methods: Meta-learners (e.g., T-learner, X-learner) using ML base models (random forests, neural nets) to estimate $\tau(x)$.

4. Heart Disease Risk Prediction

  • Common Models: Framingham Risk Score (traditional logistic regression), ML models (Random Forest, XGBoost, Neural Networks) using features like age, cholesterol, BP, ECG results.

  • AI Contribution: Handles complex interactions, uses imaging data (CT angiography) for more precise plaque analysis, integrates longitudinal EHR data for dynamic risk updating.


C. Medical Imaging & Specialized Tasks

AI in Medical Image Segmentation

  • Task: Pixel/voxel-level classification to delineate anatomical structures (tumors, organs) or pathologies in scans (MRI, CT, X-ray).

  • Key Architecture: U-Net

    • Structure: Encoder (contracting path) captures context; Decoder (expanding path) enables precise localization. Skip connections between corresponding encoder-decoder layers preserve high-resolution features.

    • Loss Function: Often Dice Loss (or Dice Coefficient) for imbalanced data:

$$\text{Dice} = \frac{2|X \cap Y|}{|X| + |Y|}$$

where $X$=prediction, $Y$=ground truth.

  • Impact: Automates tedious manual contouring, ensures consistency, aids in surgical planning and radiotherapy targeting.

D. Model Evaluation & Robustness

1. Evaluation Metrics for Model Efficiency (Healthcare Context)

Metric Definition/Formula Best For
Accuracy $(TP+TN)/(TP+TN+FP+FN)$ Balanced classes (rare in healthcare).
Precision $TP/(TP+FP)$ Minimizing false positives (e.g., cancer screening: avoid unnecessary biopsies).
Recall (Sensitivity) $TP/(TP+FN)$ Minimizing false negatives (e.g., critical disease detection: miss no cases).
F1-Score $2 \times (Precision \times Recall)/(Precision + Recall)$ Balancing P & R.
Specificity $TN/(TN+FP)$ True negative rate (important for ruling out disease).
AUC-ROC Area under ROC curve (TPR vs. FPR). Overall discriminative power across thresholds.
Calibration Agreement between predicted probability & actual frequency (e.g., reliability diagram). Trustworthy probability estimates for risk communication.

2. Techniques to Address Overfitting

  • Regularization: Add penalty to loss function.

    • L1 (Lasso): $$\displaystyle \lambda \sum |\beta_j| $$ → induces sparsity.

    • L2 (Ridge): $$\displaystyle \lambda \sum \beta_j^2 $$ → shrinks coefficients.

  • Dropout (NNs): Randomly "drop" (set to zero) a fraction of neurons during training.

  • Early Stopping: Halt training when validation performance degrades.

  • Data Augmentation: Increase training data diversity (e.g., rotate/flip medical images).

  • Cross-Validation: Robust performance estimation (k-fold CV).

  • Ensemble Methods: Bagging (e.g., Random Forest) reduces variance.

Common Pitfall: Using Accuracy on imbalanced healthcare data (e.g., 95% healthy, 5% diseased). A model predicting "all healthy" gets 95% accuracy but is useless. Always use Precision, Recall, F1, or AUC-ROC.


II. NATURAL LANGUAGE PROCESSING (NLP)

A. Text Preprocessing & Foundational Concepts

1. Tokenization

  • Definition: Splitting text into basic units (tokens: words, subwords, characters).

  • Challenges: Ambiguous boundaries (e.g., "U.S.A.", "don't"), language-specific rules (Chinese no spaces), handling punctuation/numbers.

  • Tools: Rule-based (split on whitespace/punctuation), statistical (e.g., Byte-Pair Encoding for subwords).

2. Regular Expressions (Regex)

  • Purpose: Pattern matching & text manipulation using symbolic rules.

  • Key Patterns & Examples:

    | Pattern | Matches | Example Use | | :--- | :--- | :--- | | \d+ | One or more digits | Extract ages: "Age: 45 years" | | \w+ | Word characters (letters, digits, _) | Find all words. | | \s+ | Whitespace | Normalize spaces. | | ^ / $$\displaystyle ` | Start/end of string | Validate format: `^\d{5} $$ (5-digit zip). | | [abc] | Any of a, b, c | Match "cat", "bat", "rat". | | (ing\|ed\|s)$ | Ends with ing, ed, or s | Find verb forms. |

  • Applications: Information extraction (emails, phones), text cleaning, tokenization rules.

3. Spelling Error Detection & Correction Challenges

  • Detection Challenges: Noisy text (social media), proper nouns, out-of-vocabulary words, homophones ("their" vs. "there").

  • Correction Challenges:

    • Non-word errors: "recieve" → "receive" (dictionary lookup).

    • Real-word errors: "I want to sea the movie" (context needed).

    • Candidate Generation: Finding plausible corrections within edit distance (insert, delete, substitute, transpose).

    • Disambiguation: Choosing correct candidate from many (e.g., "to", "too", "two").

  • Algorithms: Minimum Edit Distance (Levenshtein Distance) – dynamic programming to find least-cost transformation sequence (insert/delete/substitute cost=1).

4. Dictionary & Thesaurus Applications

  • Dictionary (e.g., WordNet):

    • WSD: Provides sense definitions & examples.

    • Text Expansion: Synonyms for query expansion in IR.

    • POS Disambiguation: Lists possible POS tags per word.

  • Thesaurus:

    • Feature Generation: Replace words with synonyms to reduce sparsity.

    • Semantic Similarity: Measure relatedness between words/concepts.

5. Finite-State Automata (FSA)

  • Definition: Abstract machine (states, transitions) for recognizing regular languages.

  • Components: Set of states $Q$, alphabet $\Sigma$, transition function $$\displaystyle \delta: Q \times \Sigma \rightarrow Q $$, start state $$\displaystyle q_0 $$, accept states $F$.

  • NLP Use: Efficient pattern matching (grep), morphological analysis (recognizing verb conjugations), tokenization rules.

  • Example: FSA to recognize words ending in "ing":

    
    (q0) --a-z--> (q0) | --i--> (q1) --n--> (q2) --g--> (q_accept)
    
    

    Any sequence of letters leading to q_accept is accepted.


B. Language Modeling

1. N-gram Models

  • Definition: Probabilistic model that predicts next word based on previous $n-1$ words. Assumes Markov property (next word depends only on last $n-1$ words).

  • Bigram Probability Calculation:

$$P(w_i | w_{i-1}) = \frac{C(w_{i-1}, w_i)}{C(w_{i-1})}$$

*   $$\displaystyle C(w_{i-1}, w_i) $$: Count of bigram.

*   $$\displaystyle C(w_{i-1}) $$: Count of unigram $$\displaystyle w_{i-1} $$.

*   **Example:** $$\displaystyle P(\text{"quick"}|\text{"the"}) = \frac{C(\text{"the quick"})}{C(\text{"the"})} $$.
  • Perplexity (Evaluation Metric):

    • Definition: Measures how well a probability model predicts a sample. Lower perplexity = better model.

    • Formula for test set $W$:

$$\text{PP}(W) = P(w_1 w_2 ... w_N)^{-1/N} = \sqrt[N]{\frac{1}{P(W)}}$$

*   **Interpretation:** Geometric average of inverse probability per word. Equivalent to exponential of cross-entropy.
  • Smoothing (Necessity & Example):

    • Problem: N-grams with zero count in training → $$\displaystyle P=0 $$ for entire sequence → model fails.

    • Solution: Reallocate some probability mass from seen to unseen N-grams.

    • Laplace (Add-One) Smoothing:

$$P_{\text{Laplace}}(w_i | w_{i-1}) = \frac{C(w_{i-1}, w_i) + 1}{C(w_{i-1}) + V}$$

    *   $V$: Vocabulary size.

    *   Adds 1 to all counts, denominator adjusted by $V$.

    *   **Example:** If $$\displaystyle C(\text{"the quick"})=0 $$, $$\displaystyle C(\text{"the"})=100 $$, $$\displaystyle V=10000 $$, then $$\displaystyle P_{\text{Laplace}}(\text{"quick"}|\text{"the"}) = \frac{0+1}{100+10000} \approx 0.000099 $$.

2. Grammarians Language Model

  • Components:

    1. Lexicon: Dictionary mapping words to their possible Parts-of-Speech (POS).

    2. Grammar Rules: Context-Free Grammar (CFG) rules defining how POS tags/phrases combine (e.g., $$\displaystyle S \rightarrow NP\ VP $$).

  • Function: Parses a sentence by:

    • Tokenizing & assigning all possible POS tags from lexicon.

    • Applying CFG rules to combine tags/phrases into a valid parse tree.

    • Limitation: Highly ambiguous; generates many parses without probabilities.

3. Hidden Markov Models (HMMs) – Principles

  • Principle: Probabilistic model for sequence labeling (e.g., POS tagging). Assumes:

    1. Markov assumption: Current hidden state depends only on previous state.

    2. Output independence: Current observation depends only on current hidden state.

  • Components:

    • States: Hidden POS tags (Noun, Verb, Adj...).

    • Observations: Words in the sentence.

    • Parameters:

      • Transition probabilities $$\displaystyle A = \{a_{ij}\} = P(q_j \text{ at } t | q_i \text{ at } t-1) $$.

      • Emission probabilities $$\displaystyle B = \{b_j(o)\} = P(\text{word } o | \text{ state } q_j) $$.

      • Initial state distribution $\pi$.

  • Use in NLP: Given a sentence (observations), find the most likely sequence of hidden states (POS tags) using Viterbi algorithm (dynamic programming).

  • Discussion: Simple, efficient, but limited by first-order Markov assumption and inability to model long-range dependencies.


C. Syntactic Processing

1. Part-of-Speech (POS) Tagging

  • Goal: Assign a POS tag (Noun, Verb, Adj...) to each word in a sentence.

  • Rule-Based Tagging Principles:

    • Use hand-crafted lexical and contextual rules.

    • Example Rules:

      • If word ends in "-ing" → tag as Verb (unless in exception list).

      • If previous tag is DT (Determiner) → current word likely Noun/Adj.

    • Limitation: Low accuracy (~77%), brittle, rules hard to maintain.

  • Transformation-Based Tagging (Brill Tagger):

    • Principle: Start with a baseline tagger (e.g., assign most frequent tag). Apply ordered transformation rules to correct errors.

    • Rule Format: (Trigger Condition) → (Change Tag).

      • Example: (PrevTag = "VB" & CurrWord = "ing") → Change CurrTag from "NN" to "VBG".
    • Learning: Rules are learned from a tagged corpus by comparing output to gold standard.

    • Advantage: More accurate than pure rule-based (~92-93%), rules are human-readable.

2. Context-Free Grammars (CFGs)

  • Components:

    • Terminals ($T$): Words/tokens (e.g., "cat", "runs").

    • Non-terminals ($N$): Syntactic categories (e.g., S, NP, VP).

    • Productions ($R$): Rules of the form $$\displaystyle A \rightarrow \beta $$ where $A \in N$, $$\displaystyle \beta \in (N \cup T)^* $$.

    • Start Symbol ($S$): Root non-terminal (usually S for Sentence).

  • Capturing Syntactic Structure: Generates phrase structure trees. Rules define how phrases combine.

    • Example: $$\displaystyle S \rightarrow NP\ VP $$, $$\displaystyle NP \rightarrow Det\ N $$, $$\displaystyle VP \rightarrow V\ NP $$.

    • Sentence "The cat runs" → Tree: [S [NP [Det The] [N cat]] [VP [V runs]]].

3. Syntactic Parsing

  • Process: Given a sentence and a grammar (CFG), find the parse tree(s) that generate the sentence.

  • Algorithms:

    • Top-Down (e.g., Recursive Descent): Start from $S$, expand using productions. Can get stuck on left-recursive rules.

    • Bottom-Up (e.g., Shift-Reduce): Start with words, combine using productions in reverse.

    • CYK Algorithm (Cocke-Younger-Kasami): Dynamic programming, $$\displaystyle O(n^3) $$ for CFG in Chomsky Normal Form (CNF: $$\displaystyle A \rightarrow BC $$ or $$\displaystyle A \rightarrow a $$). Fills a triangular table.

  • Ambiguity in Parse Trees:

    • Definition: A sentence has more than one valid parse tree under the grammar.

    • Example: "I saw the man with the telescope."

      • Reading 1 (Instrument): I used a telescope to see the man. → [S [NP I] [VP saw [NP [Det the] [N man] [PP with [Det the] [N telescope]]]]]

      • Reading 2 (Modifier): I saw the man who had a telescope. → [S [NP I] [VP [VP saw [NP [Det the] [N man]]] [PP with [Det the] [N telescope]]]]

    • Problem: Parser must choose or rank parses. Probabilistic CFGs (PCFGs) help.

4. Dependency Grammar

  • Idea: Represents syntactic structure as directed labeled dependencies between words (not phrases).

  • Root: Main verb of the sentence.

  • Arc: From head (governor) to dependent (modifier), labeled with relation (nsubj, dobj, amod, etc.).

  • Example (Universal Dependencies): "She eats apples."

    
    eats (root)
    
    ├── nsubj → She
    
    └── dobj → apples
    
    
  • Advantage: More directly models predicate-argument structure; useful for many languages; efficient parsing algorithms.

5. Treebanks

  • Construction Process:

    1. Annotation: Human linguists manually parse sentences (phrase structure or dependencies) from a corpus.

    2. Adjudication: Disagreements resolved by senior annotators.

    3. Format: Standardized (e.g., Penn Treebank for CFG, Universal Dependencies for deps).

  • Role:

    • Development: Provides gold-standard training data for statistical parsers (PCFGs, neural parsers).

    • Assessment: Standard test sets for evaluating parser accuracy (e.g., ParsEval metrics: labeled precision/recall/F1 for constituent spans; UAS/LAS for dependencies).

6. Probabilistic Parsing

  • Probabilistic Context-Free Grammar (PCFG):

    • Extension of CFG: Each production $$\displaystyle A \rightarrow \beta $$ has a probability $$\displaystyle P(A \rightarrow \beta) $$.

    • Constraints: $$\displaystyle \sum_{\beta} P(A \rightarrow \beta) = 1 $$ for each $A$.

    • Parse Probability: $$\displaystyle P(\text{tree}) = \prod_{\text{rules in tree}} P(\text{rule}) $$.

    • Goal: Find most probable parse (Viterbi parse).

  • Probabilistic CYK Algorithm:

    • Modification to CYK: Instead of boolean table cell $[i,j,A]$ meaning "A can span words i..j", store probability $$\displaystyle P(A \overset{w_i...w_j}{\Rightarrow}) $$.

    • Recalculation: For split point $k$:

$$P(A \rightarrow BC) \times P(B \overset{w_i...w_k}{\Rightarrow}) \times P(C \overset{w_{k+1}...w_j}{\Rightarrow})$$

*   Take **max** over all $B,C,k$ for $[i,j,A]$.

*   **Backpointers** store argmax choices to recover best parse.

D. Semantic Processing & Meaning

1. Word Sense Disambiguation (WSD)

  • Definition: Task of determining which sense (meaning) of a polysemous word is used in a given context.

    • Example: "bank" → financial institution vs. river side.
  • Methods:

    • Supervised Methods:

      • Approach: Treat as classification task. Train on sense-tagged corpus (e.g., SemCor).

      • Features: Surrounding words (bag-of-words), POS tags, dependency relations, collocations.

      • Models: SVM, Naive Bayes, Neural Networks.

    • Dictionary-Based (Knowledge-Based):

      • Approach: Use sense definitions from machine-readable dictionaries (e.g., WordNet).

      • Algorithm: Lesk Algorithm – Select sense whose definition has most word overlap with target sentence context.

      • Limitation: Definitions are short, overlap sparse.

    • Thesaurus-Based:

      • Approach: Use semantic relations (synonymy, hypernymy) from thesaurus (WordNet).

      • Algorithm: Similarity-based – For each candidate sense, compute similarity to context words' senses (using measures like Path Similarity, Wu-Palmer). Choose sense with max average similarity.

2. Compositional Semantics

  • Principle: Meaning of a complex expression is a function of the meanings of its parts and the rules used to combine them.

  • Example: "red ball"

    • Meanings: red = λx. red(x), ball = λx. ball(x).

    • Composition: red(ball) = λx. (red(x) ∧ ball(x)).

  • Contribution: Provides formal framework (often using lambda calculus) to build logical forms from parse trees, enabling logical inference and question answering.

3. First-Order Logic (FOL) in NLP

  • Why FOL? More expressive than propositional logic. Can represent objects, properties, relations, quantification.

  • Key Components:

    • Constants: john, mary, fido.

    • Predicates: Likes(john, mary), Dog(fido).

    • Variables: x, y.

    • Quantifiers:

      • Universal: $\forall x$ (for all x).

      • Existential: $\exists x$ (there exists x).

    • Connectives: $$\displaystyle \land, \lor, \rightarrow, \neg $$.

  • NLP Use: Represent meaning of sentences for natural language inference (recognizing textual entailment), semantic role labeling, and knowledge base population.


E. Advanced NLP Applications & Systems

1. Speech Recognition & NLP Enhancement

  • How it Works (Pipeline):

    1. Acoustic Model: Converts audio signal to phoneme probabilities (HMMs/DNNs).

    2. Pronunciation Lexicon: Maps phonemes to words.

    3. Language Model (N-gram/Neural): Scores word sequences for plausibility.

    4. Decoder: Finds most likely word sequence given acoustic & LM scores.

  • NLP Enhancement:

    • Better LMs: Neural LMs (RNNs, Transformers) capture longer context → lower word error rate.

    • Punctuation & Capitalization: NLP post-processing.

    • Semantic Understanding: ASR output fed to NLP for intent recognition (voice assistants).

2. Machine Translation: Transfer Model (Phases)

  • Classic Transfer-Based Approach:

    1. Analysis: Source sentence → syntactic/semantic representation (parse tree, logical form).

    2. Transfer: Map source representation to target language representation (using bilingual dictionaries & transfer rules for structural differences).

    3. Generation: Target representation → fluent target sentence (using target language grammar).

  • Limitation: Requires deep linguistic analysis for both languages; error-prone; largely superseded by statistical (SMT) and neural (NMT) methods.

3. Making Word Processors Smarter (NLP Integration)

  • Grammar & Style Check: Beyond spell-check. Detects passive voice overuse, complex sentences, jargon, plagiarism (using text similarity).

  • Contextual Suggestions: "Find synonyms" based on sentence context (WSD).

  • Auto-summarization: Extract key sentences from long documents.

  • Smart Templates: Auto-fill forms by extracting entities from emails/letters (NER).

  • Readability Scoring: Estimate grade level (Flesch-Kincaid).

4. Commercial Applications of NLP (Improving UX)

  • Search Engines: Query understanding, semantic search, featured snippets.

  • Chatbots & Virtual Assistants: Intent recognition, dialogue management (Siri, Alexa, customer service bots).

  • Sentiment Analysis: Brand monitoring on social media, product review summarization.

  • Recommendation Systems: Analyze reviews, user-generated content for better item profiling.

  • Content Moderation: Automatically detect hate speech, spam, fake news.

  • HR Tech: Resume parsing, matching JD to CVs, screening questions.

  • Legal Tech: Contract review, e-discovery (document clustering & relevance).

Exam Tip: For "commercial applications," structure answer by industry (e.g., Retail: sentiment analysis & chatbots; Healthcare: clinical note extraction; Finance: fraud detection in news). Always link to user experience: faster search, 24/7 support, personalized content.

Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in