Skip to content
CE-803 (A) · Artificial Intelligence/Quick Revision Short Notes

Artificial Intelligence (CE-803 (A)) - Unit 4 Short Notes

UNIT 4: AI Applications and Advanced Topics

I. Expert Systems

Definition and Key Characteristics

An Expert System (ES) is a computer program that emulates the decision-making ability of a human expert in a specific, narrow domain. It uses knowledge and inference rules to solve problems that typically require human expertise.

Key Characteristics:

  • High performance: Matches or exceeds human expert level in its domain.

  • Reliability: Provides consistent answers for identical inputs.

  • Explanation capability: Can justify its reasoning and conclusions.

  • Knowledge representation: Explicitly stores domain knowledge separately from the control/inference mechanism.

  • Handling uncertainty: Can work with incomplete or ambiguous information (in advanced ES).

Core Components

Component Function
Knowledge Base Stores domain-specific facts and rules (the "what" and "how").
Inference Engine The "brain"; applies logical rules to the knowledge base to derive new facts or conclusions.
User Interface Medium for interaction between user and the ES (questions, answers, explanations).
Explanation Facility Traces and explains the reasoning path (e.g., "Why?" and "How?" explanations).
Knowledge Acquisition Module Tool/interface for experts/knowledge engineers to add, update, or refine knowledge in the KB.

Inference Engines

1. Forward Chaining (Data-Driven)

  • Process: Starts with known facts (data) and repeatedly applies rules to derive new facts until a goal is reached or no new facts can be inferred.

  • Analogy: "Bottom-up" reasoning.

  • Suitable for: Situations where all initial data is available and the goal is not specified in advance (e.g., monitoring, diagnosis, classification).

  • Example: Medical diagnosis system starting from patient symptoms.

2. Backward Chaining (Goal-Driven)

  • Process: Starts with a hypothesis or goal, works backward through rules to find evidence or sub-goals that support it.

  • Analogy: "Top-down" reasoning.

  • Suitable for: Situations with a specific hypothesis to prove (e.g., planning, interpretation, "what-if" analysis).

  • Example: Legal advisor system checking if a client qualifies for a specific law.

Comparison:

Feature Forward Chaining Backward Chaining
Control Data-driven Goal-driven
Search Strategy Breadth-first (typically) Depth-first (typically)
Efficiency Can be inefficient (fires many irrelevant rules). More focused, but can get stuck in irrelevant sub-goals.
Best For Diagnosis (What is it?) Consultation (Does this apply?)

Knowledge Representation in Expert Systems

  • Production Rules (IF <condition> THEN <action>): Most common. Simple, modular, good for heuristic knowledge.

  • Predicate Logic (First-Order Logic): Highly expressive. Uses predicates, variables, quantifiers (∀, ∃).

    • Translation to Clausal Form: Convert to Conjunctive Normal Form (CNF: conjunction of disjunctions) for resolution.

    • Resolution and Refutation: A single, complete inference rule. To prove a goal G, add ¬G to the KB and show that the resulting set of clauses is unsatisfiable (leads to contradiction □).

  • Semantic Networks: Graphical representation using nodes (concepts/objects) and edges (relationships). Good for "is-a" and "has-part" hierarchies.

  • Frames and Scripts: Structured objects (frames) with slots (attributes) and default values. Scripts are frames for stereotypical event sequences (e.g., "restaurant script").

Development Process

  1. Knowledge Acquisition & Engineering: The major bottleneck. Involves interviewing domain experts, extracting, structuring, and formalizing their tacit knowledge.

  2. Knowledge Validation & Verification: Ensuring the KB is correct, complete, and consistent.

  3. Maintaining Knowledge Currency: Regular updates as domain knowledge evolves. Requires a robust acquisition module and expert involvement.

Benefits and Limitations

Advantages Limitations
Consistent, never forgets. Knowledge Acquisition Bottleneck: Hard, time-consuming to extract expertise.
Can be easily documented & replicated. Common Sense Problem: Lacks broad, everyday knowledge.
Can work in environments dangerous to humans. Brittleness: Fails on problems outside its narrow domain.
Can explain decisions (increases trust). Learning Capability: Most traditional ES cannot learn from experience.
Integrates knowledge from multiple experts. High Development & Maintenance Cost.

Case Studies

  • Medical Diagnosis (e.g., MYCIN): Used rules to diagnose bacterial infections and recommend antibiotics. Pioneered certainty factors for uncertainty.

  • Financial Planning (e.g., XCON/R1): Configured computer systems (DEC VAX). Demonstrated massive cost savings and ROI, proving ES commercial viability.

II. Natural Language Processing (NLP)

Definition and Significance

NLP is a field of AI focused on enabling computers to understand, interpret, manipulate, and generate human language. Its significance lies in bridging the human-computer interaction gap, enabling applications like search engines, translation, and conversational AI.

Core Components of NLP

Component Focus Example Tasks
Syntax (Parsing) Grammatical structure. Part-of-Speech (POS) tagging, Constituency/Dependency parsing.
Semantics Meaning of words/sentences. Word Sense Disambiguation, Semantic Role Labeling.
Pragmatics Meaning in context. Discourse analysis, Resolving pronouns ("it", "he"), Implicature.

Knowledge Representation for NLP

  • Semantic Networks: Represent word meanings and relationships (synonymy, hyponymy). Used in WordNet.

  • Frames & Scripts: Represent typical situations and participants. A script for "going to a restaurant" includes roles (customer, waiter), props (menu, money), and scenes (ordering, eating).

  • Conceptual Dependency (CD): Represents meaning using a small set of primitive acts (e.g., ATRANS for transfer of abstract relationship, PTRANS for physical transfer) and conceptual cases (actor, object, destination). Aims for language-independent representation.

  • Comparative Analysis (CD vs. Semantic Nets):

    • CD: More structured, uses primitives for deep understanding, language-independent. Complex to build.

    • Semantic Nets: Simpler, good for lexical/semantic relationships, but less procedural. Better for large-scale lexical databases.

Applications

  • Chatbots/Virtual Assistants: Use NLP for intent recognition, dialogue management (often with frames/scripts).

  • Machine Translation: Maps source language syntax/semantics to target language.

  • Sentiment Analysis: Classifies text polarity (positive/negative/neutral) using lexical and contextual features.

Role of Scripts, Schemas, and Frames in Conversational AI

They provide contextual scaffolding. A frame for a "hotel booking" has slots: [destination, check-in, check-out, room-type]. A script for the booking dialogue sequences expected turns. This helps the AI:

  1. Predict user intents and fill slots.

  2. Generate coherent, context-aware responses.

  3. Handle ellipsis and anaphora ("I want a room there" -> "there" refers to previously mentioned destination).

Ethical Implications

  • Surveillance & Monitoring: NLP enables mass analysis of communications (emails, calls), raising privacy and consent issues.

  • Bias & Fairness: Models trained on biased corpora (e.g., news, social media) perpetuate societal biases (gender, race). Mitigation: Curated datasets, bias detection algorithms, fairness-aware training.

  • Misinformation & Deepfakes: NLP can generate highly convincing fake text (GPT models), threatening information integrity.

III. Neural Networks and Deep Learning

Foundations

Concept Description
Biological Neuron Receives signals via dendrites, sums in cell body, fires if threshold crossed, signal sent via axon.
Artificial Neuron (McCulloch-Pitts) Binary threshold unit: output = 1 if Σ(w_i * x_i) >= θ, else 0. First mathematical model.
Rosenblatt's Perceptron Single-layer network with adjustable weights w_i and bias b. Learns via weight update rule: w_i(new) = w_i(old) + α*(t - y)*x_i (where t=target, y=output, α=learning rate). Limited to linearly separable problems.

Training Mechanisms

  • Backpropagation (BP): Core algorithm for training multi-layer networks (MLPs). Computes gradient of loss function w.r.t. each weight by applying chain rule backward through the network.

  • Gradient Descent: W_new = W_old - η * ∇L(W) (η = learning rate). Stochastic GD (SGD): Uses one sample per update. Mini-batch GD: Compromise, uses small batches.

  • Momentum: Accelerates SGD by adding a fraction γ of the previous update vector: ΔW = γ*ΔW_prev - η*∇L. Helps escape shallow local minima/saddle points.

  • Activation Functions:

    • Sigmoid: σ(x) = 1/(1+e^{-x}). Smooth, output (0,1). Problems: vanishing gradients, not zero-centered.

    • ReLU (Rectified Linear Unit): f(x) = max(0, x). Computationally cheap, mitigates vanishing gradient (for +ve inputs). Problem: "Dying ReLU" (neurons stuck at 0).

    • Leaky ReLU / Parametric ReLU (PReLU): Fixes dying ReLU by allowing small gradient for x<0.

ANN Architecture

  • Input Layer: Receives feature vectors. Number of neurons = number of features.

  • Hidden Layer(s): Apply non-linear transformations. Depth and width are hyperparameters.

  • Output Layer: Produces final prediction. Activation function depends on task (e.g., Softmax for multi-class).

  • Feedforward Networks (FNNs): Information flows strictly forward, no cycles. Basic MLP is a FNN.

Convolutional Neural Networks (CNNs)

Designed for grid-like data (images).

Layer Type Purpose Key Parameters
Convolutional (CONV) Extract local features (edges, textures) using learnable filters/kernels. Kernel size (FxF), Number of filters (K), Stride (S), Padding (P).
Pooling (MAX/AVG) Downsample, provide translation invariance, reduce parameters. Pool size (FxF), Stride (S).
Fully Connected (FC) Combine high-level features for final classification/regression. Standard ANN layer.

Padding & Stride:

  • Padding (P): Adding zeros around input border to control output spatial size. Output size = (W - F + 2P)/S + 1.

  • Stride (S): Step size of filter movement. Larger stride = smaller output, less computation.

Architectural Variants:

  • VGG-16: Simple, uniform architecture (only 3x3 convs, 2x2 max-pool). Deep (16 layers), but many parameters (~138M).

  • GoogLeNet (Inception Modules): Uses Inception modules with parallel convolutions (1x1, 3x3, 5x5) + pooling, concatenated. Reduces parameters, increases width/depth efficiently.

  • ResNet (Residual Learning): Introduces skip connections (identity mappings). A residual block: output = F(x) + x. Solves vanishing gradient in very deep networks (>100 layers) by allowing gradient to flow directly.

Recurrent Neural Networks (RNNs)

Designed for sequential data (time series, text).

  • Traditional RNN: Has hidden state h_t that captures past information: h_t = f(W_hh * h_{t-1} + W_xh * x_t + b). Backpropagation Through Time (BPTT): Unfolds network through time steps, then applies standard backprop.

  • Challenges:

    • Vanishing/Exploding Gradients: Gradients multiplied repeatedly by weight matrix during BPTT. If eigenvalues <1 -> vanish; >1 -> explode. Hinders learning long-range dependencies.
  • Gated Units (Solutions):

    • LSTM (Long Short-Term Memory): Has cell state C_t (highway) + gates (Input i_t, Forget f_t, Output o_t). Gates regulate information flow, mitigating vanishing gradient.

      C_t = f_t ⊙ C_{t-1} + i_t ⊙ tanh(W_c * x_t + U_c * h_{t-1} + b_c)

    • GRU (Gated Recurrent Unit): Simpler than LSTM. Combines forget and input gates into a single update gate z_t, and has a reset gate r_t. Fewer parameters, often comparable performance.

  • Bidirectional RNNs (BiRNNs): Process sequence forward and backward (with separate RNNs). Hidden states from both directions are concatenated. Provides context from both past and future (useful in NLP, e.g., POS tagging).

Evaluation Metrics for Classification

Metric Formula (from Confusion Matrix) Focus
Confusion Matrix Table of TP, TN, FP, FN. Foundation for all metrics.
Precision TP / (TP + FP) Of predicted positives, how many are correct? (Minimize FP)
Recall (Sensitivity) TP / (TP + FN) Of actual positives, how many were found? (Minimize FN)
F1-Score 2 * (Precision * Recall) / (Precision + Recall) Harmonic mean of Precision & Recall. Balances both.

Key Datasets and Applications

  • MNIST: Handwritten digit dataset (60k train, 10k test, 28x28 grayscale). Application: Benchmark for image classification, testing new architectures.

  • NLP Tasks (e.g., Language Modeling): Predict next word given previous words. Uses RNNs/LSTMs/Transformers. Datasets: Penn Treebank, WikiText. Application: Core for machine translation, text generation.

IV. Probabilistic Reasoning in AI Applications

Bayes' Theorem

Definition: A method for updating the probability of a hypothesis H given new evidence E. Mathematical Formulation:

$$ P(H|E) = \frac{P(E|H) \cdot P(H)}{P(E)} $$

Where:

  • P(H|E) = Posterior: Probability of hypothesis after seeing evidence.

  • P(E|H) = Likelihood: Probability of evidence given hypothesis is true.

  • P(H) = Prior: Initial probability of hypothesis.

  • P(E) = Marginal Likelihood: Total probability of evidence (normalizing constant).

Significance in Uncertain Reasoning:

  • Provides a principled mathematical framework for reasoning under uncertainty.

  • Combines prior knowledge (P(H)) with observed data (P(E|H)).

  • Foundation for Bayesian Networks (graphical models) and Naïve Bayes Classifier.

Applications

  • Diagnostic Reasoning in Expert Systems: E.g., in MYCIN, used certainty factors (heuristic approximation of Bayesian reasoning) to combine evidence from multiple rules.

  • Probabilistic Models in NLP: Naïve Bayes Classifier is widely used for:

    • Spam Filtering: P(Spam|Words) ∝ P(Words|Spam) * P(Spam).

    • Sentiment Analysis: P(Positive|Review).

    • Assumes feature (word) independence given the class (naïve but often effective).

V. Case Studies and Applied Problem-Solving

Game-Playing AI: Tic-Tac-Toe & Minimax

  • Tic-Tac-Toe: Small state space (~9! = 362,880 possible games). Perfect information, zero-sum, deterministic. Ideal for teaching adversarial search.

  • Minimax Search Procedure:

    1. Game Tree: Nodes = game states, Edges = moves. Players: MAX (us) and MIN (opponent).

    2. Recursive Evaluation: MAX chooses move maximizing minimum outcome (worst-case from MIN's best response).

    3. Terminal States: Assign utility (+1 win, 0 draw, -1 loss).

    4. Back up values: MIN nodes take min of children, MAX nodes take max.

    5. Goal: Choose root move leading to state with highest minimax value.

  • Alpha-Beta Pruning: Optimizes Minimax by pruning branches that cannot affect the final decision.

    • α (alpha): Best (highest) value that MAX can guarantee at that point or above.

    • β (beta): Best (lowest) value that MIN can guarantee at that point or below.

    • Prune when: α >= β at any node. The current branch cannot provide a better outcome for the player whose turn it is.

    • Effect: Reduces the number of nodes evaluated from O(b^d) to roughly O(√b^d) (with perfect ordering), where b=branching factor, d=depth. Enables deeper search in same time.

  • Extension to Complex Games (e.g., Chess): Same principles, but:

    • Huge state space (b~35, d~100). Cannot search to terminal states.

    • Use heuristic evaluation functions (e.g., material count, piece-square tables) at non-terminal leaves.

    • Use iterative deepening, transposition tables, move ordering heuristics to maximize pruning efficiency.

Robotics: Block World Problem

  • Problem Definition: A robot arm must rearrange a set of blocks on a table from an initial configuration to a goal configuration. Blocks are clear (no block on top) to be moved.

  • Significance: A classic, simplified model for robotic manipulation and spatial reasoning. Tests planning, kinematics, and perception.

  • Recent Technological Advancements:

    • Advanced Perception: 3D vision (RGB-D sensors, point clouds), deep learning for object recognition/pose estimation.

    • Motion Planning: Sampling-based planners (RRT*, PRM) for high-DOF arms in cluttered spaces.

    • Grasping: Deep learning models (e.g., GQ-CNN) for robust grasp synthesis from vision.

    • Sim-to-Real Transfer: Training policies in simulation (PyBullet, MuJoCo) and transferring to real robots via domain randomization.

    • End-to-End Learning: Learning policies directly from pixels to actions (less common for precise manipulation).

AI for Global Challenges

  • Climate Change: Optimizing energy grids (smart grids), carbon footprint tracking, climate modeling (AI-accelerated simulations), monitoring deforestation (satellite image analysis).

  • Healthcare: Diagnostics: Medical image analysis (CNNs for X-ray, MRI). Treatment: Drug discovery (generative models for molecules), personalized treatment plans (reinforcement learning). Administrative: Automating paperwork.

  • Education: Personalized Learning: Adaptive learning systems that adjust content difficulty based on student performance (recommender systems). Intelligent Tutoring Systems (ITS): Provide customized feedback and hints. Automated Grading: NLP for essay scoring.

Ethical and Social Considerations

  • Bias in AI Systems: Arises from biased training data or algorithm design. Leads to unfair outcomes (e.g., facial recognition accuracy disparity). Mitigation: Diverse datasets, fairness constraints, auditing.

  • Accountability & Transparency: "Black box" models (deep learning) lack explainability. Need: Explainable AI (XAI) techniques (LIME, SHAP). Clear liability frameworks for AI errors (e.g., autonomous vehicle accidents).

  • Societal Impact of Deployment:

    • Job Displacement: Automation of cognitive and physical tasks.

    • Economic Inequality: Potential to widen skill/premium gaps.

    • Autonomous Weapons: Ethical concerns about lethal autonomous systems (LAWS).

    • Surveillance Capitalism: Use of AI for pervasive behavioral tracking and manipulation.

VI. Supporting Knowledge Representation Schemes (Contextual)

Scripts, Schemas, and Frames

Scheme Definition Key Feature Example
Frame A data structure for stereotypical situations. Slots (attributes) with default values. Inheritance from parent frames. Frame: RESTAURANT<br>Slots: [has-part: menu, waiter, food], [customer-role: orders, eats, pays]
Script A specialized frame for temporal sequences of events in a stereotypical activity. Scenes (ordered sub-events), roles, props. Script: EAT-AT-RESTAURANT<br>Scenes: [Enter, Order, Eat, Pay, Exit]
Schema A more general, abstract mental structure for organizing knowledge. Broader than scripts/frames; can represent concepts, events, roles. Schema: CONTAINER has slots: [contains, opening, material].

Distinctions:

  • Frame vs. Script: Frame is static (what is), Script is dynamic (what happens). A script uses frames for its roles/props.

  • Schema: Often used interchangeably with frame in AI literature, but in psychology, schema is more abstract and flexible.

Application in Structuring Domain Knowledge: Provides expectations and defaults. Allows AI systems to:

  1. Fill in missing information (default slot values).

  2. Understand incomplete input by matching to known patterns.

  3. Generate coherent plans or narratives.

Semantic Networks

  • Graphical Representation: Nodes = concepts/objects/events. Edges = directed, labeled relationships (e.g., IS-A, PART-OF, INSTANCE-OF, HAS-PROPERTY).

  • Use in Knowledge Modeling:

    • Lexical Semantics: WordNet (synsets, hypernym/hyponym relations).

    • Inference: Traverse graph for inheritance (e.g., "Canary IS-A bird" + "Birds CAN fly" → "Canary CAN fly" unless overridden).

    • Natural Language Understanding: Represent meaning of sentences (e.g., [John] -subject- [eat] -object- [apple]).

  • Limitations: Cannot represent complex logical statements (negation, disjunction) easily. Lack of standard representation for procedures/actions.

Conceptual Dependency (CD)

  • Primitive Acts: Small set of language-independent, universal action primitives (e.g., ATRANS - abstract transfer, PTRANS - physical transfer, MTRANS - mental transfer, GRASP).

  • Conceptual Cases: Roles filled by concepts (e.g., ACTOR, OBJECT, DESTINATION, INSTRUMENT).

  • Representation Example: "John gave Mary a book."

    ATRANS(ACTOR=John, OBJECT=book, TO=Mary)

  • Role in NLP Understanding: Aims to capture the deep, underlying meaning of a sentence, independent of surface wording. Useful for paraphrase recognition, machine translation (interlingua approach), and question answering.

  • Comparison with Semantic Nets: CD is more procedural and action-oriented, using a fixed set of primitives. Semantic nets are more declarative and relational, better for static knowledge. CD is harder to scale due to primitive set design.


Exam Tips & Common Pitfalls:

  • Expert Systems: Be ready to differentiate forward/backward chaining with domain examples. Know the components and the knowledge bottleneck.
  • NLP: Distinguish Syntax, Semantics, Pragmatics. Know scripts vs. frames and their role in dialogue systems. Be prepared to discuss ethical issues (bias, privacy) with examples.
  • Neural Networks: Backpropagation derivation is often asked. Know ReLU vs. Sigmoid pros/cons. CNN layers (CONV, POOL, FC) and their parameters (kernel, stride, padding) are crucial. LSTM vs. GRU differences (gates, parameters).
  • Bayes' Theorem: Write the formula boxed. Apply to simple diagnostic or classification problems (Naïve Bayes).
  • Game Playing: Be able to trace Minimax and Alpha-Beta on a small tree. Know why Alpha-Beta prunes (α >= β). Understand the challenge of scaling to chess (heuristics, depth).
  • Knowledge Representation: Compare Production Rules vs. Predicate Logic (expressiveness, inference). Know resolution steps (convert to CNF, unify, resolve). Contrast Semantic Nets vs. CD.
  • Block World: Describe the problem and modern solutions (perception, planning, grasping).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in