UNIT 3: DATA WAREHOUSING & MINING - EXAM-FOCUSED SHORT NOTES
I. DATA WAREHOUSE FUNDAMENTALS & ARCHITECTURE
A. Introduction & Characteristics
Definition: A Data Warehouse (DW) is a subject-oriented, integrated, time-variant, and non-volatile collection of data in support of management's decision-making process.
Need: To support OLAP (Online Analytical Processing) for complex queries, historical analysis, and business intelligence, separate from OLTP (Online Transaction Processing) systems.
| Characteristic | Description |
|---|---|
| Subject-Oriented | Organized around key subjects (e.g., Customer, Product, Sales) rather than applications. |
| Integrated | Data from multiple, heterogeneous sources (e.g., relational DB, flat files) is consolidated with consistent naming, encoding, and formats. |
| Time-Variant | Data is stored with a time dimension, providing historical perspective (e.g., daily, monthly snapshots). |
| Non-Volatile | Data is stable; once entered, it is not updated/deleted but only appended and accessed. |
B. Data Warehouse Architecture (Three-Tier)
-
Bottom Tier (Data Source & Staging):
-
Data Sources: Operational databases, external sources, legacy systems.
-
Staging Area: Temporary workspace for data cleaning, integration, and transformation before loading into the DW.
-
-
Middle Tier (Data Warehouse & Data Marts):
-
Data Warehouse: Central repository storing integrated, subject-oriented, historical data.
-
Data Marts: Subsets of the DW tailored for specific business lines (e.g., Sales DM, Finance DM). Can be dependent (from DW) or independent.
-
Metadata: "Data about data." Stores definitions, source, transformation rules, and usage info. Crucial for DW management.
-
-
Top Tier (Access Tools):
-
OLAP Tools: For multidimensional analysis (e.g., roll-up, drill-down).
-
Query & Reporting Tools: For ad-hoc queries and standard reports.
-
Data Mining Tools: For pattern discovery.
-
Client/Server Architecture: GUI front-end connects to OLAP server/database.
-
DiagramCANVAS: Draw a three-layer pyramid. Bottom: "Data Sources & Staging Area". Middle: "Data Warehouse & Data Marts (with Metadata)". Top: "OLAP, Query, Mining Tools (Client/Server)". Arrows show flow from bottom to top.
C. Implementation Approaches & Techniques
-
Vertical Partitioning: Splitting a table vertically (by columns) into multiple tables based on access frequency.
-
Need: To improve query performance by separating frequently accessed (hot) columns from infrequently accessed (cold) columns.
-
Method:
-
Identify columns with high/low access frequency.
-
Create a main table with primary key and hot columns.
-
Create one or more extension tables with the same primary key and cold columns.
-
Queries needing only hot columns access the smaller main table.
-
-
Trade-off: Reduces I/O for common queries but adds join overhead for queries needing both hot and cold columns.
-
D. Schema Design for Multidimensional Databases
| Schema | Structure | Normalization | Advantages | Disadvantages |
|---|---|---|---|---|
| Star Schema | Single Fact Table connected to multiple, denormalized Dimension Tables. Fact table contains foreign keys to dimensions and measures. | Dimensions are denormalized (redundant). | Simple, high query performance (fewer joins), easy for end-users. | Data redundancy, update anomalies. |
| Snowflake Schema | Fact table connected to multiple, normalized Dimension Tables. Dimensions are broken into related sub-dimensions. | Dimensions are normalized (3NF). | Reduces redundancy, easier to maintain, saves storage. | Complex queries (more joins), poorer performance than Star. |
| Galaxy Schema (Fact Constellation) | Multiple Fact Tables share common Dimension Tables. A collection of star schemas. | Mixed; dimensions can be normalized or denormalized. | Handles complex subjects with multiple processes (e.g., Sales & Returns). Reuses dimensions. | Most complex design, potential for inconsistent facts, difficult to navigate. |
Key Terms:
- Fact Table: Central table containing measures (quantitative data like sales_amount) and foreign keys to dimension tables.
- Dimension Table: Contains descriptive attributes (e.g., product_name, customer_city) used to query, group, and filter facts.
II. ONLINE ANALYTICAL PROCESSING (OLAP)
A. OLAP Fundamentals
Definition: OLAP is a category of software tools that provide analysis of data stored in a DW or other data store. It enables users to analyze data from multiple perspectives.
OLAP vs. OLTP:
| Feature | OLTP | OLAP |
| :--- | :--- | :--- |
| Purpose | Day-to-day operations (insert, update, delete). | Decision support, historical analysis. |
| Data | Current, detailed, normalized. | Historical, summarized, denormalized (DW). |
| Queries | Simple, short, fast (e.g., "Update customer address"). | Complex, long-running, ad-hoc (e.g., "Sales by region, product, quarter"). |
| Users | Clerks, customers. | Managers, executives, analysts. |
Core OLAP Operations:
- Roll-up (Drill-up): Aggregating data (e.g., city → state → country).
- Drill-down: Getting finer details (e.g., country → state → city).
- Slice-and-dice: Selecting a subset (slice) and viewing it across different dimensions (dice).
- Pivot (Rotate): Rotating the data cube to view from a different dimensional perspective.
B. OLAP Architectures & Servers
| Type | Storage | Working | Advantages | Disadvantages |
|---|---|---|---|---|
| MOLAP<br>(Multidimensional OLAP) | Proprietary multidimensional array storage (cube). | Data is pre-aggregated and stored in an optimized cube format. | Fast query response (no joins), optimized for complex calculations. | Limited scalability (cube size), lengthy ETL/cube build time, data latency. |
| ROLAP<br>(Relational OLAP) | Relational Database (tables). | Maps multidimensional logic to relational tables (fact & dimension tables). Uses SQL for queries. | High scalability (handles large data volumes), uses standard RDBMS, no data latency. | Slower query response (complex SQL joins), performance depends on RDBMS optimization. |
| HOLAP<br>(Hybrid OLAP) | Combination: detailed data in RDBMS, aggregated data in MOLAP cube. | Balances between MOLAP and ROLAP. | Good balance of performance and scalability. | Complexity in management and synchronization. |
DiagramCANVAS: Draw two side-by-side diagrams.
Left (ROLAP): Box "OLAP Server" -> arrow to "Relational DB" showing tables: FACT (with FKs) and DIMENSION tables. Label: "SQL Queries".
Right (MOLAP): Box "OLAP Server" -> arrow to "Multidimensional Storage (Cube)". Label: "Pre-aggregated arrays".
C. Data Cube Computation
-
Goal: Pre-compute all possible aggregate queries (group-bys) for a set of dimensions.
-
Lattice of Cuboids: A lattice is a directed acyclic graph where nodes represent cuboids (subsets of dimensions) and edges represent roll-up/drill-down operations.
- Example for dimensions {A, B, C}: Base cuboid (A,B,C) → (A,B), (A,C), (B,C) → (A), (B), (C) → apex (∅).
-
Computation Strategies:
-
Full Cube: Compute all cuboids. Computationally expensive ($$\displaystyle 2^n $$ for n dimensions).
-
Iceberg Cube: Compute only cuboids where measure (e.g., count) meets a minimum support threshold. Prunes low-interest aggregates.
-
Closed Cube: Stores only closed itemsets (no superset with same support). More compact than full cube.
-
III. DATA PREPROCESSING FOR MINING
A. Data Cleaning
-
Missing Data:
-
Deletion: Remove tuples with missing values (only if few missing).
-
Imputation: Fill with mean/median/mode, or use predictive models (e.g., regression).
-
-
Noisy Data (Smoothing):
-
Binning: Sort values, partition into bins, smooth by bin mean/median/boundaries.
-
Regression: Fit a regression function (linear, multi-linear) to smooth data.
-
Clustering: Group similar values, treat outliers as noise.
-
-
Outlier Detection: Use boxplot (1.5*IQR), clustering, or statistical tests.
B. Data Integration & Transformation
-
Transformation Methods:
-
Smoothing: Remove noise (as above).
-
Attribute Construction: Create new attributes from existing ones (e.g., area = length * width).
-
Aggregation: Summarize data (e.g., daily sales → monthly sales).
-
Normalization: Scale attributes to a small range.
-
Min-Max: $$\displaystyle x' = \frac{x - \min}{\max - \min} $$ → [0,1]
-
Z-Score: $$\displaystyle x' = \frac{x - \mu}{\sigma} $$ → mean=0, std=1
-
Decimal Scaling: $$\displaystyle x' = \frac{x}{10^j} $$ (j smallest integer s.t. max|x'| < 1)
-
-
Discretization/Binning: Convert continuous to categorical (e.g., age → youth, adult, senior).
-
C. Data Mining Task Primitives
These define a data mining query/task:
-
Task-relevant data: The subset of the database to be mined (e.g.,
SELECT * FROM sales WHERE year > 2020). -
Background knowledge: Domain knowledge, constraints, taxonomies (e.g., "product hierarchy").
-
Interestingness measures: Metrics to evaluate patterns (e.g., support, confidence, lift for association rules; accuracy for classification).
-
Presentation/Visualization knowledge: How to display results (e.g., rules in a table, tree graph, cluster map).
IV. DATA MINING FUNDAMENTALS & KDD PROCESS
A. Introduction to Data Mining
Definition: The process of discovering interesting patterns and knowledge from large amounts of data. It is the core step of the KDD process.
KDD vs. Data Mining: KDD is the overall process (Selection → Preprocessing → Transformation → Data Mining → Interpretation/Evaluation). Data Mining is the application of algorithms to extract patterns.
Pros and Cons:
| Pros (Benefits) | Cons (Challenges) |
| :--- | :--- |
| Predictive capabilities, automated discovery, hidden pattern detection, improved decision-making. | Privacy concerns, data quality issues, misinterpretation of patterns, scalability, ethical misuse. |
Types of Data:
- Relational: Tables with rows/columns.
- Transactional: Each tuple is a transaction (e.g., market basket).
- Data Warehouse: Integrated, subject-oriented, historical (multidimensional).
- Advanced: Text, Web, Spatial, Time-series, Stream data.
B. KDD Process Model
-
Selection: Define the problem, select target data.
-
Preprocessing: Cleaning, handling missing values.
-
Transformation: Normalization, attribute construction, aggregation.
-
Data Mining: Apply core algorithms (association, classification, clustering).
-
Interpretation/Evaluation: Evaluate patterns for validity, usefulness, and visualize results.
Role of Data Mining Engine: It is the core algorithmic component of the KDD process (Step 4). It executes the chosen data mining methods (e.g., Apriori, Decision Tree, k-Means) on the prepared data to generate the initial set of patterns/models.
V. ASSOCIATION RULE MINING
A. Core Concepts
Itemset: A set of items (e.g., {bread, milk}).
- k-itemset: An itemset with k items.
Support (σ): Fraction of transactions that contain the itemset.
$$\boxed{\text{support}(X) = \frac{\sigma(X)}{N}}$$
where $N$ = total transactions.
Confidence: Conditional probability that a transaction containing X also contains Y.
$$\boxed{\text{confidence}(X \rightarrow Y) = \frac{\text{support}(X \cup Y)}{\text{support}(X)}$$
Frequent Itemset: An itemset with support ≥ min_sup threshold.
Association Rule: An implication of the form $$\displaystyle X \rightarrow Y $$, where $$\displaystyle X \cap Y = \emptyset $$.
- Strong Rule: Meets min_sup and min_conf thresholds.
B. Apriori Algorithm
Key Principle (Apriori Property): All subsets of a frequent itemset must also be frequent. (If {A,B,C} is frequent, then {A,B}, {A,C}, {B,C} must be frequent). Steps:
-
Find frequent itemsets (L_k):
-
C_k: Generate candidate k-itemsets by joining L_{k-1} with itself.
-
Prune: Remove candidates with any (k-1)-subset not in L_{k-1} (using Apriori property).
-
Count support: Scan DB, count support of remaining candidates.
-
L_k: Keep candidates with support ≥ min_sup.
-
Repeat for k=1,2,... until L_k is empty.
-
-
Generate strong rules: For each frequent itemset l (|l| ≥ 2), generate all non-empty subsets. For each subset A, rule A → (l - A) if confidence ≥ min_conf.
Example: For min_sup=50% (2/4 trans), min_conf=70%.
Transactions: {A,B,C}, {A,C}, {B,C}, {A,B,D}.
- L1 (frequent 1-itemsets): {A}(75%), {B}(50%), {C}(75%). D is infrequent.
- C2: {A,B}, {A,C}, {B,C}. Prune? All 1-subsets are frequent → keep.
- L2: {A,C}(50%), {B,C}(50%). {A,B} support=25% → pruned.
- Rule from {A,C}: A→C (conf=50%/75%=66.7% <70%), C→A (conf=50%/75%=66.7% <70%) → no strong rule.
- Rule from {B,C}: B→C (conf=50%/50%=100% ✓), C→B (conf=50%/75%=66.7% ✗). Strong rule: B → C.
C. Advanced Association Mining
-
FP-Growth Algorithm:
-
Scan DB once to get frequent items (L1).
-
Build FP-Tree:
-
Order items in each transaction by descending frequency (from L1).
-
Insert ordered transaction into a prefix tree (trie). Share common prefixes. Store count at each node.
-
-
Mine FP-Tree recursively:
-
For each item in header table (starting from least frequent):
-
Construct conditional pattern base: all paths from root to the item.
-
Build conditional FP-Tree from pattern base.
-
If tree has single path → generate all combos; else, recursively mine.
-
-
Combine suffix item with patterns from conditional tree to form frequent patterns.
-
- Advantage over Apriori: No candidate generation, only 2 DB scans.
-
-
Techniques to Improve Apriori Efficiency:
-
Hash-based technique: Use hash table to prune candidate k-itemsets during generation.
-
Transaction reduction: Remove transactions that don't contain any frequent items.
-
Partitioning: Partition DB, find local frequent itemsets, then global.
-
Sampling: Mine a random sample, then verify on full DB.
-
Dynamic item counting: Add candidate itemsets during a DB scan based on partial counts.
-
VI. CLASSIFICATION
A. Classification Process & Model
-
Learning (Training): Build a model (classifier) from training data (tuples with known class labels).
-
Testing: Evaluate model on unseen test data. Predict class labels and compare to actual.
-
Prediction (Application): Use trained model to classify new, unseen instances.
B. Decision Tree Induction (e.g., ID3/C4.5)
-
Algorithm (ID3):
-
Start with all training instances at root.
-
If all instances belong to same class, make leaf with that class.
-
Else, select the best attribute (A) to split on using Information Gain.
-
Create a branch for each value of A, partition instances.
-
Recurse on each branch with remaining attributes.
-
-
Attribute Selection Measures:
-
Information Gain (ID3): Based on Entropy.
-
Entropy(S) = $$\displaystyle -\sum_{i=1}^{c} p_i \log_2 p_i $$ (p_i = proportion of class i in S).
-
Gain(S, A) = Entropy(S) - $$\displaystyle \sum_{v \in Values(A)} \frac{|S_v|}{|S|} Entropy(S_v) $$.
-
Choose attribute with highest Gain.
-
-
Gain Ratio (C4.5): Corrects Information Gain's bias toward many-valued attributes.
-
SplitInfo(S, A) = $$\displaystyle -\sum_{v} \frac{|S_v|}{|S|} \log_2 \frac{|S_v|}{|S|} $$
-
GainRatio(S, A) = Gain(S, A) / SplitInfo(S, A).
-
Choose attribute with highest Gain Ratio.
-
-
Gini Index (CART): Measures impurity.
-
Gini(S) = $$\displaystyle 1 - \sum_{i=1}^{c} p_i^2 $$.
-
For binary split on A: Gini_A(S) = $$\displaystyle \frac{|S_1|}{|S|} Gini(S_1) + \frac{|S_2|}{|S|} Gini(S_2) $$.
-
Choose attribute/split that minimizes weighted Gini.
-
-
-
Tree Pruning:
-
Pre-pruning: Stop tree growth early (e.g., min samples per leaf, max depth).
-
Post-pruning: Grow full tree, then remove branches (e.g., error-based, cost-complexity).
-
C. Other Classification Methods
-
Rule-Based (e.g., RIPPER):
-
Sequential Covering: Learn rules one at a time.
-
Start with empty rule set.
-
Learn a rule that covers many instances of a class (e.g., using separate-and-conquer).
-
Remove covered instances.
-
Repeat until stopping condition.
-
-
RIPPER: Optimized version of sequential covering for efficiency.
-
-
Bayesian Classification (Naïve Bayes):
- Based on Bayes' Theorem:
$$P(C|X) = \frac{P(C) P(X|C)}{P(X)}$$
where C = class, X = attribute tuple.
* **"Naïve" Assumption:** Attributes are conditionally independent given the class.
$$P(X|C) = \prod_{i=1}^{n} P(x_i|C)$$
* **Classifier:** Assign class C that maximizes $$\displaystyle P(C) \prod_{i} P(x_i|C) $$.
* **Advantage:** Simple, fast, works well despite independence assumption.
D. Classifier Evaluation
-
Process of Verifying Accuracy:
-
Holdout: Split data into training set (e.g., 2/3) and test set (1/3). Train on train, test on test.
-
Cross-Validation (k-fold): Partition data into k equal folds. Train on k-1 folds, test on 1 fold. Repeat k times, average accuracy.
-
Bootstrapping: Sample n instances with replacement from original data (size n) to form training set. Test on left-out instances. Repeat.
-
-
Metrics (from Confusion Matrix):
| | Predicted Positive | Predicted Negative | | :--- | :--- | :--- | | Actual Positive | TP (True Positive) | FN (False Negative) | | Actual Negative | FP (False Positive) | TN (True Negative) |
-
Accuracy: $$\displaystyle \frac{TP+TN}{Total} $$
-
Precision (P): $$\displaystyle \frac{TP}{TP+FP} $$ (Of predicted positives, how many correct?)
-
Recall (R): $$\displaystyle \frac{TP}{TP+FN} $$ (Of actual positives, how many found?)
-
F1-Score: $$\displaystyle 2 \times \frac{P \times R}{P + R} $$ (Harmonic mean of P & R).
-
ROC Curve: Plots True Positive Rate (Recall) vs False Positive Rate ($$\displaystyle \frac{FP}{FP+TN} $$) at different classification thresholds. Area Under Curve (AUC) measures overall performance.
-
VII. CLUSTERING
A. Clustering Fundamentals
Definition: Grouping data objects into clusters such that objects within a cluster are similar to each other and dissimilar to objects in other clusters.
Types of Clusters:
- Partitional: Divides data into non-overlapping subsets (e.g., k-Means).
- Hierarchical: Creates a tree of clusters (dendrogram). Can be agglomerative (bottom-up) or divisive (top-down).
- Density-based: Clusters are dense regions separated by sparse regions (e.g., DBSCAN).
- Grid-based: Objects are mapped to a grid structure (e.g., STING).
- Model-based: Assumes data follows a probability distribution (e.g., Gaussian Mixture Models).
B. Partitioning Methods: k-Means
-
Algorithm:
-
Choose k initial centroids (randomly or heuristically).
-
Repeat:
-
Assignment: Assign each point to the nearest centroid (using Euclidean distance).
-
Update: Recalculate centroids as mean of all points in the cluster.
-
-
Until centroids no longer change (or max iterations).
-
-
Example: k=2, points: (1,1), (1,2), (2,1), (5,5), (6,5). Initial centroids: C1=(1,1), C2=(5,5). Assign → Update → Converge.
-
Limitations:
-
Requires k to be specified.
-
Sensitive to initial centroids (may converge to local optimum).
-
Assumes spherical clusters of similar size/density.
-
Sensitive to outliers.
-
C. Hierarchical Methods
-
Agglomerative (Bottom-up):
-
Start: each point as its own cluster.
-
Merge the two closest clusters.
-
Repeat until one cluster remains (or desired k).
-
-
Divisive (Top-down):
-
Start: all points in one cluster.
-
Split a cluster into two (e.g., using k-Means).
-
Repeat until each point is its own cluster (or desired k).
-
-
Linkage Criteria (for Agglomerative): Defines "closeness" between clusters.
-
Single Linkage: Min distance between points in two clusters. → Chaining effect.
-
Complete Linkage: Max distance between points. → Compact clusters.
-
Average Linkage: Average distance between all pairs.
-
Centroid Linkage: Distance between cluster centroids.
-
-
Dendrogram: Tree diagram showing the sequence of merges/splits. Cutting the dendrogram at a level gives the cluster partition.
D. Density-Based: DBSCAN
-
Core Concepts:
-
ε-neighborhood (eps): Radius around a point.
-
MinPts: Minimum number of points in ε-neighborhood to define a core point.
-
Directly density-reachable: Point p is directly reachable from q if q is core and p is in q's ε-neighborhood.
-
Density-connected: Two points are connected if there is a path of directly reachable points.
-
-
Algorithm:
-
Arbitrarily pick a point p.
-
If p is core (≥ MinPts in ε-neighborhood), create new cluster, add all points density-reachable from p.
-
If p is not core, mark as noise.
-
Repeat until all points processed.
-
-
Advantages:
-
Discovers clusters of arbitrary shape.
-
Robust to outliers (noise points not assigned to any cluster).
-
Does not require specifying number of clusters (k).
-
VIII. ADVANCED & SPECIALIZED DATA MINING APPLICATIONS
A. Web Usage Mining
-
Goal: Discover patterns from web logs to understand user behavior, improve website structure, personalize content, and for marketing.
-
Data Sources:
-
Server logs: Client requests, IP, timestamp, URL, status.
-
Client-side logs: Cookies, user profiles (if logged in).
-
Proxy logs: Requests from multiple users behind a firewall.
-
-
Major Tasks:
-
Session/Transaction Identification: Group user hits into sessions (e.g., 30-min timeout).
-
Path Analysis: Frequent navigation paths (e.g., Home → Products → Electronics → Checkout).
-
Association & Clustering of Pages: Find pages often visited together; cluster pages into topics.
-
B. Text Mining
-
Definition: Applying data mining techniques to unstructured text documents to discover patterns.
-
Difference from Data Mining: Input is text, not structured data. Requires significant preprocessing.
-
Major Tasks & Preprocessing:
-
Text Preprocessing:
-
Tokenization: Split text into words/tokens.
-
Stop-word removal: Remove common words (the, is, and).
-
Stemming: Reduce words to root form (e.g., "running" → "run").
-
Lemmatization: Reduce to dictionary form (e.g., "better" → "good").
-
-
Document Representation:
- Vector Space Model: Document = vector of term weights (e.g., TF-IDF: Term Frequency-Inverse Document Frequency).
-
Tasks:
-
Text Categorization: Assign documents to predefined classes (e.g., spam/ham, news topics).
-
Text Clustering: Group similar documents without predefined classes.
-
Text Summarization: Generate a concise summary of a document.
-
-
C. Spatial Data Mining
-
Definition: Discovering patterns from spatial data (data with geographic/location reference).
-
Characteristics: Spatial autocorrelation (nearby objects are more similar), complex data types (points, lines, polygons), spatial relationships (topological, distance, direction).
-
Major Tasks:
-
Spatial Association Rules: Rules with spatial predicates (e.g.,
near,intersects).A(x) ∧ B(y) → C(z)where x, y, z are spatial objects. -
Spatial Classification: Classify spatial objects using both non-spatial attributes and spatial attributes (e.g., classify land use type using soil type (non-spatial) and proximity to river (spatial)).
-
Spatial Clustering: Cluster objects based on spatial proximity and optionally non-spatial attributes (e.g., DBSCAN for earthquake epicenters).
-
D. Data Prediction
-
General Concept: Building a model to predict the value of a target (dependent) variable based on other (independent) variables.
-
Relationship to Classification & Regression:
-
Classification is a type of prediction where the target is categorical (discrete class labels: yes/no, spam/ham).
-
Regression is prediction where the target is continuous/real-valued (e.g., house price, temperature).
-
-
Common Techniques: Linear Regression, Polynomial Regression, Decision Trees (for both classification & regression trees - CART), Neural Networks. The evaluation metrics differ (e.g., MSE, RMSE for regression vs. Accuracy for classification).
Exam Tips & Common Pitfalls:
- Schema Design: Be ready to draw and contrast Star, Snowflake, and Galaxy schemas. Know which is most normalized (Snowflake) and which is best for query performance (Star).
- OLAP: Clearly distinguish MOLAP (cube, fast) vs ROLAP (relational, scalable). Be able to sketch the ROLAP mapping.
- Apriori: Practice a small example (4-5 transactions) showing candidate generation, pruning, and rule generation. Remember the Apriori Property.
- FP-Growth: Understand the two-step process: (1) Build FP-Tree (order by freq, share prefixes), (2) Mine recursively using conditional pattern bases. Know its main advantage: no candidate generation.
- Decision Trees: Be fluent in calculating Information Gain and Gini Index for a small dataset. Know when to use Gain Ratio (many-valued attributes).
- Clustering: Differentiate k-Means (partitional, centroid-based) vs DBSCAN (density-based, handles noise). Know single vs complete linkage effects.
- KDD: Remember the 5 steps and that Data Mining is Step 4. The Data Mining Engine is the algorithmic core.
- Vertical Partitioning: Explain it as splitting by columns to separate hot/cold data, improving I/O for common queries at the cost of joins.