Skip to content
CY-604 (C) · Dataware Housing & Mining/Quick Revision Short Notes

Dataware Housing & Mining (CY-604 (C)) - Unit 3 Short Notes

UNIT 3: DATA WAREHOUSING & MINING - EXAM-FOCUSED SHORT NOTES


I. DATA WAREHOUSE FUNDAMENTALS & ARCHITECTURE

A. Introduction & Characteristics

Definition: A Data Warehouse (DW) is a subject-oriented, integrated, time-variant, and non-volatile collection of data in support of management's decision-making process.

Need: To support OLAP (Online Analytical Processing) for complex queries, historical analysis, and business intelligence, separate from OLTP (Online Transaction Processing) systems.

Characteristic Description
Subject-Oriented Organized around key subjects (e.g., Customer, Product, Sales) rather than applications.
Integrated Data from multiple, heterogeneous sources (e.g., relational DB, flat files) is consolidated with consistent naming, encoding, and formats.
Time-Variant Data is stored with a time dimension, providing historical perspective (e.g., daily, monthly snapshots).
Non-Volatile Data is stable; once entered, it is not updated/deleted but only appended and accessed.

B. Data Warehouse Architecture (Three-Tier)

  1. Bottom Tier (Data Source & Staging):

    • Data Sources: Operational databases, external sources, legacy systems.

    • Staging Area: Temporary workspace for data cleaning, integration, and transformation before loading into the DW.

  2. Middle Tier (Data Warehouse & Data Marts):

    • Data Warehouse: Central repository storing integrated, subject-oriented, historical data.

    • Data Marts: Subsets of the DW tailored for specific business lines (e.g., Sales DM, Finance DM). Can be dependent (from DW) or independent.

    • Metadata: "Data about data." Stores definitions, source, transformation rules, and usage info. Crucial for DW management.

  3. Top Tier (Access Tools):

    • OLAP Tools: For multidimensional analysis (e.g., roll-up, drill-down).

    • Query & Reporting Tools: For ad-hoc queries and standard reports.

    • Data Mining Tools: For pattern discovery.

    • Client/Server Architecture: GUI front-end connects to OLAP server/database.

DiagramCANVAS: Draw a three-layer pyramid. Bottom: "Data Sources & Staging Area". Middle: "Data Warehouse & Data Marts (with Metadata)". Top: "OLAP, Query, Mining Tools (Client/Server)". Arrows show flow from bottom to top.

C. Implementation Approaches & Techniques

  • Vertical Partitioning: Splitting a table vertically (by columns) into multiple tables based on access frequency.

    • Need: To improve query performance by separating frequently accessed (hot) columns from infrequently accessed (cold) columns.

    • Method:

      1. Identify columns with high/low access frequency.

      2. Create a main table with primary key and hot columns.

      3. Create one or more extension tables with the same primary key and cold columns.

      4. Queries needing only hot columns access the smaller main table.

    • Trade-off: Reduces I/O for common queries but adds join overhead for queries needing both hot and cold columns.

D. Schema Design for Multidimensional Databases

Schema Structure Normalization Advantages Disadvantages
Star Schema Single Fact Table connected to multiple, denormalized Dimension Tables. Fact table contains foreign keys to dimensions and measures. Dimensions are denormalized (redundant). Simple, high query performance (fewer joins), easy for end-users. Data redundancy, update anomalies.
Snowflake Schema Fact table connected to multiple, normalized Dimension Tables. Dimensions are broken into related sub-dimensions. Dimensions are normalized (3NF). Reduces redundancy, easier to maintain, saves storage. Complex queries (more joins), poorer performance than Star.
Galaxy Schema (Fact Constellation) Multiple Fact Tables share common Dimension Tables. A collection of star schemas. Mixed; dimensions can be normalized or denormalized. Handles complex subjects with multiple processes (e.g., Sales & Returns). Reuses dimensions. Most complex design, potential for inconsistent facts, difficult to navigate.

Key Terms:

  • Fact Table: Central table containing measures (quantitative data like sales_amount) and foreign keys to dimension tables.
  • Dimension Table: Contains descriptive attributes (e.g., product_name, customer_city) used to query, group, and filter facts.

II. ONLINE ANALYTICAL PROCESSING (OLAP)

A. OLAP Fundamentals

Definition: OLAP is a category of software tools that provide analysis of data stored in a DW or other data store. It enables users to analyze data from multiple perspectives.

OLAP vs. OLTP:

| Feature | OLTP | OLAP |

| :--- | :--- | :--- |

| Purpose | Day-to-day operations (insert, update, delete). | Decision support, historical analysis. |

| Data | Current, detailed, normalized. | Historical, summarized, denormalized (DW). |

| Queries | Simple, short, fast (e.g., "Update customer address"). | Complex, long-running, ad-hoc (e.g., "Sales by region, product, quarter"). |

| Users | Clerks, customers. | Managers, executives, analysts. |

Core OLAP Operations:

  1. Roll-up (Drill-up): Aggregating data (e.g., city → state → country).
  1. Drill-down: Getting finer details (e.g., country → state → city).
  1. Slice-and-dice: Selecting a subset (slice) and viewing it across different dimensions (dice).
  1. Pivot (Rotate): Rotating the data cube to view from a different dimensional perspective.

B. OLAP Architectures & Servers

Type Storage Working Advantages Disadvantages
MOLAP<br>(Multidimensional OLAP) Proprietary multidimensional array storage (cube). Data is pre-aggregated and stored in an optimized cube format. Fast query response (no joins), optimized for complex calculations. Limited scalability (cube size), lengthy ETL/cube build time, data latency.
ROLAP<br>(Relational OLAP) Relational Database (tables). Maps multidimensional logic to relational tables (fact & dimension tables). Uses SQL for queries. High scalability (handles large data volumes), uses standard RDBMS, no data latency. Slower query response (complex SQL joins), performance depends on RDBMS optimization.
HOLAP<br>(Hybrid OLAP) Combination: detailed data in RDBMS, aggregated data in MOLAP cube. Balances between MOLAP and ROLAP. Good balance of performance and scalability. Complexity in management and synchronization.

DiagramCANVAS: Draw two side-by-side diagrams.

Left (ROLAP): Box "OLAP Server" -> arrow to "Relational DB" showing tables: FACT (with FKs) and DIMENSION tables. Label: "SQL Queries".

Right (MOLAP): Box "OLAP Server" -> arrow to "Multidimensional Storage (Cube)". Label: "Pre-aggregated arrays".

C. Data Cube Computation

  • Goal: Pre-compute all possible aggregate queries (group-bys) for a set of dimensions.

  • Lattice of Cuboids: A lattice is a directed acyclic graph where nodes represent cuboids (subsets of dimensions) and edges represent roll-up/drill-down operations.

    • Example for dimensions {A, B, C}: Base cuboid (A,B,C) → (A,B), (A,C), (B,C) → (A), (B), (C) → apex (∅).
  • Computation Strategies:

    • Full Cube: Compute all cuboids. Computationally expensive ($$\displaystyle 2^n $$ for n dimensions).

    • Iceberg Cube: Compute only cuboids where measure (e.g., count) meets a minimum support threshold. Prunes low-interest aggregates.

    • Closed Cube: Stores only closed itemsets (no superset with same support). More compact than full cube.


III. DATA PREPROCESSING FOR MINING

A. Data Cleaning

  • Missing Data:

    • Deletion: Remove tuples with missing values (only if few missing).

    • Imputation: Fill with mean/median/mode, or use predictive models (e.g., regression).

  • Noisy Data (Smoothing):

    • Binning: Sort values, partition into bins, smooth by bin mean/median/boundaries.

    • Regression: Fit a regression function (linear, multi-linear) to smooth data.

    • Clustering: Group similar values, treat outliers as noise.

  • Outlier Detection: Use boxplot (1.5*IQR), clustering, or statistical tests.

B. Data Integration & Transformation

  • Transformation Methods:

    • Smoothing: Remove noise (as above).

    • Attribute Construction: Create new attributes from existing ones (e.g., area = length * width).

    • Aggregation: Summarize data (e.g., daily sales → monthly sales).

    • Normalization: Scale attributes to a small range.

      • Min-Max: $$\displaystyle x' = \frac{x - \min}{\max - \min} $$ → [0,1]

      • Z-Score: $$\displaystyle x' = \frac{x - \mu}{\sigma} $$ → mean=0, std=1

      • Decimal Scaling: $$\displaystyle x' = \frac{x}{10^j} $$ (j smallest integer s.t. max|x'| < 1)

    • Discretization/Binning: Convert continuous to categorical (e.g., age → youth, adult, senior).

C. Data Mining Task Primitives

These define a data mining query/task:

  1. Task-relevant data: The subset of the database to be mined (e.g., SELECT * FROM sales WHERE year > 2020).

  2. Background knowledge: Domain knowledge, constraints, taxonomies (e.g., "product hierarchy").

  3. Interestingness measures: Metrics to evaluate patterns (e.g., support, confidence, lift for association rules; accuracy for classification).

  4. Presentation/Visualization knowledge: How to display results (e.g., rules in a table, tree graph, cluster map).


IV. DATA MINING FUNDAMENTALS & KDD PROCESS

A. Introduction to Data Mining

Definition: The process of discovering interesting patterns and knowledge from large amounts of data. It is the core step of the KDD process.

KDD vs. Data Mining: KDD is the overall process (Selection → Preprocessing → Transformation → Data Mining → Interpretation/Evaluation). Data Mining is the application of algorithms to extract patterns.

Pros and Cons:

| Pros (Benefits) | Cons (Challenges) |

| :--- | :--- |

| Predictive capabilities, automated discovery, hidden pattern detection, improved decision-making. | Privacy concerns, data quality issues, misinterpretation of patterns, scalability, ethical misuse. |

Types of Data:

  • Relational: Tables with rows/columns.
  • Transactional: Each tuple is a transaction (e.g., market basket).
  • Data Warehouse: Integrated, subject-oriented, historical (multidimensional).
  • Advanced: Text, Web, Spatial, Time-series, Stream data.

B. KDD Process Model

  1. Selection: Define the problem, select target data.

  2. Preprocessing: Cleaning, handling missing values.

  3. Transformation: Normalization, attribute construction, aggregation.

  4. Data Mining: Apply core algorithms (association, classification, clustering).

  5. Interpretation/Evaluation: Evaluate patterns for validity, usefulness, and visualize results.

Role of Data Mining Engine: It is the core algorithmic component of the KDD process (Step 4). It executes the chosen data mining methods (e.g., Apriori, Decision Tree, k-Means) on the prepared data to generate the initial set of patterns/models.


V. ASSOCIATION RULE MINING

A. Core Concepts

Itemset: A set of items (e.g., {bread, milk}).

  • k-itemset: An itemset with k items.

Support (σ): Fraction of transactions that contain the itemset.

$$\boxed{\text{support}(X) = \frac{\sigma(X)}{N}}$$

where $N$ = total transactions.

Confidence: Conditional probability that a transaction containing X also contains Y.

$$\boxed{\text{confidence}(X \rightarrow Y) = \frac{\text{support}(X \cup Y)}{\text{support}(X)}$$

Frequent Itemset: An itemset with support ≥ min_sup threshold.

Association Rule: An implication of the form $$\displaystyle X \rightarrow Y $$, where $$\displaystyle X \cap Y = \emptyset $$.

  • Strong Rule: Meets min_sup and min_conf thresholds.

B. Apriori Algorithm

Key Principle (Apriori Property): All subsets of a frequent itemset must also be frequent. (If {A,B,C} is frequent, then {A,B}, {A,C}, {B,C} must be frequent). Steps:

  1. Find frequent itemsets (L_k):

    • C_k: Generate candidate k-itemsets by joining L_{k-1} with itself.

    • Prune: Remove candidates with any (k-1)-subset not in L_{k-1} (using Apriori property).

    • Count support: Scan DB, count support of remaining candidates.

    • L_k: Keep candidates with support ≥ min_sup.

    • Repeat for k=1,2,... until L_k is empty.

  2. Generate strong rules: For each frequent itemset l (|l| ≥ 2), generate all non-empty subsets. For each subset A, rule A → (l - A) if confidence ≥ min_conf.

Example: For min_sup=50% (2/4 trans), min_conf=70%.

Transactions: {A,B,C}, {A,C}, {B,C}, {A,B,D}.

  • L1 (frequent 1-itemsets): {A}(75%), {B}(50%), {C}(75%). D is infrequent.
  • C2: {A,B}, {A,C}, {B,C}. Prune? All 1-subsets are frequent → keep.
  • L2: {A,C}(50%), {B,C}(50%). {A,B} support=25% → pruned.
  • Rule from {A,C}: A→C (conf=50%/75%=66.7% <70%), C→A (conf=50%/75%=66.7% <70%) → no strong rule.
  • Rule from {B,C}: B→C (conf=50%/50%=100% ✓), C→B (conf=50%/75%=66.7% ✗). Strong rule: B → C.

C. Advanced Association Mining

  • FP-Growth Algorithm:

    1. Scan DB once to get frequent items (L1).

    2. Build FP-Tree:

      • Order items in each transaction by descending frequency (from L1).

      • Insert ordered transaction into a prefix tree (trie). Share common prefixes. Store count at each node.

    3. Mine FP-Tree recursively:

      • For each item in header table (starting from least frequent):

        • Construct conditional pattern base: all paths from root to the item.

        • Build conditional FP-Tree from pattern base.

        • If tree has single path → generate all combos; else, recursively mine.

      • Combine suffix item with patterns from conditional tree to form frequent patterns.

    • Advantage over Apriori: No candidate generation, only 2 DB scans.
  • Techniques to Improve Apriori Efficiency:

    • Hash-based technique: Use hash table to prune candidate k-itemsets during generation.

    • Transaction reduction: Remove transactions that don't contain any frequent items.

    • Partitioning: Partition DB, find local frequent itemsets, then global.

    • Sampling: Mine a random sample, then verify on full DB.

    • Dynamic item counting: Add candidate itemsets during a DB scan based on partial counts.


VI. CLASSIFICATION

A. Classification Process & Model

  1. Learning (Training): Build a model (classifier) from training data (tuples with known class labels).

  2. Testing: Evaluate model on unseen test data. Predict class labels and compare to actual.

  3. Prediction (Application): Use trained model to classify new, unseen instances.

B. Decision Tree Induction (e.g., ID3/C4.5)

  • Algorithm (ID3):

    1. Start with all training instances at root.

    2. If all instances belong to same class, make leaf with that class.

    3. Else, select the best attribute (A) to split on using Information Gain.

    4. Create a branch for each value of A, partition instances.

    5. Recurse on each branch with remaining attributes.

  • Attribute Selection Measures:

    • Information Gain (ID3): Based on Entropy.

      • Entropy(S) = $$\displaystyle -\sum_{i=1}^{c} p_i \log_2 p_i $$ (p_i = proportion of class i in S).

      • Gain(S, A) = Entropy(S) - $$\displaystyle \sum_{v \in Values(A)} \frac{|S_v|}{|S|} Entropy(S_v) $$.

      • Choose attribute with highest Gain.

    • Gain Ratio (C4.5): Corrects Information Gain's bias toward many-valued attributes.

      • SplitInfo(S, A) = $$\displaystyle -\sum_{v} \frac{|S_v|}{|S|} \log_2 \frac{|S_v|}{|S|} $$

      • GainRatio(S, A) = Gain(S, A) / SplitInfo(S, A).

      • Choose attribute with highest Gain Ratio.

    • Gini Index (CART): Measures impurity.

      • Gini(S) = $$\displaystyle 1 - \sum_{i=1}^{c} p_i^2 $$.

      • For binary split on A: Gini_A(S) = $$\displaystyle \frac{|S_1|}{|S|} Gini(S_1) + \frac{|S_2|}{|S|} Gini(S_2) $$.

      • Choose attribute/split that minimizes weighted Gini.

  • Tree Pruning:

    • Pre-pruning: Stop tree growth early (e.g., min samples per leaf, max depth).

    • Post-pruning: Grow full tree, then remove branches (e.g., error-based, cost-complexity).

C. Other Classification Methods

  • Rule-Based (e.g., RIPPER):

    • Sequential Covering: Learn rules one at a time.

      1. Start with empty rule set.

      2. Learn a rule that covers many instances of a class (e.g., using separate-and-conquer).

      3. Remove covered instances.

      4. Repeat until stopping condition.

    • RIPPER: Optimized version of sequential covering for efficiency.

  • Bayesian Classification (Naïve Bayes):

    • Based on Bayes' Theorem:

$$P(C|X) = \frac{P(C) P(X|C)}{P(X)}$$

    where C = class, X = attribute tuple.

*   **"Naïve" Assumption:** Attributes are conditionally independent given the class.

$$P(X|C) = \prod_{i=1}^{n} P(x_i|C)$$

*   **Classifier:** Assign class C that maximizes $$\displaystyle P(C) \prod_{i} P(x_i|C) $$.

*   **Advantage:** Simple, fast, works well despite independence assumption.

D. Classifier Evaluation

  • Process of Verifying Accuracy:

    1. Holdout: Split data into training set (e.g., 2/3) and test set (1/3). Train on train, test on test.

    2. Cross-Validation (k-fold): Partition data into k equal folds. Train on k-1 folds, test on 1 fold. Repeat k times, average accuracy.

    3. Bootstrapping: Sample n instances with replacement from original data (size n) to form training set. Test on left-out instances. Repeat.

  • Metrics (from Confusion Matrix):

    | | Predicted Positive | Predicted Negative | | :--- | :--- | :--- | | Actual Positive | TP (True Positive) | FN (False Negative) | | Actual Negative | FP (False Positive) | TN (True Negative) |

    • Accuracy: $$\displaystyle \frac{TP+TN}{Total} $$

    • Precision (P): $$\displaystyle \frac{TP}{TP+FP} $$ (Of predicted positives, how many correct?)

    • Recall (R): $$\displaystyle \frac{TP}{TP+FN} $$ (Of actual positives, how many found?)

    • F1-Score: $$\displaystyle 2 \times \frac{P \times R}{P + R} $$ (Harmonic mean of P & R).

    • ROC Curve: Plots True Positive Rate (Recall) vs False Positive Rate ($$\displaystyle \frac{FP}{FP+TN} $$) at different classification thresholds. Area Under Curve (AUC) measures overall performance.


VII. CLUSTERING

A. Clustering Fundamentals

Definition: Grouping data objects into clusters such that objects within a cluster are similar to each other and dissimilar to objects in other clusters.

Types of Clusters:

  • Partitional: Divides data into non-overlapping subsets (e.g., k-Means).
  • Hierarchical: Creates a tree of clusters (dendrogram). Can be agglomerative (bottom-up) or divisive (top-down).
  • Density-based: Clusters are dense regions separated by sparse regions (e.g., DBSCAN).
  • Grid-based: Objects are mapped to a grid structure (e.g., STING).
  • Model-based: Assumes data follows a probability distribution (e.g., Gaussian Mixture Models).

B. Partitioning Methods: k-Means

  • Algorithm:

    1. Choose k initial centroids (randomly or heuristically).

    2. Repeat:

      • Assignment: Assign each point to the nearest centroid (using Euclidean distance).

      • Update: Recalculate centroids as mean of all points in the cluster.

    3. Until centroids no longer change (or max iterations).

  • Example: k=2, points: (1,1), (1,2), (2,1), (5,5), (6,5). Initial centroids: C1=(1,1), C2=(5,5). Assign → Update → Converge.

  • Limitations:

    • Requires k to be specified.

    • Sensitive to initial centroids (may converge to local optimum).

    • Assumes spherical clusters of similar size/density.

    • Sensitive to outliers.

C. Hierarchical Methods

  • Agglomerative (Bottom-up):

    1. Start: each point as its own cluster.

    2. Merge the two closest clusters.

    3. Repeat until one cluster remains (or desired k).

  • Divisive (Top-down):

    1. Start: all points in one cluster.

    2. Split a cluster into two (e.g., using k-Means).

    3. Repeat until each point is its own cluster (or desired k).

  • Linkage Criteria (for Agglomerative): Defines "closeness" between clusters.

    • Single Linkage: Min distance between points in two clusters. → Chaining effect.

    • Complete Linkage: Max distance between points. → Compact clusters.

    • Average Linkage: Average distance between all pairs.

    • Centroid Linkage: Distance between cluster centroids.

  • Dendrogram: Tree diagram showing the sequence of merges/splits. Cutting the dendrogram at a level gives the cluster partition.

D. Density-Based: DBSCAN

  • Core Concepts:

    • ε-neighborhood (eps): Radius around a point.

    • MinPts: Minimum number of points in ε-neighborhood to define a core point.

    • Directly density-reachable: Point p is directly reachable from q if q is core and p is in q's ε-neighborhood.

    • Density-connected: Two points are connected if there is a path of directly reachable points.

  • Algorithm:

    1. Arbitrarily pick a point p.

    2. If p is core (≥ MinPts in ε-neighborhood), create new cluster, add all points density-reachable from p.

    3. If p is not core, mark as noise.

    4. Repeat until all points processed.

  • Advantages:

    • Discovers clusters of arbitrary shape.

    • Robust to outliers (noise points not assigned to any cluster).

    • Does not require specifying number of clusters (k).


VIII. ADVANCED & SPECIALIZED DATA MINING APPLICATIONS

A. Web Usage Mining

  • Goal: Discover patterns from web logs to understand user behavior, improve website structure, personalize content, and for marketing.

  • Data Sources:

    • Server logs: Client requests, IP, timestamp, URL, status.

    • Client-side logs: Cookies, user profiles (if logged in).

    • Proxy logs: Requests from multiple users behind a firewall.

  • Major Tasks:

    1. Session/Transaction Identification: Group user hits into sessions (e.g., 30-min timeout).

    2. Path Analysis: Frequent navigation paths (e.g., Home → Products → Electronics → Checkout).

    3. Association & Clustering of Pages: Find pages often visited together; cluster pages into topics.

B. Text Mining

  • Definition: Applying data mining techniques to unstructured text documents to discover patterns.

  • Difference from Data Mining: Input is text, not structured data. Requires significant preprocessing.

  • Major Tasks & Preprocessing:

    1. Text Preprocessing:

      • Tokenization: Split text into words/tokens.

      • Stop-word removal: Remove common words (the, is, and).

      • Stemming: Reduce words to root form (e.g., "running" → "run").

      • Lemmatization: Reduce to dictionary form (e.g., "better" → "good").

    2. Document Representation:

      • Vector Space Model: Document = vector of term weights (e.g., TF-IDF: Term Frequency-Inverse Document Frequency).
    3. Tasks:

      • Text Categorization: Assign documents to predefined classes (e.g., spam/ham, news topics).

      • Text Clustering: Group similar documents without predefined classes.

      • Text Summarization: Generate a concise summary of a document.

C. Spatial Data Mining

  • Definition: Discovering patterns from spatial data (data with geographic/location reference).

  • Characteristics: Spatial autocorrelation (nearby objects are more similar), complex data types (points, lines, polygons), spatial relationships (topological, distance, direction).

  • Major Tasks:

    • Spatial Association Rules: Rules with spatial predicates (e.g., near, intersects). A(x) ∧ B(y) → C(z) where x, y, z are spatial objects.

    • Spatial Classification: Classify spatial objects using both non-spatial attributes and spatial attributes (e.g., classify land use type using soil type (non-spatial) and proximity to river (spatial)).

    • Spatial Clustering: Cluster objects based on spatial proximity and optionally non-spatial attributes (e.g., DBSCAN for earthquake epicenters).

D. Data Prediction

  • General Concept: Building a model to predict the value of a target (dependent) variable based on other (independent) variables.

  • Relationship to Classification & Regression:

    • Classification is a type of prediction where the target is categorical (discrete class labels: yes/no, spam/ham).

    • Regression is prediction where the target is continuous/real-valued (e.g., house price, temperature).

  • Common Techniques: Linear Regression, Polynomial Regression, Decision Trees (for both classification & regression trees - CART), Neural Networks. The evaluation metrics differ (e.g., MSE, RMSE for regression vs. Accuracy for classification).


Exam Tips & Common Pitfalls:

  • Schema Design: Be ready to draw and contrast Star, Snowflake, and Galaxy schemas. Know which is most normalized (Snowflake) and which is best for query performance (Star).
  • OLAP: Clearly distinguish MOLAP (cube, fast) vs ROLAP (relational, scalable). Be able to sketch the ROLAP mapping.
  • Apriori: Practice a small example (4-5 transactions) showing candidate generation, pruning, and rule generation. Remember the Apriori Property.
  • FP-Growth: Understand the two-step process: (1) Build FP-Tree (order by freq, share prefixes), (2) Mine recursively using conditional pattern bases. Know its main advantage: no candidate generation.
  • Decision Trees: Be fluent in calculating Information Gain and Gini Index for a small dataset. Know when to use Gain Ratio (many-valued attributes).
  • Clustering: Differentiate k-Means (partitional, centroid-based) vs DBSCAN (density-based, handles noise). Know single vs complete linkage effects.
  • KDD: Remember the 5 steps and that Data Mining is Step 4. The Data Mining Engine is the algorithmic core.
  • Vertical Partitioning: Explain it as splitting by columns to separate hot/cold data, improving I/O for common queries at the cost of joins.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in