Skip to content
CS-703 (B) · Data Mining and Warehousing/Quick Revision Short Notes

Data Mining and Warehousing (CS-703 (B)) - Unit 3 Short Notes

How unit 3 is examined

This unit covers data and its preparation, then what data mining does and how it relates to KDD; the marks sit in Basic data mining tasks, Data Mining vs KDD, Issues in data mining and Data Preprocessing.

Data Types

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>A data object is described by attributes, and the attribute type (interval-scaled, binary, categorical, ordinal, ratio-scaled or mixed) decides which dissimilarity measure clustering can use.</mark>

Key points.

  1. Interval-scaled variables are measured on a linear scale with no true zero (temperature in Celsius), and their dissimilarity is Euclidean or Manhattan distance after standardisation.
  2. A binary variable has only states 0 and 1; it is symmetric if both states matter (gender) and asymmetric if only 1 is important (disease present), and its dissimilarity is $d=\frac{r+s}{q+r+s+t}$ for symmetric, with $t$ dropped for asymmetric (Jaccard).
  3. A categorical (nominal) variable has more than two unordered states (colour), and its dissimilarity is $d=\frac{p-m}{p}$, where $m$ is the number of matches among $p$ variables.
  4. An ordinal variable has ranked states (small, medium, large); ranks are mapped to $[0,1]$ by $z=\frac{r-1}{M-1}$ and then treated as interval-scaled.
  5. A ratio-scaled variable has a true zero and grows exponentially (bacteria count); take the logarithm first and then treat it as interval-scaled.
  6. Mixed types are combined by a weighted average of the per-variable dissimilarities, each scaled to $[0,1]$.

Asked: [7 marks] (Jun 2025) Explain different data types used in clustering.

Quality of data

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>Data quality is its fitness for mining, judged by accuracy, completeness, consistency, timeliness and freedom from noise and outliers.</mark>

Key points.

  1. An outlier is an object that deviates so much from the rest that it seems generated by a different mechanism, and it is a data quality problem or a finding depending on the goal.
  2. Statistical (distribution-based) outlier detection assumes a distribution model (for example normal) and flags objects with very low probability under it.
  3. A block procedure tests one candidate at a time against a working hypothesis (H: all objects come from the model), and a discordancy test rejects H when the object is too improbable, so it is called an outlier.
  4. Parametric methods assume a known distribution with estimated parameters; non-parametric methods (histograms, kernels) learn the model from the data.
  5. Limitations: the distribution is often unknown, tests work mostly on one attribute, and they scale poorly to high-dimensional data.

Asked: [7 marks] (Jun 2025) Explain briefly about statistical distribution-based outlier detection.

Data Preprocessing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Data preprocessing is the set of techniques (cleaning, integration, transformation, reduction) that convert raw data into clean, consistent and compact data before mining.</mark>

Key points.

  1. Need: real data is incomplete, noisy and inconsistent, and "garbage in, garbage out" means poor data gives poor patterns.
  2. Data cleaning fills missing values (ignore tuple, global constant, attribute mean), smooths noise by binning, regression or clustering, and removes outliers and inconsistencies.
  3. Data integration merges data from several sources into one store and resolves schema conflicts, redundancy (found by correlation) and differing value representations.
  4. Data transformation rescales the data by normalisation such as min-max $v'=\frac{v-min}{max-min}$, plus smoothing, aggregation, generalisation and attribute construction.
  5. Data reduction gives a smaller representation with almost the same results, by dimensionality reduction, numerosity reduction, sampling, compression and discretisation.
  6. Impact: mining becomes faster and more accurate because the algorithm works on smaller, cleaner data.

Answer frame. Open with the definition and the need (incomplete, noisy, inconsistent data); draw the four-box flow cleaning to integration to transformation to reduction; develop one technique-rich line per form; close with the effect on quality and speed.

Asked: [7 marks] (Dec 2020, Nov 2023) What is the need of Data Preprocessing? Discuss various forms of preprocessing. Write short notes on the various preprocessing tasks.

Similarity measures

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>

Definition. <mark>A similarity measure quantifies how alike two objects are, and a dissimilarity (distance) measure quantifies how different they are; similarity is high when distance is small.</mark>

Key points.

  1. Euclidean distance is $d(x,y)=\sqrt{\sum_i (x_i-y_i)^2}$ and Manhattan distance is $d(x,y)=\sum_i |x_i-y_i|$.
  2. Cosine similarity is $\cos(x,y)=\frac{x\cdot y}{\|x\|\,\|y\|}$ and suits documents and sparse data because it ignores length.
  3. Jaccard similarity for binary data is $J=\frac{M_{11}}{M_{11}+M_{10}+M_{01}}$, ignoring 0-0 matches.
  4. Role in clustering: objects with small distance are placed in the same cluster; role in classification: a new object gets the class of its nearest neighbours (k-NN).
  5. The choice of measure changes the clusters and the neighbours, so it must match the data type and scale.

Example. $x=(1,2,3)$, $y=(4,6,3)$: Euclidean $=\sqrt{9+16+0}=5$; Manhattan $=3+4+0=7$; cosine $=\frac{25}{\sqrt{14}\sqrt{61}}=0.855$. For binary $p=10110$, $q=11010$: $M_{11}=2$, $M_{10}=1$, $M_{01}=1$, so J = 2/4 = 0.5.

Asked: [7 marks] (Dec 2025) Describe similarity measures and their role in clustering and classification.

Summary statistics

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Summary statistics are single numbers that describe the centre and spread of data.</mark>

Key points.

  1. Central tendency: mean $\bar{x}=\frac{1}{n}\sum x_i$, median (middle value) and mode (most frequent value).
  2. Spread: range, quartiles and interquartile range, variance $\sigma^2=\frac{1}{n}\sum (x_i-\bar{x})^2$ and standard deviation $\sigma$.
  3. For 2, 4, 4, 6, 9: mean 5, median 4, mode 4, variance 5.6 and $\sigma\approx 2.37$; the mean is pulled by the outlier 9 but the median is not.

Data distributions

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A data distribution shows how often each value occurs; the normal distribution is the bell-shaped, symmetric one.</mark>

Key points.

  1. The normal distribution $N(\mu,\sigma^2)$ has density $f(x)=\frac{1}{\sigma\sqrt{2\pi}}e^{-(x-\mu)^2/2\sigma^2}$, with about 68% of values within $1\sigma$ and 95% within $2\sigma$.
  2. Skewed data has a long tail on one side, so mean, median and mode differ; symmetric data has them equal.
  3. Histograms and boxplots reveal the distribution, and outlier detection uses it (see Quality of data).

Basic data mining tasks

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Data mining tasks (functionalities) specify the kind of patterns to find: descriptive tasks characterise the data, and predictive tasks use current data to predict.</mark>

Key points.

  1. Characterization summarises the general features of a target class (for example customers who spend over 1000 a year).
  2. Discrimination compares the features of a target class against contrasting classes.
  3. Association finds items that occur together, such as $\text{buys(bread)} \Rightarrow \text{buys(butter)}$ with support and confidence.
  4. Classification builds a model from class-labelled training data (decision tree, rules, neural network) to predict the class of new objects.
  5. Prediction (regression) estimates a numeric value, and clustering groups unlabelled objects so that intra-cluster similarity is high and inter-cluster similarity is low.
  6. Outlier analysis finds objects that fit no group, and evolution and trend analysis studies data that changes over time.

Diagram. Task primitives: the user specifies five inputs to the mining system, which reads the database or warehouse.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-01" viewBox="0 0 510 320.8" width="510" height="320.8" role="img" aria-label="Primitives: TRD task-relevant data, KT kind of knowledge, BK background knowledge, IM interestingness measures, PR presentation and visualisation; DM mining engine"><style>#dsfig-u3-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-01 .t{fill:#16181D;font-weight:500}#dsfig-u3-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-01 .dot{fill:#16181D}#dsfig-u3-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-01 .ah{fill:#454C5A}#dsfig-u3-01 .ah.hi{fill:#2340B8}#dsfig-u3-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-01 .e{stroke:#B1B7C3}html.dark #dsfig-u3-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-01 .t{fill:#E6E8ED}html.dark #dsfig-u3-01 .t.inv{fill:#0F1115}html.dark #dsfig-u3-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-01 .dot{fill:#E6E8ED}html.dark #dsfig-u3-01 .ann{fill:#8FA3FF}html.dark #dsfig-u3-01 .lbl{fill:#858D9C}html.dark #dsfig-u3-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-01 .ah{fill:#B1B7C3}html.dark #dsfig-u3-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh7" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M55.8,115.5 L151.5,51.6" marker-end="url(#ah7)"/><path class="e" d="M184.2,51.4 L324.2,156.4" marker-end="url(#ah7)"/><path class="e" d="M186.6,107.3 L321.5,161.2" marker-end="url(#ah7)"/><path class="e" d="M188,161.3 L320,168" marker-end="url(#ah7)"/><path class="e" d="M187.2,215.1 L320.9,175" marker-end="url(#ah7)"/><path class="e" d="M184.9,270.4 L323.4,180.4" marker-end="url(#ah7)"/><path class="e" d="M360,169 L449,169" marker-end="url(#ah7)"/><circle class="n" cx="40" cy="126" r="18"/><text class="t" x="40" y="126" dy=".35em" text-anchor="middle">DB</text><circle class="n" cx="169" cy="40" r="18"/><text class="t" x="169" y="40" dy=".35em" text-anchor="middle">TRD</text><circle class="n" cx="169" cy="100.2" r="18"/><text class="t" x="169" y="100.2" dy=".35em" text-anchor="middle">KT</text><circle class="n" cx="169" cy="160.4" r="18"/><text class="t" x="169" y="160.4" dy=".35em" text-anchor="middle">BK</text><circle class="n" cx="169" cy="220.6" r="18"/><text class="t" x="169" y="220.6" dy=".35em" text-anchor="middle">IM</text><circle class="n" cx="169" cy="280.8" r="18"/><text class="t" x="169" y="280.8" dy=".35em" text-anchor="middle">PR</text><circle class="n" cx="341" cy="169" r="18"/><text class="t" x="341" y="169" dy=".35em" text-anchor="middle">DM</text><circle class="n" cx="470" cy="169" r="18"/><text class="t" x="470" y="169" dy=".35em" text-anchor="middle">OUT</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Primitives: TRD task-relevant data, KT kind of knowledge, BK background knowledge, IM interestingness measures, PR presentation and visualisation; DM mining engine</figcaption></figure>

Task primitives.

  1. Task-relevant data selects the database, tables, attributes and conditions to mine.
  2. Kind of knowledge to mine names the task: characterization, association, classification, clustering and so on.
  3. Background knowledge such as concept hierarchies (street, city, state, country) guides the search.
  4. Interestingness measures (support, confidence, novelty, utility) filter out uninteresting patterns.
  5. Presentation and visualisation states how results are shown: rules, tables, charts, cubes.

Classification is supervised. Classification is supervised learning because the training tuples carry known class labels and the model learns from them: in step 1 the algorithm (for example a decision tree) constructs a model from the labelled training set, and in step 2 the model is tested on labelled test data and, if accurate, applied to unlabelled objects.

Feature Classification (supervised) Clustering (unsupervised)
Labels Class labels given in training data No labels
Goal Predict class of new object Discover natural groups
Training Model built from training set No training phase
Evaluation Accuracy on test set Cohesion, separation (no ground truth)
Examples Decision tree, Naive Bayes k-means, hierarchical
Application Spam or not-spam email Customer segmentation

Supervised versus unsupervised learning is the same table: supervised covers classification and regression, unsupervised covers clustering and association, and association rules are unsupervised.

Answer frame. For primitives: open with the definition (a data mining task is specified by five primitives), draw the diagram, explain the five in order with one example each, close with the point that a well-specified query gives focused, interesting results. For classification vs clustering or supervised vs unsupervised: define both, draw the six-row table, add one example application each, and close with the labelled-versus-unlabelled distinction. For functionalities: define, then list the six points with one example each.

Asked: [7 marks] (Dec 2020, Jun 2025) List out the differences between classification and clustering methods with example. Briefly explain the differences between "classification" and "Clustering" and give an informal example of an application that would benefit from each technique. Asked: [7 marks] (Nov 2023) Classification is supervised learning? Justify. (Also: explain whether association rule mining is supervised or unsupervised: it is unsupervised, since it needs no class labels.) Asked: [7 marks] (Dec 2024, Jun 2025) Explain with diagrammatic illustration the primitives for specifying a data mining task. Discuss about data mining task primitives with examples. Asked: [7 marks] (Dec 2025) Discuss the key differences between supervised and unsupervised learning with examples. Asked: [7 marks] (Jun 2025) Explain the data mining functionalities.

Data Mining vs Knowledge Discovery in Databases

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. <mark>Knowledge discovery in databases (KDD) is the whole multi-step process of turning raw data into useful knowledge, and data mining is the one step in it that applies algorithms to extract patterns.</mark>

Diagram.

<figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-02" viewBox="0 0 596 80" width="596" height="80" role="img" aria-label="KDD process: databases, selection, preprocessing, transformation, data mining, pattern evaluation, knowledge presentation"><style>#dsfig-u3-02 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-02 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-02 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-02 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-02 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-02 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-02 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-02 .t{fill:#16181D;font-weight:500}#dsfig-u3-02 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-02 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-02 .dot{fill:#16181D}#dsfig-u3-02 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-02 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-02 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-02 .ah{fill:#454C5A}#dsfig-u3-02 .ah.hi{fill:#2340B8}#dsfig-u3-02 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-02 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-02 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-02 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-02 .e{stroke:#B1B7C3}html.dark #dsfig-u3-02 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-02 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-02 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-02 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-02 .t{fill:#E6E8ED}html.dark #dsfig-u3-02 .t.inv{fill:#0F1115}html.dark #dsfig-u3-02 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-02 .dot{fill:#E6E8ED}html.dark #dsfig-u3-02 .ann{fill:#8FA3FF}html.dark #dsfig-u3-02 .lbl{fill:#858D9C}html.dark #dsfig-u3-02 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-02 .ah{fill:#B1B7C3}html.dark #dsfig-u3-02 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-02 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-02 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-02 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh8" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L105,40" marker-end="url(#ah8)"/><path class="e" d="M145,40 L191,40" marker-end="url(#ah8)"/><path class="e" d="M231,40 L277,40" marker-end="url(#ah8)"/><path class="e" d="M317,40 L363,40" marker-end="url(#ah8)"/><path class="e" d="M403,40 L449,40" marker-end="url(#ah8)"/><path class="e" d="M489,40 L535,40" marker-end="url(#ah8)"/><path class="e" d="M451,40 L147,40" marker-end="url(#ah8)"/><g class="wl"><rect x="267.3" y="31" width="61.5" height="18" rx="9"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">iterate</text></g><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">DB</text><circle class="n" cx="126" cy="40" r="18"/><text class="t" x="126" y="40" dy=".35em" text-anchor="middle">SEL</text><circle class="n" cx="212" cy="40" r="18"/><text class="t" x="212" y="40" dy=".35em" text-anchor="middle">PRE</text><circle class="n" cx="298" cy="40" r="18"/><text class="t" x="298" y="40" dy=".35em" text-anchor="middle">TRA</text><circle class="n" cx="384" cy="40" r="18"/><text class="t" x="384" y="40" dy=".35em" text-anchor="middle">DM</text><circle class="n" cx="470" cy="40" r="18"/><text class="t" x="470" y="40" dy=".35em" text-anchor="middle">EVA</text><circle class="n" cx="556" cy="40" r="18"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">KN</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">KDD process: databases, selection, preprocessing, transformation, data mining, pattern evaluation, knowledge presentation</figcaption></figure>

Steps.

  1. Data cleaning removes noise and inconsistent data, and data integration combines multiple sources.
  2. Data selection retrieves the data relevant to the analysis task from the database.
  3. Data transformation converts data into forms suitable for mining by summary or aggregation.
  4. Data mining applies intelligent methods to extract data patterns; it is the essential step.
  5. Pattern evaluation identifies the truly interesting patterns using interestingness measures.
  6. Knowledge presentation shows the mined knowledge to the user by visualisation and reports; the process loops back to earlier steps if results are poor.
Basis Data mining KDD
Scope One step, pattern extraction Whole end-to-end process
Steps Applies algorithms only Selection to presentation
Input Prepared, transformed data Raw data from databases
Output Patterns and models Useful, validated knowledge
Human role Mostly automatic Heavy user involvement, iterative
Example Running Apriori on baskets Cleaning sales data, mining, then acting on the rules

Answer frame. For KDD steps and "data mining as a step": define KDD, draw the flow with the iteration arrow, give one line per step, and close by noting mining is the core step. For "data mining vs KDD": define both, give the table, and add the retail example.

Asked: [7 marks] (Dec 2020, Nov 2023, Jun 2025) Explain the various steps involved in knowledge discovery. Explain with diagrammatic illustration data mining as a step in the process of knowledge discovery. What is KDD? Explain about data mining as a step in the process of knowledge discovery. Asked: [7 marks] (Nov 2023, Dec 2025) Explain data mining vs knowledge discovery in databases. Define data mining and differentiate it from knowledge discovery in databases.

Issues in Data mining

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. <mark>Issues in data mining are the challenges in methodology, user interaction, performance and data diversity that a mining system must handle.</mark>

Key points.

  1. Mining methodology issues: mining different kinds of knowledge, mining at multiple levels of abstraction, and handling noisy or incomplete data.
  2. User interaction issues: interactive mining, use of background knowledge, ad hoc queries, and presentation and visualisation of results.
  3. Performance issues: algorithms must be efficient and scalable to huge databases, and parallel, distributed and incremental methods are needed.
  4. Diversity of data types: relational, complex, heterogeneous and multimedia data, and data from distributed, dynamic sources.
  5. High dimensionality makes distances lose meaning and search costly.
  6. Privacy and security: mining personal data can violate privacy, so data must be protected and anonymised.
  7. Pattern evaluation: deciding which of the many discovered patterns are truly interesting.

Answer frame. Open with one line defining the challenges; group the seven points under methodology, user interaction, performance and data diversity; for the Dec 2024 question add the functionalities from Basic data mining tasks with one example each; close with privacy as the main social issue.

Asked: [7 marks] (Dec 2020, Jun 2025) Briefly explain the major issues and challenges of Data Mining. Explain major requirements and challenges in data mining. Asked: [7 marks] (Dec 2024) Explain the various data mining issues and functionalities in detail.

Introduction to Fuzzy sets and fuzzy logic

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A fuzzy set $A$ assigns each element $x$ a membership degree $\mu_A(x)\in[0,1]$ instead of only 0 or 1.</mark>

Key points.

  1. In a crisp set membership is 0 or 1, while a fuzzy set allows partial membership, for example "tall" gives $\mu(180\text{ cm})=0.8$.
  2. Operations: union $\max(\mu_A,\mu_B)$, intersection $\min(\mu_A,\mu_B)$, complement $1-\mu_A$.
  3. Membership functions are usually triangular, trapezoidal or Gaussian.
  4. Fuzzy logic reasons with degrees of truth and IF-THEN rules, and in mining it handles vague attributes such as "young" or "high income".

Last-minute revision

  • KDD is the whole process; data mining is its core pattern-extraction step.
  • KDD steps: cleaning, integration, selection, transformation, mining, evaluation, presentation.
  • Five task primitives: task-relevant data, kind of knowledge, background knowledge, interestingness, presentation.
  • Functionalities: characterization, discrimination, association, classification, prediction, clustering, outlier analysis, evolution analysis.
  • Classification is supervised (labelled training data); clustering and association are unsupervised.
  • Preprocessing forms: cleaning, integration, transformation, reduction; min-max $v'=\frac{v-min}{max-min}$.
  • Euclidean $\sqrt{\sum(x_i-y_i)^2}$, Manhattan $\sum|x_i-y_i|$, cosine $\frac{x\cdot y}{\|x\|\|y\|}$, Jaccard $\frac{M_{11}}{M_{11}+M_{10}+M_{01}}$.
  • Attribute types: interval, binary, categorical, ordinal, ratio, mixed.
  • Outlier detection by statistics: block and discordancy tests, parametric or non-parametric.
  • Issues: methodology, user interaction, performance, data diversity, privacy.
  • Fuzzy union is max, intersection is min, complement is 1 minus membership.

Memory hooks

  • KDD "CISTEP": Clean, Integrate, Select, Transform, Evaluate, Present around Mining.
  • Primitives "TKBIP": Task data, Knowledge kind, Background, Interestingness, Presentation.
  • Supervised = teacher gives labels; unsupervised = no teacher, find groups.
  • Preprocessing "CITR": Cleaning, Integration, Transformation, Reduction.
  • Issues "MUPD": Methodology, User interaction, Performance, Diversity.

Coverage checklist

  • Data Types: Jun 2025 data types used in clustering.
  • Quality of data: Jun 2025 statistical distribution-based outlier detection.
  • Data Preprocessing: Dec 2020, Nov 2023 need and forms of preprocessing.
  • Similarity measures: Dec 2025 similarity measures in clustering and classification.
  • Summary statistics: no past questions.
  • Data distributions: no past questions.
  • Basic data mining tasks: classification vs clustering, classification is supervised, task primitives, supervised vs unsupervised, functionalities.
  • Data Mining V/s knowledge discovery in databases.: KDD steps, mining as a KDD step, data mining vs KDD.
  • Issues in Data mining.: major issues and challenges, issues and functionalities.
  • Introduction to Fuzzy sets and fuzzy logic: no past questions.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in