How unit 3 is examined
This unit covers data types, pre-processing, similarity, statistics, the KDD process, mining tasks, issues and fuzzy logic; KDD with mining tasks carries the most marks, then pre-processing and fuzzy sets.
Data Types, Quality of data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. A data type says what kind of values an attribute holds and which operations make sense on it; <mark>data quality is the fitness of data for mining, judged by accuracy, completeness, consistency, timeliness and believability.</mark>
Key points.
- Interval-scaled variables are measured on a linear scale with no true zero (temperature in Celsius), so differences are meaningful but ratios are not; distance is Euclidean or Manhattan after standardising.
- Ratio-scaled variables have a true zero (weight, age, income), so ratios are meaningful; they are compared by Euclidean distance or by a log transform when growth is exponential.
- Binary variables take only 0 or 1; symmetric binary uses the simple matching coefficient $d=\frac{b+c}{a+b+c+d}$, while asymmetric binary (rare positive outcome) uses the Jaccard-style $d=\frac{b+c}{a+b+c}$.
- Categorical (nominal) variables have unordered labels such as colour; dissimilarity is $d=\frac{p-m}{p}$, where $m$ is the number of matching attributes and $p$ the total.
- Ordinal variables have a meaningful order but unknown gaps (low, medium, high); ranks are mapped to $z=\frac{r-1}{M-1}$ and then treated as interval values.
- Clustering needs the correct type because the dissimilarity formula depends on it, and a wrong measure puts unlike objects in one cluster.
- Poor quality shows as noise, missing values, outliers, duplicates and inconsistency, and it makes the mined patterns unreliable.
Answer frame. Open with the definition of a data type; list the five types with one example each; give the dissimilarity for each; close with why clustering needs the right measure.
Asked: [7 marks] (May 2024) Explain different data types used in clustering.
Data Pre-processing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>Data pre-processing is the set of techniques that clean, integrate, transform and reduce raw data so that it is fit for mining.</mark>
Key points.
- Real data is noisy, incomplete and inconsistent, so mining it directly gives poor patterns: garbage in, garbage out.
- Data cleaning fills missing values (mean, most probable value, ignoring the tuple), smooths noise (binning, regression, clustering) and removes outliers and inconsistencies.
- Data integration merges data from several sources into one store, resolving schema conflicts, redundant attributes (found by correlation) and duplicate records.
- Data transformation converts data to suitable forms by smoothing, aggregation, generalisation, normalisation (min-max, z-score) and attribute construction.
- Data reduction gives a smaller data set with nearly the same analytical result, through dimensionality reduction, numerosity reduction, data cube aggregation, compression and discretisation.
- Justification: quality mining needs quality data, and pre-processing improves accuracy, shortens mining time and makes results interpretable.
Min-max normalisation: $v'=\frac{v-\min}{\max-\min}$; z-score: $v'=\frac{v-\mu}{\sigma}$.
Answer frame. Open with the definition and the problem of noisy, incomplete and inconsistent data; list the four tasks with two techniques each; close with the effect on mining accuracy and speed, which is the justification.
Pitfall: Only naming the four tasks without linking each to a data problem gives no justification marks.
Asked: [7 marks] (May 2023, Jun 2025) "Data preprocessing is necessary before data mining process" Justify your answer. Why data preprocessing is needed and explain data preprocessing tasks?
Similarity measures
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>A similarity measure gives a number showing how alike two objects are; a distance (dissimilarity) measure gives how far apart they are, and the two are inversely related.</mark>
Key points.
- Euclidean distance is $d(x,y)=\sqrt{\sum_{i}(x_i-y_i)^2}$ and suits continuous numeric attributes such as coordinates.
- Manhattan distance is $d(x,y)=\sum_i |x_i-y_i|$, the sum of absolute differences along each axis.
- Cosine similarity is $\cos(x,y)=\frac{x\cdot y}{\|x\|\,\|y\|}$; it measures the angle, so it suits text documents where length does not matter.
- Jaccard similarity is $J(A,B)=\frac{|A\cap B|}{|A\cup B|}$ and suits sets and asymmetric binary data such as market baskets.
Example. For $x=(1,2,3)$ and $y=(4,6,3)$: Euclidean $=\sqrt{9+16+0}=5$; Manhattan $=3+4+0=7$; cosine $=\frac{25}{\sqrt{14}\sqrt{61}}=0.855$. For $A=\{a,b,c,d\}$, $B=\{c,d,e\}$: $J=\frac{2}{5}=0.4$. Euclidean = 5, cosine = 0.855, Jaccard = 0.4.
Asked: [7 marks] (May 2023) Discuss about any two measures of similarity.
Summary statistics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Summary statistics describe a data set by a few numbers for its centre and spread.
Key points.
- Mean $\bar x=\frac{1}{n}\sum x_i$ is the centre but is sensitive to outliers; the median is the middle value and resists them.
- Mode is the most frequent value; variance $\sigma^2=\frac{1}{n}\sum(x_i-\bar x)^2$ and standard deviation $\sigma=\sqrt{\sigma^2}$ measure spread.
- For 2, 4, 4, 4, 5, 5, 7, 9: mean $=5$, median $=4.5$, variance $=4$, $\sigma=2$.
Data distributions
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A data distribution shows how often each value or range of values occurs in the data.
Key points.
- A histogram plots the frequency of each range and shows the shape of the data.
- The normal distribution is bell shaped and symmetric with $f(x)=\frac{1}{\sigma\sqrt{2\pi}}e^{-(x-\mu)^2/2\sigma^2}$; about 68% of values lie within $\mu\pm\sigma$.
- Skewed data has a long tail on one side, so mean and median differ; knowing the distribution helps choose normalisation and detect outliers.
Basic data mining tasks, Data Mining Vs knowledge discovery in databases
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Knowledge discovery in databases (KDD) is the whole iterative process of extracting useful, valid and novel patterns from data, and data mining is the single step of that process where the pattern-finding algorithms are applied.</mark>
Diagram. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u3-01" viewBox="0 0 603 209" width="603" height="209" role="img" aria-label="KDD process. DB data sources, Sel selection (target data), Pre pre-processing (clean data), Trn transformation, Mine data mining (patterns), Eval interpretation and evaluation, K knowledge; the back arrow shows iteration."><style>#dsfig-u3-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u3-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u3-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u3-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u3-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u3-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u3-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u3-01 .t{fill:#16181D;font-weight:500}#dsfig-u3-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u3-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u3-01 .dot{fill:#16181D}#dsfig-u3-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u3-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u3-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u3-01 .ah{fill:#454C5A}#dsfig-u3-01 .ah.hi{fill:#2340B8}#dsfig-u3-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u3-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u3-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u3-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u3-01 .e{stroke:#B1B7C3}html.dark #dsfig-u3-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u3-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u3-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u3-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u3-01 .t{fill:#E6E8ED}html.dark #dsfig-u3-01 .t.inv{fill:#0F1115}html.dark #dsfig-u3-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u3-01 .dot{fill:#E6E8ED}html.dark #dsfig-u3-01 .ann{fill:#8FA3FF}html.dark #dsfig-u3-01 .lbl{fill:#858D9C}html.dark #dsfig-u3-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u3-01 .ah{fill:#B1B7C3}html.dark #dsfig-u3-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u3-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u3-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u3-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh4" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,40 L122.2,40" marker-end="url(#ah4)"/><path class="e" d="M162.2,40 L225.4,40" marker-end="url(#ah4)"/><path class="e" d="M265.4,40 L328.6,40" marker-end="url(#ah4)"/><path class="e" d="M368.6,40 L424.8,40" marker-end="url(#ah4)"/><path class="e" d="M478.8,40 L528,40" marker-end="url(#ah4)"/><path class="e" d="M556,66 L556,148" marker-end="url(#ah4)"/><path class="e" d="M530,40 L164.2,40" marker-end="url(#ah4)"/><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">DB</text><circle class="n" cx="143.2" cy="40" r="18"/><text class="t" x="143.2" y="40" dy=".35em" text-anchor="middle">Sel</text><circle class="n" cx="246.4" cy="40" r="18"/><text class="t" x="246.4" y="40" dy=".35em" text-anchor="middle">Pre</text><circle class="n" cx="349.6" cy="40" r="18"/><text class="t" x="349.6" y="40" dy=".35em" text-anchor="middle">Trn</text><rect class="n" x="427.8" y="25" width="50" height="30" rx="15"/><text class="t" x="452.8" y="40" dy=".35em" text-anchor="middle">Mine</text><rect class="n" x="531" y="25" width="50" height="30" rx="15"/><text class="t" x="556" y="40" dy=".35em" text-anchor="middle">Eval</text><circle class="n" cx="556" cy="169" r="18"/><text class="t" x="556" y="169" dy=".35em" text-anchor="middle">K</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">KDD process. DB data sources, Sel selection (target data), Pre pre-processing (clean data), Trn transformation, Mine data mining (patterns), Eval interpretation and evaluation, K knowledge; the back arrow shows iteration.</figcaption></figure>
Key points (phases of KDD).
- Data selection retrieves the task-relevant data from the databases, giving the target data.
- Pre-processing cleans the target data by removing noise and duplicates and handling missing values.
- Transformation converts the data into forms suitable for mining, by aggregation, normalisation and dimensionality reduction.
- Data mining applies algorithms such as classification, clustering or association to extract patterns.
- Interpretation and evaluation identifies the truly interesting patterns using measures such as support and confidence, and discards the rest.
- Knowledge presentation shows the result through visualisation and reports, and the process is repeated with new choices if the result is poor.
- Data mining is therefore only one step of KDD, although the two terms are often used loosely for each other.
Data mining tasks. Tasks are predictive (use current data to predict unknown or future values) or descriptive (find human-interpretable patterns in the data).
| Task | Type | Purpose and example |
|---|---|---|
| Classification | Predictive | Assign objects to predefined classes; loan applicants as safe or risky |
| Regression | Predictive | Predict a continuous value; house price from area |
| Time series analysis | Predictive | Study values over time; forecast next month's sales |
| Prediction | Predictive | Estimate future or unknown values; next year's demand |
| Clustering | Descriptive | Group similar objects with no predefined classes; customer segments |
| Association rules | Descriptive | Find items that occur together; bread implies butter |
| Summarisation | Descriptive | Give a compact description of data; average sales per region |
| Sequence discovery | Descriptive | Find ordered patterns; buying a phone then a cover |
Answer frame. For the phases question, open with the KDD definition, draw the flow diagram with the feedback arrow, describe points 1-6 in order with one line each, and close by stating that mining is one step of KDD. For the tasks question, open with predictive versus descriptive, then give each task with its example from the table, and close with a line that the choice of task depends on the goal.
Asked: [7 marks] (May 2023, May 2024, Jun 2025) Describe the various phases in knowledge discovery process with a neat diagram. Explain data mining as a step in knowledge discovery process. Asked: [7 marks] (May 2023) Explain about various Data Mining Tasks with appropriate examples.
Issues in Data mining
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Issues in data mining are the challenges in methodology, performance, data diversity and society that limit how well mining works in practice.</mark>
Key points.
- Mining methodology issues include mining different kinds of knowledge, handling noisy and incomplete data, and judging the interestingness of patterns.
- Performance issues are efficiency and scalability, since algorithms must handle huge databases in reasonable time, using parallel and incremental methods.
- Diversity of data types raises the problem of mining relational, text, multimedia, spatial and web data from distributed heterogeneous sources.
- Data quality issues such as missing, noisy and inconsistent data reduce the accuracy of the patterns found.
- Privacy and security is a serious issue, because mining personal data can expose individuals and lead to misuse.
- The result must be presented in an understandable way, so visualisation and user interaction remain open challenges.
Asked: [7 marks] (May 2023) Discuss about issues in data mining.
Introduction to Fuzzy sets and fuzzy logic
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>A fuzzy set is a set in which each element has a degree of membership between 0 and 1 given by a membership function $\mu_A(x)$, whereas a crisp set allows only 0 (not a member) or 1 (member).</mark>
Key points.
- In a crisp set "tall" is a sharp cut such as height above 180 cm; in a fuzzy set a person of 175 cm may be tall with membership 0.6.
- The membership function $\mu_A:X\to[0,1]$ maps each element to its degree and is often triangular, trapezoidal or Gaussian.
- Fuzzy operations are: union $\mu_{A\cup B}=\max(\mu_A,\mu_B)$, intersection $\mu_{A\cap B}=\min(\mu_A,\mu_B)$ and complement $\mu_{A'}=1-\mu_A$.
- Fuzzy logic is a many-valued logic where truth is any value in $[0,1]$, using linguistic variables such as "temperature" with values cold, warm and hot.
- Fuzzy rules have the form IF temperature is high THEN fan speed is fast, and reasoning combines the rules and defuzzifies the result to a crisp output.
- In data mining it handles uncertainty: fuzzy clustering lets an object belong to several clusters, fuzzy classification gives soft class labels, and fuzzy association rules work on quantitative attributes turned into terms like "young" or "high income".
Example. For $A=\{x_1:0.2,\,x_2:0.7,\,x_3:1\}$ and $B=\{x_1:0.5,\,x_2:0.4,\,x_3:0.6\}$: union $=\{0.5,0.7,1\}$, intersection $=\{0.2,0.4,0.6\}$, complement of $A=\{0.8,0.3,0\}$.
Answer frame. Open by contrasting crisp and fuzzy sets; give the membership function and the three operations with the example; then explain linguistic variables and rules; close with the data mining applications.
Asked: [7 marks] (May 2024, Jun 2025) Discuss about fuzzy sets and fuzzy logic. Elaborate the concept of "fuzzy sets" and discuss how fuzzy logic can be applied to data mining.
Last-minute revision
- KDD is the whole process; data mining is one step in it.
- KDD phases: selection, pre-processing, transformation, mining, interpretation and evaluation.
- Predictive tasks: classification, regression, time series, prediction; descriptive tasks: clustering, association, summarisation, sequence discovery.
- Pre-processing tasks: cleaning, integration, transformation, reduction.
- Min-max: $v'=(v-\min)/(\max-\min)$; z-score: $v'=(v-\mu)/\sigma$.
- Data types for clustering: interval, ratio, binary, categorical, ordinal.
- Euclidean $=\sqrt{\sum(x_i-y_i)^2}$; cosine $=x\cdot y/(\|x\|\|y\|)$; Jaccard $=|A\cap B|/|A\cup B|$.
- Fuzzy: union max, intersection min, complement $1-\mu$.
- Crisp membership is 0 or 1; fuzzy membership is any value in $[0,1]$.
- Mining issues: methodology, performance, data diversity, privacy.
Memory hooks
- KDD phases: "Some People Take Many Estimates" (Selection, Pre-processing, Transformation, Mining, Evaluation).
- Predictive = predict the future; descriptive = describe the present.
- Fuzzy: union is Max, intersection is Min, complement is One minus.
- Cleaning, Integration, Transformation, Reduction: "CITR".
Coverage checklist
- Data Types, Quality of data: data types used in clustering (Q4).
- Data Pre-processing: necessity and tasks of pre-processing (Q3).
- Similarity measures: any two measures of similarity (Q7).
- Summary statistics: mean, median, variance (not asked).
- Data distributions: histogram, normal distribution (not asked).
- Basic data mining tasks, Data Mining Vs knowledge discovery in databases: KDD phases with diagram, data mining as a KDD step (Q1), mining tasks with examples (Q2).
- Issues in Data mining: issues in data mining (Q6).
- Introduction to Fuzzy sets and fuzzy logic: fuzzy sets, fuzzy logic and its use in data mining (Q5).