How unit 4 is examined
This unit covers data mining, its process and methods, data preparation (cleaning, missing values, recoding, scaling, normalizing) and visualization; no topic has been asked recently, so learn each definition and its core points.
Definition and applications of data mining
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data mining is the process of discovering useful, previously unknown patterns and relationships in large data sets using statistics, machine learning and database techniques.</mark>
Key points.
- It is the core step of Knowledge Discovery in Databases (KDD), which turns raw data into knowledge.
- It finds patterns such as associations, clusters, classes and outliers that ordinary queries cannot reveal.
- Applications include market-basket analysis in retail, fraud detection in banking and insurance, and customer churn prediction in telecom.
- It is also used for medical diagnosis, credit scoring and targeted marketing.
Data mining process
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The data mining process is the ordered set of steps that takes a business problem through data preparation and modelling to a deployed, evaluated result.</mark>
Key points.
- CRISP-DM has six phases: business understanding, data understanding, data preparation, modelling, evaluation and deployment.
- KDD gives a similar chain: selection, pre-processing, transformation, data mining, then interpretation and evaluation.
- Data preparation usually takes the largest share of the effort, often more than half.
- The process is iterative, so a poor evaluation sends the analyst back to earlier phases.
Analysis methodologies
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Analysis methodologies are the families of techniques used to model data and extract knowledge, chosen by whether the goal is prediction or description.</mark>
Key points.
- Supervised (predictive) methods learn from labelled data; classification and regression are the main examples.
- Unsupervised (descriptive) methods such as clustering find structure in unlabelled data.
- Association analysis finds items that occur together, for example bread and butter in a basket.
- Time-series analysis and forecasting study data ordered in time to predict future values.
Typical pre-processing operations: combining values into one
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Pre-processing is the preparation of raw data so that it is clean, consistent and suitable for mining; combining values into one merges several values or attributes into a single one.</mark>
Key points.
- Typical operations are cleaning, integration, transformation and reduction.
- Combining values replaces several detailed values by one, such as merging day, month and year fields into one date.
- It also merges rare categories into one class such as "Other", which reduces the number of distinct values.
- Combining reduces dimensionality but loses some detail, so it should keep what the analysis needs.
Handling incomplete or incorrect data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Incomplete data lacks values or attributes, and incorrect data holds wrong or out-of-range values; both must be found and treated before mining.</mark>
Key points.
- Incomplete data arises from unrecorded fields, equipment faults or fields that were not applicable.
- Incorrect data arises from entry errors, faulty sensors or transmission problems.
- Detection uses range checks, validation rules and outlier tests on each attribute.
- Treatment is to correct the value from the source, estimate it, or drop the record.
Handling missing values
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Missing values are attribute values that are absent from a record, and they are handled by deleting or filling them (imputation).</mark>
Key points.
- Ignoring or deleting the record is simple but wasteful when many values are missing.
- A constant such as "Unknown" can be filled in by hand or by rule.
- The attribute mean or median fills numeric gaps, and the mode fills categorical gaps.
- The most probable value can be predicted by regression or a decision tree, which is accurate but costly.
Recoding values
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Recoding replaces the existing values of an attribute with new codes or categories, for example turning numeric ages into age groups.</mark>
Key points.
- Binning converts a continuous value into intervals, such as ages 0-18, 19-60 and above 60.
- Categories can be converted to numbers, for example Male as 1 and Female as 0.
- Recoding gives inconsistent labels one standard form, such as "M", "male" and "Male" all becoming "M".
- It simplifies analysis but loses detail, so keep the original column.
Sub setting, sorting
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Sub setting selects the rows or columns needed from a data set, and sorting arranges records in ascending or descending order of one or more attributes.</mark>
Key points.
- Row subsetting filters records by a condition, such as sales in one region.
- Column subsetting keeps only the relevant attributes and reduces data volume.
- Sampling draws a random subset when the full data is too large to mine.
- Sorting orders the data, which helps ranking, finding extremes and removing duplicates.
Transforming scale, determining percentiles
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Scale transformation changes the range or units of values, and the percentile of a value is the percentage of observations that lie at or below it.</mark>
Key points.
- Rescaling puts attributes on a common scale, such as converting rupees to lakhs or marks to a 0-100 range.
- A log transform compresses a skewed range of values.
- The $p$th percentile is the value below which $p\%$ of the sorted data falls.
- The 50th percentile is the median, and the 25th and 75th are the quartiles $Q_1$ and $Q_3$.
Formula. Position of the $p$th percentile in $n$ sorted values: $i = \dfrac{p}{100}(n+1)$.
Data manipulation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data manipulation is the process of restructuring, combining and deriving data so that it is in the form the analysis needs.</mark>
Key points.
- It includes filtering, sorting, grouping and aggregating records, for example total sales per month.
- Joining or merging tables combines data from several sources on a common key.
- Derived attributes are computed from existing ones, such as profit equal to revenue minus cost.
- It is also called data wrangling, and it comes before modelling.
Removing noise, removing inconsistencies
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Noise is random error or variance in a measured value, and an inconsistency is a conflict between values that should agree; both are removed during cleaning.</mark>
Key points.
- Binning smooths noisy data by replacing each value with its bin mean, median or boundary.
- Regression smooths data by fitting a function to it, and clustering exposes outliers that lie outside groups.
- Inconsistencies, such as a wrong date format or an age that contradicts a birth date, are fixed by validation rules and cross-checks.
- Duplicate records are also detected and removed so they do not distort counts.
Transformations, standardizing, normalizing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Transformation converts data into forms suitable for mining, and standardizing or normalizing rescales attributes to a common range or distribution.</mark>
Key points.
- Attributes with large ranges, such as salary, would otherwise dominate small ones, such as age, in distance-based methods.
- Normalization scales values into a fixed range such as 0 to 1.
- Standardization rescales values to mean 0 and standard deviation 1.
- Other transformations are smoothing, aggregation, generalization and discretization.
Min-max normalization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Min-max normalization linearly rescales values into a new range, usually 0 to 1, keeping the relations among the original values.</mark>
Formula. $$v' = \frac{v - \min}{\max - \min}\,(\text{new\_max} - \text{new\_min}) + \text{new\_min}$$
Example. Values 200, 300, 400, 600, 1000 with min 200 and max 1000 give $v'=(v-200)/800$, so the results are 0, 0.125, 0.25, 0.5, 1. The value 400 becomes 0.25.
Key points.
- The minimum maps to 0 and the maximum to 1.
- It is sensitive to outliers, because one extreme value squeezes the rest into a narrow band.
Z-score standardization
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Z-score standardization rescales each value by subtracting the mean and dividing by the standard deviation, so the result has mean 0 and standard deviation 1.</mark>
Formula. $$z = \frac{v - \mu}{\sigma}$$
Example. For 200, 300, 400, 600, 1000 the mean is $\mu=500$ and the population standard deviation is $\sigma=\sqrt{80000}\approx 282.84$. Then 200 gives $z=(200-500)/282.84$, so $z \approx -1.06$.
Key points.
- A positive $z$ means the value is above the mean, and a negative $z$ means it is below.
- It is less affected by outliers than min-max and does not need a known minimum and maximum.
Rules of standardizing data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The rules of standardizing data are the guidelines that keep scaling consistent, so that all attributes are comparable and no information leaks.</mark>
Key points.
- Standardize only numeric attributes, and never scale identifiers or category codes.
- Compute the mean, standard deviation, minimum and maximum on the training data only, then reuse them on new data.
- Choose z-score when the data has outliers, and min-max when a fixed range is needed.
- Keep the original values and the scaling parameters, so results can be converted back.
Role of visualization in analytics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data visualization is the graphical presentation of data through charts, graphs and maps so that patterns, trends and outliers can be seen and understood quickly.</mark>
Key points.
- Visuals let the eye spot trends and outliers that tables of numbers hide.
- They support exploration of the data before modelling and help check results afterwards.
- They communicate findings to non-technical decision makers.
- Dashboards combine charts so that managers can monitor KPIs and decide faster.
Different techniques for visualizing data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Visualization techniques are the chart types chosen to match the data and the question, such as comparison, distribution, relationship or composition.</mark>
Key points.
- Bar and column charts compare categories, and line charts show trends over time.
- Pie charts show parts of a whole, and histograms and box plots show distributions.
- Scatter plots show the relationship between two numeric variables.
- Heat maps show intensity across a matrix, and geographic maps show location data.
Last-minute revision
- Data mining discovers useful, unknown patterns in large data; it is the core step of KDD.
- CRISP-DM phases: business understanding, data understanding, data preparation, modelling, evaluation, deployment.
- Supervised methods use labelled data (classification, regression); unsupervised methods (clustering) do not.
- Missing values are handled by deleting, filling a constant, mean, median or mode, or predicting the value.
- Binning smooths noise by replacing values with the bin mean, median or boundary.
- Recoding replaces values with new codes or groups, such as age bands.
- Percentile position: $i = \frac{p}{100}(n+1)$; the 50th percentile is the median.
- Min-max: $v'=\frac{v-\min}{\max-\min}(\text{new\_max}-\text{new\_min})+\text{new\_min}$, range 0 to 1 by default.
- Z-score: $z=\frac{v-\mu}{\sigma}$, giving mean 0 and standard deviation 1.
- Compute scaling parameters on training data only.
- Bar for comparison, line for trend, scatter for relationship, histogram for distribution.
Memory hooks
- CRISP-DM: "Business Data Data Model Eval Deploy" is BD-DMED.
- Missing values: "Delete, Default, Mean, Predict".
- Min-max squeezes into 0 to 1; z-score centres at 0.
- Chart choice: Bar compares, Line trends, Scatter relates, Histogram distributes.
Coverage checklist
- Definition and applications of data mining: definition, KDD role, applications (no past questions).
- Data mining process: CRISP-DM and KDD steps (no past questions).
- Analysis methodologies: supervised, unsupervised, association, time series (no past questions).
- Typical pre-processing operations: combining values into one: pre-processing operations and merging values (no past questions).
- Handling incomplete or incorrect data: causes, detection, treatment (no past questions).
- Handling missing values: deletion and imputation methods (no past questions).
- Recoding values: binning, coding, standard labels (no past questions).
- Sub setting, sorting: row and column subsets, sampling, sorting (no past questions).
- Transforming scale, determining percentiles: rescaling, percentile formula (no past questions).
- Data manipulation: filtering, joining, aggregating, derived attributes (no past questions).
- Removing noise, removing inconsistencies: binning, regression, clustering, validation (no past questions).
- Transformations, standardizing, normalizing: purpose and kinds of transformation (no past questions).
- Min-max normalization: formula and worked example (no past questions).
- Z-score standardization: formula and worked example (no past questions).
- Rules of standardizing data: guidelines for consistent scaling (no past questions).
- Role of visualization in analytics: exploration, communication, dashboards (no past questions).
- Different techniques for visualizing data: chart types and their uses (no past questions).