Skip to content
AD-702 (B) · Business Intelligence/Quick Revision Short Notes

Business Intelligence (AD-702 (B)) - Unit 4 Short Notes

How unit 4 is examined

This unit covers data mining, its process and methods, data preparation (cleaning, missing values, recoding, scaling, normalizing) and visualization; no topic has been asked recently, so learn each definition and its core points.

Definition and applications of data mining

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data mining is the process of discovering useful, previously unknown patterns and relationships in large data sets using statistics, machine learning and database techniques.</mark>

Key points.

  1. It is the core step of Knowledge Discovery in Databases (KDD), which turns raw data into knowledge.
  2. It finds patterns such as associations, clusters, classes and outliers that ordinary queries cannot reveal.
  3. Applications include market-basket analysis in retail, fraud detection in banking and insurance, and customer churn prediction in telecom.
  4. It is also used for medical diagnosis, credit scoring and targeted marketing.

Data mining process

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>The data mining process is the ordered set of steps that takes a business problem through data preparation and modelling to a deployed, evaluated result.</mark>

Key points.

  1. CRISP-DM has six phases: business understanding, data understanding, data preparation, modelling, evaluation and deployment.
  2. KDD gives a similar chain: selection, pre-processing, transformation, data mining, then interpretation and evaluation.
  3. Data preparation usually takes the largest share of the effort, often more than half.
  4. The process is iterative, so a poor evaluation sends the analyst back to earlier phases.

Analysis methodologies

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Analysis methodologies are the families of techniques used to model data and extract knowledge, chosen by whether the goal is prediction or description.</mark>

Key points.

  1. Supervised (predictive) methods learn from labelled data; classification and regression are the main examples.
  2. Unsupervised (descriptive) methods such as clustering find structure in unlabelled data.
  3. Association analysis finds items that occur together, for example bread and butter in a basket.
  4. Time-series analysis and forecasting study data ordered in time to predict future values.

Typical pre-processing operations: combining values into one

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Pre-processing is the preparation of raw data so that it is clean, consistent and suitable for mining; combining values into one merges several values or attributes into a single one.</mark>

Key points.

  1. Typical operations are cleaning, integration, transformation and reduction.
  2. Combining values replaces several detailed values by one, such as merging day, month and year fields into one date.
  3. It also merges rare categories into one class such as "Other", which reduces the number of distinct values.
  4. Combining reduces dimensionality but loses some detail, so it should keep what the analysis needs.

Handling incomplete or incorrect data

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Incomplete data lacks values or attributes, and incorrect data holds wrong or out-of-range values; both must be found and treated before mining.</mark>

Key points.

  1. Incomplete data arises from unrecorded fields, equipment faults or fields that were not applicable.
  2. Incorrect data arises from entry errors, faulty sensors or transmission problems.
  3. Detection uses range checks, validation rules and outlier tests on each attribute.
  4. Treatment is to correct the value from the source, estimate it, or drop the record.

Handling missing values

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Missing values are attribute values that are absent from a record, and they are handled by deleting or filling them (imputation).</mark>

Key points.

  1. Ignoring or deleting the record is simple but wasteful when many values are missing.
  2. A constant such as "Unknown" can be filled in by hand or by rule.
  3. The attribute mean or median fills numeric gaps, and the mode fills categorical gaps.
  4. The most probable value can be predicted by regression or a decision tree, which is accurate but costly.

Recoding values

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Recoding replaces the existing values of an attribute with new codes or categories, for example turning numeric ages into age groups.</mark>

Key points.

  1. Binning converts a continuous value into intervals, such as ages 0-18, 19-60 and above 60.
  2. Categories can be converted to numbers, for example Male as 1 and Female as 0.
  3. Recoding gives inconsistent labels one standard form, such as "M", "male" and "Male" all becoming "M".
  4. It simplifies analysis but loses detail, so keep the original column.

Sub setting, sorting

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Sub setting selects the rows or columns needed from a data set, and sorting arranges records in ascending or descending order of one or more attributes.</mark>

Key points.

  1. Row subsetting filters records by a condition, such as sales in one region.
  2. Column subsetting keeps only the relevant attributes and reduces data volume.
  3. Sampling draws a random subset when the full data is too large to mine.
  4. Sorting orders the data, which helps ranking, finding extremes and removing duplicates.

Transforming scale, determining percentiles

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Scale transformation changes the range or units of values, and the percentile of a value is the percentage of observations that lie at or below it.</mark>

Key points.

  1. Rescaling puts attributes on a common scale, such as converting rupees to lakhs or marks to a 0-100 range.
  2. A log transform compresses a skewed range of values.
  3. The $p$th percentile is the value below which $p\%$ of the sorted data falls.
  4. The 50th percentile is the median, and the 25th and 75th are the quartiles $Q_1$ and $Q_3$.

Formula. Position of the $p$th percentile in $n$ sorted values: $i = \dfrac{p}{100}(n+1)$.

Data manipulation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data manipulation is the process of restructuring, combining and deriving data so that it is in the form the analysis needs.</mark>

Key points.

  1. It includes filtering, sorting, grouping and aggregating records, for example total sales per month.
  2. Joining or merging tables combines data from several sources on a common key.
  3. Derived attributes are computed from existing ones, such as profit equal to revenue minus cost.
  4. It is also called data wrangling, and it comes before modelling.

Removing noise, removing inconsistencies

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Noise is random error or variance in a measured value, and an inconsistency is a conflict between values that should agree; both are removed during cleaning.</mark>

Key points.

  1. Binning smooths noisy data by replacing each value with its bin mean, median or boundary.
  2. Regression smooths data by fitting a function to it, and clustering exposes outliers that lie outside groups.
  3. Inconsistencies, such as a wrong date format or an age that contradicts a birth date, are fixed by validation rules and cross-checks.
  4. Duplicate records are also detected and removed so they do not distort counts.

Transformations, standardizing, normalizing

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Transformation converts data into forms suitable for mining, and standardizing or normalizing rescales attributes to a common range or distribution.</mark>

Key points.

  1. Attributes with large ranges, such as salary, would otherwise dominate small ones, such as age, in distance-based methods.
  2. Normalization scales values into a fixed range such as 0 to 1.
  3. Standardization rescales values to mean 0 and standard deviation 1.
  4. Other transformations are smoothing, aggregation, generalization and discretization.

Min-max normalization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Min-max normalization linearly rescales values into a new range, usually 0 to 1, keeping the relations among the original values.</mark>

Formula. $$v' = \frac{v - \min}{\max - \min}\,(\text{new\_max} - \text{new\_min}) + \text{new\_min}$$

Example. Values 200, 300, 400, 600, 1000 with min 200 and max 1000 give $v'=(v-200)/800$, so the results are 0, 0.125, 0.25, 0.5, 1. The value 400 becomes 0.25.

Key points.

  1. The minimum maps to 0 and the maximum to 1.
  2. It is sensitive to outliers, because one extreme value squeezes the rest into a narrow band.

Z-score standardization

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Z-score standardization rescales each value by subtracting the mean and dividing by the standard deviation, so the result has mean 0 and standard deviation 1.</mark>

Formula. $$z = \frac{v - \mu}{\sigma}$$

Example. For 200, 300, 400, 600, 1000 the mean is $\mu=500$ and the population standard deviation is $\sigma=\sqrt{80000}\approx 282.84$. Then 200 gives $z=(200-500)/282.84$, so $z \approx -1.06$.

Key points.

  1. A positive $z$ means the value is above the mean, and a negative $z$ means it is below.
  2. It is less affected by outliers than min-max and does not need a known minimum and maximum.

Rules of standardizing data

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>The rules of standardizing data are the guidelines that keep scaling consistent, so that all attributes are comparable and no information leaks.</mark>

Key points.

  1. Standardize only numeric attributes, and never scale identifiers or category codes.
  2. Compute the mean, standard deviation, minimum and maximum on the training data only, then reuse them on new data.
  3. Choose z-score when the data has outliers, and min-max when a fixed range is needed.
  4. Keep the original values and the scaling parameters, so results can be converted back.

Role of visualization in analytics

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Data visualization is the graphical presentation of data through charts, graphs and maps so that patterns, trends and outliers can be seen and understood quickly.</mark>

Key points.

  1. Visuals let the eye spot trends and outliers that tables of numbers hide.
  2. They support exploration of the data before modelling and help check results afterwards.
  3. They communicate findings to non-technical decision makers.
  4. Dashboards combine charts so that managers can monitor KPIs and decide faster.

Different techniques for visualizing data

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Visualization techniques are the chart types chosen to match the data and the question, such as comparison, distribution, relationship or composition.</mark>

Key points.

  1. Bar and column charts compare categories, and line charts show trends over time.
  2. Pie charts show parts of a whole, and histograms and box plots show distributions.
  3. Scatter plots show the relationship between two numeric variables.
  4. Heat maps show intensity across a matrix, and geographic maps show location data.

Last-minute revision

  • Data mining discovers useful, unknown patterns in large data; it is the core step of KDD.
  • CRISP-DM phases: business understanding, data understanding, data preparation, modelling, evaluation, deployment.
  • Supervised methods use labelled data (classification, regression); unsupervised methods (clustering) do not.
  • Missing values are handled by deleting, filling a constant, mean, median or mode, or predicting the value.
  • Binning smooths noise by replacing values with the bin mean, median or boundary.
  • Recoding replaces values with new codes or groups, such as age bands.
  • Percentile position: $i = \frac{p}{100}(n+1)$; the 50th percentile is the median.
  • Min-max: $v'=\frac{v-\min}{\max-\min}(\text{new\_max}-\text{new\_min})+\text{new\_min}$, range 0 to 1 by default.
  • Z-score: $z=\frac{v-\mu}{\sigma}$, giving mean 0 and standard deviation 1.
  • Compute scaling parameters on training data only.
  • Bar for comparison, line for trend, scatter for relationship, histogram for distribution.

Memory hooks

  • CRISP-DM: "Business Data Data Model Eval Deploy" is BD-DMED.
  • Missing values: "Delete, Default, Mean, Predict".
  • Min-max squeezes into 0 to 1; z-score centres at 0.
  • Chart choice: Bar compares, Line trends, Scatter relates, Histogram distributes.

Coverage checklist

  • Definition and applications of data mining: definition, KDD role, applications (no past questions).
  • Data mining process: CRISP-DM and KDD steps (no past questions).
  • Analysis methodologies: supervised, unsupervised, association, time series (no past questions).
  • Typical pre-processing operations: combining values into one: pre-processing operations and merging values (no past questions).
  • Handling incomplete or incorrect data: causes, detection, treatment (no past questions).
  • Handling missing values: deletion and imputation methods (no past questions).
  • Recoding values: binning, coding, standard labels (no past questions).
  • Sub setting, sorting: row and column subsets, sampling, sorting (no past questions).
  • Transforming scale, determining percentiles: rescaling, percentile formula (no past questions).
  • Data manipulation: filtering, joining, aggregating, derived attributes (no past questions).
  • Removing noise, removing inconsistencies: binning, regression, clustering, validation (no past questions).
  • Transformations, standardizing, normalizing: purpose and kinds of transformation (no past questions).
  • Min-max normalization: formula and worked example (no past questions).
  • Z-score standardization: formula and worked example (no past questions).
  • Rules of standardizing data: guidelines for consistent scaling (no past questions).
  • Role of visualization in analytics: exploration, communication, dashboards (no past questions).
  • Different techniques for visualizing data: chart types and their uses (no past questions).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in