How unit 1 is examined
This unit covers what predictive analytics is, the three kinds of analytics and models (predictive, descriptive, decision), the analytical techniques used, and how data is classified, prepared and explored; none of the ten topics was asked in the supplied papers, so each is kept short but complete enough to answer if it appears.
Introduction to predictive analytics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Predictive analytics uses historical data, statistics and machine learning to estimate the probability or value of a future outcome.</mark>
Key points.
- It learns a pattern from past records and applies it to new cases, for example predicting customer churn, loan default or next month's sales.
- The output is a score, probability or forecast, never a certainty, so it supports a decision rather than replacing it.
- The usual flow is: define the goal, collect data, prepare it, build a model, validate it on unseen data and deploy it.
- Its accuracy depends on data quality and on the future resembling the past.
Business analytics: types, applications
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Business analytics is the use of data, statistics and quantitative methods to support and improve business decisions.</mark>
Key points.
- Descriptive analytics answers "what happened" using reports, dashboards and summaries.
- Predictive analytics answers "what will happen" using models and forecasts.
- Prescriptive analytics answers "what should we do" using optimization and simulation.
- Applications include retail demand forecasting, credit scoring, fraud detection, targeted marketing, churn prevention and supply-chain planning.
Predictive models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A predictive model maps input variables to a target variable so that the target can be estimated for unseen data.</mark>
Key points.
- Regression models predict a numeric target such as sales, for example $\hat{y}=\beta_0+\beta_1x_1+\dots+\beta_kx_k$.
- Classification models predict a category such as fraud or not fraud, and time-series models forecast values over time.
- They are trained on labelled historical data, which makes them supervised.
- They must be judged on held-out test data, not on the training data.
Descriptive models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A descriptive model summarizes and finds structure in existing data without predicting a target variable.</mark>
Key points.
- Clustering groups similar customers into segments, which is called customer profiling.
- Association rules find items that occur together, such as bread and butter in market-basket analysis.
- They are usually unsupervised because no target label is used.
- Their output explains the present data and often becomes an input to later predictive models.
Decision models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A decision model uses objectives, constraints and predictions to recommend the best action.</mark>
Key points.
- It is the prescriptive step and typically uses optimization, such as linear programming, or simulation.
- Its inputs are decision variables, an objective (maximize profit or minimize cost) and constraints such as budget or capacity.
- It often consumes predictions, for example forecast demand, to choose stock levels or prices.
- Its output is a recommended action, not a forecast.
Analytical techniques
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Analytical techniques are the statistical and data-mining methods used to extract patterns and build models from data.</mark>
Key points.
- Statistical techniques include regression, hypothesis testing, correlation and time-series analysis.
- Data-mining techniques include classification, clustering, association rules and decision trees.
- Machine-learning methods such as neural networks and ensembles learn complex non-linear patterns from data.
- The technique is chosen by the goal (predict, group or optimize) and by the type of data available.
Data types and associated techniques
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data is structured, semi-structured or unstructured, and each type needs different techniques.</mark>
Key points.
- Structured data sits in tables with fixed fields and suits SQL, regression and classification.
- Semi-structured data such as JSON or XML has tags but no fixed schema and needs parsing before analysis.
- Unstructured data such as text, images and audio needs text mining, computer vision or speech processing.
- Variables are also numeric (continuous or discrete) or categorical (nominal or ordinal), and this decides which statistics and encodings are valid.
Complexities of data
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data complexity is the set of properties, such as large volume, wide variety and poor quality, that make data hard to analyze.</mark>
Key points.
- Volume, velocity and variety of big data strain storage and processing.
- Missing values, noise, outliers and duplicates reduce the accuracy of models.
- Data from many sources has inconsistent formats and units and needs integration.
- High dimensionality and imbalanced classes make models overfit or ignore rare events such as fraud.
Data preparation, pre-processing
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Data preparation converts raw data into a clean, consistent form suitable for modelling.</mark>
Key points.
- Cleaning fills or removes missing values, corrects errors and removes duplicates and outliers.
- Integration merges data from several sources into one consistent set.
- Transformation applies scaling, normalization $x'=\frac{x-x_{min}}{x_{max}-x_{min}}$, and encoding of categories as numbers.
- Reduction lowers dimensions or rows through feature selection or sampling, and the whole step usually takes most of the project time.
Exploratory data analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Exploratory data analysis (EDA) is the visual and numerical examination of data to understand its structure, spot anomalies and form hypotheses before modelling.</mark>
Key points.
- Summary statistics such as mean, median, standard deviation and quartiles describe each variable.
- Histograms, box plots and scatter plots show distribution, outliers and relationships.
- A correlation matrix reveals which variables move together and which are redundant.
- Findings guide cleaning, feature choice and model selection.
Last-minute revision
- Predictive analytics estimates future outcomes from historical data using statistics and machine learning.
- Its flow: define goal, collect, prepare, model, validate, deploy.
- Three types of analytics: descriptive (what happened), predictive (what will happen), prescriptive (what to do).
- Predictive models are supervised: regression for numbers, classification for categories.
- Regression form: $\hat{y}=\beta_0+\beta_1x_1+\dots+\beta_kx_k$.
- Descriptive models (clustering, association rules) are usually unsupervised.
- Decision models use optimization or simulation with an objective and constraints.
- Data is structured, semi-structured (JSON, XML) or unstructured (text, images, audio).
- Variables are numeric (continuous, discrete) or categorical (nominal, ordinal).
- Complexities: volume, variety, velocity, missing values, noise, outliers, imbalance.
- Preparation steps: clean, integrate, transform, reduce.
- Min-max normalization is $x'=\frac{x-x_{min}}{x_{max}-x_{min}}$, giving values between 0 and 1.
- EDA uses summary statistics, plots and correlation before any modelling.
Memory hooks
- D-P-P: Describe, Predict, Prescribe.
- CITR for preparation: Clean, Integrate, Transform, Reduce.
- Regression = number, classification = label.
- Descriptive = no target, predictive = has target, decision = has constraints.
- SSU for data: Structured, Semi-structured, Unstructured.
- EDA = look before you model.
Coverage checklist
- Introduction to predictive analytics: definition, flow and limits (no past questions).
- Business analytics: types, applications: descriptive, predictive, prescriptive and uses (no past questions).
- Predictive models: regression and classification (no past questions).
- Descriptive models: clustering and association rules (no past questions).
- Decision models: optimization and simulation (no past questions).
- Analytical techniques: statistical, data mining and machine learning (no past questions).
- Data types and associated techniques: structured, semi-structured, unstructured, numeric, categorical (no past questions).
- Complexities of data: volume, quality, integration, imbalance (no past questions).
- Data preparation, pre-processing: cleaning, integration, transformation, reduction (no past questions).
- Exploratory data analysis: summary statistics, plots, correlation (no past questions).