How unit 2 is examined
This unit covers the standard model types (propensity, cluster, recommender), statistical analysis, and the discipline of building a trustworthy model: selection, validation, overfitting, bias-variance, balancing and baselines. No topic has appeared in recent papers, so learn each definition and its key points.
Propensity models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A propensity model estimates the probability that a customer will take a specific action, such as buy, churn or default, and gives each customer a score between 0 and 1.</mark>
Key points.
- It is a supervised classification model, usually logistic regression or a tree, trained on past customers whose outcome is known.
- The output is a score, so customers can be ranked and only the top decile is targeted, which saves campaign cost.
- Common types are propensity to buy, to churn and to respond.
- Logistic form: $P(y=1)=\dfrac{1}{1+e^{-(b_0+b_1x_1+\dots+b_kx_k)}}$.
Cluster models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A cluster model groups records into segments so that records in one cluster are similar and records in different clusters are dissimilar, without using any labels.</mark>
Key points.
- Clustering is unsupervised, because no target variable is given.
- K-means places $k$ centroids and repeatedly assigns points to the nearest centroid, then recomputes the centroids.
- Similarity is measured by distance, for example Euclidean: $d=\sqrt{\sum_i (x_i-y_i)^2}$.
- Typical use is customer segmentation, so that each segment gets its own offer.
Collaborative filtering
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Collaborative filtering recommends items to a user based on the past ratings and behaviour of similar users or similar items, without needing item content.</mark>
Key points.
- User-based filtering finds users with similar tastes and recommends what they liked.
- Item-based filtering recommends items similar to those the user already liked.
- Similarity is often cosine: $\text{sim}(u,v)=\dfrac{u\cdot v}{\|u\|\,\|v\|}$.
- Its weaknesses are the cold-start problem for new users or items and sparse rating data.
Univariate Statistical analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Univariate analysis describes a single variable at a time, using its centre, spread and distribution, without studying relationships with other variables.</mark>
Key points.
- Central tendency is measured by mean, median and mode, with mean $\bar{x}=\frac{1}{n}\sum x_i$.
- Spread is measured by range, variance $s^2=\frac{\sum (x_i-\bar{x})^2}{n-1}$ and standard deviation.
- Histograms, box plots and frequency tables show the distribution and outliers.
- It is the first step of exploratory analysis and checks skewness before modelling.
Multivariate Statistical analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Multivariate analysis studies three or more variables together to understand the relationships and joint behaviour among them.</mark>
Key points.
- It examines correlations and dependencies between variables, which univariate analysis cannot show.
- Common techniques are multiple regression, principal component analysis, factor analysis, MANOVA and cluster analysis.
- Principal component analysis reduces many correlated variables to a few uncorrelated components.
- Correlated predictors (multicollinearity) must be checked, because they make coefficients unstable.
Model Selection
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Model selection is choosing, from several candidate models, the one that generalises best to unseen data, judged on validation performance and complexity.</mark>
Key points.
- Candidates are compared on the same held-out or cross-validated data.
- Simpler models are preferred when performance is similar (Occam's razor).
- AIC and BIC penalise complexity: $AIC=2k-2\ln L$ and $BIC=k\ln n-2\ln L$, where $k$ is the number of parameters and $L$ the likelihood. The lower value wins.
- Never select using the test set, because that leaks information.
Supervised versus unsupervised methods
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Supervised methods learn from labelled data to predict a known target, while unsupervised methods find structure in unlabelled data.</mark>
| Basis | Supervised | Unsupervised |
|---|---|---|
| Data | Labelled | Unlabelled |
| Goal | Predict target | Discover patterns |
| Tasks | Classification, regression | Clustering, association |
| Examples | Logistic regression, trees | K-means, PCA |
| Evaluation | Accuracy, error | Cohesion, interpretation |
Statistical and data mining methodology
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A data mining methodology is a standard, repeatable process for turning a business problem into a deployed model; CRISP-DM is the most common.</mark>
Key points.
- CRISP-DM has six phases: business understanding, data understanding, data preparation, modelling, evaluation and deployment.
- The process is iterative, so results of evaluation often send the team back to earlier phases.
- The statistical approach starts from a hypothesis and tests it, while data mining explores large data for patterns.
- Business understanding comes first, because a model that answers the wrong question has no value.
Cross-validation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Cross-validation estimates how well a model generalises by repeatedly training on part of the data and testing on the remaining part.</mark>
Key points.
- In k-fold cross-validation the data is split into $k$ equal folds; each fold is used once as the test set while the other $k-1$ train.
- The final estimate is the average of the $k$ scores, which is more stable than a single split.
- Holdout uses one train-test split, which is fast but depends on the split chosen.
- Leave-one-out is the case $k=n$, and stratified folds keep class proportions.
Overfitting
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Overfitting occurs when a model learns the noise in the training data, so it scores very well on training data but poorly on new data.</mark>
Key points.
- The sign is low training error with much higher validation error.
- Causes are an over-complex model, too little data and too many features.
- Remedies are more data, feature reduction, pruning, regularisation and cross-validation.
- Underfitting is the opposite: the model is too simple and errors are high on both sets.
Bias-variance trade-off
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The bias-variance trade-off is the tension that lowering a model's bias (error from wrong assumptions) usually raises its variance (sensitivity to the training sample), and vice versa.</mark>
Key points.
- Expected error $=\text{Bias}^2+\text{Variance}+\text{Irreducible noise}$.
- Simple models have high bias and low variance, so they underfit.
- Complex models have low bias and high variance, so they overfit.
- The best model sits at the complexity where total error is lowest, and regularisation moves a model toward it.
Balancing the training dataset
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Balancing adjusts the class proportions of the training data so that a rare class, such as fraud, is not ignored by the model.</mark>
Key points.
- On imbalanced data a model can reach high accuracy by always predicting the majority class.
- Undersampling removes majority records, while oversampling duplicates minority records.
- SMOTE creates synthetic minority records by interpolating between neighbouring minority points.
- Balance only the training set, never the test set, and judge with precision, recall or F1 instead of accuracy.
Establishing baseline performance
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A baseline is the performance of a very simple model or rule, against which every proposed model must show a real improvement.</mark>
Key points.
- For classification the naive baseline predicts the majority class; for regression it predicts the mean.
- A complex model that cannot beat the baseline is not worth deploying.
- Other baselines are the current business rule or the previous model.
- Example: if 95 percent of customers do not churn, the majority baseline has 95 percent accuracy.
Last-minute revision
- A propensity model outputs the probability of an action, from logistic regression or trees.
- Clustering is unsupervised; K-means assigns points to the nearest centroid and recomputes centroids.
- Collaborative filtering uses similar users or items, with cosine similarity; cold start is its weakness.
- Univariate studies one variable; multivariate studies several together, for example PCA.
- AIC $=2k-2\ln L$ and BIC $=k\ln n-2\ln L$; the lower value is better.
- Supervised uses labels; unsupervised does not.
- CRISP-DM phases: business, data understanding, preparation, modelling, evaluation, deployment.
- K-fold cross-validation averages $k$ test scores.
- Overfitting means low training error and high test error.
- Error $=\text{Bias}^2+\text{Variance}+\text{noise}$.
- SMOTE synthesises minority samples; balance only the training set.
- The baseline is the majority class or the mean.
Memory hooks
- Propensity = "chance score" for each customer.
- Cluster = birds of a feather, no labels.
- Bias underfits, variance overfits.
- CRISP-DM: "Big Dogs Play Music, Eat Dinner" (Business, Data, Prepare, Model, Evaluate, Deploy).
- Baseline = beat the lazy guess first.
Coverage checklist
- Propensity models: covered; no past questions.
- Cluster models: covered; no past questions.
- Collaborative filtering: covered; no past questions.
- Univariate Statistical analysis: covered; no past questions.
- Multivariate Statistical analysis: covered; no past questions.
- Model Selection: covered; no past questions.
- Supervised versus unsupervised methods: covered; no past questions.
- Statistical and data mining methodology: covered; no past questions.
- Cross-validation: covered; no past questions.
- Overfitting: covered; no past questions.
- Bias-variance trade-off: covered; no past questions.
- Balancing the training dataset: covered; no past questions.
- Establishing baseline performance: covered; no past questions.