Skip to content
AD-702 (D) · Predictive Analytics/Quick Revision Short Notes

Predictive Analytics (AD-702 (D)) - Unit 2 Short Notes

How unit 2 is examined

This unit covers the standard model types (propensity, cluster, recommender), statistical analysis, and the discipline of building a trustworthy model: selection, validation, overfitting, bias-variance, balancing and baselines. No topic has appeared in recent papers, so learn each definition and its key points.

Propensity models

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A propensity model estimates the probability that a customer will take a specific action, such as buy, churn or default, and gives each customer a score between 0 and 1.</mark>

Key points.

  1. It is a supervised classification model, usually logistic regression or a tree, trained on past customers whose outcome is known.
  2. The output is a score, so customers can be ranked and only the top decile is targeted, which saves campaign cost.
  3. Common types are propensity to buy, to churn and to respond.
  4. Logistic form: $P(y=1)=\dfrac{1}{1+e^{-(b_0+b_1x_1+\dots+b_kx_k)}}$.

Cluster models

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A cluster model groups records into segments so that records in one cluster are similar and records in different clusters are dissimilar, without using any labels.</mark>

Key points.

  1. Clustering is unsupervised, because no target variable is given.
  2. K-means places $k$ centroids and repeatedly assigns points to the nearest centroid, then recomputes the centroids.
  3. Similarity is measured by distance, for example Euclidean: $d=\sqrt{\sum_i (x_i-y_i)^2}$.
  4. Typical use is customer segmentation, so that each segment gets its own offer.

Collaborative filtering

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Collaborative filtering recommends items to a user based on the past ratings and behaviour of similar users or similar items, without needing item content.</mark>

Key points.

  1. User-based filtering finds users with similar tastes and recommends what they liked.
  2. Item-based filtering recommends items similar to those the user already liked.
  3. Similarity is often cosine: $\text{sim}(u,v)=\dfrac{u\cdot v}{\|u\|\,\|v\|}$.
  4. Its weaknesses are the cold-start problem for new users or items and sparse rating data.

Univariate Statistical analysis

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Univariate analysis describes a single variable at a time, using its centre, spread and distribution, without studying relationships with other variables.</mark>

Key points.

  1. Central tendency is measured by mean, median and mode, with mean $\bar{x}=\frac{1}{n}\sum x_i$.
  2. Spread is measured by range, variance $s^2=\frac{\sum (x_i-\bar{x})^2}{n-1}$ and standard deviation.
  3. Histograms, box plots and frequency tables show the distribution and outliers.
  4. It is the first step of exploratory analysis and checks skewness before modelling.

Multivariate Statistical analysis

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Multivariate analysis studies three or more variables together to understand the relationships and joint behaviour among them.</mark>

Key points.

  1. It examines correlations and dependencies between variables, which univariate analysis cannot show.
  2. Common techniques are multiple regression, principal component analysis, factor analysis, MANOVA and cluster analysis.
  3. Principal component analysis reduces many correlated variables to a few uncorrelated components.
  4. Correlated predictors (multicollinearity) must be checked, because they make coefficients unstable.

Model Selection

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Model selection is choosing, from several candidate models, the one that generalises best to unseen data, judged on validation performance and complexity.</mark>

Key points.

  1. Candidates are compared on the same held-out or cross-validated data.
  2. Simpler models are preferred when performance is similar (Occam's razor).
  3. AIC and BIC penalise complexity: $AIC=2k-2\ln L$ and $BIC=k\ln n-2\ln L$, where $k$ is the number of parameters and $L$ the likelihood. The lower value wins.
  4. Never select using the test set, because that leaks information.

Supervised versus unsupervised methods

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Supervised methods learn from labelled data to predict a known target, while unsupervised methods find structure in unlabelled data.</mark>

Basis Supervised Unsupervised
Data Labelled Unlabelled
Goal Predict target Discover patterns
Tasks Classification, regression Clustering, association
Examples Logistic regression, trees K-means, PCA
Evaluation Accuracy, error Cohesion, interpretation

Statistical and data mining methodology

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A data mining methodology is a standard, repeatable process for turning a business problem into a deployed model; CRISP-DM is the most common.</mark>

Key points.

  1. CRISP-DM has six phases: business understanding, data understanding, data preparation, modelling, evaluation and deployment.
  2. The process is iterative, so results of evaluation often send the team back to earlier phases.
  3. The statistical approach starts from a hypothesis and tests it, while data mining explores large data for patterns.
  4. Business understanding comes first, because a model that answers the wrong question has no value.

Cross-validation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Cross-validation estimates how well a model generalises by repeatedly training on part of the data and testing on the remaining part.</mark>

Key points.

  1. In k-fold cross-validation the data is split into $k$ equal folds; each fold is used once as the test set while the other $k-1$ train.
  2. The final estimate is the average of the $k$ scores, which is more stable than a single split.
  3. Holdout uses one train-test split, which is fast but depends on the split chosen.
  4. Leave-one-out is the case $k=n$, and stratified folds keep class proportions.

Overfitting

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Overfitting occurs when a model learns the noise in the training data, so it scores very well on training data but poorly on new data.</mark>

Key points.

  1. The sign is low training error with much higher validation error.
  2. Causes are an over-complex model, too little data and too many features.
  3. Remedies are more data, feature reduction, pruning, regularisation and cross-validation.
  4. Underfitting is the opposite: the model is too simple and errors are high on both sets.

Bias-variance trade-off

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>The bias-variance trade-off is the tension that lowering a model's bias (error from wrong assumptions) usually raises its variance (sensitivity to the training sample), and vice versa.</mark>

Key points.

  1. Expected error $=\text{Bias}^2+\text{Variance}+\text{Irreducible noise}$.
  2. Simple models have high bias and low variance, so they underfit.
  3. Complex models have low bias and high variance, so they overfit.
  4. The best model sits at the complexity where total error is lowest, and regularisation moves a model toward it.

Balancing the training dataset

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Balancing adjusts the class proportions of the training data so that a rare class, such as fraud, is not ignored by the model.</mark>

Key points.

  1. On imbalanced data a model can reach high accuracy by always predicting the majority class.
  2. Undersampling removes majority records, while oversampling duplicates minority records.
  3. SMOTE creates synthetic minority records by interpolating between neighbouring minority points.
  4. Balance only the training set, never the test set, and judge with precision, recall or F1 instead of accuracy.

Establishing baseline performance

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A baseline is the performance of a very simple model or rule, against which every proposed model must show a real improvement.</mark>

Key points.

  1. For classification the naive baseline predicts the majority class; for regression it predicts the mean.
  2. A complex model that cannot beat the baseline is not worth deploying.
  3. Other baselines are the current business rule or the previous model.
  4. Example: if 95 percent of customers do not churn, the majority baseline has 95 percent accuracy.

Last-minute revision

  • A propensity model outputs the probability of an action, from logistic regression or trees.
  • Clustering is unsupervised; K-means assigns points to the nearest centroid and recomputes centroids.
  • Collaborative filtering uses similar users or items, with cosine similarity; cold start is its weakness.
  • Univariate studies one variable; multivariate studies several together, for example PCA.
  • AIC $=2k-2\ln L$ and BIC $=k\ln n-2\ln L$; the lower value is better.
  • Supervised uses labels; unsupervised does not.
  • CRISP-DM phases: business, data understanding, preparation, modelling, evaluation, deployment.
  • K-fold cross-validation averages $k$ test scores.
  • Overfitting means low training error and high test error.
  • Error $=\text{Bias}^2+\text{Variance}+\text{noise}$.
  • SMOTE synthesises minority samples; balance only the training set.
  • The baseline is the majority class or the mean.

Memory hooks

  • Propensity = "chance score" for each customer.
  • Cluster = birds of a feather, no labels.
  • Bias underfits, variance overfits.
  • CRISP-DM: "Big Dogs Play Music, Eat Dinner" (Business, Data, Prepare, Model, Evaluate, Deploy).
  • Baseline = beat the lazy guess first.

Coverage checklist

  • Propensity models: covered; no past questions.
  • Cluster models: covered; no past questions.
  • Collaborative filtering: covered; no past questions.
  • Univariate Statistical analysis: covered; no past questions.
  • Multivariate Statistical analysis: covered; no past questions.
  • Model Selection: covered; no past questions.
  • Supervised versus unsupervised methods: covered; no past questions.
  • Statistical and data mining methodology: covered; no past questions.
  • Cross-validation: covered; no past questions.
  • Overfitting: covered; no past questions.
  • Bias-variance trade-off: covered; no past questions.
  • Balancing the training dataset: covered; no past questions.
  • Establishing baseline performance: covered; no past questions.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in