How unit 3 is examined
This unit covers how regression and classification models are built and scored; no topic was asked in the supplied papers, so every topic is short but complete.
Measuring Performance in Regression Models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Regression performance is measured by comparing predicted values with observed values using error metrics such as RMSE, MAE and $R^2$.</mark>
Key points.
- $\text{RMSE}=\sqrt{\frac{1}{n}\sum (y_i-\hat y_i)^2}$ is in the units of the response and penalises large errors heavily.
- $\text{MAE}=\frac{1}{n}\sum |y_i-\hat y_i|$ is less sensitive to outliers than RMSE.
- $R^2=1-\frac{SS_{res}}{SS_{tot}}$ is the share of variance explained, but it measures correlation-like fit and not accuracy of prediction.
- Metrics must be computed on held-out data, not on the training set.
Linear Regression and Its Cousins
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==Linear regression models the response as a linear combination of predictors, $y=\beta_0+\beta_1x_1+\dots+\beta_px_p+\varepsilon$, with coefficients found by ordinary least squares (OLS).==
Key points.
- OLS chooses $\beta$ to minimise $\sum(y_i-\hat y_i)^2$, giving $\hat\beta=(X^TX)^{-1}X^Ty$.
- It is simple and interpretable but fails with correlated predictors or more predictors than samples.
- Cousins add a penalty: ridge ($L_2$) shrinks coefficients, lasso ($L_1$) sets some to zero, and elastic net mixes both.
- Partial least squares and principal component regression handle collinearity by using derived components.
Non-Linear Regression Models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Non-linear regression models capture curved relationships between predictors and response that a straight line cannot fit.</mark>
Key points.
- Neural networks fit flexible curves through hidden layers of non-linear units.
- Multivariate adaptive regression splines (MARS) build the fit from piecewise linear hinge functions.
- Support vector machines for regression fit a tube around the data and ignore errors inside it.
- K-nearest neighbours predicts the average response of the $k$ closest training points, so $k$ controls flexibility.
Regression Trees and Rule-Based Models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A regression tree repeatedly splits the data on predictor values so that each leaf predicts the mean response of its samples.</mark>
Key points.
- Each split is chosen to minimise the sum of squared errors within the two child nodes.
- Trees handle non-linearity, interactions and mixed data types without scaling, but are unstable and overfit if grown deep.
- Model trees fit a linear regression in each leaf instead of a constant.
- Rule-based models turn each path into an if-then rule; bagging, random forests and boosting combine many trees to reduce variance.
Case Study: Compressive Strength of Concrete Mixtures
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The concrete case study predicts the compressive strength of a concrete mix from its ingredients and age, and compares many regression models on it.</mark>
Key points.
- Predictors are cement, slag, fly ash, water, superplasticizer, coarse and fine aggregate, and age in days.
- The response is compressive strength in MPa, a continuous value, so this is a regression task.
- Models compared include linear regression, lasso, MARS, neural networks, SVM, regression trees and ensembles, using resampled RMSE and $R^2$.
- Tree ensembles and boosting usually perform best because strength depends non-linearly on ingredient ratios.
Measuring Performance in Classification Models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Classification performance is judged from the confusion matrix of predicted against actual classes, along with class probabilities and AUC.</mark>
Key points.
- $\text{Accuracy}=\frac{TP+TN}{TP+TN+FP+FN}$ misleads when classes are imbalanced.
- $\text{Precision}=\frac{TP}{TP+FP}$ and $\text{Recall}=\frac{TP}{TP+FN}$; specificity is $\frac{TN}{TN+FP}$.
- $F_1=\frac{2PR}{P+R}$ balances precision and recall.
- The Kappa statistic corrects accuracy for chance agreement.
Discriminant Analysis and Other Linear Classification Models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Linear discriminant analysis (LDA) classifies a sample to the class with the highest posterior probability, assuming Gaussian classes with a common covariance, which gives linear boundaries.</mark>
Key points.
- LDA finds the linear combination of predictors that best separates class means relative to within-class spread.
- Logistic regression models the log-odds $\ln\frac{p}{1-p}=\beta_0+\beta^Tx$ and outputs class probabilities directly.
- Partial least squares discriminant analysis and penalised versions (glmnet) handle correlated or many predictors.
- Nearest shrunken centroids shrink class centroids to select useful predictors.
Non-Linear Classification Models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Non-linear classifiers form curved decision boundaries, for example with kernels, neighbours, networks or Bayes rule.</mark>
Key points.
- Support vector machines maximise the margin and use kernels (polynomial, RBF) to separate classes non-linearly.
- K-nearest neighbours assigns the majority class among the $k$ closest points and needs scaled predictors.
- Neural networks learn non-linear boundaries through hidden layers.
- Naive Bayes applies Bayes rule assuming predictors are independent given the class; quadratic and flexible discriminant analysis also give curved boundaries.
Classification Trees and Rule-Based Models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A classification tree splits data on predictors to make leaves as pure as possible, each leaf predicting its majority class.</mark>
Key points.
- Splits are chosen by reducing impurity: Gini $=1-\sum p_k^2$ or entropy $=-\sum p_k\log_2 p_k$.
- C4.5 uses gain ratio and C5.0 adds boosting and rule extraction; CART builds binary splits.
- Pruning cuts back the tree to avoid overfitting.
- Rule sets (if-then) are readable, and bagging, random forests and boosting improve accuracy.
Model Evaluation Techniques
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Model evaluation estimates how well a model will perform on unseen data using validation and threshold-independent tools such as the ROC curve.</mark>
Key points.
- Hold-out, k-fold cross-validation and bootstrap resampling estimate future performance without touching the test set.
- The ROC curve plots true positive rate (sensitivity) against false positive rate ($1-$specificity) as the threshold varies.
- AUC is the area under the ROC curve: 0.5 means random guessing and 1 means perfect separation.
- Lift and gain charts show how well a model ranks positives at the top.
Last-minute revision
- RMSE $=\sqrt{\text{mean}(y-\hat y)^2}$; MAE $=\text{mean}|y-\hat y|$; $R^2=1-SS_{res}/SS_{tot}$.
- OLS: $\hat\beta=(X^TX)^{-1}X^Ty$; ridge uses $L_2$, lasso uses $L_1$.
- Non-linear regressors: neural network, MARS, SVM, kNN.
- Regression tree leaf predicts the mean; model tree fits a linear model in the leaf.
- Concrete strength depends on cement, water, age and other ingredients; ensembles win.
- Accuracy $=(TP+TN)/N$; precision $=TP/(TP+FP)$; recall $=TP/(TP+FN)$.
- $F_1=2PR/(P+R)$; Kappa corrects for chance.
- LDA assumes Gaussian classes with equal covariance; logistic regression models log-odds.
- Gini $=1-\sum p^2$; entropy $=-\sum p\log_2 p$.
- ROC plots TPR against FPR; AUC 0.5 is random.
Memory hooks
- RMSE squares, MAE absolutes: R for "ruthless" on big errors.
- Ridge shrinks, lasso drops.
- SVM = Sharpest Vector Margin; kernels bend the line.
- Gini and entropy both mean "how mixed is the leaf".
- ROC: True positives up the side, False positives along the bottom.
Coverage checklist
- Measuring Performance in Regression Models: RMSE, MAE, $R^2$ (no past questions).
- Linear Regression and Its Cousins: OLS, ridge, lasso (no past questions).
- Non-Linear Regression Models: neural network, MARS, SVM, kNN (no past questions).
- Regression Trees and Rule-Based Models: regression and model trees (no past questions).
- Case Study: Compressive Strength of Concrete Mixtures: predictors, response, model comparison (no past questions).
- Measuring Performance in Classification Models: confusion matrix metrics (no past questions).
- Discriminant Analysis and Other Linear Classification Models: LDA, logistic regression (no past questions).
- Non-Linear Classification Models: SVM, kNN, neural network, naive Bayes (no past questions).
- Classification Trees and Rule-Based Models: Gini, entropy, pruning, C5.0 (no past questions).
- Model Evaluation Techniques: cross-validation, ROC, AUC (no past questions).