How unit 2 is examined
This unit covers correlation, least-squares regression and its variants, and ANOVA; no past questions have been asked, so every topic is short and formula-led.
Correlation, Scatter diagram
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Correlation measures the strength and direction of the linear association between two variables, and a scatter diagram plots the paired observations $(x_i, y_i)$ to show it.</mark>
Key points.
- Points rising from left to right show positive correlation, and points falling show negative correlation.
- A shapeless cloud shows no correlation, and points on a straight line show perfect correlation.
- Correlation shows association only; it does not prove that one variable causes the other.
Karl Pearson's coefficient of correlation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Pearson's $r$ is the covariance of $x$ and $y$ divided by the product of their standard deviations.</mark>
Formula.
$$r=\frac{\operatorname{Cov}(x,y)}{\sigma_x\sigma_y}=\frac{n\sum xy-\sum x\sum y}{\sqrt{n\sum x^2-(\sum x)^2}\,\sqrt{n\sum y^2-(\sum y)^2}}$$
Key points.
- The value always lies between $-1$ and $+1$, where $+1$ is perfect positive and $-1$ is perfect negative correlation.
- It is unchanged by a change of origin or a positive change of scale.
- It measures only linear relationship, so $r=0$ does not rule out a curved one.
Spearman's Rank correlation coefficient
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Spearman's coefficient is Pearson's $r$ computed on the ranks of the data, and it suits ordinal or non-normal data.</mark>
Formula.
$$\rho=1-\frac{6\sum d^2}{n(n^2-1)}$$
where $d$ is the difference between the two ranks of each item.
Key points.
- It lies between $-1$ and $+1$, like Pearson's $r$.
- For tied ranks, give each tied item the average rank and add $\frac{m(m^2-1)}{12}$ to $\sum d^2$ for each group of $m$ ties.
- It measures any monotonic relationship, not only linear.
Methods of least square
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>The method of least squares fits the line or curve that minimises the sum of the squared vertical deviations $\sum e_i^2$ between observed and fitted values.</mark>
Key points.
- For the line $y=a+bx$, minimising $\sum(y-a-bx)^2$ gives the normal equations $\sum y=na+b\sum x$ and $\sum xy=a\sum x+b\sum x^2$.
- Solving them gives $b=\frac{n\sum xy-\sum x\sum y}{n\sum x^2-(\sum x)^2}$ and $a=\bar y-b\bar x$.
- The fitted line always passes through the point $(\bar x,\bar y)$, and the residuals sum to zero.
Simple linear Regression model, SLR assumptions and prediction
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==Simple linear regression models one response as $y=\beta_0+\beta_1x+\varepsilon$, where $\varepsilon$ is a random error.==
Key points.
- The relationship between $x$ and the mean of $y$ is linear.
- Errors are independent, have mean zero, have constant variance (homoscedasticity) and are normally distributed.
- The slope $\beta_1$ is the change in $y$ for a unit change in $x$, estimated by $b=r\,\frac{s_y}{s_x}$.
- Prediction is $\hat y=a+bx$, reliable only for $x$ inside the observed range.
Multiple linear Regression, MLR assumption and prediction
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==Multiple linear regression models one response on several predictors as $y=\beta_0+\beta_1x_1+\dots+\beta_kx_k+\varepsilon$.==
Key points.
- In matrix form the least-squares estimate is $\hat\beta=(X^TX)^{-1}X^Ty$.
- It keeps the SLR assumptions of linearity, independent errors, constant variance and normal errors.
- It adds no perfect multicollinearity, meaning no predictor is an exact linear combination of the others.
- Each $\beta_j$ is the change in $y$ per unit of $x_j$ with the other predictors held fixed, and prediction substitutes the new $x$ values.
Polynomial Regression
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==Polynomial regression fits a curve $y=\beta_0+\beta_1x+\beta_2x^2+\dots+\beta_kx^k+\varepsilon$ to curved data.==
Key points.
- It is still linear in the coefficients, so ordinary least squares fits it by treating $x,x^2,\dots,x^k$ as separate predictors.
- A higher degree fits the sample better but overfits, so use the lowest degree that captures the curve.
- Powers of $x$ are highly correlated, so centring $x$ reduces multicollinearity.
Logistics Regression
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==Logistic regression models the probability of a binary outcome through the logit, $\ln\frac{p}{1-p}=\beta_0+\beta_1x_1+\dots+\beta_kx_k$.==
Key points.
- The predicted probability is $p=\frac{1}{1+e^{-(\beta_0+\beta_1x)}}$, always between 0 and 1.
- Coefficients are estimated by maximum likelihood, not least squares.
- $e^{\beta_j}$ is the odds ratio for a unit rise in $x_j$.
- It is used for classification, such as pass or fail.
Poisson Regression
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==Poisson regression models a count response $Y\sim\text{Poisson}(\mu)$ with the log link $\ln\mu=\beta_0+\beta_1x_1+\dots+\beta_kx_k$.==
Key points.
- It is a generalised linear model for counts such as accidents or calls per day.
- The log link keeps the predicted mean $\mu=e^{\beta_0+\beta_1x}$ positive.
- It assumes the mean equals the variance, and overdispersion breaks this.
- Coefficients are fitted by maximum likelihood, and $e^{\beta_j}$ is the multiplicative change in the mean.
Non-Linear Regression
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==Non-linear regression fits a model that is non-linear in its parameters, such as $y=ae^{bx}$, by minimising the sum of squared errors.==
Key points.
- Some models linearise by a transformation; for example $y=ae^{bx}$ becomes $\ln y=\ln a+bx$.
- Truly non-linear models have no closed-form solution and are fitted by iterative methods such as Gauss-Newton.
- Iteration needs good starting values and may reach a local minimum.
Analysis of Variance (One way & Two Way)
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>ANOVA tests whether the means of three or more groups are equal by comparing the variance between groups with the variance within groups using the F-ratio.</mark>
Formula. One way: $SST=SSB+SSW$ and $F=\frac{SSB/(k-1)}{SSW/(N-k)}$.
Key points.
- One way uses one factor, and two way uses two factors and can also test their interaction.
- $H_0$ says all group means are equal, and it is rejected when $F$ exceeds the table value.
- It assumes normal populations, equal variances and independent samples.
Analysis of Covariance
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>ANCOVA combines ANOVA with regression to compare group means after adjusting for a continuous covariate.</mark>
Key points.
- The covariate is a nuisance variable, such as pre-test score, that affects the response.
- It removes the covariate's effect from the error, which raises power and gives adjusted means.
- It also assumes the regression slope is the same in every group.
Multivariate Analysis of Variance
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>MANOVA extends ANOVA to test whether group mean vectors differ on several dependent variables at once.</mark>
Key points.
- It avoids the inflated Type I error of running a separate ANOVA on each variable.
- Common test statistics are Wilks' lambda, Pillai's trace and Roy's largest root.
- It assumes multivariate normality and equal covariance matrices across groups.
Last-minute revision
- Pearson $r=\frac{\operatorname{Cov}(x,y)}{\sigma_x\sigma_y}$, always between $-1$ and $+1$.
- Spearman $\rho=1-\frac{6\sum d^2}{n(n^2-1)}$.
- Least squares minimises $\sum e^2$, and the line passes through $(\bar x,\bar y)$.
- Slope $b=\frac{n\sum xy-\sum x\sum y}{n\sum x^2-(\sum x)^2}$ and intercept $a=\bar y-b\bar x$.
- Also $b=r\,s_y/s_x$.
- SLR assumptions are linearity, independence, constant variance and normal errors.
- MLR estimate is $\hat\beta=(X^TX)^{-1}X^Ty$, and it needs no multicollinearity.
- Polynomial regression is linear in its coefficients.
- Logistic regression uses the logit link for binary outcomes, and Poisson regression uses the log link for counts.
- One-way ANOVA $F=\frac{SSB/(k-1)}{SSW/(N-k)}$.
- ANCOVA adjusts means for a covariate, and MANOVA handles several dependent variables.
Memory hooks
- Pearson is for measured data, Spearman is for ranks.
- Least squares means the smallest squared misses.
- Logit for Yes/No, log for Count.
- ANOVA compares Between with Within, and the ratio is F.
- ANCOVA is ANOVA plus a Covariate, and MANOVA is ANOVA with Many outcomes.
Coverage checklist
- Correlation, Scatter diagram: no past questions.
- Karl Pearson's coefficient of correlation: no past questions.
- Spearman's Rank correlation coefficient: no past questions.
- Methods of least square: no past questions.
- Simple linear Regression model, SLR assumptions and prediction: no past questions.
- Multiple linear Regression, MLR assumption and prediction: no past questions.
- Polynomial Regression: no past questions.
- Logistics Regression: no past questions.
- Poisson Regression: no past questions.
- Non-Linear Regression: no past questions.
- Analysis of Variance (One way & Two Way): no past questions.
- Analysis of Covariance: no past questions.
- Multivariate Analysis of Variance: no past questions.