Skip to content
AD-703 (C) · Advanced Statistical Analytics/Quick Revision Short Notes

Advanced Statistical Analytics (AD-703 (C)) - Unit 2 Short Notes

How unit 2 is examined

This unit covers correlation, least-squares regression and its variants, and ANOVA; no past questions have been asked, so every topic is short and formula-led.

Correlation, Scatter diagram

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Correlation measures the strength and direction of the linear association between two variables, and a scatter diagram plots the paired observations $(x_i, y_i)$ to show it.</mark>

Key points.

  1. Points rising from left to right show positive correlation, and points falling show negative correlation.
  2. A shapeless cloud shows no correlation, and points on a straight line show perfect correlation.
  3. Correlation shows association only; it does not prove that one variable causes the other.

Karl Pearson's coefficient of correlation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Pearson's $r$ is the covariance of $x$ and $y$ divided by the product of their standard deviations.</mark>

Formula.

$$r=\frac{\operatorname{Cov}(x,y)}{\sigma_x\sigma_y}=\frac{n\sum xy-\sum x\sum y}{\sqrt{n\sum x^2-(\sum x)^2}\,\sqrt{n\sum y^2-(\sum y)^2}}$$

Key points.

  1. The value always lies between $-1$ and $+1$, where $+1$ is perfect positive and $-1$ is perfect negative correlation.
  2. It is unchanged by a change of origin or a positive change of scale.
  3. It measures only linear relationship, so $r=0$ does not rule out a curved one.

Spearman's Rank correlation coefficient

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Spearman's coefficient is Pearson's $r$ computed on the ranks of the data, and it suits ordinal or non-normal data.</mark>

Formula.

$$\rho=1-\frac{6\sum d^2}{n(n^2-1)}$$

where $d$ is the difference between the two ranks of each item.

Key points.

  1. It lies between $-1$ and $+1$, like Pearson's $r$.
  2. For tied ranks, give each tied item the average rank and add $\frac{m(m^2-1)}{12}$ to $\sum d^2$ for each group of $m$ ties.
  3. It measures any monotonic relationship, not only linear.

Methods of least square

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>The method of least squares fits the line or curve that minimises the sum of the squared vertical deviations $\sum e_i^2$ between observed and fitted values.</mark>

Key points.

  1. For the line $y=a+bx$, minimising $\sum(y-a-bx)^2$ gives the normal equations $\sum y=na+b\sum x$ and $\sum xy=a\sum x+b\sum x^2$.
  2. Solving them gives $b=\frac{n\sum xy-\sum x\sum y}{n\sum x^2-(\sum x)^2}$ and $a=\bar y-b\bar x$.
  3. The fitted line always passes through the point $(\bar x,\bar y)$, and the residuals sum to zero.

Simple linear Regression model, SLR assumptions and prediction

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==Simple linear regression models one response as $y=\beta_0+\beta_1x+\varepsilon$, where $\varepsilon$ is a random error.==

Key points.

  1. The relationship between $x$ and the mean of $y$ is linear.
  2. Errors are independent, have mean zero, have constant variance (homoscedasticity) and are normally distributed.
  3. The slope $\beta_1$ is the change in $y$ for a unit change in $x$, estimated by $b=r\,\frac{s_y}{s_x}$.
  4. Prediction is $\hat y=a+bx$, reliable only for $x$ inside the observed range.

Multiple linear Regression, MLR assumption and prediction

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==Multiple linear regression models one response on several predictors as $y=\beta_0+\beta_1x_1+\dots+\beta_kx_k+\varepsilon$.==

Key points.

  1. In matrix form the least-squares estimate is $\hat\beta=(X^TX)^{-1}X^Ty$.
  2. It keeps the SLR assumptions of linearity, independent errors, constant variance and normal errors.
  3. It adds no perfect multicollinearity, meaning no predictor is an exact linear combination of the others.
  4. Each $\beta_j$ is the change in $y$ per unit of $x_j$ with the other predictors held fixed, and prediction substitutes the new $x$ values.

Polynomial Regression

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==Polynomial regression fits a curve $y=\beta_0+\beta_1x+\beta_2x^2+\dots+\beta_kx^k+\varepsilon$ to curved data.==

Key points.

  1. It is still linear in the coefficients, so ordinary least squares fits it by treating $x,x^2,\dots,x^k$ as separate predictors.
  2. A higher degree fits the sample better but overfits, so use the lowest degree that captures the curve.
  3. Powers of $x$ are highly correlated, so centring $x$ reduces multicollinearity.

Logistics Regression

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==Logistic regression models the probability of a binary outcome through the logit, $\ln\frac{p}{1-p}=\beta_0+\beta_1x_1+\dots+\beta_kx_k$.==

Key points.

  1. The predicted probability is $p=\frac{1}{1+e^{-(\beta_0+\beta_1x)}}$, always between 0 and 1.
  2. Coefficients are estimated by maximum likelihood, not least squares.
  3. $e^{\beta_j}$ is the odds ratio for a unit rise in $x_j$.
  4. It is used for classification, such as pass or fail.

Poisson Regression

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==Poisson regression models a count response $Y\sim\text{Poisson}(\mu)$ with the log link $\ln\mu=\beta_0+\beta_1x_1+\dots+\beta_kx_k$.==

Key points.

  1. It is a generalised linear model for counts such as accidents or calls per day.
  2. The log link keeps the predicted mean $\mu=e^{\beta_0+\beta_1x}$ positive.
  3. It assumes the mean equals the variance, and overdispersion breaks this.
  4. Coefficients are fitted by maximum likelihood, and $e^{\beta_j}$ is the multiplicative change in the mean.

Non-Linear Regression

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==Non-linear regression fits a model that is non-linear in its parameters, such as $y=ae^{bx}$, by minimising the sum of squared errors.==

Key points.

  1. Some models linearise by a transformation; for example $y=ae^{bx}$ becomes $\ln y=\ln a+bx$.
  2. Truly non-linear models have no closed-form solution and are fitted by iterative methods such as Gauss-Newton.
  3. Iteration needs good starting values and may reach a local minimum.

Analysis of Variance (One way & Two Way)

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>ANOVA tests whether the means of three or more groups are equal by comparing the variance between groups with the variance within groups using the F-ratio.</mark>

Formula. One way: $SST=SSB+SSW$ and $F=\frac{SSB/(k-1)}{SSW/(N-k)}$.

Key points.

  1. One way uses one factor, and two way uses two factors and can also test their interaction.
  2. $H_0$ says all group means are equal, and it is rejected when $F$ exceeds the table value.
  3. It assumes normal populations, equal variances and independent samples.

Analysis of Covariance

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>ANCOVA combines ANOVA with regression to compare group means after adjusting for a continuous covariate.</mark>

Key points.

  1. The covariate is a nuisance variable, such as pre-test score, that affects the response.
  2. It removes the covariate's effect from the error, which raises power and gives adjusted means.
  3. It also assumes the regression slope is the same in every group.

Multivariate Analysis of Variance

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>MANOVA extends ANOVA to test whether group mean vectors differ on several dependent variables at once.</mark>

Key points.

  1. It avoids the inflated Type I error of running a separate ANOVA on each variable.
  2. Common test statistics are Wilks' lambda, Pillai's trace and Roy's largest root.
  3. It assumes multivariate normality and equal covariance matrices across groups.

Last-minute revision

  • Pearson $r=\frac{\operatorname{Cov}(x,y)}{\sigma_x\sigma_y}$, always between $-1$ and $+1$.
  • Spearman $\rho=1-\frac{6\sum d^2}{n(n^2-1)}$.
  • Least squares minimises $\sum e^2$, and the line passes through $(\bar x,\bar y)$.
  • Slope $b=\frac{n\sum xy-\sum x\sum y}{n\sum x^2-(\sum x)^2}$ and intercept $a=\bar y-b\bar x$.
  • Also $b=r\,s_y/s_x$.
  • SLR assumptions are linearity, independence, constant variance and normal errors.
  • MLR estimate is $\hat\beta=(X^TX)^{-1}X^Ty$, and it needs no multicollinearity.
  • Polynomial regression is linear in its coefficients.
  • Logistic regression uses the logit link for binary outcomes, and Poisson regression uses the log link for counts.
  • One-way ANOVA $F=\frac{SSB/(k-1)}{SSW/(N-k)}$.
  • ANCOVA adjusts means for a covariate, and MANOVA handles several dependent variables.

Memory hooks

  • Pearson is for measured data, Spearman is for ranks.
  • Least squares means the smallest squared misses.
  • Logit for Yes/No, log for Count.
  • ANOVA compares Between with Within, and the ratio is F.
  • ANCOVA is ANOVA plus a Covariate, and MANOVA is ANOVA with Many outcomes.

Coverage checklist

  • Correlation, Scatter diagram: no past questions.
  • Karl Pearson's coefficient of correlation: no past questions.
  • Spearman's Rank correlation coefficient: no past questions.
  • Methods of least square: no past questions.
  • Simple linear Regression model, SLR assumptions and prediction: no past questions.
  • Multiple linear Regression, MLR assumption and prediction: no past questions.
  • Polynomial Regression: no past questions.
  • Logistics Regression: no past questions.
  • Poisson Regression: no past questions.
  • Non-Linear Regression: no past questions.
  • Analysis of Variance (One way & Two Way): no past questions.
  • Analysis of Covariance: no past questions.
  • Multivariate Analysis of Variance: no past questions.
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in