How unit 4 is examined
Covers how two variables move together (correlation) and how one is predicted from the other (regression). Marks sit in Karl Pearson's $r$, Spearman's rank correlation with ties, and the lines of regression; every one is a 7-mark numerical or short explain.
Correlation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Low weight</span>
Definition. <mark>Correlation is a statistical measure of the direction and degree of the linear relationship between two variables.</mark>
Key points.
- Correlation is positive when both variables rise and fall together, negative when one rises as the other falls, and zero when they are unrelated.
- The coefficient $r$ measures the degree of association and always lies between $-1$ and $+1$.
- Regression uses the relationship to predict the dependent variable from the independent one, so the two are "two sides of the same coin": both study the same relationship between variables.
- The link is $r=\pm\sqrt{b_{yx}\,b_{xy}}$, and $r$ carries the common sign of both regression coefficients.
- They differ in purpose: $r$ is symmetric in $X$ and $Y$ and has no unit, while regression is asymmetric (Y on X differs from X on Y) and gives an equation with units.
- Correlation shows association, not causation.
Answer frame. Open with the definition of both; draw a two-column table (correlation vs regression: purpose, symmetry, unit, output) and the scatter diagram with both regression lines crossing at $(\bar x,\bar y)$ from Lines of Regression; state $r=\pm\sqrt{b_{yx}b_{xy}}$; close that each complements the other, since $r$ tells how strong the relation is and regression uses it to predict.
Asked: [7 marks] (Nov 2022) "Correlation and Regression are two sides of the same coin". Explain.
Coefficient of Correlation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Karl Pearson's coefficient of correlation $r$ is the ratio of the covariance of $X$ and $Y$ to the product of their standard deviations, and measures the strength and direction of their linear relationship.</mark>
Formula. $$r=\frac{\sum (X-\bar X)(Y-\bar Y)}{\sqrt{\sum (X-\bar X)^2\,\sum (Y-\bar Y)^2}}=\frac{\operatorname{Cov}(X,Y)}{\sigma_X\sigma_Y}$$
Key points.
- $r$ lies between $-1$ and $+1$; $+1$ is perfect positive, $-1$ perfect negative and $0$ means no linear relation.
- $r$ is independent of change of origin and scale (for positive scale), so deviations from an assumed mean may be used.
- $r$ is the geometric mean of the two regression coefficients: $r^2=b_{yx}b_{xy}$.
- Independent variables have $r=0$, but $r=0$ does not imply independence, since a non-linear relation can exist.
- $|r|>0.7$ is usually called strong, $0.3$ to $0.7$ moderate and below $0.3$ weak.
Steps.
Step 1: Find the means of X and Y.
Step 2: Form deviations u = X - mean(X), v = Y - mean(Y).
Step 3: Find sum u^2, sum v^2 and sum uv.
Step 4: r = sum uv / sqrt(sum u^2 * sum v^2); comment on sign and size.
Example. Price $X$ and quantity $Y$: $X=10,10,11,12,12$; $Y=5,6,4,3,3$.
| $X$ | $Y$ | $u=X-11$ | $v=Y-4.2$ | $u^2$ | $v^2$ | $uv$ |
|---|---|---|---|---|---|---|
| 10 | 5 | -1 | 0.8 | 1 | 0.64 | -0.8 |
| 10 | 6 | -1 | 1.8 | 1 | 3.24 | -1.8 |
| 11 | 4 | 0 | -0.2 | 0 | 0.04 | 0 |
| 12 | 3 | 1 | -1.2 | 1 | 1.44 | -1.2 |
| 12 | 3 | 1 | -1.2 | 1 | 1.44 | -1.2 |
$\bar X=55/5=11$, $\bar Y=21/5=4.2$, $\sum u^2=4$, $\sum v^2=6.8$, $\sum uv=-5$. $$r=\frac{-5}{\sqrt{4\times 6.8}}=\frac{-5}{5.215}=-0.96$$ r = -0.96: the sign is negative, so quantity falls as price rises, and the magnitude is close to 1, so the inverse relation is very strong.
Second numerical (father and son heights). $\bar X=68$, $\bar Y=69$; with $u=X-68$, $v=Y-69$: $\sum u^2=36$, $\sum v^2=44$, $\sum uv=24$, so $r=24/\sqrt{36\times44}=24/39.8=$ 0.60, a moderate positive correlation.
Third numerical. $x=11,10,9,8,7,6,5$; $y=20,18,12,8,10,5,4$: $\bar x=8$, $\bar y=11$, $\sum u^2=28$, $\sum v^2=226$, $\sum uv=76$, so $r=76/\sqrt{28\times226}=$ 0.96, strong positive.
Answer frame. Open with the formula and definition; draw the deviation table with all seven columns; show the sums, substitute, and box $r$; close with a comment on sign (direction) and magnitude (strength).
Pitfall: Take deviations from the true mean before squaring; using raw values or forgetting the sign of $\sum uv$ flips or inflates $r$.
Asked: [7 marks] (Nov 2022, Jun 2023, Dec 2024, Dec 2025) Price and quantity of a commodity for 5 months: prices 10, 10, 11, 12, 12; quantity 5, 6, 4, 3, 3. Find Karl Pearson's coefficient of correlation and comment on its sign and magnitude. (Same slot also asked with father-son heights X = 65, 66, 67, 67, 68, 69, 70, 72 and Y = 67, 68, 65, 68, 72, 72, 69, 71, and with x = 11 to 5, y = 20, 18, 12, 8, 10, 5, 4.)
Rank Correlation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Spearman's rank correlation coefficient $\rho$ measures the degree of agreement between two rankings of the same $n$ individuals.</mark>
Formula. $$\rho=1-\frac{6\sum d^2}{n(n^2-1)}$$ With tied ranks (tied items get the average of their ranks), add a correction to $\sum d^2$: $$\rho=1-\frac{6\left[\sum d^2+\sum \frac{m(m^2-1)}{12}\right]}{n(n^2-1)}$$ where $m$ is the number of items in each tie group, one term per tie group in $X$ and in $Y$.
Key points.
- $d=R_X-R_Y$ is the difference of the two ranks of each individual, and $\sum d=0$ is a check on the ranking.
- $\rho$ lies between $-1$ (ranks fully opposite) and $+1$ (ranks identical).
- It is used when data are qualitative or only ranked (judges, beauty, honesty), or when actual values are given and ranks are easier.
- Rank the values from largest as 1 (or smallest as 1) but do the same for both variables.
- Tied values take the mean of the ranks they would occupy, for example two items tied at ranks 2 and 3 each get 2.5.
- If the data are given as ranks by several judges, find $\rho$ for each pair; the pair with the largest positive $\rho$ has the nearest common taste.
Example (tied ranks). $X=68,64,75,50,64,80,75,40,55,64$; $Y=62,58,68,45,81,60,68,48,50,70$.
| $R_X$ | 4 | 6 | 2.5 | 9 | 6 | 1 | 2.5 | 10 | 8 | 6 |
|---|---|---|---|---|---|---|---|---|---|---|
| $R_Y$ | 5 | 7 | 3.5 | 10 | 1 | 6 | 3.5 | 9 | 8 | 2 |
| $d$ | -1 | -1 | -1 | -1 | 5 | -5 | -1 | 1 | 0 | 4 |
| $d^2$ | 1 | 1 | 1 | 1 | 25 | 25 | 1 | 1 | 0 | 16 |
$\sum d^2=72$. Ties: $X$ has 75 twice and 64 three times; $Y$ has 68 twice. Correction $=\frac{2\cdot3}{12}+\frac{3\cdot8}{12}+\frac{2\cdot3}{12}=0.5+2+0.5=3$. $$\rho=1-\frac{6(72+3)}{10\times99}=1-\frac{450}{990}=\textbf{0.545}$$
Example (judges). $n=10$, $n(n^2-1)=990$.
| Competitor | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| A | 1 | 6 | 5 | 10 | 3 | 2 | 4 | 9 | 7 | 8 |
| B | 3 | 5 | 8 | 4 | 7 | 10 | 2 | 1 | 6 | 9 |
| C | 6 | 4 | 9 | 8 | 1 | 2 | 3 | 10 | 5 | 7 |
| $d_{AB}$ | -2 | 1 | -3 | 6 | -4 | -8 | 2 | 8 | 1 | -1 |
| $d_{BC}$ | -3 | 1 | -1 | -4 | 6 | 8 | -1 | -9 | 1 | 2 |
| $d_{AC}$ | -5 | 2 | -4 | 2 | 2 | 0 | 1 | -1 | 2 | 1 |
Squaring: $\sum d_{AB}^2=4+1+9+36+16+64+4+64+1+1=200$, $\sum d_{BC}^2=9+1+1+16+36+64+1+81+1+4=214$, $\sum d_{AC}^2=25+4+16+4+4+0+1+1+4+1=60$.
| Pair | $\sum d^2$ | $\rho=1-6\sum d^2/990$ |
|---|---|---|
| A, B | 200 | -0.21 |
| B, C | 214 | -0.30 |
| A, C | 60 | 0.64 |
Judges A and C have the nearest approach to common likings, as $\rho(A,C)=0.64$ is the highest and positive.
Example (no ties). Chemistry and Physics marks of 10 students, ranked from the highest: $\sum d^2=26$, so $\rho=1-\frac{6\times26}{990}=$ 0.84, a high positive agreement.
Answer frame. Open with the definition and formula; write the rank table with $R_X,R_Y,d,d^2$ and mention how ties are averaged; add the tie correction only if values repeat; box $\rho$; close with the meaning of its sign and size.
Asked: [7 marks] (Jun 2023, Dec 2024, Dec 2025) Obtain the rank correlation coefficient for X = 68, 64, 75, 50, 64, 80, 75, 40, 55, 64 and Y = 62, 58, 68, 45, 81, 60, 68, 48, 50, 70. (Also asked as ten students' marks in Chemistry 78, 36, 98, 25, 75, 82, 90, 62, 65, 39 and Physics 84, 51, 91, 60, 68, 62, 86, 58, 63, 47.) Asked: [7 marks] (Dec 2023) Ten competitors were ranked by three judges A, B and C. Using the rank correlation method, discuss which pair of judges has the nearest approach to common likings in music.
Lines of Regression
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Regression is the statistical method of estimating the average value of a dependent variable from the known value of an independent variable; the line of regression is the best-fitting straight line obtained by the method of least squares.</mark>
Formula. Regression of $Y$ on $X$ predicts $y$ from $x$; regression of $X$ on $Y$ predicts $x$ from $y$. $$y-\bar y=b_{yx}(x-\bar x),\quad b_{yx}=\frac{\operatorname{Cov}(X,Y)}{\sigma_x^2}=r\frac{\sigma_y}{\sigma_x}$$ $$x-\bar x=b_{xy}(y-\bar y),\quad b_{xy}=\frac{\operatorname{Cov}(X,Y)}{\sigma_y^2}=r\frac{\sigma_x}{\sigma_y}$$ Derivation: minimise $\sum(y-a-bx)^2$ by setting its derivatives in $a$ and $b$ to zero; this gives $\sum y=na+b\sum x$, hence the line passes through $(\bar x,\bar y)$, and $b=\operatorname{Cov}(X,Y)/\sigma_x^2$.
Key points.
- Both regression lines pass through the point of means $(\bar x,\bar y)$, which is where they intersect.
- $b_{yx}$ and $b_{xy}$ always have the same sign as $r$ (and as $\operatorname{Cov}$), so they cannot have opposite signs.
- $b_{yx}\,b_{xy}=r^2$, so $r=\pm\sqrt{b_{yx}b_{xy}}$ and the product of the two coefficients cannot exceed 1.
- The arithmetic mean of $b_{yx}$ and $b_{xy}$ is at least $r$ (in magnitude).
- Regression coefficients are independent of change of origin but not of scale.
- If $r=\pm1$ the two lines coincide, and if $r=0$ they are perpendicular to each other (parallel to the axes).
- The dependent variable is the one predicted (on the left), the independent variable is the one used to predict (on the right), so the two lines are not interchangeable.
Diagram. Scatter diagram with the two lines of regression crossing at the point of means $M=(\bar x,\bar y)$. <figure class="ds-fig" style="margin:1.4rem 0;overflow-x:auto"><svg xmlns="http://www.w3.org/2000/svg" id="dsfig-u4-01" viewBox="0 0 596 424" width="596" height="424" role="img" aria-label="Scatter of points with line Y on X (A-B) and line X on Y (C-D) meeting at the means M"><style>#dsfig-u4-01 .e{stroke:#454C5A;stroke-width:1.4;fill:none}#dsfig-u4-01 .e.hi{stroke:#2340B8;stroke-width:2.6}#dsfig-u4-01 .n{fill:#FFFFFF;stroke:#16181D;stroke-width:1.4}#dsfig-u4-01 .n.hi{fill:#E3E9FC;stroke:#2340B8;stroke-width:2.2}#dsfig-u4-01 .n.rb-b{fill:#16181D;stroke:#16181D}#dsfig-u4-01 .n.rb-r{fill:#BD3227;stroke:#BD3227}#dsfig-u4-01 text{font-family:"JetBrains Mono",ui-monospace,Menlo,Consolas,monospace;font-size:13px}#dsfig-u4-01 .t{fill:#16181D;font-weight:500}#dsfig-u4-01 .t.inv{fill:#FFFFFF;font-weight:700}#dsfig-u4-01 .kd{stroke:#16181D;stroke-width:1.2}#dsfig-u4-01 .dot{fill:#16181D}#dsfig-u4-01 .ann{fill:#2340B8;font-size:11px;font-weight:700}#dsfig-u4-01 .lbl{fill:#6F7787;font-family:system-ui,-apple-system,sans-serif;font-size:12px;font-weight:700}#dsfig-u4-01 .ptr{fill:#2340B8;font-size:12px;font-weight:700}#dsfig-u4-01 .ah{fill:#454C5A}#dsfig-u4-01 .ah.hi{fill:#2340B8}#dsfig-u4-01 .wl rect{fill:#FFFFFF;stroke:#DCE0E7}#dsfig-u4-01 .wl .t{font-size:12px;font-weight:700}#dsfig-u4-01 .wl.hi rect{fill:#2340B8;stroke:#2340B8}#dsfig-u4-01 .wl.hi .t{fill:#FFFFFF}html.dark #dsfig-u4-01 .e{stroke:#B1B7C3}html.dark #dsfig-u4-01 .e.hi{stroke:#8FA3FF}html.dark #dsfig-u4-01 .n{fill:#161920;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.hi{fill:#1E2748;stroke:#8FA3FF}html.dark #dsfig-u4-01 .n.rb-b{fill:#E6E8ED;stroke:#E6E8ED}html.dark #dsfig-u4-01 .n.rb-r{fill:#FF7E71;stroke:#FF7E71}html.dark #dsfig-u4-01 .t{fill:#E6E8ED}html.dark #dsfig-u4-01 .t.inv{fill:#0F1115}html.dark #dsfig-u4-01 .kd{stroke:#E6E8ED}html.dark #dsfig-u4-01 .dot{fill:#E6E8ED}html.dark #dsfig-u4-01 .ann{fill:#8FA3FF}html.dark #dsfig-u4-01 .lbl{fill:#858D9C}html.dark #dsfig-u4-01 .ptr{fill:#8FA3FF}html.dark #dsfig-u4-01 .ah{fill:#B1B7C3}html.dark #dsfig-u4-01 .ah.hi{fill:#8FA3FF}html.dark #dsfig-u4-01 .wl rect{fill:#161920;stroke:#2A2E37}html.dark #dsfig-u4-01 .wl.hi rect{fill:#8FA3FF;stroke:#8FA3FF}html.dark #dsfig-u4-01 .wl.hi .t{fill:#0F1115}</style><defs><marker id="ah1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah" d="M0,1 L9,5 L0,9 z"/></marker><marker id="ahh1" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="7" markerHeight="7" orient="auto-start-reverse"><path class="ah hi" d="M0,1 L9,5 L0,9 z"/></marker></defs><path class="e" d="M59,384 L535,384" marker-end="url(#ah1)"/><path class="e" d="M40,365 L40,61" marker-end="url(#ah1)"/><path class="e" d="M99.3,331.2 L281.7,221.8"/><path class="e" d="M314.3,202.2 L496.7,92.8"/><path class="e" d="M180.4,368.8 L286.6,227.2"/><path class="e" d="M309.4,196.8 L415.6,55.2"/><g class="wl"><rect x="170.1" y="267.5" width="40.8" height="18" rx="9"/><text class="t" x="190.5" y="276.5" dy=".35em" text-anchor="middle">YonX</text></g><g class="wl"><rect x="213.1" y="289" width="40.8" height="18" rx="9"/><text class="t" x="233.5" y="298" dy=".35em" text-anchor="middle">XonY</text></g><circle class="n" cx="40" cy="384" r="18"/><text class="t" x="40" y="384" dy=".35em" text-anchor="middle">O</text><circle class="n" cx="556" cy="384" r="18"/><text class="t" x="556" y="384" dy=".35em" text-anchor="middle">X</text><circle class="n" cx="40" cy="40" r="18"/><text class="t" x="40" y="40" dy=".35em" text-anchor="middle">Y</text><circle class="n" cx="83" cy="341" r="18"/><text class="t" x="83" y="341" dy=".35em" text-anchor="middle">A</text><circle class="n" cx="513" cy="83" r="18"/><text class="t" x="513" y="83" dy=".35em" text-anchor="middle">B</text><circle class="n" cx="169" cy="384" r="18"/><text class="t" x="169" y="384" dy=".35em" text-anchor="middle">C</text><circle class="n" cx="427" cy="40" r="18"/><text class="t" x="427" y="40" dy=".35em" text-anchor="middle">D</text><circle class="n" cx="298" cy="212" r="18"/><text class="t" x="298" y="212" dy=".35em" text-anchor="middle">M</text><circle class="n" cx="126" cy="289.4" r="18"/><text class="t" x="126" y="289.4" dy=".35em" text-anchor="middle">p1</text><circle class="n" cx="212" cy="289.4" r="18"/><text class="t" x="212" y="289.4" dy=".35em" text-anchor="middle">p2</text><circle class="n" cx="263.6" cy="237.8" r="18"/><text class="t" x="263.6" y="237.8" dy=".35em" text-anchor="middle">p3</text><circle class="n" cx="349.6" cy="151.8" r="18"/><text class="t" x="349.6" y="151.8" dy=".35em" text-anchor="middle">p4</text><circle class="n" cx="418.4" cy="160.4" r="18"/><text class="t" x="418.4" y="160.4" dy=".35em" text-anchor="middle">p5</text><circle class="n" cx="470" cy="91.6" r="18"/><text class="t" x="470" y="91.6" dy=".35em" text-anchor="middle">p6</text></svg><figcaption style="font-size:.82em;opacity:.72;margin-top:.45rem">Scatter of points with line Y on X (A-B) and line X on Y (C-D) meeting at the means M</figcaption></figure>
Role in business and economics.
- Demand estimation: regressing quantity demanded on price or income gives a demand function, and its slope is the price (or income) effect on demand.
- Sales forecasting: regressing sales on advertising spend or on time gives a line that projects future sales for a planned budget or period.
- Pricing: regressing revenue or quantity on price estimates how a price change will move revenue, which guides the choice of price.
- Cost control: regressing total cost on output separates fixed cost (intercept) from variable cost per unit (slope) for budgeting.
- Econometrics: the consumption function $C=a+bY$ is a regression of consumption on income, where the slope $b$ is the marginal propensity to consume (MPC).
- Policy impact: regression estimates the effect of a tax, subsidy or interest-rate change on a dependent variable such as demand, investment or output.
Business example. Advertising $x$ (Rs. lakh) $=1,2,3,4,5$; sales $y$ (Rs. lakh) $=3,5,6,8,8$. $\bar x=3$, $\bar y=6$, $\sum uv=13$, $\sum u^2=10$, so $b_{yx}=1.3$ and $y=6+1.3(x-3)=1.3x+2.1$. Each extra Rs. 1 lakh of advertising raises sales by Rs. 1.3 lakh on average; a Rs. 6 lakh budget forecasts sales of Rs. 9.9 lakh.
Example 1 (given summary). $\bar x=65$, $\bar y=67$, $\sigma_x=2.5$, $\sigma_y=3.5$, $r=0.8$; find $y$ at $x=70$. $$b_{yx}=0.8\times\frac{3.5}{2.5}=1.12,\quad y-67=1.12(x-65)$$ At $x=70$: $y=67+1.12\times5=$ Rs. 72.6.
Example 2 (raw data). $x=2,4,6,8,10$; $y=5,7,9,8,11$. $\bar x=6$, $\bar y=8$; $u=-4,-2,0,2,4$; $v=-3,-1,1,0,3$; $\sum u^2=40$, $\sum v^2=20$, $\sum uv=26$. $b_{yx}=26/40=0.65$, $b_{xy}=26/20=1.3$. y on x: $y-8=0.65(x-6)$, i.e. $y=0.65x+4.1$. x on y: $x-6=1.3(y-8)$, i.e. $x=1.3y-4.4$. Slope: $y$ rises 0.65 unit per unit of $x$; $r=\sqrt{0.65\times1.3}=0.92$.
Example 3 (validity check). $Y=5+2.8X$ gives $b_{yx}=2.8$; $X=3-0.5Y$ gives $b_{xy}=-0.5$. The signs differ, which is impossible because both must share the sign of $r$. Also $r^2=2.8\times(-0.5)=-1.4<0$, but $0\le r^2\le1$. Hence they cannot be the regression equations.
Example 4 (joint pdf). $f(x,y)=\frac13(x+y)$, $0\le x\le1$, $0\le y\le2$. Marginals: $f_X(x)=\int_0^2 f\,dy=\frac{2(x+1)}{3}$, $f_Y(y)=\int_0^1 f\,dx=\frac{1+2y}{6}$.
| Quantity | Working | Value |
|---|---|---|
| $E(X)$ | $\frac23\int_0^1(x^2+x)dx=\frac23\left(\frac13+\frac12\right)$ | $5/9$ |
| $E(Y)$ | $\frac16\int_0^2(y+2y^2)dy=\frac16\left(2+\frac{16}3\right)$ | $11/9$ |
| $E(X^2)$ | $\frac23\int_0^1(x^3+x^2)dx=\frac23\left(\frac14+\frac13\right)$ | $7/18$ |
| $\operatorname{Var}X$ | $7/18-25/81$ | $13/162$ |
| $E(Y^2)$ | $\frac16\int_0^2(y^2+2y^3)dy=\frac16\left(\frac83+8\right)$ | $16/9$ |
| $\operatorname{Var}Y$ | $16/9-121/81$ | $23/81$ |
| $E(XY)$ | $\frac13\int_0^1\!\int_0^2(x^2y+xy^2)dy\,dx=\frac13\int_0^1\left(2x^2+\frac{8x}3\right)dx=\frac13\left(\frac23+\frac43\right)$ | $2/3$ |
| $\operatorname{Cov}$ | $2/3-55/81$ | $-1/81$ |
(i) $r=\dfrac{-1/81}{\sqrt{(13/162)(23/81)}}=-\sqrt{2/299}=$ -0.082. (ii) $b_{yx}=\frac{-1/81}{13/162}=-\frac2{13}$, $b_{xy}=\frac{-1/81}{23/81}=-\frac1{23}$. y on x: $y-\frac{11}{9}=-\frac{2}{13}\left(x-\frac59\right)$; x on y: $x-\frac59=-\frac1{23}\left(y-\frac{11}9\right)$. (iii) Regression curves of the means (conditional expectations): $E(Y|X=x)=\dfrac{3x+4}{3(x+1)}$ and $E(X|Y=y)=\dfrac{2+3y}{3(1+2y)}$; they are curves, not lines, because the pdf is not bivariate normal.
Answer frame. Definition question: open with the definition, draw the scatter diagram with both lines meeting at $(\bar x,\bar y)$, distinguish dependent and independent variables, then develop role points 1-6, give the advertising-sales example with its slope, and close that it turns a relationship into a usable prediction. Numerical: write the formula for the required line, substitute the means, $\sigma$ or sums, simplify to the equation, then read off the asked value; close with the check $r^2=b_{yx}b_{xy}$. Validity question: state both coefficients, test sign, then test $b_{yx}b_{xy}\le1$, and conclude.
Pitfall: Do not interchange the lines: to predict $y$ use $y$ on $x$ (with $r\sigma_y/\sigma_x$), never invert the line of $x$ on $y$.
Asked: [7 marks] (Nov 2022) What do you understand by regression? What role does it play in business and economic analysis? Asked: [7 marks] (Dec 2023) Variables X and Y have joint p.d.f. $f(x,y)=\frac13(x+y)$, $0\le x\le1$, $0\le y\le2$. Find (i) $r(X,Y)$ (ii) the two lines of regression (iii) the two regression curves for the means. Asked: [7 marks] (Dec 2023) Find the most likely price in Bombay corresponding to the price of Rs. 70 at Calcutta: average price 65 and 67, standard deviation 2.5 and 3.5, $r=0.8$. Asked: [7 marks] (Jun 2023) Can $Y=5+2.8X$ and $X=3-0.5Y$ be the estimated regression equations of Y on X and X on Y? Explain with suitable theoretical arguments. Asked: [7 marks] (Dec 2024) From the data x = 2, 4, 6, 8, 10 and y = 5, 7, 9, 8, 11 determine the lines of regression.
Multiple and Partial Correlation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Multiple correlation $R_{1.23}$ measures the relationship of one variable $X_1$ with two or more other variables $X_2,X_3$ taken together; partial correlation $r_{12.3}$ measures the relationship of $X_1$ and $X_2$ after removing the linear effect of $X_3$.
Key points.
- $R_{1.23}=\sqrt{\dfrac{r_{12}^2+r_{13}^2-2r_{12}r_{13}r_{23}}{1-r_{23}^2}}$ and it always lies between $0$ and $1$.
- $r_{12.3}=\dfrac{r_{12}-r_{13}r_{23}}{\sqrt{(1-r_{13}^2)(1-r_{23}^2)}}$ and it lies between $-1$ and $+1$.
- The multiple correlation coefficient is never smaller than any simple $|r_{1j}|$, and $R_{1.23}^2$ is the proportion of variation in $X_1$ explained by $X_2$ and $X_3$.
- Partial correlation is used to test whether a link between two variables is real or caused by a third.
Last-minute revision
- $r=\dfrac{\sum uv}{\sqrt{\sum u^2\sum v^2}}$ with $u=X-\bar X$, $v=Y-\bar Y$; $-1\le r\le1$.
- Price-quantity data: $\bar X=11$, $\bar Y=4.2$, $\sum uv=-5$, $r=-0.96$.
- Father-son heights: $\sum u^2=36$, $\sum v^2=44$, $\sum uv=24$, $r=0.60$.
- $\rho=1-6\sum d^2/[n(n^2-1)]$; ties add $\sum m(m^2-1)/12$ to $\sum d^2$; tied items share the mean rank.
- Tied-rank question: $\sum d^2=72$, correction 3, $n=10$, $\rho=0.545$.
- Judges A, B, C: $\rho_{AB}=-0.21$, $\rho_{BC}=-0.30$, $\rho_{AC}=0.64$, so A and C agree most.
- $b_{yx}=r\sigma_y/\sigma_x$, $b_{xy}=r\sigma_x/\sigma_y$, $b_{yx}b_{xy}=r^2$; both share the sign of $r$.
- Regression lines meet at $(\bar x,\bar y)$; identical if $r=\pm1$, perpendicular if $r=0$.
- Calcutta-Bombay: $y=67+1.12(x-65)$, so $y(70)=72.6$.
- $x=2,4,6,8,10$; $y=5,7,9,8,11$: $y=0.65x+4.1$, $x=1.3y-4.4$.
- $2.8$ and $-0.5$ cannot be regression coefficients: opposite signs and product $-1.4$.
- Joint pdf $\frac13(x+y)$: $r=-0.082$, $b_{yx}=-2/13$, $b_{xy}=-1/23$.
Memory hooks
- Pearson = Product of deviations over root of sums of squares; Spearman = Sixty ("6") over $n(n^2-1)$.
- "Regress on the right": the variable after "on" is the known one used for prediction.
- Same sign, small product: $b_{yx}$ and $b_{xy}$ agree in sign and multiply to $r^2\le1$.
- Ties: average the ranks, then add $m(m^2-1)/12$ for each tie group.
- Both lines cross at the means; $r=0$ gives a right angle, $r=\pm1$ one line.
Coverage checklist
- Correlation: "Correlation and Regression are two sides of the same coin" (Nov 2022).
- Coefficient of Correlation: Karl Pearson's $r$ for price-quantity, father-son heights and the $x,y$ table (Nov 2022, Jun 2023, Dec 2024, Dec 2025).
- Rank Correlation: tied-rank and Chemistry-Physics data (Jun 2023, Dec 2024, Dec 2025); three judges (Dec 2023).
- Lines of Regression: meaning and role (Nov 2022); joint pdf (Dec 2023); Calcutta-Bombay price (Dec 2023); validity of given equations (Jun 2023); lines from data (Dec 2024).
- Multiple and Partial Correlation: no recent questions; definition and formulas covered.