Skip to content
AD-302 · Probability and Statistics for Data Science/Quick Revision Short Notes

Probability and Statistics for Data Science (AD-302) - Unit 5 Short Notes

How unit 5 is examined

Hypothesis testing basics, errors, chi-square tests and the t, F and Z tests; marks come from the t/F/Z numericals (high) and the chi-square variance and goodness-of-fit numericals (medium).

Testing of hypothesis: Null and Alternative hypothesis

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Testing of hypothesis is a rule that uses sample data to decide whether a stated claim about a population parameter should be accepted or rejected.</mark>

Key points.

  1. The null hypothesis $H_0$ is the claim of no difference or no effect, for example $H_0:\mu=\mu_0$; it is assumed true until the data contradict it.
  2. The alternative hypothesis $H_1$ is what we accept if $H_0$ is rejected, for example $\mu\ne\mu_0$ (two-tailed), $\mu>\mu_0$ (right-tailed) or $\mu<\mu_0$ (left-tailed).
  3. The form of $H_1$ decides whether the critical region lies in both tails or in one tail.
  4. Steps: state $H_0$ and $H_1$, fix $\alpha$, compute the test statistic, compare it with the table value, and conclude.

Two types of errors

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A Type I error is rejecting a true $H_0$, and a Type II error is accepting a false $H_0$.</mark>

Key points.

  1. $P(\text{Type I})=\alpha$, the level of significance; $P(\text{Type II})=\beta$.
  2. For a fixed sample size, reducing $\alpha$ increases $\beta$, so both cannot be made small together except by increasing $n$.
  3. Decision table: reject a true $H_0$ is Type I; accept a false $H_0$ is Type II; the other two decisions are correct.

Level of significance and power of the test

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>The level of significance $\alpha$ is the maximum probability of rejecting a true $H_0$; the power of a test is $1-\beta$, the probability of correctly rejecting a false $H_0$.</mark>

Key points.

  1. Common levels are 5% ($\alpha=0.05$) and 1%; the critical region is the set of values of the statistic with total probability $\alpha$ under $H_0$.
  2. Two-tailed 5% critical Z is 1.96 and 1% is 2.58; one-tailed 5% is 1.645 and 1% is 2.33.
  3. If the calculated value falls in the critical region, $H_0$ is rejected; otherwise it is accepted (not rejected).
  4. Power rises with a larger sample size and a larger true difference; the p-value is the smallest $\alpha$ at which $H_0$ would be rejected.

Tests of significance: Chi-square distribution

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==If $Z_1,\dots,Z_n$ are independent standard normal variates, then $\chi^2=\sum Z_i^2$ follows the chi-square distribution with $n$ degrees of freedom.==

Key points.

  1. The chi-square distribution is positively skewed, takes only non-negative values, and its mean is $n$ and variance is $2n$.
  2. It tends to the normal distribution as $n$ becomes large, and the sum of independent chi-squares is chi-square with the summed degrees of freedom.
  3. It is used for the test of a population variance and for goodness of fit; the test is one-tailed to the right except for a two-sided variance test.
  4. The calculated $\chi^2$ is compared with the table value at the given $\alpha$ and degrees of freedom.

Test of population variance and goodness of fit

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>

Definition. ==The chi-square test of variance checks $H_0:\sigma^2=\sigma_0^2$ using $\chi^2=(n-1)s^2/\sigma_0^2$, and the goodness-of-fit test checks whether observed frequencies agree with expected frequencies using $\chi^2=\sum (O-E)^2/E$.==

Formula. $$\chi^2=\frac{(n-1)s^2}{\sigma_0^2}=\frac{\sum(x-\bar x)^2}{\sigma_0^2}\ (\text{df}=n-1),\qquad \chi^2=\sum\frac{(O_i-E_i)^2}{E_i}\ (\text{df}=k-1)$$

Key points.

  1. For the variance test $H_0:\sigma=\sigma_0$ against $H_1:\sigma\ne\sigma_0$ (two-tailed), or $\sigma^2>\sigma_0^2$ (right-tailed) when the claim is "no more than".
  2. Here $s^2=\sum(x-\bar x)^2/(n-1)$ is the unbiased sample variance.
  3. For large $n$ ($>30$) $s$ is approximately normal, so $Z=\dfrac{s-\sigma_0}{\sigma_0/\sqrt{2n}}$ can be used instead.
  4. In goodness of fit, $E_i$ comes from $H_0$ (equal frequency gives $E=N/k$), and each expected frequency should be at least 5.
  5. Degrees of freedom are $k-1$, less one more for each parameter estimated from the data.

Example (Jun 2023, Dec 2023). $H_0:\sigma=10$, $H_1:\sigma\ne10$, $n=50$, $s=15$. $\chi^2=\dfrac{49\times225}{100}=110.25$, df $=49$, and $Z=\dfrac{15-10}{10/\sqrt{100}}=5$. Since $|Z|=5>1.96$ (also $\chi^2=110.25>70.22$, the upper 5% two-tailed limit), reject $H_0$: the population SD is not 10.

Example (Dec 2023 variant). $H_0:\sigma^2\le0.16$, $H_1:\sigma^2>0.16$, $n=11$, $\sum x=27.6$, $\bar x=2.509$, $\sum(x-\bar x)^2=0.1891$. $\chi^2=0.1891/0.16=1.18$, df $=10$, table value at 1% is $23.21$. Since $1.18<23.21$, accept $H_0$: the variance is not more than 0.16.

Example (Jun 2023, digits). $H_0$: digits occur equally often, so $E=10000/10=1000$.

Digit 0 1 2 3 4 5 6 7 8 9
$O-E$ 26 107 -3 -34 75 -67 107 -28 -36 -147

$\chi^2=\dfrac{1}{1000}[26^2+107^2+3^2+34^2+75^2+67^2+107^2+28^2+36^2+147^2]=58.542$, df $=9$, table value $16.92$. $58.542>16.92$, so reject $H_0$: the digits do not occur equally frequently.

Answer frame. Open with the null hypothesis and the statistic used; for variance write hypotheses, $\chi^2$, df, table value, decision; for fit write $E=N/k$, the $(O-E)^2/E$ working, df $=k-1$, table value; close with the decision in words about the original claim.

Asked: [7 marks] (Jun 2023, Dec 2023) Test the hypothesis $\sigma=10$, given $s=15$ for a random sample of size 50 from a normal population. / Instrument variance no more than 0.16: write hypotheses and test at 1% with 11 measurements 2.5, 2.3, 2.4, 2.3, 2.5, 2.7, 2.5, 2.6, 2.6, 2.7, 2.5. Asked: [7 marks] (Jun 2023) Digits in a telephone directory (0 to 9: 1026, 1107, 997, 966, 1075, 933, 1107, 972, 964, 853; total 10000): test whether the digits occur equally frequently. Pitfall: use the unbiased $s^2$ with $n-1$ in the variance test, and do not forget df $=k-1$ (not $k$) in goodness of fit.

t, F, Z distribution and tests based on them

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>

Definition. ==The Z test uses the standard normal statistic for large samples or known $\sigma$, the t test uses Student's t with $n-1$ df for small samples with unknown $\sigma$, and the F test compares two variances using $F=s_1^2/s_2^2$.==

Formula. $$Z=\frac{\bar x-\mu}{\sigma/\sqrt n},\quad Z=\frac{\bar x_1-\bar x_2}{\sigma\sqrt{\frac1{n_1}+\frac1{n_2}}},\quad t=\frac{\bar x-\mu}{s/\sqrt n}\ (s^2=\tfrac{\sum(x-\bar x)^2}{n-1}),\quad F=\frac{s_1^2}{s_2^2}\ (s_1^2>s_2^2)$$

Key points.

  1. Use Z when $n>30$ or $\sigma$ is known, t when $n\le30$ and $\sigma$ is unknown (population normal), and F to test equality of two population variances.
  2. The t distribution is symmetric about 0 like the normal but has heavier tails, and it approaches the normal as df grows.
  3. For t, df $=n-1$; for two-sample F, df are $(n_1-1,n_2-1)$ with the larger variance on top.
  4. Critical Z at 5% is 1.96 (two-tailed) and 1.645 (one-tailed).
  5. A $(1-\alpha)$ confidence interval for the mean is $\bar x\pm z_{\alpha/2}\,\sigma/\sqrt n$; the 95% interval uses 1.96.
  6. Reject $H_0$ when the calculated value exceeds the table value in absolute terms, else accept.
  7. Assumptions: random samples, normal population, and independent samples for the two-sample tests.

Example 1 (t test, Jun 2023, Dec 2023, Dec 2024, Dec 2025). $H_0:\mu=64$, $H_1:\mu>64$ (one-tailed), $n=10$. $\sum x=654$, $\bar x=65.4$, $\sum(x-\bar x)^2=122.4$, $s^2=122.4/9=13.6$, $s=3.688$. $t=\dfrac{65.4-64}{3.688/\sqrt{10}}=\dfrac{1.4}{1.166}=1.20$, table $t_{0.05}(9)=1.833$. $1.20<1.833$, so accept $H_0$: the average height cannot be said to exceed 64 inches.

Example 2 (Z test, Dec 2024). $H_0:\mu_1=\mu_2$, $n_1=100,\ \bar x_1=210,\ n_2=150,\ \bar x_2=220,\ \sigma=11$. $Z=\dfrac{210-220}{11\sqrt{1/100+1/150}}=\dfrac{-10}{11\times0.1291}=-7.04$. $|Z|=7.04>1.96$, so reject $H_0$: the average incomes differ significantly.

Example 3 (F test, Dec 2025). $H_0:\sigma_1^2=\sigma_2^2$. Sample I: $\bar x=11.75$, $\sum(x-\bar x)^2=33.5$, $s_1^2=33.5/7=4.786$. Sample II: $\bar x=10.43$, $\sum(x-\bar x)^2=23.71$, $s_2^2=23.71/6=3.952$. $F=4.786/3.952=1.21$, table $F_{0.05}(7,6)=4.20$. $1.21<4.20$, so accept $H_0$: the variance estimates do not differ significantly.

Example 4 (confidence interval, Dec 2025). $\bar x=1250$, $\sigma=150$, $n=100$. Margin $=1.96\times150/10=29.4$. 95% interval: $1220.6$ to $1279.4$ units.

Other t-test variants asked.

  • I.Q. sample of 10 (Dec 2023): $\bar x=97.2$, $s=14.27$, $t=\dfrac{97.2-100}{14.27/\sqrt{10}}=-0.62$; $|t|<2.262$ (df 9, two-tailed), so accept $\mu=100$. Range of sample means: $97.2\pm2.262\times4.51$, i.e. $87.0$ to $107.4$.
  • 9 items 45, 47, 50, 52, 48, 47, 49, 53, 51: $\bar x=49.11$, $s=2.619$ (df 8), $t=\dfrac{1.611}{2.619/3}=1.85<2.31$, so the mean does not differ significantly from 47.5.
  • Mica washers ($\mu=10$, $\bar x=9.52$, $s=0.6$, $n=10$): with $s$ as the sample SD (divisor $n$), $t=\dfrac{9.52-10}{0.6/\sqrt{9}}=-2.4$; $|t|$ exceeds the 9-df table value 2.26, so the machine setting differs.

Answer frame. Open with "Given the sample, we test whether ..."; state $H_0$ and $H_1$ with tail; write the formula and choose t, Z or F with its reason; show $\bar x$, $s$ (with $n-1$) in a small working, then the statistic; compare with the table value at the given df and close with the decision in words about the original claim. For the CI question skip hypotheses and finish with the interval.

Asked: [7 marks] (Jun 2023, Dec 2023, Dec 2024, Dec 2025) Heights of 10 males 70, 67, 62, 68, 61, 68, 70, 64, 64, 60: is the average height greater than 64 inches at 5% ($P(t>1.833)=0.05$, 9 df)? / 10 boys' I.Q. 70, 120, 110, 101, 88, 83, 95, 98, 107, 100: support mean 100? Find a range for sample mean I.Q. / 9 items 45, 47, 50, 52, 48, 47, 49, 53, 51 vs assumed mean 47.5 ($t_{0.05}=2.31$) / Mica washers, 10 mils, sample of 10 with mean 9.52 and SD 0.6: find t. Asked: [7 marks] (Dec 2024) Average income 210 (SD 10, n=100) versus 220 (SD 12, n=150), city SD 11: is the difference significant? Asked: [7 marks] (Dec 2025) Samples of 8 (9, 11, 13, 11, 15, 9, 12, 14) and 7 (10, 12, 10, 14, 9, 8, 10) items: do the variance estimates differ significantly (F = 4.20 for 7, 6 df)? Asked: [7 marks] (Dec 2025) Mean monthly consumption of 100 families is 1250 units, SD 150: construct a 95% confidence interval for the actual mean. Pitfall: put the larger variance in the numerator of F, use $n-1$ in $s$, and quote a one-tailed table value when $H_1$ is one-sided.

Last-minute revision

  • $H_0$ is the no-difference claim; rejecting a true $H_0$ is Type I ($\alpha$), accepting a false $H_0$ is Type II ($\beta$); power is $1-\beta$.
  • Critical Z at 5%: 1.96 (two-tailed), 1.645 (one-tailed); at 1%: 2.58 and 2.33.
  • Small sample, $\sigma$ unknown: $t=(\bar x-\mu)/(s/\sqrt n)$ with df $=n-1$.
  • Two means with known $\sigma$: $Z=(\bar x_1-\bar x_2)/[\sigma\sqrt{1/n_1+1/n_2}]$.
  • $F=s_1^2/s_2^2$ (larger on top), df $(n_1-1,n_2-1)$.
  • Variance test: $\chi^2=(n-1)s^2/\sigma_0^2$; goodness of fit: $\sum(O-E)^2/E$, df $=k-1$.
  • Heights: $\bar x=65.4$, $s^2=13.6$, $t=1.20<1.833$, accept.
  • Incomes: $Z=-7.04$, reject. Variances: $F=1.21<4.20$, accept.
  • Consumption 95% CI: $1250\pm29.4=(1220.6,\,1279.4)$.
  • Digits: $\chi^2=58.542>16.92$ (df 9), reject. $\sigma=10$, $s=15$: $Z=5$, reject.

Memory hooks

  • "Small and unknown, take t; large or known, take Z; variances need F."
  • Type I = "false alarm" (reject a true $H_0$); Type II = "miss" (accept a false $H_0$).
  • Chi-square fit: "Observed minus Expected, square, over Expected, add up."
  • Df ladder: t uses $n-1$, F uses $(n_1-1,n_2-1)$, fit uses $k-1$.

Coverage checklist

  • Testing of hypothesis: Null and Alternative hypothesis: definition, $H_0$/$H_1$, tails, steps.
  • two types of errors: Type I and II, $\alpha$, $\beta$.
  • level of significance and power of the test: $\alpha$, critical values, power $1-\beta$, p-value.
  • Tests of significance: Chi-square distribution: definition, properties, uses.
  • test of popular variance and test of goodness of fit: sigma = 10 with s = 15 (Jun 2023, Dec 2023), variance 0.16 test, telephone digits (Jun 2023).
  • t, F ,Z distribution and tests based on them: heights, I.Q., 9-item, mica t-tests (Jun 2023, Dec 2023, Dec 2024, Dec 2025), income Z test (Dec 2024), F test (Dec 2025), 95% CI (Dec 2025).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in