How unit 1 is examined
This unit covers probability distributions (the normal-distribution and discrete-distribution numericals carry the marks), inferential statistics, hypothesis tests, and regression with ANOVA (asked once as a short note).
Probability Distributions
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Medium weight</span>
Definition. <mark>A probability distribution lists every value a random variable can take together with its probability, and the probabilities add up to 1.</mark>
Key points.
- A discrete random variable has a probability function $P(x)$ with $P(x)\ge 0$ and $\sum P(x)=1$; this condition is used to find any unknown constant such as $k$.
- The mean (expected value) of a discrete variable is $E[X]=\sum xP(x)$.
- The variance is $Var(X)=E[X^2]-(E[X])^2$, where $E[X^2]=\sum x^2P(x)$; the standard deviation is its square root.
- The normal distribution is continuous, bell-shaped and symmetric about its mean $\mu$, so the area on each side of the mean is 0.5 and the total area is 1.
- Any normal variable is standardised with $Z=(X-\mu)/\sigma$, so the standard normal table (mean 0, SD 1) can be used.
- The number of items in a range is $N\times$ (area for that range), and a percentage is the area $\times 100$.
Formula. $Z=\dfrac{X-\mu}{\sigma}$, $E[X]=\sum xP(x)$, $Var(X)=E[X^2]-(E[X])^2$.
Example 1 (normal, 800 students, $\mu=28.8$, $\sigma=2.06$).
| Limit | Z | Table area from mean |
|---|---|---|
| 28.4 | $-0.19$ | 0.0753 |
| 30.4 | $0.78$ | 0.2823 |
| 31.3 | $1.21$ | 0.3869 |
i) $P(28.4<X<30.4)=0.0753+0.2823=0.3576$, so students $=800\times0.3576\approx$ 286. ii) $P(X>31.3)=0.5-0.3869=0.1131$, so students $=800\times0.1131\approx$ 90.
Example 2 (500 businesses, $\mu=36000$, $\sigma=10000$). For 40,000, $Z=0.4$, and for 30,000, $Z=-0.6$. i) $P(X>40000)=0.5-0.1554=0.3446$, so businesses $=500\times0.3446\approx$ 172. ii) $P(30000<X<40000)=0.2257+0.1554=0.3811$, that is about 38.1%.
Example 3 (find $k$, mean, variance). Values $x=0..7$ have $P(x)=0,k,2k,2k,3k,k^2,2k^2,7k^2+k$. Sum $=10k^2+9k=1$, so $10k^2+9k-1=0$, giving $(10k-1)(k+1)=0$; reject $k=-1$ because a probability cannot be negative, so $k=0.1$. Then $P(x)=0,0.1,0.2,0.2,0.3,0.01,0.02,0.17$.
| Quantity | Working | Value |
|---|---|---|
| $E[X]$ | $0.1+0.4+0.6+1.2+0.05+0.12+1.19$ | 3.66 |
| $E[X^2]$ | $0.1+0.8+1.8+4.8+0.25+0.72+8.33$ | 16.8 |
| $Var(X)$ | $16.8-3.66^2$ | 3.4044 |
Answer: $k=0.1$, mean $=3.66$, variance $=3.4044$.
Answer frame. Open by stating the given mean, SD, N and the normal-distribution assumption; convert each limit to Z, read the table area and multiply by N; close with the boxed count or percentage. For the $k$ problem, open with $\sum P(x)=1$, solve the quadratic and reject the negative root, then compute $E[X]$, $E[X^2]$ and the variance in a table.
Pitfall: Forgetting to reject the negative root of $k$, or reading the table area from the left instead of from the mean, loses marks.
Asked: [7 marks] (Nov 2023) Weights of 800 male students are normally distributed with mean 28.8 kg and SD 2.06 kg; find the number of students weighing i) between 28.4 kg and 30.4 kg, ii) more than 31.3 kg. Also, 500 businesses have average sales Rs. 36,000 with SD Rs. 10,000 (normal); find i) the number with sales above Rs. 40,000, ii) the percentage with sales between Rs. 30,000 and Rs. 40,000. Asked: [7 marks] (Nov 2023) A random variable has probability function P(x) = 0, K, 2K, 2K, 3K, K^2, 2K^2, 7K^2+K for x = 0 to 7; determine i) k, ii) mean, iii) variance.
Inferential Statistics
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Inferential statistics draws conclusions about a population from a sample, using probability to state how reliable the conclusion is.
Key points.
- Descriptive statistics only summarises the data in hand, whereas inferential statistics generalises from a sample to the whole population.
- Estimation gives a point estimate (a single value such as the sample mean $\bar x$) or an interval estimate (a confidence interval).
- A 95% confidence interval for the mean is $\bar x\pm z\,\sigma/\sqrt n$, with $z=1.96$.
- Hypothesis testing is the second main tool of inference.
Inferential Statistics through hypothesis tests
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. A hypothesis test uses sample data to decide whether to reject a null hypothesis $H_0$ in favour of an alternative $H_1$.
Key points.
- The null hypothesis $H_0$ states no effect or no difference, and the alternative $H_1$ states the claim being tested.
- The test statistic, for example the t-test $t=\dfrac{\bar x-\mu}{s/\sqrt n}$, measures how far the sample is from $H_0$.
- The p-value is the probability of a result at least as extreme as observed if $H_0$ is true; if $p<\alpha$ (usually 0.05), reject $H_0$.
- A Type I error is rejecting a true $H_0$ (probability $\alpha$), and a Type II error is accepting a false $H_0$.
Regression & ANOVA
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. Regression fits a line $y=a+bx$ that predicts a dependent variable from an independent one, and ANOVA compares group means through variances.
Key points.
- The least-squares slope is $b=\dfrac{\sum(x-\bar x)(y-\bar y)}{\sum(x-\bar x)^2}$ and the intercept is $a=\bar y-b\bar x$.
- The coefficient of determination $R^2$ is the fraction of variation in $y$ explained by the line.
- ANOVA tests whether three or more group means are equal, using the F-ratio of between-group variance to within-group variance.
- A large F (small p-value) means at least one group mean differs.
Regression ANOVA (Analysis of Variance)
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">High weight</span>
Definition. <mark>Regression ANOVA splits the total variation in the dependent variable into the part explained by the regression line and the residual part, and uses the F-test to judge whether the regression is significant.</mark>
Key points.
- The total sum of squares is $SST=\sum(y-\bar y)^2$, which measures the whole variation in $y$.
- The regression sum of squares is $SSR=\sum(\hat y-\bar y)^2$, the variation explained by the fitted line.
- The error (residual) sum of squares is $SSE=\sum(y-\hat y)^2$, the variation left unexplained, and $SST=SSR+SSE$.
- Each sum of squares is divided by its degrees of freedom to give a mean square: $MSR=SSR/1$ for simple regression and $MSE=SSE/(n-2)$.
- The test statistic is $F=MSR/MSE$, compared with the F-table value at $(1,\,n-2)$ degrees of freedom.
- The null hypothesis is $H_0$: slope $=0$ (no linear relationship); if the calculated F exceeds the table value, reject $H_0$ and the regression is significant.
- The goodness of fit is $R^2=SSR/SST$, and a value near 1 means the line explains most of the variation.
ANOVA table.
| Source | SS | df | MS | F |
|---|---|---|---|---|
| Regression | SSR | 1 | $SSR/1$ | $MSR/MSE$ |
| Error | SSE | $n-2$ | $SSE/(n-2)$ | |
| Total | SST | $n-1$ |
Example. If a regression of sales on advertising gives $SSR=90$, $SSE=10$ and $n=12$, then $SST=100$, $MSR=90$, $MSE=10/10=1$, $F=90$ and $R^2=0.9$; since 90 far exceeds the table value of 4.96 at (1,10) and 5%, the regression is significant.
Answer frame. Open with the definition; draw the ANOVA table; develop points 1-3 (the three sums of squares), then 4-5 (mean squares and F), then 6-7 (decision and $R^2$); close by stating that a significant F means the line predicts $y$ better than the mean alone. In a three-of-five short note, give the definition, the table and one line on the F-test, about half a page.
Asked: [14 marks] (Dec 2020) Write short note on any three: i) Capacity scheduler in Map Reduce, ii) 5P's of Big Data, iii) Metastore in Hive, iv) Regression ANOVA, v) Term frequency.
Last-minute revision
- A probability distribution has $\sum P(x)=1$ and every $P(x)\ge0$.
- $E[X]=\sum xP(x)$ and $Var(X)=E[X^2]-(E[X])^2$.
- $Z=(X-\mu)/\sigma$; the normal curve is symmetric, with area 0.5 on each side of the mean.
- Count $=N\times$ area; percentage $=$ area $\times100$.
- Weights problem: 286 students between 28.4 and 30.4 kg, 90 above 31.3 kg.
- Sales problem: 172 businesses above Rs. 40,000, and 38.1% between Rs. 30,000 and Rs. 40,000.
- Probability-function problem: $k=0.1$ (reject $-1$), mean 3.66, variance 3.4044.
- Confidence interval for the mean: $\bar x\pm z\sigma/\sqrt n$.
- Reject $H_0$ when $p<\alpha$; Type I is rejecting a true $H_0$.
- Regression line: $y=a+bx$, $b=\sum(x-\bar x)(y-\bar y)/\sum(x-\bar x)^2$.
- $SST=SSR+SSE$; $F=MSR/MSE$; $R^2=SSR/SST$.
Memory hooks
- Probabilities always sum to one: "total is 1, or it is not a distribution".
- Z means "how many SDs from the mean": subtract mean, divide by SD.
- SST = SSR + SSE: "Total = Regression explained + Error left".
- F is a fraction of explained over unexplained; big F means the line matters.
- Type I is a false alarm, Type II is a missed alarm.
Coverage checklist
- Probability Distributions: Nov 2023 normal-distribution numerical (students and businesses), Nov 2023 find k, mean, variance.
- Inferential Statistics: no past questions.
- Inferential Statistics through hypothesis tests: no past questions.
- Regression & ANOVA: no past questions.
- Regression ANOVA(Analysis of Variance): Dec 2020 short note on Regression ANOVA.