How unit 4 is examined
This unit covers estimating an unknown population parameter from a sample: what makes an estimator good, two methods to build one (moments and maximum likelihood), and large-sample confidence intervals. No topic was asked in the supplied papers, so each is short.
Problem of point estimation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Point estimation uses a single sample value, the estimator $T(X_1,\dots,X_n)$, to estimate an unknown population parameter $\theta$.</mark>
Key points.
- A parameter is a fixed unknown constant of the population, while an estimator is a statistic, that is, a random variable that changes from sample to sample.
- The number an estimator gives for one particular sample is called the estimate.
- For example, $\bar{x}$ estimates $\mu$ and $s^2$ estimates $\sigma^2$.
- Different estimators can exist for the same parameter, which is why criteria to choose among them are needed.
Criteria of a good Estimator
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A good estimator is one that is unbiased, consistent, efficient and sufficient.</mark>
Key points.
- Unbiasedness means the estimator is correct on average, $E(T)=\theta$.
- Consistency means the estimator gets closer to $\theta$ as the sample size $n$ grows.
- Efficiency means that among unbiased estimators it has the smallest variance.
- Sufficiency means it captures all the information the sample holds about $\theta$.
Unbiasedness
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==An estimator $T$ is unbiased for $\theta$ if $E(T)=\theta$; its bias is $E(T)-\theta$.==
Key points.
- The sample mean is unbiased for $\mu$ because $E(\bar{X})=\mu$.
- The sample variance with divisor $n-1$, $s^2=\frac{1}{n-1}\sum(x_i-\bar{x})^2$, is unbiased for $\sigma^2$, whereas the divisor $n$ gives $E=\frac{n-1}{n}\sigma^2$, which is biased.
- Unbiased estimators need not be unique, and an unbiased estimator may not exist for every parameter.
Consistency
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>$T_n$ is a consistent estimator of $\theta$ if $T_n \to \theta$ in probability, that is, $P(|T_n-\theta|<\epsilon)\to 1$ as $n\to\infty$.</mark>
Key points.
- A sufficient condition is $E(T_n)\to\theta$ and $\mathrm{Var}(T_n)\to 0$ as $n\to\infty$.
- The sample mean $\bar{X}$ is consistent for $\mu$ because its variance $\sigma^2/n$ tends to 0.
- Consistency is a large-sample property, so a consistent estimator can still be biased for small $n$.
Efficiency, Sufficiency, Minimum Variance and Unbiasedness (Small sample)
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Of two unbiased estimators, the more efficient one is the one with smaller variance; the minimum variance unbiased estimator (MVUE) has the least variance among all unbiased estimators.</mark>
Key points.
- Relative efficiency of $T_1$ with respect to $T_2$ is $\mathrm{Var}(T_2)/\mathrm{Var}(T_1)$.
- The Cramer-Rao lower bound is $\mathrm{Var}(T)\ge \dfrac{1}{nI(\theta)}$, and an unbiased estimator that reaches it is the MVUE.
- A statistic is sufficient if the conditional distribution of the sample given it does not depend on $\theta$; by the factorisation theorem, $L=g(T,\theta)\,h(x)$.
- For a normal population, $\bar{X}$ is sufficient and MVUE for $\mu$.
Method of moments
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==In the method of moments, the population moments $E(X^r)$ are equated to the corresponding sample moments $m_r'=\frac{1}{n}\sum x_i^r$ and solved for the parameters.==
Key points.
- For $k$ unknown parameters, equate the first $k$ moments and solve the $k$ equations.
- Example: for a Poisson($\lambda$) population, $E(X)=\lambda$, so $\hat\lambda=\bar{x}$; for a normal population, $\hat\mu=\bar{x}$ and $\hat\sigma^2=\frac{1}{n}\sum(x_i-\bar{x})^2$.
- The method is simple, but its estimators are not always efficient or unbiased.
Method of Maximum Likelihood
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==The maximum likelihood estimator (MLE) is the value of $\theta$ that maximises the likelihood $L(\theta)=\prod_{i=1}^{n} f(x_i;\theta)$.==
Steps.
Step 1: Write L(theta) = product of f(x_i; theta).
Step 2: Take the log, ln L.
Step 3: Solve d(ln L)/d(theta) = 0 for theta.
Step 4: Check that d2(ln L)/d(theta)^2 < 0.
Key points.
- For a normal population, $\hat\mu=\bar{x}$ and $\hat\sigma^2=\frac{1}{n}\sum(x_i-\bar{x})^2$.
- The MLE is consistent, asymptotically normal and efficient, and it has the invariance property: the MLE of $g(\theta)$ is $g(\hat\theta)$.
Consistency & Efficiency (Large sample)
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>An estimator is asymptotically efficient if, for large $n$, it is consistent and asymptotically normal with variance equal to the Cramer-Rao bound $\frac{1}{nI(\theta)}$.</mark>
Key points.
- Under regularity conditions, the MLE satisfies $\hat\theta\sim N\!\left(\theta,\frac{1}{nI(\theta)}\right)$ for large $n$.
- Asymptotic efficiency is judged by comparing asymptotic variances, and the smaller one is better.
- The estimator is consistent when its bias and variance both vanish as $n\to\infty$.
Interval Estimation: Confidence Intervals of mean and proportion in large samples
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A $(1-\alpha)100\%$ confidence interval is a random interval that contains the parameter with probability $1-\alpha$.</mark>
Formula. $$\bar{x}\pm z_{\alpha/2}\frac{\sigma}{\sqrt{n}}\quad(\text{use } s \text{ if } \sigma \text{ is unknown}),\qquad \hat{p}\pm z_{\alpha/2}\sqrt{\frac{\hat{p}\hat{q}}{n}}$$
Key points.
- A sample is large when $n\ge 30$, so the CLT makes $\bar{X}$ approximately normal.
- The critical values are $z_{0.025}=1.96$ for 95% and $z_{0.005}=2.58$ for 99%.
- A higher confidence level gives a wider interval, and a larger $n$ gives a narrower one.
Last-minute revision
- A point estimator is a statistic used to estimate a parameter; a single value it gives is the estimate.
- Unbiased means $E(T)=\theta$; the divisor $n-1$ makes $s^2$ unbiased for $\sigma^2$.
- Consistent means $T_n\to\theta$ in probability; $E(T_n)\to\theta$ and $\mathrm{Var}(T_n)\to0$ is sufficient.
- Efficiency compares variances of unbiased estimators; relative efficiency is $\mathrm{Var}(T_2)/\mathrm{Var}(T_1)$.
- Cramer-Rao bound: $\mathrm{Var}(T)\ge 1/(nI(\theta))$.
- Sufficiency is tested by the factorisation $L=g(T,\theta)h(x)$.
- Method of moments equates population moments to sample moments.
- MLE maximises $L(\theta)$, usually through $d\ln L/d\theta=0$.
- Normal MLE: $\hat\mu=\bar{x}$, $\hat\sigma^2=\frac1n\sum(x_i-\bar{x})^2$ (biased).
- Large-sample CI for the mean is $\bar{x}\pm z_{\alpha/2}\sigma/\sqrt{n}$; $z=1.96$ for 95%.
- Large-sample CI for a proportion is $\hat{p}\pm z_{\alpha/2}\sqrt{\hat{p}\hat{q}/n}$.
Memory hooks
- UCES: Unbiased, Consistent, Efficient, Sufficient are the four properties of a good estimator.
- Divide by $n-1$ for an unbiased sample variance, and by $n$ for the (biased) MLE.
- MOM matches moments, MLE maximises likelihood.
- 1.96 for 95% and 2.58 for 99%.
Coverage checklist
- Problem of point estimation: definition, estimator versus estimate (no past questions).
- Criteria of a good Estimator: the four criteria (no past questions).
- Unbiasedness: definition, bias, $s^2$ (no past questions).
- Consistency: definition and sufficient condition (no past questions).
- Efficiency, Sufficiency Minimum Variance and Unbiasedness (Small sample): efficiency, Cramer-Rao, factorisation (no past questions).
- Method of moments: procedure and examples (no past questions).
- Method of Maximum Likelihood: steps, normal MLE, properties (no past questions).
- Consistency & Efficiency (Large sample): asymptotic normality and efficiency (no past questions).
- Interval Estimation: Confidence Intervals of mean and proportion in large samples: both formulas and $z$ values (no past questions).