Skip to content
AD-703 (C) · Advanced Statistical Analytics/Quick Revision Short Notes

Advanced Statistical Analytics (AD-703 (C)) - Unit 4 Short Notes

How unit 4 is examined

This unit covers estimating an unknown population parameter from a sample: what makes an estimator good, two methods to build one (moments and maximum likelihood), and large-sample confidence intervals. No topic was asked in the supplied papers, so each is short.

Problem of point estimation

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Point estimation uses a single sample value, the estimator $T(X_1,\dots,X_n)$, to estimate an unknown population parameter $\theta$.</mark>

Key points.

  1. A parameter is a fixed unknown constant of the population, while an estimator is a statistic, that is, a random variable that changes from sample to sample.
  2. The number an estimator gives for one particular sample is called the estimate.
  3. For example, $\bar{x}$ estimates $\mu$ and $s^2$ estimates $\sigma^2$.
  4. Different estimators can exist for the same parameter, which is why criteria to choose among them are needed.

Criteria of a good Estimator

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A good estimator is one that is unbiased, consistent, efficient and sufficient.</mark>

Key points.

  1. Unbiasedness means the estimator is correct on average, $E(T)=\theta$.
  2. Consistency means the estimator gets closer to $\theta$ as the sample size $n$ grows.
  3. Efficiency means that among unbiased estimators it has the smallest variance.
  4. Sufficiency means it captures all the information the sample holds about $\theta$.

Unbiasedness

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==An estimator $T$ is unbiased for $\theta$ if $E(T)=\theta$; its bias is $E(T)-\theta$.==

Key points.

  1. The sample mean is unbiased for $\mu$ because $E(\bar{X})=\mu$.
  2. The sample variance with divisor $n-1$, $s^2=\frac{1}{n-1}\sum(x_i-\bar{x})^2$, is unbiased for $\sigma^2$, whereas the divisor $n$ gives $E=\frac{n-1}{n}\sigma^2$, which is biased.
  3. Unbiased estimators need not be unique, and an unbiased estimator may not exist for every parameter.

Consistency

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>$T_n$ is a consistent estimator of $\theta$ if $T_n \to \theta$ in probability, that is, $P(|T_n-\theta|<\epsilon)\to 1$ as $n\to\infty$.</mark>

Key points.

  1. A sufficient condition is $E(T_n)\to\theta$ and $\mathrm{Var}(T_n)\to 0$ as $n\to\infty$.
  2. The sample mean $\bar{X}$ is consistent for $\mu$ because its variance $\sigma^2/n$ tends to 0.
  3. Consistency is a large-sample property, so a consistent estimator can still be biased for small $n$.

Efficiency, Sufficiency, Minimum Variance and Unbiasedness (Small sample)

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Of two unbiased estimators, the more efficient one is the one with smaller variance; the minimum variance unbiased estimator (MVUE) has the least variance among all unbiased estimators.</mark>

Key points.

  1. Relative efficiency of $T_1$ with respect to $T_2$ is $\mathrm{Var}(T_2)/\mathrm{Var}(T_1)$.
  2. The Cramer-Rao lower bound is $\mathrm{Var}(T)\ge \dfrac{1}{nI(\theta)}$, and an unbiased estimator that reaches it is the MVUE.
  3. A statistic is sufficient if the conditional distribution of the sample given it does not depend on $\theta$; by the factorisation theorem, $L=g(T,\theta)\,h(x)$.
  4. For a normal population, $\bar{X}$ is sufficient and MVUE for $\mu$.

Method of moments

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==In the method of moments, the population moments $E(X^r)$ are equated to the corresponding sample moments $m_r'=\frac{1}{n}\sum x_i^r$ and solved for the parameters.==

Key points.

  1. For $k$ unknown parameters, equate the first $k$ moments and solve the $k$ equations.
  2. Example: for a Poisson($\lambda$) population, $E(X)=\lambda$, so $\hat\lambda=\bar{x}$; for a normal population, $\hat\mu=\bar{x}$ and $\hat\sigma^2=\frac{1}{n}\sum(x_i-\bar{x})^2$.
  3. The method is simple, but its estimators are not always efficient or unbiased.

Method of Maximum Likelihood

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==The maximum likelihood estimator (MLE) is the value of $\theta$ that maximises the likelihood $L(\theta)=\prod_{i=1}^{n} f(x_i;\theta)$.==

Steps.

Step 1: Write L(theta) = product of f(x_i; theta).
Step 2: Take the log, ln L.
Step 3: Solve d(ln L)/d(theta) = 0 for theta.
Step 4: Check that d2(ln L)/d(theta)^2 < 0.

Key points.

  1. For a normal population, $\hat\mu=\bar{x}$ and $\hat\sigma^2=\frac{1}{n}\sum(x_i-\bar{x})^2$.
  2. The MLE is consistent, asymptotically normal and efficient, and it has the invariance property: the MLE of $g(\theta)$ is $g(\hat\theta)$.

Consistency & Efficiency (Large sample)

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>An estimator is asymptotically efficient if, for large $n$, it is consistent and asymptotically normal with variance equal to the Cramer-Rao bound $\frac{1}{nI(\theta)}$.</mark>

Key points.

  1. Under regularity conditions, the MLE satisfies $\hat\theta\sim N\!\left(\theta,\frac{1}{nI(\theta)}\right)$ for large $n$.
  2. Asymptotic efficiency is judged by comparing asymptotic variances, and the smaller one is better.
  3. The estimator is consistent when its bias and variance both vanish as $n\to\infty$.

Interval Estimation: Confidence Intervals of mean and proportion in large samples

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A $(1-\alpha)100\%$ confidence interval is a random interval that contains the parameter with probability $1-\alpha$.</mark>

Formula. $$\bar{x}\pm z_{\alpha/2}\frac{\sigma}{\sqrt{n}}\quad(\text{use } s \text{ if } \sigma \text{ is unknown}),\qquad \hat{p}\pm z_{\alpha/2}\sqrt{\frac{\hat{p}\hat{q}}{n}}$$

Key points.

  1. A sample is large when $n\ge 30$, so the CLT makes $\bar{X}$ approximately normal.
  2. The critical values are $z_{0.025}=1.96$ for 95% and $z_{0.005}=2.58$ for 99%.
  3. A higher confidence level gives a wider interval, and a larger $n$ gives a narrower one.

Last-minute revision

  • A point estimator is a statistic used to estimate a parameter; a single value it gives is the estimate.
  • Unbiased means $E(T)=\theta$; the divisor $n-1$ makes $s^2$ unbiased for $\sigma^2$.
  • Consistent means $T_n\to\theta$ in probability; $E(T_n)\to\theta$ and $\mathrm{Var}(T_n)\to0$ is sufficient.
  • Efficiency compares variances of unbiased estimators; relative efficiency is $\mathrm{Var}(T_2)/\mathrm{Var}(T_1)$.
  • Cramer-Rao bound: $\mathrm{Var}(T)\ge 1/(nI(\theta))$.
  • Sufficiency is tested by the factorisation $L=g(T,\theta)h(x)$.
  • Method of moments equates population moments to sample moments.
  • MLE maximises $L(\theta)$, usually through $d\ln L/d\theta=0$.
  • Normal MLE: $\hat\mu=\bar{x}$, $\hat\sigma^2=\frac1n\sum(x_i-\bar{x})^2$ (biased).
  • Large-sample CI for the mean is $\bar{x}\pm z_{\alpha/2}\sigma/\sqrt{n}$; $z=1.96$ for 95%.
  • Large-sample CI for a proportion is $\hat{p}\pm z_{\alpha/2}\sqrt{\hat{p}\hat{q}/n}$.

Memory hooks

  • UCES: Unbiased, Consistent, Efficient, Sufficient are the four properties of a good estimator.
  • Divide by $n-1$ for an unbiased sample variance, and by $n$ for the (biased) MLE.
  • MOM matches moments, MLE maximises likelihood.
  • 1.96 for 95% and 2.58 for 99%.

Coverage checklist

  • Problem of point estimation: definition, estimator versus estimate (no past questions).
  • Criteria of a good Estimator: the four criteria (no past questions).
  • Unbiasedness: definition, bias, $s^2$ (no past questions).
  • Consistency: definition and sufficient condition (no past questions).
  • Efficiency, Sufficiency Minimum Variance and Unbiasedness (Small sample): efficiency, Cramer-Rao, factorisation (no past questions).
  • Method of moments: procedure and examples (no past questions).
  • Method of Maximum Likelihood: steps, normal MLE, properties (no past questions).
  • Consistency & Efficiency (Large sample): asymptotic normality and efficiency (no past questions).
  • Interval Estimation: Confidence Intervals of mean and proportion in large samples: both formulas and $z$ values (no past questions).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in