How unit 5 is examined
This unit covers Bayesian inference and estimation, MCMC, hierarchical models, survival analysis, causal inference and high-dimensional data; no topic has been asked in the supplied papers, so learn each definition and formula.
Introduction to Bayesian inference
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Bayesian inference updates the prior belief about a parameter $\theta$ with the observed data to give the posterior distribution, using Bayes theorem.</mark>
$$p(\theta\mid x)=\frac{p(x\mid\theta)\,p(\theta)}{p(x)}\propto \text{likelihood}\times\text{prior}$$
Key points.
- The prior $p(\theta)$ states belief about the parameter before seeing data.
- The likelihood $p(x\mid\theta)$ measures how well each value of $\theta$ explains the data.
- The posterior $p(\theta\mid x)$ combines both and is the final answer of the analysis.
- The denominator $p(x)=\int p(x\mid\theta)p(\theta)\,d\theta$ is a normalising constant.
Bayesian parameter estimation
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Bayesian estimation summarises the posterior distribution by a single value, such as the posterior mean, median or mode (MAP), or by an interval.</mark>
Key points.
- The posterior mean $E[\theta\mid x]$ is the Bayes estimate under squared-error loss.
- The MAP estimate is the posterior mode, $\hat\theta=\arg\max_\theta\, p(x\mid\theta)p(\theta)$.
- A credible interval is an interval containing $\theta$ with stated posterior probability, for example 95%.
- For a Beta$(a,b)$ prior and $k$ successes in $n$ trials, the posterior is Beta$(a+k,,b+n-k)$.
Markov Chain Monte Carlo (MCMC) methods
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>MCMC draws samples from a posterior that cannot be solved in closed form by building a Markov chain whose stationary distribution is that posterior.</mark>
Key points.
- Metropolis-Hastings proposes a new value and accepts it with probability $\min\!\left(1,\frac{p(\theta')q(\theta\mid\theta')}{p(\theta)q(\theta'\mid\theta)}\right)$.
- Gibbs sampling updates one parameter at a time from its full conditional distribution, and every draw is accepted.
- The early draws, called burn-in, are discarded before the chain reaches its stationary distribution.
- Posterior means and intervals are computed from the retained samples, and convergence is checked with trace plots.
Bayesian hierarchical models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A Bayesian hierarchical model places priors on the parameters of the prior, so that group-level parameters are drawn from a common population distribution controlled by hyperparameters.</mark>
Key points.
- The structure is data $y_{ij}\sim p(y\mid\theta_j)$, group parameters $\theta_j\sim p(\theta\mid\phi)$ and hyperparameters $\phi\sim p(\phi)$.
- Groups share information, so small groups are pulled towards the overall mean; this is called shrinkage.
- It suits multilevel data such as students within schools.
- It is fitted usually by MCMC.
Survival analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Survival analysis studies the time until an event such as failure or death occurs, and handles censored observations whose event time is not fully seen.</mark>
Key points.
- Censoring means the event had not happened when observation ended, so only a lower bound on the time is known.
- The survival function is $S(t)=P(T>t)$.
- The hazard function is $h(t)=\dfrac{f(t)}{S(t)}$, the instantaneous event rate at time $t$ given survival to $t$.
- The Kaplan-Meier estimator estimates $S(t)$ non-parametrically, and the Cox model $h(t)=h_0(t)e^{\beta^{T}x}$ relates covariates to hazard.
Causal inference
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Causal inference estimates the effect of a treatment on an outcome, by comparing what happened with the counterfactual of what would have happened without it.</mark>
Key points.
- Each unit has potential outcomes $Y(1)$ and $Y(0)$, but only one of them is ever observed.
- The average treatment effect is $ATE=E[Y(1)-Y(0)]$.
- Randomised experiments remove confounding, so the difference in group means estimates the ATE.
- In observational data, confounders are handled by matching, propensity scores or regression adjustment.
High-dimensional data analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>High-dimensional data has a number of variables $p$ that is large compared with the number of observations $n$, often $p>n$, so classical methods fail.</mark>
Key points.
- With $p>n$ ordinary least squares has no unique solution and overfits badly.
- Sparsity assumes that only a few variables truly matter.
- Lasso minimises $\sum(y_i-\hat y_i)^2+\lambda\sum|\beta_j|$ and shrinks some coefficients exactly to zero, so it selects variables.
- Dimension reduction, such as PCA, and regularisation such as ridge address the curse of dimensionality.
Last-minute revision
- Posterior $\propto$ likelihood $\times$ prior.
- MAP is the posterior mode; the posterior mean is the Bayes estimate under squared-error loss.
- Beta$(a,b)$ prior with $k$ successes in $n$ trials gives Beta$(a+k,b+n-k)$.
- Metropolis-Hastings accepts with probability $\min(1,\text{ratio})$; Gibbs samples full conditionals.
- Burn-in draws are discarded in MCMC.
- Hierarchical models give shrinkage through shared hyperparameters.
- $S(t)=P(T>t)$ and $h(t)=f(t)/S(t)$.
- Censored data give only a lower bound on the survival time.
- $ATE=E[Y(1)-Y(0)]$; randomisation removes confounding.
- Lasso uses an $L_1$ penalty and selects variables; ridge uses $L_2$.
Memory hooks
- Posterior = Prior x Likelihood, then normalise.
- MAP = Mode; Mean = squared loss.
- Gibbs gives one coordinate at a time; Metropolis proposes and accepts or rejects.
- Hazard is the risk now; survival is the chance of lasting past now.
- Lasso lassoes coefficients to zero.
Coverage checklist
- Introduction to Bayesian inference: no past questions; definition, Bayes theorem formula.
- Bayesian parameter estimation: no past questions; posterior mean, MAP, credible interval.
- Markov Chain Monte Carlo (MCMC) methods: no past questions; Metropolis-Hastings, Gibbs.
- Bayesian hierarchical models: no past questions; hyperparameters, shrinkage.
- Survival analysis: no past questions; censoring, survival and hazard functions.
- Causal inference: no past questions; potential outcomes, ATE.
- High-dimensional data analysis: no past questions; sparsity, lasso.