How unit 5 is examined
This unit extends ARMA to nonstationary and seasonal data (ARIMA, SARIMA), then covers regression with ARMA errors, VAR, state-space models and deep learning forecasters. No topic was asked in the supplied papers, so each is short.
ARIMA models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>An ARIMA(p,d,q) process is one whose d-th difference is a stationary ARMA(p,q) process.</mark>
Key points.
- Differencing removes trend: $\nabla X_t = X_t - X_{t-1} = (1-B)X_t$, where $B$ is the backshift operator.
- The model is $\phi(B)(1-B)^d X_t = \theta(B) Z_t$, with $Z_t$ white noise.
- $p$ is the autoregressive order, $d$ the number of differences and $q$ the moving-average order.
- ARIMA(0,1,0) is the random walk; one or two differences usually suffice.
- If $d=0$, ARIMA reduces to ARMA($p,q$), so ARMA is the stationary special case.
- The number of differences lost is $d$, because each difference shortens the series by one observation.
Example. ARIMA(1,1,0): $W_t = X_t - X_{t-1}$ follows $W_t = 0.5\,W_{t-1} + Z_t$, so $X_t = 1.5X_{t-1} - 0.5X_{t-2} + Z_t$.
Identification techniques
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Identification means choosing $d$, $p$ and $q$ from the data using the sample ACF, PACF and information criteria (Box-Jenkins method).</mark>
Key points.
- A slowly decaying ACF signals nonstationarity, so difference until the ACF dies quickly.
- For the differenced series, the PACF cutting off after lag $p$ suggests AR($p$), and the ACF cutting off after lag $q$ suggests MA($q$).
- If both tail off, the model is mixed ARMA; compare candidates with AIC or BIC and choose the smallest.
- Then estimate the parameters and check that residuals look like white noise.
- Overdifferencing is a warning sign: the ACF at lag 1 becomes close to $-0.5$ and the variance of the series increases, so use the smallest $d$ that gives stationarity.
- The Box-Jenkins cycle has three stages: identify the model, estimate the parameters, and run diagnostic checks, repeating if the residuals fail.
| Pattern in differenced series | ACF | PACF | Model |
|---|---|---|---|
| AR($p$) | tails off | cuts off after $p$ | AR |
| MA($q$) | cuts off after $q$ | tails off | MA |
| ARMA | tails off | tails off | mixed |
Unit roots in time series
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A unit root exists when the AR polynomial has a root on the unit circle, so the series is nonstationary and shocks never die out.</mark>
Key points.
- The random walk $X_t = X_{t-1} + Z_t$ has a unit root ($\phi = 1$).
- The Augmented Dickey-Fuller (ADF) test regresses $\nabla X_t$ on $X_{t-1}$ and lagged differences.
- The null hypothesis is that a unit root exists; a very negative test statistic rejects it and indicates stationarity.
- If the null is not rejected, difference the series.
- For $X_t=\phi X_{t-1}+Z_t$, the ADF regression is $\nabla X_t = \gamma X_{t-1} + \sum_i \delta_i \nabla X_{t-i} + Z_t$ with $\gamma=\phi-1$, and the null is $\gamma=0$.
- The ADF statistic is compared with Dickey-Fuller critical values, not the normal table; at 5% with a constant the value is about $-2.86$.
- A series that becomes stationary after $d$ differences is called integrated of order $d$, written I($d$).
Forecasting ARIMA models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==An ARIMA forecast is the conditional expectation $\hat X_{n+h} = E[X_{n+h} \mid X_1,\dots,X_n]$, obtained by forecasting the differenced series and summing back (integrating).==
Key points.
- Rewrite the model as a difference equation for $X_t$, replace future noise by 0 and future values by their forecasts.
- Forecasts of ARIMA models with $d \ge 1$ do not converge to a constant mean.
- The forecast error variance grows with horizon $h$, so prediction intervals widen.
- An approximate 95% interval is $\hat X_{n+h} \pm 1.96\,\sigma_h$.
- The random walk forecast is flat, $\hat X_{n+h}=X_n$, with error variance $h\sigma^2$.
Example. ARIMA(1,1,0) with $\phi=0.5$, $X_n=100$, $X_{n-1}=96$. The last difference is 4, and future differences shrink by half each step.
| $h$ | forecast difference | $\hat X_{n+h}$ |
|---|---|---|
| 1 | $0.5\times4=2$ | 102 |
| 2 | $0.5\times2=1$ | 103 |
| 3 | $0.5\times1=0.5$ | 103.5 |
Seasonal ARIMA models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>SARIMA$(p,d,q)\times(P,D,Q)_s$ combines non-seasonal and seasonal AR, differencing and MA terms, where $s$ is the season length.</mark>
Key points.
- The model is $\Phi(B^s)\phi(B)(1-B)^d(1-B^s)^D X_t = \Theta(B^s)\theta(B) Z_t$.
- Seasonal differencing $(1-B^s)$ removes a seasonal pattern; $s=12$ for monthly and $s=4$ for quarterly data.
- Seasonal ACF and PACF spikes at lags $s, 2s, \dots$ identify $P$ and $Q$.
- The airline model ARIMA$(0,1,1)\times(0,1,1)_{12}$ is the classic example.
- Seasonal differencing for monthly data is $Y_t = X_t - X_{t-12}$, which removes the repeating yearly pattern.
- The airline model expands to $(1-B)(1-B^{12})X_t = (1+\theta B)(1+\Theta B^{12})Z_t$; it needs only two parameters.
- Forecasting uses the same difference-equation method as ARIMA, and the seasonal pattern continues in the forecasts.
Regression with ARMA errors
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. ==Regression with ARMA errors is $Y_t = \beta^T x_t + \eta_t$, where the error $\eta_t$ follows an ARMA(p,q) process instead of being white noise.==
Key points.
- Ordinary least squares ignores correlated errors, so its standard errors and tests become unreliable.
- Coefficients $\beta$ and ARMA parameters are estimated together by maximum likelihood (or generalised least squares).
- The model is also called dynamic regression; with an AR error part it is often written ARMAX.
- Check that the final residuals are white noise.
- Method: fit a regression, examine the residual ACF and PACF, identify an ARMA model for the residuals, then refit everything jointly.
- Example: sales $Y_t = \beta_0 + \beta_1 (\text{advertising})_t + \eta_t$ with $\eta_t = 0.6\,\eta_{t-1} + Z_t$, so today's unexplained sales carry over to tomorrow.
Multivariate Time Series analysis
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A multivariate time series is a vector of series $\mathbf{X}_t$ observed together; the vector autoregression VAR($p$) models each series on past values of all series.</mark>
Key points.
- VAR($p$) is $\mathbf{X}_t = c + A_1\mathbf{X}_{t-1} + \dots + A_p\mathbf{X}_{t-p} + \mathbf{Z}_t$.
- Each equation is fitted by ordinary least squares, and the lag $p$ is chosen by AIC or BIC.
- The cross-correlation function measures the lead-lag relation between two series.
- Cointegration means individually nonstationary series have a stationary linear combination, so they share a long-run equilibrium.
- A VAR with $k$ series and $p$ lags has $k^2 p$ slope coefficients; for $k=3$, $p=2$ this is 18.
- Granger causality asks whether past values of one series improve the forecast of another, and impulse responses show the effect of a shock over time.
- Differencing cointegrated series loses long-run information; the vector error-correction model (VECM) keeps it.
State-Space Models
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>A state-space model describes an observed series through a hidden state that evolves over time, using a state equation and an observation equation.</mark>
Key points.
- State equation: $\mathbf{S}_t = F\,\mathbf{S}_{t-1} + \mathbf{w}_t$; observation equation: $Y_t = H\,\mathbf{S}_t + v_t$.
- The Kalman filter updates the state estimate recursively in two steps, predict and then correct with each new observation.
- ARMA, ARIMA and structural (trend plus seasonal) models can all be written in state-space form.
- It handles missing values naturally.
- Scalar Kalman example: prior state 10 with variance $P=4$, observation $y=12$ with noise variance $R=1$, and $H=1$. The gain is $K=P/(P+R)=0.8$, the updated state is $10+0.8(12-10)=11.6$, and the variance falls to $(1-K)P=0.8$.
- The local level model $S_t=S_{t-1}+w_t$, $Y_t=S_t+v_t$ is the simplest structural model and is equivalent to ARIMA(0,1,1).
Deep Learning techniques of time series forecasting
<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>
Definition. <mark>Deep learning forecasters are neural networks that learn temporal patterns directly from windows of past values, without a fixed statistical form.</mark>
Key points.
- An LSTM has input, forget and output gates and a memory cell, so it captures long-range dependence and avoids vanishing gradients.
- A temporal CNN uses dilated causal convolutions, which see far into the past without looking at the future.
- Input is a sliding window of lagged values; the network predicts the next value or several steps ahead.
- They need much data and tuning, and are less interpretable than ARIMA.
- Recurrent networks feed the hidden state back into themselves, and the LSTM forget gate decides which memory to keep: $c_t = f_t \odot c_{t-1} + i_t \odot \tilde c_t$.
- Data must be scaled and split in time order, never shuffled, to avoid leaking the future into training.
- Hybrid approaches, an ARIMA for the linear part and a network for the residual, often beat either alone.
Last-minute revision
- ARIMA(p,d,q): the d-th difference is ARMA(p,q); $\nabla = 1-B$.
- Random walk is ARIMA(0,1,0).
- Slowly decaying ACF means differencing is needed.
- PACF cut-off gives $p$; ACF cut-off gives $q$; smallest AIC or BIC picks the model.
- ADF null hypothesis: unit root present.
- ARIMA forecast intervals widen with horizon.
- SARIMA$(p,d,q)\times(P,D,Q)_s$; $s=12$ monthly.
- Airline model: ARIMA$(0,1,1)\times(0,1,1)_{12}$.
- Regression with ARMA errors: $Y_t = \beta^T x_t + \eta_t$, $\eta_t$ is ARMA.
- VAR($p$) uses all series' lags; cointegration gives a stationary combination.
- Kalman filter: predict, then update.
Memory hooks
- I-for-Integrated: I means "difference d times to get stationary".
- PACF-AR, ACF-MA: PACF cuts for AR, ACF cuts for MA.
- ADF null = trouble: the null is a unit root, so you want to reject it.
- Season in the exponent: seasonal terms use $B^s$.
- Kalman = guess, then correct.
Coverage checklist
- ARIMA models: definition, differencing, model equation (no past questions).
- Identification techniques: ACF/PACF, AIC/BIC (no past questions).
- Unit roots in time series: ADF test (no past questions).
- Forecasting ARIMA models: point forecasts and intervals (no past questions).
- Seasonal ARIMA models: SARIMA equation, airline model (no past questions).
- Regression with ARMA errors: definition and estimation (no past questions).
- Multivariate Time Series analysis: VAR, cointegration (no past questions).
- State-Space Models: state and observation equations, Kalman filter (no past questions).
- Deep Learning techniques of time series forecasting: LSTM, temporal CNN (no past questions).