Skip to content
AD-702 (D) · Predictive Analytics/Quick Revision Short Notes

Predictive Analytics (AD-702 (D)) - Unit 5 Short Notes

How unit 5 is examined

This unit extends ARMA to nonstationary and seasonal data (ARIMA, SARIMA), then covers regression with ARMA errors, VAR, state-space models and deep learning forecasters. No topic was asked in the supplied papers, so each is short.

ARIMA models

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>An ARIMA(p,d,q) process is one whose d-th difference is a stationary ARMA(p,q) process.</mark>

Key points.

  1. Differencing removes trend: $\nabla X_t = X_t - X_{t-1} = (1-B)X_t$, where $B$ is the backshift operator.
  2. The model is $\phi(B)(1-B)^d X_t = \theta(B) Z_t$, with $Z_t$ white noise.
  3. $p$ is the autoregressive order, $d$ the number of differences and $q$ the moving-average order.
  4. ARIMA(0,1,0) is the random walk; one or two differences usually suffice.
  5. If $d=0$, ARIMA reduces to ARMA($p,q$), so ARMA is the stationary special case.
  6. The number of differences lost is $d$, because each difference shortens the series by one observation.

Example. ARIMA(1,1,0): $W_t = X_t - X_{t-1}$ follows $W_t = 0.5\,W_{t-1} + Z_t$, so $X_t = 1.5X_{t-1} - 0.5X_{t-2} + Z_t$.

Identification techniques

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Identification means choosing $d$, $p$ and $q$ from the data using the sample ACF, PACF and information criteria (Box-Jenkins method).</mark>

Key points.

  1. A slowly decaying ACF signals nonstationarity, so difference until the ACF dies quickly.
  2. For the differenced series, the PACF cutting off after lag $p$ suggests AR($p$), and the ACF cutting off after lag $q$ suggests MA($q$).
  3. If both tail off, the model is mixed ARMA; compare candidates with AIC or BIC and choose the smallest.
  4. Then estimate the parameters and check that residuals look like white noise.
  5. Overdifferencing is a warning sign: the ACF at lag 1 becomes close to $-0.5$ and the variance of the series increases, so use the smallest $d$ that gives stationarity.
  6. The Box-Jenkins cycle has three stages: identify the model, estimate the parameters, and run diagnostic checks, repeating if the residuals fail.
Pattern in differenced series ACF PACF Model
AR($p$) tails off cuts off after $p$ AR
MA($q$) cuts off after $q$ tails off MA
ARMA tails off tails off mixed

Unit roots in time series

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A unit root exists when the AR polynomial has a root on the unit circle, so the series is nonstationary and shocks never die out.</mark>

Key points.

  1. The random walk $X_t = X_{t-1} + Z_t$ has a unit root ($\phi = 1$).
  2. The Augmented Dickey-Fuller (ADF) test regresses $\nabla X_t$ on $X_{t-1}$ and lagged differences.
  3. The null hypothesis is that a unit root exists; a very negative test statistic rejects it and indicates stationarity.
  4. If the null is not rejected, difference the series.
  5. For $X_t=\phi X_{t-1}+Z_t$, the ADF regression is $\nabla X_t = \gamma X_{t-1} + \sum_i \delta_i \nabla X_{t-i} + Z_t$ with $\gamma=\phi-1$, and the null is $\gamma=0$.
  6. The ADF statistic is compared with Dickey-Fuller critical values, not the normal table; at 5% with a constant the value is about $-2.86$.
  7. A series that becomes stationary after $d$ differences is called integrated of order $d$, written I($d$).

Forecasting ARIMA models

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==An ARIMA forecast is the conditional expectation $\hat X_{n+h} = E[X_{n+h} \mid X_1,\dots,X_n]$, obtained by forecasting the differenced series and summing back (integrating).==

Key points.

  1. Rewrite the model as a difference equation for $X_t$, replace future noise by 0 and future values by their forecasts.
  2. Forecasts of ARIMA models with $d \ge 1$ do not converge to a constant mean.
  3. The forecast error variance grows with horizon $h$, so prediction intervals widen.
  4. An approximate 95% interval is $\hat X_{n+h} \pm 1.96\,\sigma_h$.
  5. The random walk forecast is flat, $\hat X_{n+h}=X_n$, with error variance $h\sigma^2$.

Example. ARIMA(1,1,0) with $\phi=0.5$, $X_n=100$, $X_{n-1}=96$. The last difference is 4, and future differences shrink by half each step.

$h$ forecast difference $\hat X_{n+h}$
1 $0.5\times4=2$ 102
2 $0.5\times2=1$ 103
3 $0.5\times1=0.5$ 103.5

Seasonal ARIMA models

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>SARIMA$(p,d,q)\times(P,D,Q)_s$ combines non-seasonal and seasonal AR, differencing and MA terms, where $s$ is the season length.</mark>

Key points.

  1. The model is $\Phi(B^s)\phi(B)(1-B)^d(1-B^s)^D X_t = \Theta(B^s)\theta(B) Z_t$.
  2. Seasonal differencing $(1-B^s)$ removes a seasonal pattern; $s=12$ for monthly and $s=4$ for quarterly data.
  3. Seasonal ACF and PACF spikes at lags $s, 2s, \dots$ identify $P$ and $Q$.
  4. The airline model ARIMA$(0,1,1)\times(0,1,1)_{12}$ is the classic example.
  5. Seasonal differencing for monthly data is $Y_t = X_t - X_{t-12}$, which removes the repeating yearly pattern.
  6. The airline model expands to $(1-B)(1-B^{12})X_t = (1+\theta B)(1+\Theta B^{12})Z_t$; it needs only two parameters.
  7. Forecasting uses the same difference-equation method as ARIMA, and the seasonal pattern continues in the forecasts.

Regression with ARMA errors

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. ==Regression with ARMA errors is $Y_t = \beta^T x_t + \eta_t$, where the error $\eta_t$ follows an ARMA(p,q) process instead of being white noise.==

Key points.

  1. Ordinary least squares ignores correlated errors, so its standard errors and tests become unreliable.
  2. Coefficients $\beta$ and ARMA parameters are estimated together by maximum likelihood (or generalised least squares).
  3. The model is also called dynamic regression; with an AR error part it is often written ARMAX.
  4. Check that the final residuals are white noise.
  5. Method: fit a regression, examine the residual ACF and PACF, identify an ARMA model for the residuals, then refit everything jointly.
  6. Example: sales $Y_t = \beta_0 + \beta_1 (\text{advertising})_t + \eta_t$ with $\eta_t = 0.6\,\eta_{t-1} + Z_t$, so today's unexplained sales carry over to tomorrow.

Multivariate Time Series analysis

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A multivariate time series is a vector of series $\mathbf{X}_t$ observed together; the vector autoregression VAR($p$) models each series on past values of all series.</mark>

Key points.

  1. VAR($p$) is $\mathbf{X}_t = c + A_1\mathbf{X}_{t-1} + \dots + A_p\mathbf{X}_{t-p} + \mathbf{Z}_t$.
  2. Each equation is fitted by ordinary least squares, and the lag $p$ is chosen by AIC or BIC.
  3. The cross-correlation function measures the lead-lag relation between two series.
  4. Cointegration means individually nonstationary series have a stationary linear combination, so they share a long-run equilibrium.
  5. A VAR with $k$ series and $p$ lags has $k^2 p$ slope coefficients; for $k=3$, $p=2$ this is 18.
  6. Granger causality asks whether past values of one series improve the forecast of another, and impulse responses show the effect of a shock over time.
  7. Differencing cointegrated series loses long-run information; the vector error-correction model (VECM) keeps it.

State-Space Models

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>A state-space model describes an observed series through a hidden state that evolves over time, using a state equation and an observation equation.</mark>

Key points.

  1. State equation: $\mathbf{S}_t = F\,\mathbf{S}_{t-1} + \mathbf{w}_t$; observation equation: $Y_t = H\,\mathbf{S}_t + v_t$.
  2. The Kalman filter updates the state estimate recursively in two steps, predict and then correct with each new observation.
  3. ARMA, ARIMA and structural (trend plus seasonal) models can all be written in state-space form.
  4. It handles missing values naturally.
  5. Scalar Kalman example: prior state 10 with variance $P=4$, observation $y=12$ with noise variance $R=1$, and $H=1$. The gain is $K=P/(P+R)=0.8$, the updated state is $10+0.8(12-10)=11.6$, and the variance falls to $(1-K)P=0.8$.
  6. The local level model $S_t=S_{t-1}+w_t$, $Y_t=S_t+v_t$ is the simplest structural model and is equivalent to ARIMA(0,1,1).

Deep Learning techniques of time series forecasting

<span style="display:inline-block;padding:.16em .6em;border:1.5px solid currentColor;border-radius:999px;font-size:.68em;font-weight:700;letter-spacing:.06em;text-transform:uppercase;opacity:.75">Not asked since 2022</span>

Definition. <mark>Deep learning forecasters are neural networks that learn temporal patterns directly from windows of past values, without a fixed statistical form.</mark>

Key points.

  1. An LSTM has input, forget and output gates and a memory cell, so it captures long-range dependence and avoids vanishing gradients.
  2. A temporal CNN uses dilated causal convolutions, which see far into the past without looking at the future.
  3. Input is a sliding window of lagged values; the network predicts the next value or several steps ahead.
  4. They need much data and tuning, and are less interpretable than ARIMA.
  5. Recurrent networks feed the hidden state back into themselves, and the LSTM forget gate decides which memory to keep: $c_t = f_t \odot c_{t-1} + i_t \odot \tilde c_t$.
  6. Data must be scaled and split in time order, never shuffled, to avoid leaking the future into training.
  7. Hybrid approaches, an ARIMA for the linear part and a network for the residual, often beat either alone.

Last-minute revision

  • ARIMA(p,d,q): the d-th difference is ARMA(p,q); $\nabla = 1-B$.
  • Random walk is ARIMA(0,1,0).
  • Slowly decaying ACF means differencing is needed.
  • PACF cut-off gives $p$; ACF cut-off gives $q$; smallest AIC or BIC picks the model.
  • ADF null hypothesis: unit root present.
  • ARIMA forecast intervals widen with horizon.
  • SARIMA$(p,d,q)\times(P,D,Q)_s$; $s=12$ monthly.
  • Airline model: ARIMA$(0,1,1)\times(0,1,1)_{12}$.
  • Regression with ARMA errors: $Y_t = \beta^T x_t + \eta_t$, $\eta_t$ is ARMA.
  • VAR($p$) uses all series' lags; cointegration gives a stationary combination.
  • Kalman filter: predict, then update.

Memory hooks

  • I-for-Integrated: I means "difference d times to get stationary".
  • PACF-AR, ACF-MA: PACF cuts for AR, ACF cuts for MA.
  • ADF null = trouble: the null is a unit root, so you want to reject it.
  • Season in the exponent: seasonal terms use $B^s$.
  • Kalman = guess, then correct.

Coverage checklist

  • ARIMA models: definition, differencing, model equation (no past questions).
  • Identification techniques: ACF/PACF, AIC/BIC (no past questions).
  • Unit roots in time series: ADF test (no past questions).
  • Forecasting ARIMA models: point forecasts and intervals (no past questions).
  • Seasonal ARIMA models: SARIMA equation, airline model (no past questions).
  • Regression with ARMA errors: definition and estimation (no past questions).
  • Multivariate Time Series analysis: VAR, cointegration (no past questions).
  • State-Space Models: state and observation equations, Kalman filter (no past questions).
  • Deep Learning techniques of time series forecasting: LSTM, temporal CNN (no past questions).
Go to where you left off?

Quick Add to Notes

Save questions, your own notes and screenshots into notes filed by unit. It takes a free account.

Create free account

Have an account? Log in