The Minnesota Prior#
The Minnesota prior [Doan et al., 1984, Litterman, 1986] is the most widely used prior for Bayesian VARs, and it is what Impulso applies by default. By default it encodes the belief that each variable follows a random walk, with coefficients on other variables’ lags shrunk toward zero.
This page is the reference statement of the prior. For a worked introduction with figures, prior predictive checks, and a bias–variance experiment, see The Minnesota Prior, From Scratch.
Setup#
Impulso estimates the VAR(\(p\))
with \(y_t \in \mathbb{R}^n\) and each \(A_l \in \mathbb{R}^{n \times n}\). The lag matrices are stacked into the single coefficient matrix that VAR.fit samples,
Write \(\beta_{ij}^{(l)}\) for the entry of \(A_l\) in row \(i\) and column \(j\): the coefficient on lag \(l\) of variable \(j\) in the equation for variable \(i\). Columns of \(B\) are ordered lag-major — all \(n\) variables at lag 1, then all \(n\) at lag 2, and so on — so \(\beta_{ij}^{(l)}\) sits at column \((l-1)n + j\).
The prior#
Every coefficient gets an independent normal prior,
with mean
where \(\delta_i\) is variable \(i\)’s own-lag mean (own_lag_mean): 1 (the default) for a random walk, 0 for a stationary series. The standard deviation is built from four multiplicative factors,
where \(\sigma_i\) is the AR(1) residual standard deviation of variable \(i\) (impulso._conjugate.ar1_residual_sd) — the classical Litterman cross-variable scale, and the same \(\sigma\) the conjugate NIWPrior already applies via minnesota_dummies. On own lags (\(i = j\)) the ratio is 1, so it changes nothing there; it only rescales cross-lag entries, converting a coefficient’s prior from coefficient space into contribution space so it no longer depends on the units the two variables happen to be measured in. See ADR-0015 for the rationale.
MinnesotaPrior.build_priors(n_vars, n_lags, sigma) returns Eq. 12 and Eq. 13 as the arrays B_mu and B_sigma, both of shape (n_vars, n_vars * n_lags). sigma is required and keyword-only; VAR.fit computes it once via ar1_residual_sd(data.endog) and passes it here. VAR.fit then passes B_mu/B_sigma straight to PyMC as the mu and sigma of a normal prior on B.
Hyperparameters#
Parameter |
Symbol |
Default |
Meaning |
|---|---|---|---|
|
\(\lambda\) |
|
Overall shrinkage. \(\lambda \to 0\) freezes the model at the random walk; \(\lambda \to \infty\) recovers OLS. Must be \(> 0\). |
|
\(d(l)\) |
|
How fast the prior tightens on longer lags. |
|
\(\kappa\) |
|
Shrinkage on other variables’ lags relative to own lags. \(\kappa = 0\) reduces the VAR to \(n\) independent AR(\(p\)) models; \(\kappa = 1\) treats own and cross lags alike. Must lie in \([0, 1]\). |
|
\(\delta_i\) |
|
Prior mean of each variable’s own first lag. A scalar applies to every variable; a tuple gives one entry per variable, in |
\(\sigma_i / \sigma_j\) is not a hyperparameter — it is always on, derived from the data, and has no opt-out.
Rough guidance: tighten \(\lambda\) as the system grows, since the number of coefficients rises quadratically in \(n\) while the sample does not — Giannone et al. [2015] show the optimal tightness falls with dimension. Prefer "geometric" decay when you have many lags and no reason to expect long-cycle dynamics.
What the prior implies#
With the default \(\delta_i = 1\) the prior mean is a random walk, whose companion matrix has spectral radius exactly one. The Minnesota prior therefore sits on the boundary of stationarity by construction, and \(\lambda\) controls how far around that boundary the prior mass spreads. It is not a stationarity prior — most prior draws are technically explosive at any \(\lambda\), because the largest of several near-unit eigenvalues is biased upward. What small \(\lambda\) buys is that they are only barely so: at \(\lambda = 0.05\) in a 3-variable VAR(4), the median spectral radius is 1.05 and only about an eighth of draws exceed 1.10. At \(\lambda = 1\) the median draw doubles a shock every period.
Scope and caveats#
Scale adjustment is always on. Eq. 13 includes the classical Litterman \(\sigma_i / \sigma_j\) term (ADR-0015), so a cross-lag coefficient’s prior is expressed in contribution space rather than raw coefficient space, regardless of the units your variables happen to be measured in. There is no flag to switch it off. Standardising your variables before fitting is no longer necessary for the prior to make sense, though it remains a reasonable habit for reading raw coefficient values. NIWPrior has scaled by the same \(\sigma\) since it was written, via minnesota_dummies.
Levels, not growth rates. The default mean in Eq. 12 is a statement about levels. On differenced, de-meaned, or standardised series a prior mean of one on the own first lag is too persistent; the honest centre is nearer zero. Set \(\delta_i\) for those series with own_lag_mean, e.g. MinnesotaPrior(own_lag_mean=(1.0, 0.0, 0.0)) for one series in levels followed by two growth rates.
Coefficients only. MinnesotaPrior governs the lag coefficients. Intercepts receive a fixed \(\mathcal{N}(0, 1)\) prior, and the residual covariance \(\Sigma\) is handled by the volatility process (Constant by default).
Not conjugate. Eq. 11 places independent normal priors on the coefficients with a separately parameterised covariance, so the posterior has no closed form and is sampled with NUTS. That independence is what buys per-equation own-versus-cross asymmetry via \(\kappa\). The conjugate alternative, NIWPrior, gives up that asymmetry in exchange for a closed-form posterior and marginal likelihood — which in turn lets the data select \(\lambda\). See The Conjugate VAR.
Usage#
from impulso import VAR
from impulso.priors import MinnesotaPrior
# Use defaults
spec = VAR(lags=4, prior="minnesota")
# Customize hyperparameters
prior = MinnesotaPrior(tightness=0.2, decay="geometric", cross_shrinkage=0.3)
spec = VAR(lags=4, prior=prior)
# A stationary prior mean for the second and third series
prior = MinnesotaPrior(own_lag_mean=(1.0, 0.0, 0.0))
spec = VAR(lags=4, prior=prior)