Bayesian and hierarchical stacking
bayesian_stacking.RdModel weights with a posterior, sampled by NUTS from pointwise held-out
log densities, the log score of the mixture as the likelihood (Yao et
al., 2018 and 2022). bayes_stacking() puts a Dirichlet prior on one
weight vector. hierarchical_stacking() lets the weights vary with
covariates, w = softmax(alpha + B x) against the last model as
reference, with the priors of BayesBlend's HierarchicalBayesStacking:
a model can be trusted in one part of the portfolio and not another.
Scale continuous covariates (BayesBlend divides by twice their standard
deviation) and dummy-code discrete ones first, with the dummies as the
first discrete columns. Without pooling, posterior moments match exact
grid integration (validation/scripts/stacking_grid.py).
Usage
bayes_stacking(
lpd,
concentration = NULL,
chains = 4,
tune = 1000,
draws = 1000,
seed = 0
)
hierarchical_stacking(
lpd,
covariates,
discrete = 0,
alpha_loc = 0,
alpha_scale = 1,
beta_loc = 0,
beta_scale = 1,
partial_pooling = FALSE,
pooling = pooling_scales(),
adaptive = NULL,
chains = 4,
tune = 1000,
draws = 1000,
seed = 0
)
pooling_scales(
tau_mu_global = 1,
tau_mu_discrete = 1,
tau_mu_continuous = 1,
tau_sigma_discrete = 1,
tau_sigma_continuous = 1
)Arguments
- lpd
A matrix of held-out log densities, one row per observation and one column per model (for example each model's
bayes_loo(m)$pointwise, as columns).- concentration
Dirichlet concentration, one per model;
NULLis uniform.- chains, tune, draws, seed
Sampler settings.
- covariates
A matrix or data frame of numeric covariates, one row per observation.
- discrete
Number of leading columns of
covariatesthat are dummy codes of discrete covariates.- alpha_loc, alpha_scale, beta_loc, beta_scale
Normal priors on the intercepts and (without pooling) slopes.
- partial_pooling
Pool the slopes (BayesBlend's pooling model).
- pooling
Pooling scales, from
pooling_scales().- adaptive
NULL, or the rate of the exponential prior onlambda(BayesBlend's default is 4).- tau_mu_global, tau_mu_discrete, tau_mu_continuous, tau_sigma_discrete, tau_sigma_continuous
Scales of the global mean, the model-level means and the slopes around them; BayesBlend's defaults are all 1.
Value
A stacking_fit with properties weights (posterior mean
weights: one row per observation for hierarchical stacking, one row
otherwise), alpha_draws, beta_draws and divergences. Use
stats::predict() with new covariates for their weights, and
blend_predictive() with the weight matrix.
Details
With partial_pooling = TRUE, each model's slopes on the discrete
columns, and separately on the continuous ones, are drawn around a
model-level mean, which is drawn around a global mean; pooling_scales()
sets the scales (0 removes a level: tau_mu_global = 0 fixes the global
mean at 0, tau_mu_* = 0 pools completely, tau_sigma_* = 0 sets every
slope to its model's mean). BayesBlend warns that pooling needs at least
three covariates. adaptive, the rate of an exponential prior on
lambda, multiplies the prior scales by N^lambda, which weakens them
as the data grow.
Examples
x <- seq(-0.5, 0.5, length.out = 100)
lpd <- cbind(a = ifelse(x < 0, -0.5, -2), b = ifelse(x < 0, -2, -0.5))
h <- hierarchical_stacking(lpd, data.frame(x = x), chains = 2, tune = 300, draws = 300)
predict(h, data.frame(x = c(-0.4, 0.4)))
#> a b
#> [1,] 0.8186926 0.1813074
#> [2,] 0.1756024 0.8243976
bayes_stacking(lpd, chains = 2, tune = 300, draws = 300)@weights
#> a b
#> [1,] 0.5001511 0.4998489