Skip to contents

Gradient boosting through the lightgbm or xgboost package, behind the same protocol as glm_fit(): stats::predict() gives means, predict_distribution() joint draws, and compare_models() and cross_validate() take it like any other fit.

Usage

booster_fit(
  formula,
  data,
  family = "poisson",
  engine = c("lightgbm", "xgboost"),
  power = NULL,
  n_rounds = 200,
  learning_rate = 0.05,
  params = list(),
  n_boot = 0,
  seed = 0,
  offset = NULL,
  weights = NULL,
  alpha = NULL,
  dispersion_model = FALSE
)

Arguments

formula

A model formula; offset(...) terms are honoured.

data

A data frame.

family

"poisson", "gamma", "tweedie" (needs power), "gaussian" or "quantile" (needs alpha); a log link for the first three.

engine

"lightgbm" or "xgboost".

power

Tweedie variance power in (1, 2).

n_rounds

Boosting rounds.

learning_rate

Shrinkage per round.

params

A named list of further engine parameters (num_leaves, max_depth, monotone_constraints, ...), passed as they are.

n_boot

Bootstrap refits for parameter uncertainty in predict_distribution(); 0 gives process noise only.

seed

Seeds the engine and the bootstrap resamples; R's own generator state is left as it was.

offset

Optional offset added to any offset() terms.

weights

Optional prior weights.

alpha

Quantile level in (0, 1), for family = "quantile".

dispersion_model

Fit a dispersion per row (gamma, Tweedie and Gaussian families).

Value

A booster_model with properties model (the engine's booster), boots (the refits), base, dispersion (1 for the Poisson, otherwise Pearson's estimate sum(w (y - mu)^2 / V(mu)) / n, with no degrees-of-freedom correction since trees have no fixed parameter count), family, engine and power.

Details

  • The formula's offset() terms and offset are the engine's starting score (lightgbm init_score, xgboost base_margin) on the link scale: log exposure for a frequency. A constant base is added so the trees start from the weighted mean response, as the engines' own boost-from-average does without an offset: log(sum(w y) / sum(w exp(offset))) for the log link, sum(w (y - offset)) / sum(w) for the identity.

  • The design is R's model.matrix() without the intercept, so factors enter as treatment-coded columns.

  • xgboost computes margins in single precision, so its means agree with a double-precision calculation to about 1e-6 relative.

  • predict_distribution() draws each row's response from the family around its mean (process noise); with n_boot bootstrap refits, each simulation also uses one refit's means (parameter uncertainty).

  • dispersion_model = TRUE adds a distributional head: a second booster (gamma objective, log link) fitted to the Pearson residuals w (y - mu)^2 / V(mu), whose expectation is each row's dispersion. The residuals come from 5-fold cross-fitted means, so a mean model that fits the training rows closely does not shrink them. predict_distribution() then uses one dispersion per row (predict_dispersion()), as a double GLM does (Smyth, 1989).

  • family = "quantile" fits the alpha quantile instead of the mean (lightgbm quantile, xgboost reg:quantileerror), starting from the weighted alpha quantile of the response; it takes no offset. Score it with pinball_loss(); predict_quantiles() puts several fits together as non-crossing quantile sets.

The lightgbm and xgboost packages are optional: install the one you use.

Examples

if (requireNamespace("lightgbm", quietly = TRUE)) {
  set.seed(1)
  d <- data.frame(x = rep(1:10 / 10, 20), exposure = rep(c(0.5, 1), 100))
  d$claims <- rpois(200, d$exposure * (0.1 + 0.3 * d$x))
  m <- booster_fit(claims ~ x + offset(log(exposure)), d, n_rounds = 20)
  head(predict(m, d))
}
#> [1] 0.09450206 0.18900412 0.09450206 0.24074162 0.11351360 0.23349051