Fit gradient-boosted trees
booster_fit.RdGradient boosting through the lightgbm or xgboost package, behind the
same protocol as glm_fit(): stats::predict() gives means,
predict_distribution() joint draws, and compare_models() and
cross_validate() take it like any other fit.
Arguments
- formula
A model formula;
offset(...)terms are honoured.- data
A data frame.
- family
"poisson","gamma","tweedie"(needspower),"gaussian"or"quantile"(needsalpha); a log link for the first three.- engine
"lightgbm"or"xgboost".- power
Tweedie variance power in
(1, 2).- n_rounds
Boosting rounds.
- learning_rate
Shrinkage per round.
- params
A named list of further engine parameters (
num_leaves,max_depth,monotone_constraints, ...), passed as they are.- n_boot
Bootstrap refits for parameter uncertainty in
predict_distribution(); 0 gives process noise only.- seed
Seeds the engine and the bootstrap resamples; R's own generator state is left as it was.
- offset
Optional offset added to any
offset()terms.- weights
Optional prior weights.
- alpha
Quantile level in
(0, 1), forfamily = "quantile".- dispersion_model
Fit a dispersion per row (gamma, Tweedie and Gaussian families).
Value
A booster_model with properties model (the engine's
booster), boots (the refits), base, dispersion (1 for the
Poisson, otherwise Pearson's estimate sum(w (y - mu)^2 / V(mu)) / n,
with no degrees-of-freedom correction since trees have no fixed
parameter count), family, engine and power.
Details
The formula's
offset()terms andoffsetare the engine's starting score (lightgbminit_score, xgboostbase_margin) on the link scale: log exposure for a frequency. A constantbaseis added so the trees start from the weighted mean response, as the engines' own boost-from-average does without an offset:log(sum(w y) / sum(w exp(offset)))for the log link,sum(w (y - offset)) / sum(w)for the identity.The design is R's
model.matrix()without the intercept, so factors enter as treatment-coded columns.xgboost computes margins in single precision, so its means agree with a double-precision calculation to about 1e-6 relative.
predict_distribution()draws each row's response from the family around its mean (process noise); withn_bootbootstrap refits, each simulation also uses one refit's means (parameter uncertainty).dispersion_model = TRUEadds a distributional head: a second booster (gamma objective, log link) fitted to the Pearson residualsw (y - mu)^2 / V(mu), whose expectation is each row's dispersion. The residuals come from 5-fold cross-fitted means, so a mean model that fits the training rows closely does not shrink them.predict_distribution()then uses one dispersion per row (predict_dispersion()), as a double GLM does (Smyth, 1989).family = "quantile"fits thealphaquantile instead of the mean (lightgbmquantile, xgboostreg:quantileerror), starting from the weightedalphaquantile of the response; it takes no offset. Score it withpinball_loss();predict_quantiles()puts several fits together as non-crossing quantile sets.
The lightgbm and xgboost packages are optional: install the one you use.
Examples
if (requireNamespace("lightgbm", quietly = TRUE)) {
set.seed(1)
d <- data.frame(x = rep(1:10 / 10, 20), exposure = rep(c(0.5, 1), 100))
d$claims <- rpois(200, d$exposure * (0.1 + 0.3 * d$x))
m <- booster_fit(claims ~ x + offset(log(exposure)), d, n_rounds = 20)
head(predict(m, d))
}
#> [1] 0.09450206 0.18900412 0.09450206 0.24074162 0.11351360 0.23349051