Skip to contents

Fits every model on each split's training rows and scores it on the test rows with every metric: one table across engines, comparing like with like. The paired difference_std_error (the standard error of each split's score minus the best model's) is much less noisy than either mean; a model within about two of them of the best is not clearly worse.

Usage

compare_models(models, data, splits, scores)

Arguments

models

A named list of fit(train) functions, each returning a fitted model (from glm_fit(), gam_fit(), elastic_net_fit() or anything score accepts).

data

A data frame.

splits

Splits from k_fold(), group_k_fold() or time_ordered().

scores

A named list of score(model, test) losses (lower is better).

Value

A data frame with one row per model and metric: model, metric, mean, std_error (the split scores' standard deviation over the square root of their number) and difference_std_error, with attribute split_scores, an array indexed by model, metric and split.

Examples

d <- data.frame(x = 1:40 / 10)
d$y <- 1 + 2 * d$x + sin(1:40)
mse <- function(m, test) mean((test$y - predict(m, test))^2)
compare_models(
  list(linear = function(train) glm_fit(y ~ x, train, family = "gaussian"),
       flat = function(train) glm_fit(y ~ 1, train, family = "gaussian")),
  d, k_fold(nrow(d), 4, seed = 1), list(mse = mse)
)
#>    model metric     mean  std_error difference_std_error
#> 1 linear    mse 0.606725 0.08202587             0.000000
#> 2   flat    mse 5.988080 1.32917848             1.294539