models.Comparison

The result of compare: every model’s score on every metric and

Usage

Source

models.Comparison(
    models,
    metrics,
    split_scores,
)

split.

Attributes

models: list of str
metrics: list of str
split_scores: dict
split_scores[(model, metric)] is the list of per-split scores.

Methods

Name Description
best() The model with the lowest mean metric.
difference_std_error() Standard error of the per-split difference from the best model.
mean() Mean score over the splits.
std_error() Standard deviation of the split scores over the square root of
table() One row per model and metric: model, metric, mean,

best()

The model with the lowest mean metric.

Usage

Source

best(metric)

difference_std_error()

Standard error of the per-split difference from the best model.

Usage

Source

difference_std_error(model, metric)

The splits are shared, so the paired difference is much less noisy than either mean: a model within about two of these of the best is not clearly worse.


mean()

Mean score over the splits.

Usage

Source

mean(model, metric)

std_error()

Standard deviation of the split scores over the square root of

Usage

Source

std_error(model, metric)

their number.


table()

One row per model and metric: model, metric, mean,

Usage

Source

table()

std_error and difference_std_error, as a list of dicts (pass it to pandas.DataFrame for a frame).