models.GlmFit

A fitted GLM, from Glm.fit.

Usage

models.GlmFit()

Attributes

Name Description
aic AIC, -2 loglik + 2 p.
coefficients Estimated coefficients.
covariance Covariance of the coefficients, as a list of rows.
deviance Residual deviance.
df_resid Residual degrees of freedom.
dispersion Dispersion.
fitted Fitted means on the training data.
input_hash Hash of the training data (design, offset, weights, response).
iterations IRLS iterations used.
log_likelihood Log-likelihood.
names Coefficient names.
null_deviance Deviance of the intercept-and-offset model.
p_values Two-sided p-values (normal for a fixed dispersion, Student’s t when
std_errors Standard errors.

aic

AIC, -2 loglik + 2 p.

aic: float


coefficients

Estimated coefficients.

coefficients: list[float]


covariance

Covariance of the coefficients, as a list of rows.

covariance: list[list[float]]


deviance

Residual deviance.

deviance: float


df_resid

Residual degrees of freedom.

df_resid: float


dispersion

Dispersion.

dispersion: float


fitted

Fitted means on the training data.

fitted: list[float]


input_hash

Hash of the training data (design, offset, weights, response).

input_hash: str


iterations

IRLS iterations used.

iterations: int


log_likelihood

Log-likelihood.

log_likelihood: float


names

Coefficient names.

names: list[str]


null_deviance

Deviance of the intercept-and-offset model.

null_deviance: float


p_values

Two-sided p-values (normal for a fixed dispersion, Student’s t when

p_values: list[float]

it is estimated).


std_errors

Standard errors.

std_errors: list[float]

Methods

Name Description
from_json() Reads an artifact written by to_json.
predict() Expected response for each row.
predict_distribution() Joint predictive distribution across the rows, with parameter and
robust_covariance() Sandwich (heteroskedasticity- or cluster-robust) covariance of the
robust_std_errors() Square roots of the diagonal of robust_covariance, with the same
to_json() The fit as a versioned JSON artifact: spec, estimates, covariance,

from_json()

Reads an artifact written by to_json.

Usage

from_json(text)
Parameters
text: str
Returns
GlmFit
Raises
ValueError
For malformed JSON, another format, a newer format version or inconsistent fields.

predict()

Expected response for each row.

Usage

predict(design)
Parameters
design: Design
Same columns as the training design.
Returns
list of float

predict_distribution()

Joint predictive distribution across the rows, with parameter and

Usage

predict_distribution(design, n_sims, seed, parameters="normal")

process uncertainty, keyed row = 0, 1, ....

Parameters
design: Design
n_sims: int
seed: int
parameters: (normal, mean_preserving, fixed) = "normal"
How the coefficients are drawn: beta ~ N(beta_hat, Sigma) (through a log link the draws’ mean is mu_hat * exp(x' Sigma x / 2)); the same with each row’s linear predictor shifted so its draws average the fitted mean exactly (log or identity link); or fixed at beta_hat (process uncertainty only).
Returns
PredictiveDistribution

robust_covariance()

Sandwich (heteroskedasticity- or cluster-robust) covariance of the

Usage

robust_covariance(design, y, kind="HC0", groups=None)

coefficients, as statsmodels’ cov_type="HC0" and "cluster".

It stays valid when the variance function or dispersion is wrong, as long as the mean is right. The dispersion cancels. For a non-canonical link it uses the observed information, as statsmodels does (R’s sandwich uses the expected).

Parameters
design: Design

The design the model was fitted on.

y: list of float

The response the model was fitted on.

kind: (HC0, HC1, cluster) = "HC0"

"HC1" scales HC0 by n / (n - p); "cluster" sums the scores within each cluster and scales by G / (G - 1) * (n - 1) / (n - p).

groups: list of int or str = None
One cluster label per row, for kind="cluster": a policy or an event, say.
Returns
list of list of float
Examples
>>> from prospicio.models import Design, Glm
>>> d = Design([[1.0] * 6, [0.0, 0.0, 0.0, 1.0, 1.0, 1.0]], ["(Intercept)", "x"])
>>> y = [1.0, 2.0, 6.0, 1.0, 4.0, 2.0]
>>> fit = Glm("poisson").fit(d, y)
>>> round(fit.robust_covariance(d, y)[0][0] * 81, 10)

14.0

>>> se = fit.robust_std_errors(d, y, "cluster", groups=[1, 1, 2, 2, 3, 3])

robust_std_errors()

Square roots of the diagonal of robust_covariance, with the same

Usage

robust_std_errors(design, y, kind="HC0", groups=None)

arguments.

Parameters
design: Design
y: list of float
kind: (HC0, HC1, cluster) = "HC0"
groups: list of int or str = None
Returns
list of float

to_json()

The fit as a versioned JSON artifact: spec, estimates, covariance,

Usage

to_json()

fit statistics, fitted values and provenance (crate version and a hash of the training data). GlmFit.from_json reads it back exactly; pickling uses it too.

Returns
str
Examples
>>> import pickle
>>> from prospicio.models import Design, Glm, GlmFit
>>> d = Design([[1.0] * 4, [0.0, 1.0, 2.0, 3.0]], ["(Intercept)", "x"])
>>> fit = Glm("poisson").fit(d, [1.0, 2.0, 2.0, 5.0])
>>> GlmFit.from_json(fit.to_json()).coefficients == fit.coefficients

True

>>> pickle.loads(pickle.dumps(fit)).input_hash == fit.input_hash

True