models.GlmFit
A fitted GLM, from Glm.fit.
Usage
models.GlmFit()Attributes
| Name | Description |
|---|---|
| aic |
AIC, -2 loglik + 2 p.
|
| coefficients | Estimated coefficients. |
| covariance | Covariance of the coefficients, as a list of rows. |
| deviance | Residual deviance. |
| df_resid | Residual degrees of freedom. |
| dispersion | Dispersion. |
| fitted | Fitted means on the training data. |
| input_hash | Hash of the training data (design, offset, weights, response). |
| iterations | IRLS iterations used. |
| log_likelihood | Log-likelihood. |
| names | Coefficient names. |
| null_deviance | Deviance of the intercept-and-offset model. |
| p_values | Two-sided p-values (normal for a fixed dispersion, Student’s t when |
| std_errors | Standard errors. |
aic
AIC, -2 loglik + 2 p.
aic: float
coefficients
Estimated coefficients.
coefficients: list[float]
covariance
Covariance of the coefficients, as a list of rows.
covariance: list[list[float]]
deviance
Residual deviance.
deviance: float
df_resid
Residual degrees of freedom.
df_resid: float
dispersion
Dispersion.
dispersion: float
fitted
Fitted means on the training data.
fitted: list[float]
input_hash
Hash of the training data (design, offset, weights, response).
input_hash: str
iterations
IRLS iterations used.
iterations: int
log_likelihood
Log-likelihood.
log_likelihood: float
names
Coefficient names.
names: list[str]
null_deviance
Deviance of the intercept-and-offset model.
null_deviance: float
p_values
Two-sided p-values (normal for a fixed dispersion, Student’s t when
p_values: list[float]
it is estimated).
std_errors
Standard errors.
std_errors: list[float]
Methods
| Name | Description |
|---|---|
| from_json() | Reads an artifact written by to_json. |
| predict() | Expected response for each row. |
| predict_distribution() | Joint predictive distribution across the rows, with parameter and |
| robust_covariance() | Sandwich (heteroskedasticity- or cluster-robust) covariance of the |
| robust_std_errors() | Square roots of the diagonal of robust_covariance, with the same |
| to_json() | The fit as a versioned JSON artifact: spec, estimates, covariance, |
from_json()
Reads an artifact written by to_json.
Usage
from_json(text)Parameters
text: str
Returns
GlmFit
Raises
ValueError- For malformed JSON, another format, a newer format version or inconsistent fields.
predict()
Expected response for each row.
Usage
predict(design)Parameters
design: Design- Same columns as the training design.
Returns
list of float
predict_distribution()
Joint predictive distribution across the rows, with parameter and
Usage
predict_distribution(design, n_sims, seed, parameters="normal")process uncertainty, keyed row = 0, 1, ....
Parameters
design: Designn_sims: intseed: intparameters: (normal, mean_preserving, fixed) = "normal"-
How the coefficients are drawn:
beta ~ N(beta_hat, Sigma)(through a log link the draws’ mean ismu_hat * exp(x' Sigma x / 2)); the same with each row’s linear predictor shifted so its draws average the fitted mean exactly (log or identity link); or fixed atbeta_hat(process uncertainty only).
Returns
PredictiveDistribution
robust_covariance()
Sandwich (heteroskedasticity- or cluster-robust) covariance of the
Usage
robust_covariance(design, y, kind="HC0", groups=None)coefficients, as statsmodels’ cov_type="HC0" and "cluster".
It stays valid when the variance function or dispersion is wrong, as long as the mean is right. The dispersion cancels. For a non-canonical link it uses the observed information, as statsmodels does (R’s sandwich uses the expected).
Parameters
design: Design-
The design the model was fitted on.
y: list of float-
The response the model was fitted on.
kind: (HC0, HC1, cluster) = "HC0"-
"HC1"scales HC0 byn / (n - p);"cluster"sums the scores within each cluster and scales byG / (G - 1) * (n - 1) / (n - p). groups: list of int or str = None-
One cluster label per row, for
kind="cluster": a policy or an event, say.
Returns
list of list of float
Examples
>>> from prospicio.models import Design, Glm
>>> d = Design([[1.0] * 6, [0.0, 0.0, 0.0, 1.0, 1.0, 1.0]], ["(Intercept)", "x"])
>>> y = [1.0, 2.0, 6.0, 1.0, 4.0, 2.0]
>>> fit = Glm("poisson").fit(d, y)
>>> round(fit.robust_covariance(d, y)[0][0] * 81, 10)14.0
>>> se = fit.robust_std_errors(d, y, "cluster", groups=[1, 1, 2, 2, 3, 3])robust_std_errors()
Square roots of the diagonal of robust_covariance, with the same
Usage
robust_std_errors(design, y, kind="HC0", groups=None)arguments.
Parameters
design: Designy: list of floatkind: (HC0, HC1, cluster) = "HC0"groups: list of int or str = None
Returns
list of float
to_json()
The fit as a versioned JSON artifact: spec, estimates, covariance,
Usage
to_json()fit statistics, fitted values and provenance (crate version and a hash of the training data). GlmFit.from_json reads it back exactly; pickling uses it too.
Returns
str
Examples
>>> import pickle
>>> from prospicio.models import Design, Glm, GlmFit
>>> d = Design([[1.0] * 4, [0.0, 1.0, 2.0, 3.0]], ["(Intercept)", "x"])
>>> fit = Glm("poisson").fit(d, [1.0, 2.0, 2.0, 5.0])
>>> GlmFit.from_json(fit.to_json()).coefficients == fit.coefficientsTrue
>>> pickle.loads(pickle.dumps(fit)).input_hash == fit.input_hashTrue