Skip to contents

The lasso (alpha = 1), ridge (alpha = 0) and everything between, for every family and link of glm_fit(), by coordinate descent inside IRLS. Minimizes glmnet's objective sum(w * d) / (2 * sum(w)) + lambda * sum(pf * ((1 - alpha) / 2 * b^2 + alpha * abs(b))), where b are the coefficients of the columns standardized to unit standard deviation (when standardize = TRUE). The intercept is not penalized; coefficients are reported on the design's scale. Results match glmnet (validation/scripts/r_glmnet.R), except that for the Gaussian glmnet divides the ridge part of its penalty by sd(y).

Usage

elastic_net_fit(
  formula,
  data,
  family = "poisson",
  link = NULL,
  alpha = 1,
  lambda = NULL,
  nlambda = 100,
  lambda_min_ratio = 1e-04,
  standardize = TRUE,
  penalty_factor = NULL,
  offset = NULL,
  weights = NULL,
  theta = NULL,
  power = NULL,
  link_power = NULL
)

Arguments

formula

A model formula; offset(...) terms are honoured.

data

A data frame.

family

"gaussian", "poisson", "gamma", "inverse_gaussian", "binomial" (response a proportion, weights the trials), "negative_binomial" (needs theta) or "tweedie" (needs power).

"identity", "log", "logit", "probit", "cloglog", "inverse", "inverse_squared" or "power" (needs link_power); NULL for the family's canonical link.

alpha

Mixing between ridge (0) and the lasso (1).

lambda

Penalty strengths; NULL for a path of nlambda values log-spaced from the smallest lambda that zeroes every coefficient down to lambda_min_ratio times it.

nlambda, lambda_min_ratio

The default path.

standardize

Penalize standardized coefficients.

penalty_factor

Optional penalty factors: a vector with one per design column, or a named vector for some columns (the rest get 1); 0 leaves a column unpenalized.

offset

Optional offset added to any offset() terms.

weights

Optional prior weights.

theta

Negative binomial theta (variance mu + mu^2 / theta).

power

Tweedie power in (1, 2).

Exponent of the power link.

Value

An elastic_net_model with properties lambda, coefficients (a matrix, one column per lambda), deviance, deviance_ratio, df (non-zero coefficients), null_deviance and alpha. Use stats::coef(), stats::predict() and predict_distribution(), each with a lambda from the path. Tune lambda and alpha with k_fold() and family_deviance().

Examples

set.seed(1)
d <- data.frame(x1 = rnorm(50), x2 = rnorm(50))
d$y <- 1 + 2 * d$x1 + rnorm(50)
m <- elastic_net_fit(y ~ x1 + x2, d, family = "gaussian", alpha = 1)
m@df[c(1, 20, 100)]
#> [1] 0 1 2
coef(m, lambda = m@lambda[20])
#> (Intercept)          x1          x2 
#>    0.879445    1.682121    0.000000