dataframe-learn-2.4.1.0: Interpretable, expression-returning machine learning for the dataframe ecosystem.
Safe HaskellNone
LanguageHaskell2010

DataFrame.LinearSolver

Description

Proximal-gradient (FISTA) solver for L1/L2-regularized generalized linear models. fitL1Logistic is the binary logistic split solver; fitProx generalizes it to any SmoothLoss. Features are standardized internally.

Synopsis

Model

data LinearModel Source #

A fitted linear classifier: predicts the positive class when sum (weights .* features) + intercept > 0. Weights of exactly 0 mark features dropped by the L1 penalty (filtered out by modelToExpr).

Constructors

LinearModel 

Fields

Instances

Instances details
Show LinearModel Source # 
Instance details

Defined in DataFrame.LinearSolver

Eq LinearModel Source # 
Instance details

Defined in DataFrame.LinearSolver

Configuration

data SolverConfig Source #

Hyper-parameters for the FISTA solver.

Constructors

SolverConfig 

Fields

  • scL1Lambda :: !Double

    Strength of the L1 penalty on weights (intercept is not regularized).

  • scL2Lambda :: !Double

    Strength of the L2 penalty (λ₂/2)·|w|² (Elastic Net; Zou & Hastie 2005). Combined with scL1Lambda this is the elastic-net objective; 0 reduces the solver to pure L1.

  • scMaxIter :: !Int

    Maximum number of FISTA iterations.

  • scTol :: !Double

    Convergence tolerance on the weight delta (L-inf norm).

  • scSampleWeights :: !(Maybe (Vector Double))

    Optional per-row sample weights, length n (Nothing is uniform). Weights should have mean 1 (i.e. Σ w_i = N) so the Lipschitz bound stays valid; see fitLinearCandidate for the class-balanced construction.

Instances

Instances details
Show SolverConfig Source # 
Instance details

Defined in DataFrame.LinearSolver

Eq SolverConfig Source # 
Instance details

Defined in DataFrame.LinearSolver

Solvers

fitL1Logistic :: SolverConfig -> Vector (Vector Double) -> Vector Double -> Vector Text -> LinearModel Source #

Fit L1-regularized binary logistic regression by FISTA. Rows are feature vectors of equal length; labels are in {-1,+1}. Features are standardized internally and weights de-standardized, so the model applies to raw values.

fitProx :: SmoothLoss -> SolverConfig -> Vector (Vector Double) -> Vector Double -> Vector Text -> LinearModel Source #

Fit any SmoothLoss with the elastic-net proximal-gradient engine. The Lipschitz constant uses the spectral norm of the standardized Gram matrix (power iteration), tight for squared and squared-hinge losses.

Expr conversion

modelToExpr :: LinearModel -> Expr Bool Source #

Convert a fitted model to an 'Expr Bool' over its feature columns, dropping zero-weight features. With no non-zero weights it returns the constant Lit (intercept > 0).

Internals (exposed for testing)

standardize :: Vector (Vector Double) -> (Vector (Vector Double), Vector Double, Vector Double, Vector Double) Source #

Standardize each column to zero mean and unit variance, also returning (means, stds, variances). Near-constant columns get std 1; callers use the raw variances to detect and drop them (see fitL1Logistic).

columnStats :: Vector (Vector Double) -> (Vector Double, Vector Double, Vector Double) Source #

Per-column (means, stds, variances) of a feature matrix. Cheaper than standardize when only the statistics are needed. unsafeIndex within is safe: all rows share width d.

softThreshold :: Double -> Double -> Double Source #

Proximal operator for the L1 norm: shrink v toward zero by lambda, clamping at zero.

sigmoid :: Double -> Double Source #

Numerically stable logistic sigmoid.

dotProduct :: Vector Double -> Vector Double -> Double Source #

Dot product of two unboxed vectors. Caller must ensure equal length; lengths are not checked.