dataframe-learn-2.4.1.0: Interpretable, expression-returning machine learning for the dataframe ecosystem.
Safe HaskellNone
LanguageHaskell2010

DataFrame.ModelSelection

Description

Cross-validation and grid search for hyperparameter tuning. The model fitters have heterogeneous types, so these helpers are parameterized by a user-supplied train -> test -> score closure; the search maximizes the mean cross-validated score (use a negated error metric to minimize). Splitting reuses the deterministic kFolds from dataframe-operations.

Synopsis

Documentation

crossValScore :: Int -> Int -> (DataFrame -> DataFrame -> Double) -> DataFrame -> [Double] Source #

Per-fold scores from k-fold cross-validation. scoreFn train test fits on the training rows and returns a score on the held-out fold.

crossValidate :: Int -> Int -> Metric -> Expr Double -> (DataFrame -> Expr Double) -> DataFrame -> [Double] Source #

scikit-learn cross_val_score: fit a model on each training fold and score its prediction expression against a truth column on the held-out fold.

fitPredict train fits on the training frame and returns the prediction expression; truth is the target column. Returns the per-fold metric values.

crossValidate 5 0 rmse (F.col @Double "target")
  (\tr -> predict (fit defaultLinearConfig (F.col @Double "target") tr)) df

data GridSearchResult c Source #

The outcome of a grid search: the best config, its score, and all results.

Constructors

GridSearchResult 

Fields

Instances

Instances details
Show c => Show (GridSearchResult c) Source # 
Instance details

Defined in DataFrame.ModelSelection

gridSearch :: Int -> Int -> [c] -> (c -> DataFrame -> DataFrame -> Double) -> DataFrame -> GridSearchResult c Source #

Search configurations by mean cross-validated score, returning the maximizer. scoreFn cfg train test fits cfg on train and scores on test.