dataframe-operations-2.4.0.0: Column operations, expression DSL, and statistics for the dataframe ecosystem.
Safe HaskellNone
LanguageHaskell2010

DataFrame.Typed.Sampling

Description

Typed sampling and splitting. All operations are schema-preserving: they change which rows are present, never the columns, so every result reuses the input schema cols.

Synopsis

Documentation

randomSplit :: forall g (cols :: [(Symbol, Type)]). RandomGen g => g -> Double -> TypedDataFrame cols -> (TypedDataFrame cols, TypedDataFrame cols) Source #

Split rows into two DataFrames by a fraction.

kFolds :: forall g (cols :: [(Symbol, Type)]). RandomGen g => g -> Int -> TypedDataFrame cols -> [TypedDataFrame cols] Source #

Partition rows into k folds.

selectRows :: forall (cols :: [(Symbol, Type)]). [Int] -> TypedDataFrame cols -> TypedDataFrame cols Source #

Select rows by index. | This may fail if the indices are out of bounds; | use with caution or use filter to select rows by a predicate instead.

stratifiedSample :: forall g a (cols :: [(Symbol, Type)]). (SplittableGen g, Columnable a) => g -> Double -> TExpr cols a -> TypedDataFrame cols -> TypedDataFrame cols Source #

Sample a fraction of rows, preserving the distribution of a strata column.

stratifiedSplit :: forall g a (cols :: [(Symbol, Type)]). (SplittableGen g, Columnable a) => g -> Double -> TExpr cols a -> TypedDataFrame cols -> (TypedDataFrame cols, TypedDataFrame cols) Source #

Split rows by a fraction, preserving the distribution of a strata column.