dataframe-operations-2.4.0.0: Column operations, expression DSL, and statistics for the dataframe ecosystem.
Safe HaskellNone
LanguageHaskell2010

DataFrame.Operations.Aggregation

Synopsis

Documentation

aggregate :: [NamedExpr] -> GroupedDataFrame -> DataFrame Source #

Aggregate a grouped dataframe using the expressions given. All ungrouped columns will be dropped.

computeRowHashes :: [Int] -> DataFrame -> Vector Int Source #

Per-row key hash over the selected key columns. Delegates to the shared computeRowHashesIO kernel, which forks over contiguous row ranges for large frames (the hashing of a wide 1e7-row text/factor join key dominates that join) and is bit-for-bit identical to a single sequential pass at any capability count.

distinct :: DataFrame -> DataFrame Source #

Filter out all non-unique values in a dataframe.

interpretNamed :: GroupedDataFrame -> NamedExpr -> Column Source #

The fall-back path: evaluate one named aggregation via the interpreter.

groupBy :: [Text] -> DataFrame -> GroupedDataFrame #

O(k * n) group the dataframe by the given key columns, bucketing rows with an open-addressing hash table that re-verifies keys on each hash hit. Groups are numbered in first-appearance order; valueIndices/offsets follow by counting sort.

buildRowToGroup :: Int -> Vector Int -> Vector Int -> Vector Int #

Build the rowToGroup lookup vector from valueIndices and offsets. rowToGroup[i] = k means row i belongs to group k.