| Safe Haskell | None |
|---|---|
| Language | Haskell2010 |
DataFrame.Operations.Aggregation
Synopsis
- aggregate :: [NamedExpr] -> GroupedDataFrame -> DataFrame
- computeRowHashes :: [Int] -> DataFrame -> Vector Int
- distinct :: DataFrame -> DataFrame
- selectIndices :: Vector Int -> DataFrame -> DataFrame
- interpretNamed :: GroupedDataFrame -> NamedExpr -> Column
- groupBy :: [Text] -> DataFrame -> GroupedDataFrame
- buildRowToGroup :: Int -> Vector Int -> Vector Int -> Vector Int
- changingPoints :: Vector (Int, Int) -> Vector Int
Documentation
aggregate :: [NamedExpr] -> GroupedDataFrame -> DataFrame Source #
Aggregate a grouped dataframe using the expressions given. All ungrouped columns will be dropped.
computeRowHashes :: [Int] -> DataFrame -> Vector Int Source #
Per-row key hash over the selected key columns. Delegates to the shared
computeRowHashesIO kernel, which forks over contiguous row ranges for large
frames (the hashing of a wide 1e7-row text/factor join key dominates that join)
and is bit-for-bit identical to a single sequential pass at any capability count.
interpretNamed :: GroupedDataFrame -> NamedExpr -> Column Source #
The fall-back path: evaluate one named aggregation via the interpreter.
groupBy :: [Text] -> DataFrame -> GroupedDataFrame #
O(k * n) group the dataframe by the given key columns, bucketing rows with an
open-addressing hash table that re-verifies keys on each hash hit. Groups are
numbered in first-appearance order; valueIndices/offsets follow by counting sort.