| Safe Haskell | None |
|---|---|
| Language | Haskell2010 |
DataFrame.Internal.GroupingDirect
Description
Low-cardinality direct-indexed grouping fast path: when the key is a single
clean unboxed Int column of small value range, the value itself indexes a dense
accumulator (no hashing/probing). Emits groups in ascending value order.
Synopsis
- directGroupThreshold :: Int
- tryDirectGroupColumn :: Column -> Maybe DirectGrouping
- data DirectGrouping = DirectGrouping {
- dgRowToGroup :: !(Vector Int)
- dgValueIndices :: !(Vector Int)
- dgOffsets :: !(Vector Int)
- dgNGroups :: !Int
Documentation
directGroupThreshold :: Int Source #
Largest key value RANGE (max - min + 1) the direct grouping path accepts. A
2^20-slot histogram is 8MB; the low-cardinality questions sit far below it
(id4 range 100, id6 range 1e5). Wider ranges fall back to the hash group-by.
tryDirectGroupColumn :: Column -> Maybe DirectGrouping Source #
Take the direct path if the (single) key column is a clean non-null unboxed
Int column with a small value range. Returns Nothing to fall back to the
hash group-by on anything else (boxed/text keys, nullable, wide ranges, empty).
data DirectGrouping Source #
The grouping layout the hash path also produces: rowToGroup, the
group-sorted valueIndices, the offsets prefix array, and the group count.
Constructors
| DirectGrouping | |
Fields
| |