| Safe Haskell | None |
|---|---|
| Language | Haskell2010 |
DataFrame.Internal.GroupingPar
Description
Parallel partitioned group-by: rows are counting-sorted into partitions by the
top hash bits, then one task per capability groups its partitions independently.
Output is bit-for-bit identical to the sequential groupBy.
Synopsis
- parallelAssignGroups :: Int -> Vector Int -> (Int -> Int -> Bool) -> IO (Vector Int, Vector Int, Vector Int)
- shouldParallelize :: Int -> Bool
- parThreshold :: Int
- numPartitionsFor :: Int -> Int
Documentation
parallelAssignGroups :: Int -> Vector Int -> (Int -> Int -> Bool) -> IO (Vector Int, Vector Int, Vector Int) Source #
Parallel group assignment. parallelAssignGroups n hashes eqRow returns
(rowToGroup, valueIndices, offsets) in canonical group order. eqRow a b must
report whether rows a and b share all key columns (null-aware).
shouldParallelize :: Int -> Bool Source #
Whether groupBy should take the parallel path: more than one capability
and at least parThreshold rows.
parThreshold :: Int Source #
Below this many rows the partition/fork overhead is not worth it; groupBy
uses its sequential ST path instead.
numPartitionsFor :: Int -> Int Source #
Number of partitions: a power of two, at least 4 * caps (P >> cores for
skew tolerance), floored at 256.