| Safe Haskell | None |
|---|---|
| Language | Haskell2010 |
DataFrame.Internal.RowHash
Description
Row-hash kernels with a parallel driver, feeding grouping and the join build/probe. Each row's hash depends only on its own bytes, so hashing disjoint ranges in parallel is race-free and bit-identical to the sequential pass.
Synopsis
- computeRowHashesIO :: Int -> [Column] -> IO (Vector Int)
- hashRowRange :: IOVector Int -> Int -> Int -> [Column] -> IO ()
- parRowHashThreshold :: Int
Documentation
computeRowHashesIO :: Int -> [Column] -> IO (Vector Int) Source #
Compute the per-row key hash over the selected key columns of an n-row
frame. Forks one worker per capability over disjoint row ranges when the row
count justifies it, else hashes the single full range; output is capability-independent.
hashRowRange :: IOVector Int -> Int -> Int -> [Column] -> IO () Source #
Mix every selected column over the row range [lo, hi) into mv, seeding
each slot with fnvOffset. Must match the sequential grouping hash byte-for-byte
so grouping and joins bucket identically.
parRowHashThreshold :: Int Source #
At least this many rows make the fork/coordination overhead of the parallel hash worth it. Below it the sequential single range is used. Matches the grouping/join parallel thresholds so the whole pipeline switches together.