dataframe-core-2.4.0.0: Core data structures for the dataframe library.
Safe HaskellNone
LanguageHaskell2010

DataFrame.Internal.RowHash

Description

Row-hash kernels with a parallel driver, feeding grouping and the join build/probe. Each row's hash depends only on its own bytes, so hashing disjoint ranges in parallel is race-free and bit-identical to the sequential pass.

Synopsis

Documentation

computeRowHashesIO :: Int -> [Column] -> IO (Vector Int) Source #

Compute the per-row key hash over the selected key columns of an n-row frame. Forks one worker per capability over disjoint row ranges when the row count justifies it, else hashes the single full range; output is capability-independent.

hashRowRange :: IOVector Int -> Int -> Int -> [Column] -> IO () Source #

Mix every selected column over the row range [lo, hi) into mv, seeding each slot with fnvOffset. Must match the sequential grouping hash byte-for-byte so grouping and joins bucket identically.

parRowHashThreshold :: Int Source #

At least this many rows make the fork/coordination overhead of the parallel hash worth it. Below it the sequential single range is used. Matches the grouping/join parallel thresholds so the whole pipeline switches together.