dataframe-operations-2.4.0.0: Column operations, expression DSL, and statistics for the dataframe ecosystem.
Safe HaskellNone
LanguageHaskell2010

DataFrame.Operations.Join

Synopsis

Join types

data JoinType Source #

Equivalent to SQL join types.

Constructors

INNER 
LEFT 
RIGHT 
FULL_OUTER 

Instances

Instances details
Show JoinType Source # 
Instance details

Defined in DataFrame.Operations.Join

Joins

join :: JoinType -> [Text] -> DataFrame -> DataFrame -> DataFrame Source #

Join two dataframes using SQL join semantics.

innerJoin :: [Text] -> DataFrame -> DataFrame -> DataFrame Source #

Performs an inner join on two dataframes using the specified key columns. Returns only rows where the key values exist in both dataframes.

Example

Expand
ghci> df = D.fromNamedColumns [("key", D.fromList [K0, K1, K2, K3]), (A, D.fromList [A0, A1, A2, A3])]
ghci> other = D.fromNamedColumns [("key", D.fromList [K0, K1, K2]), (B, D.fromList [B0, B1, B2])]
ghci> D.innerJoin ["key"] df other

-----------------
 key  |  A  |  B
------|-----|----
 Text | Text| Text
------|-----|----
 K0   | A0  | B0
 K1   | A1  | B1
 K2   | A2  | B2

leftJoin :: [Text] -> DataFrame -> DataFrame -> DataFrame Source #

Performs a left join on two dataframes using the specified key columns. Returns all rows from the left dataframe, with matching rows from the right dataframe. Non-matching rows will have Nothing/null values for columns from the right dataframe.

Example

Expand
ghci> df = D.fromNamedColumns [("key", D.fromList [K0, K1, K2, K3]), (A, D.fromList [A0, A1, A2, A3])]
ghci> other = D.fromNamedColumns [("key", D.fromList [K0, K1, K2]), (B, D.fromList [B0, B1, B2])]
ghci> D.leftJoin ["key"] df other

------------------------
 key  |  A  |     B
------|-----|----------
 Text | Text| Maybe Text
------|-----|----------
 K0   | A0  | Just B0
 K1   | A1  | Just B1
 K2   | A2  | Just B2
 K3   | A3  | Nothing

rightJoin :: [Text] -> DataFrame -> DataFrame -> DataFrame Source #

Performs a right join on two dataframes using the specified key columns. Returns all rows from the right dataframe, with matching rows from the left dataframe. Non-matching rows will have Nothing/null values for columns from the left dataframe.

Example

Expand
ghci> df = D.fromNamedColumns [("key", D.fromList [K0, K1, K2, K3]), (A, D.fromList [A0, A1, A2, A3])]
ghci> other = D.fromNamedColumns [("key", D.fromList [K0, K1]), (B, D.fromList [B0, B1])]
ghci> D.rightJoin ["key"] df other

-----------------
 key  |  A  |  B
------|-----|----
 Text | Text| Text
------|-----|----
 K0   | A0  | B0
 K1   | A1  | B1

Low-level join kernels

Reused by the lazy executor and the parallel-join tests; not a stable public API (candidates for a future Join.Internal split).

buildHashColumn :: [Text] -> DataFrame -> Vector Int Source #

Compute hashes for the given key column names in a DataFrame.

buildCompactIndex :: Vector Int -> CompactIndex Source #

Build a compact index from a vector of row hashes: sort by hash, scan for contiguous runs, insert each run into an open-addressing table. Capacity is sized for the worst case (every row distinct) so building never resizes.

hashProbeKernel Source #

Arguments

:: CompactIndex

Built once from the full right/build side.

-> Vector Int

Probe hashes (one batch).

-> (Vector Int, Vector Int) 

Probe one batch of rows against a pre-built CompactIndex, returning (probeExpandedIxs, buildExpandedIxs). Unlike hashInnerKernel it neither builds the index nor guards cross-product size — the caller sizes batches.

hashLeftKernel :: Vector Int -> Vector Int -> (Vector Int, Vector Int) Source #

Hash-based left join kernel, returning (leftExpandedIndices, rightExpandedIndices) with -1 marking unmatched right rows. Grows its output buffer dynamically rather than pre-allocating the full cross product.

innerKernel :: Bool -> Vector Int -> Vector Int -> (Vector Int, Vector Int) Source #

Select the inner-join kernel: parallel chunked-probe hash join when the caller decided the probe side is large enough and multi-core, otherwise the sequential single-threaded hash kernel.

parInnerKernel :: Vector Int -> Vector Int -> (Vector Int, Vector Int) Source #

Parallel inner-join kernel: build the CompactIndex on buildHashes once, then probe probeHashes in parallel. Output is bit-for-bit identical to hashInnerKernel (probe-row order preserved).

parLeftKernel :: Vector Int -> Vector Int -> (Vector Int, Vector Int) Source #

Parallel left-join kernel: build the CompactIndex on rightHashes once, then probe leftHashes in parallel. Bit-for-bit identical output to hashLeftKernel (left-row order preserved, -1 sentinel for misses).

assembleInner :: Set Text -> DataFrame -> DataFrame -> Vector Int -> Vector Int -> DataFrame Source #

Assemble the result DataFrame for an inner join from expanded index vectors.

assembleLeft :: Set Text -> DataFrame -> DataFrame -> Vector Int -> Vector Int -> DataFrame Source #

Assemble the result DataFrame for a left join. Right index vectors use -1 sentinel, gathered via gatherWithSentinel.