| Safe Haskell | None |
|---|---|
| Language | Haskell2010 |
DataFrame.Operations.Join
Synopsis
- data JoinType
- = INNER
- | LEFT
- | RIGHT
- | FULL_OUTER
- join :: JoinType -> [Text] -> DataFrame -> DataFrame -> DataFrame
- innerJoin :: [Text] -> DataFrame -> DataFrame -> DataFrame
- leftJoin :: [Text] -> DataFrame -> DataFrame -> DataFrame
- rightJoin :: [Text] -> DataFrame -> DataFrame -> DataFrame
- fullOuterJoin :: [Text] -> DataFrame -> DataFrame -> DataFrame
- buildHashColumn :: [Text] -> DataFrame -> Vector Int
- buildCompactIndex :: Vector Int -> CompactIndex
- hashProbeKernel :: CompactIndex -> Vector Int -> (Vector Int, Vector Int)
- hashInnerKernel :: Vector Int -> Vector Int -> (Vector Int, Vector Int)
- hashLeftKernel :: Vector Int -> Vector Int -> (Vector Int, Vector Int)
- innerKernel :: Bool -> Vector Int -> Vector Int -> (Vector Int, Vector Int)
- parInnerKernel :: Vector Int -> Vector Int -> (Vector Int, Vector Int)
- parLeftKernel :: Vector Int -> Vector Int -> (Vector Int, Vector Int)
- assembleInner :: Set Text -> DataFrame -> DataFrame -> Vector Int -> Vector Int -> DataFrame
- assembleLeft :: Set Text -> DataFrame -> DataFrame -> Vector Int -> Vector Int -> DataFrame
Join types
Equivalent to SQL join types.
Constructors
| INNER | |
| LEFT | |
| RIGHT | |
| FULL_OUTER |
Joins
join :: JoinType -> [Text] -> DataFrame -> DataFrame -> DataFrame Source #
Join two dataframes using SQL join semantics.
innerJoin :: [Text] -> DataFrame -> DataFrame -> DataFrame Source #
Performs an inner join on two dataframes using the specified key columns. Returns only rows where the key values exist in both dataframes.
Example
ghci> df = D.fromNamedColumns [("key", D.fromList [K0, K1, K2, K3]), (A, D.fromList [A0, A1, A2, A3])]
ghci> other = D.fromNamedColumns [("key", D.fromList [K0, K1, K2]), (B, D.fromList [B0, B1, B2])]
ghci> D.innerJoin ["key"] df other
-----------------
key | A | B
------|-----|----
Text | Text| Text
------|-----|----
K0 | A0 | B0
K1 | A1 | B1
K2 | A2 | B2
leftJoin :: [Text] -> DataFrame -> DataFrame -> DataFrame Source #
Performs a left join on two dataframes using the specified key columns. Returns all rows from the left dataframe, with matching rows from the right dataframe. Non-matching rows will have Nothing/null values for columns from the right dataframe.
Example
ghci> df = D.fromNamedColumns [("key", D.fromList [K0, K1, K2, K3]), (A, D.fromList [A0, A1, A2, A3])]
ghci> other = D.fromNamedColumns [("key", D.fromList [K0, K1, K2]), (B, D.fromList [B0, B1, B2])]
ghci> D.leftJoin ["key"] df other
------------------------
key | A | B
------|-----|----------
Text | Text| Maybe Text
------|-----|----------
K0 | A0 | Just B0
K1 | A1 | Just B1
K2 | A2 | Just B2
K3 | A3 | Nothing
rightJoin :: [Text] -> DataFrame -> DataFrame -> DataFrame Source #
Performs a right join on two dataframes using the specified key columns. Returns all rows from the right dataframe, with matching rows from the left dataframe. Non-matching rows will have Nothing/null values for columns from the left dataframe.
Example
ghci> df = D.fromNamedColumns [("key", D.fromList [K0, K1, K2, K3]), (A, D.fromList [A0, A1, A2, A3])]
ghci> other = D.fromNamedColumns [("key", D.fromList [K0, K1]), (B, D.fromList [B0, B1])]
ghci> D.rightJoin ["key"] df other
-----------------
key | A | B
------|-----|----
Text | Text| Text
------|-----|----
K0 | A0 | B0
K1 | A1 | B1
Low-level join kernels
Reused by the lazy executor and the parallel-join tests; not a
stable public API (candidates for a future Join.Internal split).
buildHashColumn :: [Text] -> DataFrame -> Vector Int Source #
Compute hashes for the given key column names in a DataFrame.
buildCompactIndex :: Vector Int -> CompactIndex Source #
Build a compact index from a vector of row hashes: sort by hash, scan for contiguous runs, insert each run into an open-addressing table. Capacity is sized for the worst case (every row distinct) so building never resizes.
Arguments
| :: CompactIndex | Built once from the full right/build side. |
| -> Vector Int | Probe hashes (one batch). |
| -> (Vector Int, Vector Int) |
Probe one batch of rows against a pre-built CompactIndex, returning
(probeExpandedIxs, buildExpandedIxs). Unlike hashInnerKernel it neither
builds the index nor guards cross-product size — the caller sizes batches.
hashLeftKernel :: Vector Int -> Vector Int -> (Vector Int, Vector Int) Source #
Hash-based left join kernel, returning (leftExpandedIndices,
rightExpandedIndices) with -1 marking unmatched right rows. Grows its output
buffer dynamically rather than pre-allocating the full cross product.
innerKernel :: Bool -> Vector Int -> Vector Int -> (Vector Int, Vector Int) Source #
Select the inner-join kernel: parallel chunked-probe hash join when the caller decided the probe side is large enough and multi-core, otherwise the sequential single-threaded hash kernel.
parInnerKernel :: Vector Int -> Vector Int -> (Vector Int, Vector Int) Source #
Parallel inner-join kernel: build the CompactIndex on buildHashes once,
then probe probeHashes in parallel. Output is bit-for-bit identical to
hashInnerKernel (probe-row order preserved).
parLeftKernel :: Vector Int -> Vector Int -> (Vector Int, Vector Int) Source #
Parallel left-join kernel: build the CompactIndex on rightHashes once,
then probe leftHashes in parallel. Bit-for-bit identical output to
hashLeftKernel (left-row order preserved, -1 sentinel for misses).