dataframe-core-2.2.0.0: Core data structures for the dataframe library.
Safe HaskellNone
LanguageHaskell2010

DataFrame.Internal.PackedText

Description

Packed-text payload + byte-slice primitives. A PackedTextData shares one UTF-8 byte buffer across all rows of a string column, with n+1 row offsets, so no per-row Text header is materialized until decode is demanded.

Synopsis

Documentation

data PackedTextData Source #

A shared UTF-8 byte buffer plus n+1 row offsets (base row r spans bytes [offsets!r, offsets!(r+1))); validity lives in the column's bitmap. ptSel is an optional selection layer letting a gatherjoinsort result share the buffer.

Constructors

PackedTextData 

Fields

mkPackedContiguous :: Array -> Vector Int -> PackedTextData Source #

Build a contiguous packed payload (no selection): the freeze-path shape.

packedGather :: Vector Int -> PackedTextData -> PackedTextData Source #

Reindex a packed payload by a selection vector, sharing the byte buffer; logical row i becomes base row indices!i. A negative or out-of-range index decodes to the empty slice. Composes with an existing selection.

packedTake :: Int -> PackedTextData -> PackedTextData Source #

Take the first k logical rows, sharing the byte buffer via a capped selection layer. O(k), no byte copy or decode — cheap take/display on a large packed column.

packedRowOffsetVec :: PackedTextData -> Maybe (Array, Vector Int) Source #

The shared buffer + contiguous n+1 offsets when the payload is the unselected base; a selected (gathered) payload returns Nothing (its rows are non-contiguous). Lets the boxed-Text fallback take the fast contiguous path.

packedLength :: PackedTextData -> Int Source #

Row count: length sel when selected, else length offsets - 1.

packedSlice :: PackedTextData -> Int -> (Array, Int, Int) Source #

Raw byte slice for logical row i: (buffer, offset, length). The hot accessor.

packedIndexText :: PackedTextData -> Int -> Text Source #

On-demand single Text for row i, using the same validate-or-lenient decode as sliceTextVector so output is bit-identical.

sliceEqBytes :: Array -> Int -> Int -> Array -> Int -> Int -> Bool Source #

Byte-wise equality of two slices. UTF-8 is injective on valid scalar sequences and lenient decode is deterministic, so this agrees with Text's == on the decoded values.

sliceCmpBytes :: Array -> Int -> Int -> Array -> Int -> Int -> Ordering Source #

Unsigned byte-lexicographic comparison (memcmp semantics). For well-formed UTF-8 this matches compare exactly, since UTF-8 byte order equals codepoint order for all valid scalars.