dataframe-parquet-1.5.0.0: Parquet reader and writer for the dataframe ecosystem.
Safe HaskellNone
LanguageHaskell2010

DataFrame.IO.Parquet.Page

Synopsis

Documentation

type PageDecoder a = Maybe DictVals -> Encoding -> Int -> ByteString -> Vector a Source #

A type-specific page decoder. Given the optional dictionary, the page encoding, the number of present values, and the decompressed value bytes, returns exactly nPresent values.

foldColumnPagesM :: (RandomAccess m, MonadIO m, Vector v a) => ColumnDescription -> (Maybe DictVals -> Encoding -> Int -> ByteString -> v a) -> [ColumnChunk] -> (acc -> (v a, Vector Int, Vector Int) -> m acc) -> acc -> m acc Source #

Left-fold over per-page value triples, decoding each page with decoder. A thin wrapper over foldColumnDataPagesM for the typed (numeric/boxed) column paths.

foldColumnDataPagesM :: (RandomAccess m, MonadIO m) => ColumnDescription -> [ColumnChunk] -> (acc -> RawPage -> m acc) -> acc -> m acc Source #

Left-fold a monadic step over every DATA page (V1 or V2) of every column chunk, in order, threading the running dictionary internally and handing each page to step as a RawPage (no value decoding — the step decides how to materialize, e.g. into a typed vector or straight into a text buffer).

Dictionary pages update the running dictionary; INDEX_PAGEs are skipped. Only one page's bytes are live at a time.

  • - TODO: when a page index is available, use it here to compute which page
  • - byte ranges to request from the RandomAccess layer instead of reading the
  • - entire column chunk in one contiguous read.
  • - TODO: accept an optional row-range and use the column/offset page index
  • - (when present in file metadata) to skip pages whose row range does not
  • - overlap the requested range, avoiding decompression of irrelevant pages
  • - entirely.

type RawPage = (Maybe DictVals, Encoding, Int, ByteString, Vector Int, Vector Int) Source #

A decoded DATA page handed to a fold step: the running dictionary, the page encoding, the present-value count, the (decompressed) value bytes, and the definition/repetition level vectors.

appendStringPageIO :: TextBuilder RealWorld -> Maybe DictVals -> Encoding -> Int -> ByteString -> IO () Source #

Append a non-nullable BYTE_ARRAY page's strings straight into a TextBuilder, no intermediate boxed Text vector.

PLAIN pages copy each value's UTF-8 bytes directly from the page buffer pointer (appendTextSliceFromPtr), skipping the per-value decodeUtf8Lenient + ByteString slicing + Text allocation that readNTexts would do. Dictionary pages append the shared dictionary Texts by reference.

appendNullableStringPageIO :: TextBuilder RealWorld -> Int -> Maybe DictVals -> Encoding -> Int -> ByteString -> Vector Int -> IO () Source #

Append a NULLABLE BYTE_ARRAY page into a TextBuilder, interleaving nulls by walking the definition levels: a present row (def == maxDef) consumes the next stored value (dictionary entry by reference, or the next length-prefixed PLAIN slice by memcpy); a null row (appendNull) writes an offset and clears a validity bit. Builds PackedText + validity bitmap with no per-row Text or [Maybe a] allocation. Only present values are stored in the page payload.