dataframe-parquet-1.5.0.0: Parquet reader and writer for the dataframe ecosystem.
Safe HaskellNone
LanguageHaskell2010

DataFrame.IO.Parquet

Synopsis

Documentation

data ParquetReadOptions Source #

Options for reading Parquet data.

These options are applied in this order:

  1. predicate filtering
  2. column projection
  3. row range
  4. safe column promotion

Column selection for selectedColumns uses leaf column names only.

Constructors

ParquetReadOptions 

Fields

  • selectedColumns :: Maybe [Text]

    Columns to keep in the final dataframe. If set, only these columns are returned. Predicate-referenced columns are read automatically when needed and projected out after filtering.

  • predicate :: Maybe (Expr Bool)

    Optional row filter expression applied before projection.

  • rowRange :: Maybe (Int, Int)

    Optional row slice (start, end) with start-inclusive/end-exclusive semantics.

  • safeColumns :: Bool

    When True, every column is promoted to OptionalColumn after read, regardless of nullability in the schema.

Instances

Instances details
Show ParquetReadOptions Source # 
Instance details

Defined in DataFrame.IO.Parquet

defaultParquetReadOptions :: ParquetReadOptions Source #

Default Parquet read options.

Equivalent to:

ParquetReadOptions
    { selectedColumns = Nothing
    , predicate = Nothing
    , rowRange = Nothing
    , safeColumns = False
    }

readParquet :: FilePath -> IO DataFrame Source #

Read a parquet file from path and load it into a dataframe.

Example

Expand
ghci> D.readParquet "./data/mtcars.parquet"

readParquetWithOpts :: ParquetReadOptions -> FilePath -> IO DataFrame Source #

Read a Parquet file using explicit read options.

Example

Expand
ghci> D.readParquetWithOpts
ghci|   (D.defaultParquetReadOptions{D.selectedColumns = Just ["id"], D.rowRange = Just (0, 10)})
ghci|   ".testsdata/alltypes_plain.parquet"

When selectedColumns is set and predicate references other columns, those predicate columns are auto-included for decoding, then projected back to the requested output columns.

_readParquetWithOpts :: ForceNonSeekable -> ParquetReadOptions -> FilePath -> IO DataFrame Source #

Internal entry point used by tests to force non-seekable mode.

readParquetFiles :: FilePath -> IO DataFrame Source #

Read Parquet files from a directory or glob path.

This is equivalent to calling readParquetFilesWithOpts with defaultParquetReadOptions.

readParquetFilesWithOpts :: ParquetReadOptions -> FilePath -> IO DataFrame Source #

Read multiple Parquet files (directory or glob) using explicit options.

If path is a directory, all non-directory entries are read. If path is a glob, matching files are read.

For multi-file reads, rowRange is applied once after concatenation (global range semantics).

Example

Expand
ghci> D.readParquetFilesWithOpts
ghci|   (D.defaultParquetReadOptions{D.selectedColumns = Just ["id"], D.rowRange = Just (0, 5)})
ghci|   ".testsdata/alltypes_plain*.parquet"

parseParquetWithOpts :: (RandomAccess m, MonadIO m) => ParquetReadOptions -> m DataFrame Source #

Parse a Parquet file via the RandomAccess handle, applying all read options. This is the central parsing entry point used by _readParquetWithOpts.

parseFileMetadata :: RandomAccess m => m FileMetadata Source #

Parse the file-level Thrift metadata from the Parquet file footer. Validates the trailing 4-byte magic marker ("PAR1") before decoding.

readMetadataFromPath :: FilePath -> IO FileMetadata Source #

Read the file metadata from a Parquet file at the given path.

readMetadataFromHandle :: FileBufferedOrSeekable -> IO FileMetadata Source #

Read only the file metadata from an open FileBufferedOrSeekable handle.

columnChunksForAll :: FileMetadata -> [[ColumnChunk]] Source #

Collect column chunks per column (transposed across all row groups).

parseColumnChunks :: (RandomAccess m, MonadIO m) => Int -> [ColumnChunk] -> ColumnDescription -> m Column Source #

Dispatch a column's chunks to the correct decoder path.

getNonNullableColumn :: (RandomAccess m, MonadIO m) => Int -> ColumnDescription -> [ColumnChunk] -> m Column Source #

Decode a required (non-nullable, non-repeated) column.

getNullableColumn :: (RandomAccess m, MonadIO m) => Int -> ColumnDescription -> [ColumnChunk] -> m Column Source #

Decode an optional (nullable) column.

getRepeatedColumn :: (RandomAccess m, MonadIO m) => ColumnDescription -> [ColumnChunk] -> m Column Source #

Decode a repeated (list/nested) column.

applyDescLogicalType :: ColumnDescription -> Column -> Column Source #

Apply a column-description's logical type annotation to convert raw decoded values (e.g. millisecond integers → UTCTime).

epochToUTCTime :: Int64 -> Integer -> Int64 -> UTCTime Source #

Convert an epoch timestamp expressed as ticksPerSecond ticks/second (each tick = psPerTick picoseconds) to UTCTime, at full precision.

Splits into days + within-day picoseconds with integer divMod (which floors, so the split is correct for pre-epoch negative values too), and never forms picoseconds-since-epoch — that would overflow Int64 for modern dates — only the bounded within-day picosecond count. 40587 is the Modified Julian Day of the Unix epoch (1970-01-01).