dataframe-parquet-1.5.0.0: Parquet reader and writer for the dataframe ecosystem.
Safe HaskellNone
LanguageHaskell2010

DataFrame.Typed.IO.Parquet

Description

Typed Parquet reading.

The reader validates the file against a type-level schema as it loads, supplied by type application:

type Trips = '[ '("id", Int), '("fare", Double)]

trips <- readParquet @Trips "trips.parquet"   -- IO (TypedDataFrame Trips)

readParquet (and readParquetWithOpts/readParquetFiles) throw a DataFrameException on schema mismatch; readParquetWithError returns the mismatch as an Either instead.

Synopsis

Documentation

readParquet :: forall (cols :: [(Symbol, Type)]). KnownSchema cols => FilePath -> IO (TypedDataFrame cols) Source #

Read a Parquet file into a typed DataFrame, throwing on schema mismatch. Reads only the columns cols names.

Example

Expand
ghci> trips <- readParquet @Trips "trips.parquet"

readParquetWithError :: forall (cols :: [(Symbol, Type)]). KnownSchema cols => FilePath -> IO (Either Text (TypedDataFrame cols)) Source #

Read a Parquet file, returning a descriptive error on schema mismatch or a missing column instead of throwing.

Example

Expand
ghci> readParquetWithError @Trips "trips.parquet"
Right (TDF ...)

readParquetWithOpts :: forall (cols :: [(Symbol, Type)]). KnownSchema cols => ParquetReadOptions -> FilePath -> IO (TypedDataFrame cols) Source #

Read a Parquet file with custom options, throwing on schema mismatch. The schema still supplies the column selection; an explicit selectedColumns takes precedence.

Example

Expand
ghci> trips <- readParquetWithOpts @Trips defaultParquetReadOptions{rowRange = Just (0, 10)} "trips.parquet"

readParquetFiles :: forall (cols :: [(Symbol, Type)]). KnownSchema cols => FilePath -> IO (TypedDataFrame cols) Source #

Read a directory/glob of Parquet files into a typed DataFrame.

Example

Expand
ghci> trips <- readParquetFiles @Trips "./data/trips/*.parquet"