dataframe-lazy-2.4.0.0: Lazy query engine for the dataframe ecosystem.
Safe HaskellNone
LanguageHaskell2010

DataFrame.Lazy

Synopsis

Documentation

filter :: Expr Bool -> LazyDataFrame -> LazyDataFrame Source #

Keep rows that satisfy the predicate.

join Source #

Arguments

:: JoinType 
-> Text

Left join key column name

-> Text

Right join key column name

-> LazyDataFrame

Left sub-query

-> LazyDataFrame

Right sub-query

-> LazyDataFrame 

Join two lazy queries on the given key columns.

take :: Int -> LazyDataFrame -> LazyDataFrame Source #

Retain at most n rows.

groupBy Source #

Arguments

:: [Text]

Group-by key columns

-> [(Text, UExpr)]
[(outputName, aggregateExpr)]
-> LazyDataFrame 
-> LazyDataFrame 

Group by a set of columns and compute aggregate expressions.

Each aggregate expression should use an Agg node (e.g. sumOf, meanOf).

sortBy :: [(Text, SortOrder)] -> LazyDataFrame -> LazyDataFrame Source #

Sort the result by the given (column, direction) pairs.

derive :: Columnable a => Text -> Expr a -> LazyDataFrame -> LazyDataFrame Source #

Add a computed column (or overwrite an existing one).

select :: [Text] -> LazyDataFrame -> LazyDataFrame Source #

Retain only the listed columns.

scanCsv :: Schema -> Text -> LazyDataFrame Source #

Scan a CSV file with the default comma separator and the in-tree strict reader. For the SIMD reader use scanCsvWith.

The Schema both types and selects: only the columns it names are read, matching scanParquet.

Example

Expand
ghci> schema = D.makeSchema [("id", D.schemaType @Int), ("name", D.schemaType @Text)]
ghci> L.runDataFrame (L.scanCsv schema "customers.csv")

scanSeparated :: Char -> Schema -> Text -> LazyDataFrame Source #

Scan a character-separated file with the default strict reader.

Example

Expand
ghci> L.runDataFrame (L.scanSeparated ';' schema "customers.txt")

scanParquet :: Schema -> Text -> LazyDataFrame Source #

Scan a Parquet file, directory of files, or glob pattern.

fromDataFrame :: DataFrame -> LazyDataFrame Source #

Lift an already-loaded eager DataFrame into the lazy plan.

data LazyDataFrame Source #

A lazy query that has not been executed yet: a LogicalPlan tree whose execution is deferred until runDataFrame is called.

Constructors

LazyDataFrame 

Fields

Instances

Instances details
Show LazyDataFrame Source # 
Instance details

Defined in DataFrame.Lazy.Internal.DataFrame

runDataFrame :: LazyDataFrame -> IO DataFrame Source #

Execute the lazy query: optimise the logical plan, then stream-execute the resulting physical plan into a fully-materialised DataFrame.

scanCsvWith :: CsvReader -> Schema -> Text -> LazyDataFrame Source #

Like scanCsv but with an explicit CSV reader (e.g. the SIMD reader fastReadCsvWithOpts from dataframe-fastcsv). The scan derives the reader's ReadOptions from the schema and separator, so any CsvReader projects.

Example

Expand
ghci> import qualified DataFrame.IO.CSV.Fast as Fast
ghci> L.runDataFrame (L.scanCsvWith Fast.fastReadCsvWithOpts schema "customers.csv")

scanCsvStreamingWith :: CsvReader -> Schema -> Text -> LazyDataFrame Source #

Like scanCsvWith, but the file is read in bounded-memory windows instead of one pass per chunk — for files too large to hold in memory even after the schema's projection.

Example

Expand
ghci> L.runDataFrame (L.scanCsvStreamingWith Fast.fastReadCsvWithOpts schema "huge.csv")

scanSeparatedWith :: CsvReader -> Char -> Schema -> Text -> LazyDataFrame Source #

Like scanSeparated but with an explicit CSV reader.

Example

Expand
ghci> L.runDataFrame (L.scanSeparatedWith Fast.fastReadCsvWithOpts ';' schema "customers.txt")

data SortOrder Source #

Sort direction used in Sort nodes and the public API.

Constructors

Ascending 
Descending