dataframe-operations-2.4.0.0: Column operations, expression DSL, and statistics for the dataframe ecosystem.
Safe HaskellNone
LanguageHaskell2010

DataFrame.Typed.Aggregate

Synopsis

Typed groupBy

groupBy :: forall (keys :: [Symbol]) (cols :: [(Symbol, Type)]). (AllKnownSymbol keys, AssertAllPresent keys cols) => TypedDataFrame cols -> TypedGrouped keys cols Source #

Group a typed DataFrame by one or more key columns.

grouped = groupBy @'["department"] employees

Naming an aggregation

as :: forall (name :: Symbol) a (keys :: [Symbol]) (cols :: [(Symbol, Type)]) (aggs :: [(Symbol, Type)]). (KnownSymbol name, Columnable a) => TExpr cols a -> TAgg keys cols aggs -> TAgg keys cols ('(name, a) ': aggs) Source #

Build a named aggregation entry. The result column name is supplied via TypeApplications; the underlying expression is validated against the source schema at compile time.

as produces a transformer on the aggregation chain — entries compose with plain (.) from Prelude (or via (|>) for SQL-like postfix reading). aggregate applies the composed transformer to the empty chain internally, so no terminator is needed.

Prefix form

Expand
result = grouped |> aggregate
    ( as @"total"  (sum   (col @"amount"))
    . as @"orders" (count (col @"order_id"))
    . as @"avg"    (mean  (col @"amount"))
    )

Postfix form (SQL-like)

Expand
result = grouped |> aggregate
    ( (sum   (col @"amount")   |> as @"total")
    . (count (col @"order_id") |> as @"orders")
    . (mean  (col @"amount")   |> as @"avg")
    )

Per-entry parentheses are required in the postfix form because (.) binds tighter than (|>).

Running aggregations

aggregate :: forall (keys :: [Symbol]) (cols :: [(Symbol, Type)]) (aggs :: [(Symbol, Type)]). (TAgg keys cols ('[] :: [(Symbol, Type)]) -> TAgg keys cols aggs) -> TypedGrouped keys cols -> TypedDataFrame (Append (GroupKeyColumns keys cols) (Reverse aggs)) Source #

Run a typed aggregation against a grouped DataFrame.

The first argument is a chain of as entries composed with (.). The empty composition (id) yields just the group keys. The result schema is the group-key columns followed by the aggregation columns in declaration order.

result = grouped |> aggregate
    ( as @"total"  (sum (col @"amount"))
    . as @"orders" (count (col @"order_id"))
    )
-- result :: TypedDataFrame
--     '[ '("region", Text)
--      , '("total", Double)
--      , '("orders", Int)
--      ]

Escape hatch

aggregateUntyped :: forall (keys :: [Symbol]) (cols :: [(Symbol, Type)]). [NamedExpr] -> TypedGrouped keys cols -> DataFrame Source #

Escape hatch: run an untyped aggregation and return a raw DataFrame.