Kwker

Quickstart: DataFrames

Run group-by, sorts and "top N rows" on your tables faster. kwker.frame takes a pandas DataFrame, a Polars DataFrame or a pyarrow Table, and returns the same type with the results that library would give. About five minutes.

Install

ShellOn your machine.
pip install kwker pandas          # or polars, or pyarrow
python -m kwker doctor            # your CPU, the engine Kwker picked, any warnings

Group-by

One row per key, in key order, with the aggregations you name:

PythonNeeds pandas: runs on your machine.
import pandas as pd
import kwker.frame as kf

orders = pd.DataFrame({"customer": ["b", "a", "b", "a", "c"], "amount": [10.0, 2.0, 5.0, 4.0, 1.0]})
print(kf.group_by(orders, "customer", {"orders": "size", "revenue": ("amount", "sum")}))

The aggregations are sum, mean, min, max, count, size, first, last and median.

Sort and top rows

PythonNeeds pandas: runs on your machine.
import pandas as pd
import kwker.frame as kf

orders = pd.DataFrame({"customer": ["b", "a", "b", "a", "c"], "amount": [10.0, 2.0, 5.0, 4.0, 1.0]})
print(kf.sort(orders, ["customer", "amount"], descending=[False, True]))   # ORDER BY customer, amount DESC
print(kf.top_k(orders, "amount", 2, descending=True))                     # ORDER BY amount DESC LIMIT 2

The same calls take Polars and pyarrow tables. Pass threads=4 for large tables.

Check the results

Compare with the library's own answer once, on your data:

PythonNeeds pandas: runs on your machine.
import numpy as np
import pandas as pd
import kwker.frame as kf

rng = np.random.default_rng(0)
df = pd.DataFrame({"k": rng.integers(0, 100, 10_000), "v": rng.random(10_000)})
mine = kf.group_by(df, "k", {"total": ("v", "sum")})
theirs = df.groupby("k", as_index=False).agg(total=("v", "sum"))
print(np.allclose(mine["total"].to_numpy(), theirs["total"].to_numpy()))
Output
True

Measure it on your machine

ShellOn your machine.
python -m kwker.bench --frames      # group-by, table sorts and ORDER BY ... LIMIT vs pandas, Polars and DuckDB

Next steps