Kwker

DataFrames, Arrow and DuckDB

kwker.frame sorts, ranks and groups pandas DataFrames, Polars DataFrames and pyarrow Tables. You get back the type you passed in, with the results that library would give.

Sort, top-k and argsort

PythonNeeds pandas: runs on your machine.
import pandas as pd
import kwker.frame as sf

df = pd.DataFrame({"region": ["eu", "us", "eu", "us", "eu"],
                   "price": [3.5, 1.0, None, 7.25, 2.0]})
print(sf.sort(df, ["region", "price"], descending=[False, True]))   # ORDER BY region, price DESC; the NaN row last
print(sf.top_k(df, "price", 2, descending=True))                     # ORDER BY price DESC LIMIT 2
order = sf.argsort(df, "price")                                      # the row order, int64

The same calls work on Polars and pyarrow:

PythonNeeds polars, pyarrow: runs on your machine.
import polars as pl, pyarrow as pa
import kwker.frame as sf

pdf = pl.DataFrame({"k": [3, None, 1, 2], "v": ["c", "x", "a", "b"]})
print(sf.sort(pdf, "k"))                                   # Polars: nulls first by default
tbl = pa.table({"k": [3, None, 1, 2]})
print(sf.sort(tbl, "k", nulls_last=True).column("k"))      # pyarrow: nulls last by default

Key columns can be:

Pass threads=4 for large frames.

Group-by

PythonNeeds pandas: runs on your machine.
import pandas as pd
import kwker.frame as sf

df = pd.DataFrame({"user": ["b", "a", "b", "a", "c"], "amount": [10.0, 2.0, 5.0, 4.0, 1.0]})
print(sf.group_by(df, "user", {"n": "size", "total": ("amount", "sum"), "p50": ("amount", "median")}))

This computes SQL's GROUP BY ... ORDER BY the keys: one row per distinct key tuple, in ascending key order. The operations are sum, mean, min, max, count (non-missing values), size (rows), first, last and median (exact).

Arrow arrays

PythonNeeds pyarrow: runs on your machine.
import pyarrow as pa
import kwker

arr = pa.array(["pear", None, "apple", "fig"])
print(kwker.arrow_argsort(arr))                    # nulls at the end, like pyarrow.compute.sort_indices
print(kwker.arrow_top_k(arr, 2, order="descending"))

These calls read Arrow memory directly through the C Data Interface, with no conversion. They accept:

DuckDB

PythonNeeds duckdb, pyarrow: runs on your machine.
import duckdb
import kwker.duck as sd

con = duckdb.connect()
rel = con.sql("SELECT range % 7 AS g, random() AS x FROM range(10000)")
top = sd.top_k(rel, ["x"], 5, descending=True, connection=con)       # a DuckDB relation back
grp = sd.group_by(rel, ["g"], {"n": "size", "mean_x": ("x", "mean")}, connection=con)
print(grp.fetchall()[:2])

kwker.duck returns a DuckDB relation. Pass arrow=True to get a pyarrow Table instead.

Missing values, NaN and equal keys

Each library keeps its own rules:

Library Missing values NaN
pandas NaN, None, NaT and pd.NA last in both directions (na_position) missing
Polars nulls first unless nulls_last=True a value above every number
pyarrow nulls last after every number, in both directions
DuckDB NULLs last unless you ask otherwise the largest value