Polars
kwker.frame takes a Polars DataFrame and returns a Polars DataFrame with the rows Polars' own call gives, in the same
order. Swap the call; keep the rest of your code.
import polars as pl
import kwker.frame as sf
df = pl.DataFrame({"user": ["b", "a", "b", "a", "c"], "amount": [10.0, 2.0, 5.0, 4.0, 1.0]})
print(sf.sort(df, ["user", "amount"], descending=[False, True]))
print(sf.top_k(df, "amount", 2, descending=True))
print(sf.group_by(df, "user", {"n": "size", "total": ("amount", "sum")}))
print(sf.sort(df, "amount").equals(df.sort("amount", maintain_order=True)))
Output
shape: (5, 2) ┌──────┬────────┐ │ user ┆ amount │ │ --- ┆ --- │ │ str ┆ f64 │ ╞══════╪════════╡ │ a ┆ 4.0 │ │ a ┆ 2.0 │ │ b ┆ 10.0 │ │ b ┆ 5.0 │ │ c ┆ 1.0 │ └──────┴────────┘ shape: (2, 2) ┌──────┬────────┐ │ user ┆ amount │ │ --- ┆ --- │ │ str ┆ f64 │ ╞══════╪════════╡ │ b ┆ 10.0 │ │ b ┆ 5.0 │ └──────┴────────┘ shape: (3, 3) ┌──────┬─────┬───────┐ │ user ┆ n ┆ total │ │ --- ┆ --- ┆ --- │ │ str ┆ i64 ┆ f64 │ ╞══════╪═════╪═══════╡ │ a ┆ 2 ┆ 6.0 │ │ b ┆ 2 ┆ 15.0 │ │ c ┆ 1 ┆ 1.0 │ └──────┴─────┴───────┘ True
From Polars calls to Kwker calls
| Polars | Kwker |
|---|---|
df.sort("a") |
sf.sort(df, "a") |
df.sort(["a", "b"], descending=[False, True]) |
sf.sort(df, ["a", "b"], descending=[False, True]) |
df.sort("a", nulls_ |
sf.sort(df, "a", nulls_ |
df.top_ |
sf.top_ |
df.bottom_ |
sf.top_ |
df["a"].arg_ |
sf.argsort(df, "a") |
df.group_ |
sf.group_ |
The group-by operations are sum, mean, min, max, count, size, first, last and median. Sort keys can
be numbers, booleans, dates and times, strings, categoricals and Enums.
Threads
Polars uses every core by default; kwker.frame uses one thread unless you pass threads. Give both the same cores
when you compare them:
import polars as pl
import kwker.frame as sf
df = pl.DataFrame({"x": list(range(100_000, 0, -1))})
out = sf.sort(df, "x", threads=4)
print(out["x"][:3].to_list())
Output
[1, 2, 3]
python -m kwker.bench --frames times Kwker against your Polars install on the same tables (Evaluate).
Notes
- A LazyFrame needs
.collect()first:kwker.frameworks on DataFrames. - Nulls come first unless
nulls_last=True, and NaN sorts above every number, as in Polars. Rows with equal keys keep their input order, likemaintain_order=True. group_byreturns one row per key in ascending key order, like.sort("g")after Polars' group-by. NaN propagates through sums and means, as in Polars.- Small tables (under about 4,000 rows) run Polars' own sort, which is faster at that size.
Related
- DataFrames, Arrow and DuckDB: the same calls on pandas and pyarrow tables.
- Group-by on a table: a step-by-step group-by.
- Groups, merges and sets: totals per key without a DataFrame.