Kwker
DataFrames and analytics

Faster group-by, sorting and top-k on the tables you have

Reports, feature pipelines, leaderboards and batch jobs spend their time grouping, sorting and picking the top rows. Kwker does those steps on pandas, Polars, pyarrow and DuckDB data and hands back the same kind of table, with the same results.

Against the libraries you use today

Speed-up of Kwker over each library on the same table of 1 million rows, every result checked against the library's own.

QuerypandasPolarspyarrowDuckDB
Group-by, 1 key, 5 aggregations (4 threads)4.23×2.11×2.11×5.54×
Group-by with an exact median (4 threads)4.25×1.21×–5.40×
Group-by, 2 keys, 3 aggregations (4 threads)7.40×1.69×1.49×4.02×
Sort a 5-column table by a string (1 thread)3.32×7.83×2.47×–
Sort a 5-column table by an int64 (1 thread)2.30×1.59×2.88×–
ORDER BY a, b LIMIT 1000 (1 thread)–9.1×–2.43×

pandas 3.0, Polars 1.44, pyarrow 25, DuckDB 1.5 with the same thread count as Kwker; Intel Xeon (Cascade Lake), October 2026. A dash: not measured. Queries and method · Run it yourself

How it works

  • Same table type back: pandas in, pandas out; Polars and pyarrow the same. DuckDB relations and Arrow arrays work too.
  • Each library's rules: where nulls and NaN go and what a sum of nothing is follow the library you called it on; every sort is stable.
  • One pass where it counts: group-by uses hashing for few distinct keys and sorting for many; top-k never sorts every row.
  • Strings, dates, categoricals: string keys, string views, timestamps, dictionaries and enums sort natively, without casting.
PythonRuns on your machine.
import pandas as pd
import kwker.frame as sf

orders = pd.DataFrame({
    "customer": ["ana", "ben", "ana", "cy", "ben", "ana"],
    "region": ["eu", "us", "eu", "us", "us", "eu"],
    "amount": [120.0, 35.5, 80.0, 220.0, 15.0, 60.0]})

per_customer = sf.group_by(orders, "customer", {
    "orders": "size",
    "revenue": ("amount", "sum"),
    "median": ("amount", "median")})
print(per_customer)
print(sf.top_k(orders, "amount", 2, descending=True))
Output
  customer  orders  revenue  median
0      ana       3    260.0   80.00
1      ben       2     50.5   25.25
2       cy       1    220.0  220.00
  customer region  amount
3       cy     us   220.0
0      ana     eu   120.0
The DataFrames guide, with runnable examples

Requirements

  • Python 3.9 or newer with the library you already use: pandas, Polars, pyarrow or DuckDB.
  • Any 64-bit processor: AVX-512 and AVX2 on x86-64, NEON and SVE on ARM64, and a portable engine elsewhere. The fastest engine is picked when your program starts.
  • Outside Python: the same group-by, sorting and top-k are in the C, C++, Rust, Java, Go and C# APIs.

DataFrames guide, groups and sets, top-k.

Free to start

Free for companies under 100 people and about $1.3M revenue (PolyForm Small Business). Every product line is in every tier.