Data and analytics
Kwker speeds up the sorting, ranking, grouping and searching inside data work. Most of it needs no new code: NumPy and
DataFrame calls keep their results, and Kwker runs underneath. Group-by on a million rows ran 1.10-3.51x faster than
the fastest of Polars, DuckDB, pyarrow and pandas on one thread, and ORDER BY a, b LIMIT k 1.52-21.7x faster than
DuckDB.
Examples
Without changing your code
- NumPy drop-in:
np.sort,np.argsort,np.uniqueand friends on Kwker, the same results. - SciPy drop-in: sparse conversions,
rankdataand the rank tests, distances and median filters on Kwker, the same results.
Tables
- DataFrames, Arrow and DuckDB: sort, top-k and group-by on pandas, Polars and pyarrow frames.
- Polars: Polars calls and their Kwker equivalents.
- Group-by on a table: a lesson from a small table to millions of rows.
- ORDER BY in DuckDB: DuckDB queries with Kwker's sorts.
Larger data
- Sort a file larger than memory: a fixed memory budget, runs on disk, one merge.
- Sparse matrices: COO to CSR and CSC, with duplicates summed, as SciPy does.
- Statistics and data helpers: ranks, quantiles, distances between distributions.
Next steps
- Quickstart: DataFrames, DuckDB and NumPy.
- Evaluate on your machine:
python -m kwker.bench --framesagainst your own pandas, Polars and DuckDB.