Kwker 0.1.0 release notes
Note
Draft. The version number, package names and download links take effect with the release.
Kwker 0.1.0 is the first public release. It sorts, selects, ranks, searches, merges and groups keys of 4 to 128 bits, strings and tables, with indices, values or records. SIMD engines for AVX-512, AVX2, SSE4.2, ARM NEON and ARM SVE and a portable fallback are chosen at run time. For PyTorch on x86-64 Linux it also replaces CPU kernels and runs CNNs, BERT-family encoders and Llama-family, mixture-of-experts and GPT-2-family decoders natively.
Speed
Kwker is compared against the fastest public library for each operation: x86-simd-sort and VQSort for sorts, their key-value and selection forms for those operations. Every ratio below is Kwker's speed over the faster of the two, one thread, measured on Intel Xeon cloud VMs with AVX-512 (the per-call grid on a Cascade Lake Xeon). The AVX2 figures run the AVX2 engine and the competitors' AVX2 builds on the same CPUs.
| Benchmark | AVX-512 | AVX2 |
|---|---|---|
| Sort sweep: 6 key types x 40 input families x 1K / 100K / 1M keys (720 cells), geomean | 1.99x | 2.14x |
| Sort sweep cells below 1.0x | 12 of 720 | 23 of 720 |
| Per-call sorts of 2 to 900 keys, 7 input patterns x 6 key types (672 cells), geomean | 2.84x | 3.01x |
| Per-call cells below 1.0x | 2 of 672 | 1 of 672 |
On real data never used for tuning (token ids, model weights, file-system metadata; 1K, 100K and 1M keys, 17 cells per
engine) no AVX-512 cell was below 1.25x and no AVX2 cell below 1.0x. The cells where Kwker trails are listed in Known limitations: mostly
near-bandwidth inputs (all equal, three values) and duplicate-heavy 64-bit floats at 100K-1M keys. Your hardware will
differ: python -m kwker.bench measures your own machine, and python -m kwker.bench --suite --native repeats the
sort sweep's input patterns there against x86-simd-sort and VQSort built for your CPU.
What is in it
- Key types: 8-, 16-, 32-, 64- and 128-bit integers, float16, bfloat16, FP8 (E4M3, E5M2), float32, float64, packed int4, strings (byte, ASCII caseless and natural order) and records by a key field.
- Operations: sort and stable sort, argsort, partial sort, selection and multi-selection, top-k (batched, masked, per group, streaming), ranking, sorted search, merges, set operations, unique / run-length encoding, group-by reductions, lexicographic sorts over several columns, key-value sorts (stable or not), rows, segments and N-D axes.
- Orders: ascending or descending, NaNs first or last; -0.0 before +0.0; results are deterministic.
- Scale: multithreaded sorts, caller-owned workspaces, scratch limits (including no allocation at all), cancellation and progress, out-of-core file sorts.
- Data: Apache Arrow arrays read in place, chunked arrays included; pandas, Polars, pyarrow and DuckDB through
kwker.frameandkwker.duck. - PyTorch and JAX: drop-in CPU kernels for sorting, selection, indexing and random numbers; a
torch.compilebackend; compiled operators; the CPU backend (AMX / AVX-512 GEMMs and attention, int8 / int4 modes); KwkCNN, KwkEncoder and KwkDecoder. Prebuilt for torch 2.12, 2.13 and 2.14. - Languages: Rust, C, C++ and Python, plus bindings for Go, Java, C#, JavaScript (Node and WebAssembly), Ruby, PHP, Perl, R, MATLAB / Octave, Swift, Objective-C, Zig, Fortran, COBOL and assembly.
Platforms
Linux x86-64 and ARM64, Windows x64 and macOS ARM64 have prebuilt packages; Windows ARM64 and macOS x86-64 build from source. The installation guide lists each package; Compatibility lists the toolchains, Python and torch releases each one is tested with.
Compatibility
0.1.x releases keep the C ABI, the C++ header, the Rust crate's public API, the Python API and the torch operator schemas compatible. Sorts, selections and stable key-value sorts return the same output on every engine and in every 0.1.x release; index operations break ties by index. A public name is deprecated for at least one minor release before it is removed; nothing is deprecated in 0.1.0. The details are in Compatibility.
Known limitations
See Known limitations. In short:
- The PyTorch extensions and the CPU backend are Linux x86-64 only.
- There is no locale-aware string collation.
- 128-bit integer keys are not exposed in Python.
- The ARM engines and the SSE4.2 engine (x86 without AVX2) are less tuned than the AVX-512 and AVX2 ones.