Known limitations
This page lists what Kwker does not do, and where it is not the fastest. Benchmarks in this documentation are compared against the fastest public library for each operation, and losses are reported along with wins.
Speed
- Kwker is not the fastest on every input at every size. The benchmark reports list each cell where it trails
x86-simd-sort or VQSort. To check your own case, run
python -m kwker benchon your machine. - All-equal and already-sorted inputs are memory-bound. At best Kwker matches the one read pass any sort needs.
- Below about 100 elements per call from Python, call overhead dominates. Batch rows into one call (
axis=,sort_segments), or use the C / Rust API directly. - The x86 engines (AVX-512, AVX2) are the most tuned. The ARM engines (NEON, SVE) are newer. Measure on your hardware.
- On x86 CPUs without AVX2 the SSE4.2 engine runs. Against VQSort's SSE4 build (or the standard library, where faster) it is about 1.15x on 1K keys and 1.3x on 100K-1M keys. It is still slower on large random inputs (0.8-0.85x at 1M keys). Other CPUs without a SIMD engine (RISC-V and others) run the portable engine, which has no SIMD kernels.
Platforms and packages
- The PyTorch extensions and the CPU backend are Linux x86-64 only. On other platforms the Python package works without them.
- No prebuilt packages exist for Windows ARM64 or macOS x86-64; both build from source. macOS has no AVX-512 engine.
- The CPU backend's AMX kernels need Linux 5.16 or newer, so the kernel can grant AMX state.
- Under WebAssembly, 64-bit keys run scalar code: SIMD128 has no 64-bit min / max.
- The WebAssembly module holds at most 4 GB, so about 4 GB of keys per call. Larger arrays need the 64-bit module, which works in browsers with Memory64 (Chrome 133+, Firefox 134+) and holds up to 16 GB. See Arrays over 4 GB.
- Some package registries do not accept Kwker's license. Where that applies, packages come from Kwker's own channels instead.
Operations
- Strings sort by bytes, ASCII caseless, natural or a 256-entry weight table. Locale-aware collation (ICU) is not built in: sort by keys that ICU generates.
- 128-bit integer keys are available in Rust and C, not in Python (Arrow decimal128 / 256 columns do sort there).
sort_kvis unstable by default; passstable=Truewhen equal keys must keep their values' order.- Distributed sorting comes as primitives (sampling, splitters, partitioning, merging). There is no distributed runtime.
- Comparator-based sorts of arbitrary objects are much slower than key sorts. Extract a key column when you can.
PyTorch and the CPU backend
-
CPU tensors only. Graphs on CUDA, XPU or MPS go to PyTorch / Inductor unchanged.
-
The
"medium", int8 and int4 modes change results by design. Check each model withcheck_accuracyfirst. -
int8 activations use one scale per tensor, so a few large activations cost resolution for the rest.
-
The eager convolution override covers only 2-D convolutions with
groups=1. KwkCNN covers depthwise and grouped convolutions itself. -
The runners cover fixed architecture families:
- KwkCNN: ResNet, MobileNet, EfficientNet, RegNet, ResNeXt, ConvNeXt.
- KwkEncoder: BERT-family encoders, on x86 CPUs with AVX2 or AVX-512 (int8 needs AMX, AVX-512 VNNI or AVX-VNNI).
- KwkDecoder: Llama-family, mixture-of-experts and GPT-2-family decoders (the list is in PyTorch).
For anything else,
runner_unsupported(model)explains why; usetorch.compile(backend="kwker")orDecoderinstead. -
KwkCNN in float32 on ResNeXt and RegNet runs at about 0.93x of OpenVINO float32 on the AMX test machine. Its int8 mode leads OpenVINO int8 on both (1.21x on ResNeXt-50, 1.52x on RegNetX-1.6GF, AMX-INT8 on both sides).
-
int4 decoding quality depends on the model. On SmolLM2-135M, int4 loses more perplexity than llama.cpp's Q4_K_M (+8.4% against +6.1%; on that model Q4_K_M is mostly a 5-bit format). On SmolLM2-1.7B it loses less (+6.0% against +8.4%). Use
calib=andgptq=Trueto narrow the gap, and measure on your own text.