Kwker

Evaluate Kwker on your machine

You don't have to take our numbers on trust. Three commands show what Kwker does on your hardware and for your code. None of them uploads anything.

Step Command What you learn Takes
1. Check the machine python -m kwker doctor your CPU's features, the engine Kwker picked, warnings seconds
2. Benchmark your libraries python -m kwker.bench --quick Kwker against NumPy, PyTorch, pyarrow and Polars on this CPU about 10 s
3. Run your own program python -m kwker audit -- train.py end-to-end time, CPU time, and whether the results match 6 runs of your program

1. Check the machine

ShellOn your machine.
python -m kwker doctor
python -m kwker doctor --json > machine.json      # the same report as JSON

doctor prints the Kwker version, the CPU and its instruction sets, the engine Kwker runs (AVX-512, AVX2, NEON, SVE or portable) and the state of the PyTorch extensions. It also warns about anything that slows Kwker down, such as a torch release the wheel was not built for.

2. Benchmark the libraries you use

ShellOn your machine.
python -m kwker.bench --quick                      # 100K keys, about 10 seconds
python -m kwker.bench                              # 1K, 100K and 1M keys, 4 key types, 3 input shapes
python -m kwker.bench --md report.md --json report.json
python -m kwker.bench --native                     # also x86-simd-sort and VQSort, built on this machine
python -m kwker.bench --suite --native             # 43 input patterns x 6 key types, as on kwker.io/benchmarks
python -m kwker.bench --frames                     # tables: group-by, sorts, ORDER BY ... LIMIT
python -m kwker.bench --models --threads 4         # image classifiers: KwkCNN, PyTorch, OpenVINO
python -m kwker.bench --encoders --threads 4       # text encoders (BERT, MiniLM): KwkEncoder, OpenVINO, PyTorch
python -m kwker.bench --llm HuggingFaceTB/SmolLM2-135M --text wiki.test.raw   # a language model: KwkDecoder, transformers

The benchmark times Kwker against every library you have installed: NumPy, PyTorch, pyarrow and Polars. They all get the same input on one thread (--threads changes that), and every result is checked against NumPy's. The report names the CPU, the engine and every library version, so you can share it as it is.

With --native, it also times Intel's x86-simd-sort and Google Highway's VQSort, the two libraries kwker.io/benchmarks compares against. It downloads both at the commits behind the published numbers and builds them once for your CPU, which takes about a minute. It needs g++ or clang++ and does not run on Windows yet.

With --suite, it sorts the 43 input patterns behind the sort results on kwker.io/benchmarks instead: random keys, skewed and duplicate-heavy keys, presorted and reversed runs, timestamps, string-like keys and more, each as six key types (uint32, int32, float32, uint64, int64, float64) at 1K, 100K and 1M keys. The report has one table per size: a row per pattern, a column per key type, and Kwker's speed-up over the fastest other library in each cell. --families uniform,mod-8 picks patterns, --types and --sizes narrow it further, and --ops sort,argsort adds argsort. The full suite takes a few minutes; add --native to include x86-simd-sort and VQSort.

With --frames, it times the data-frame cases of kwker.io/benchmarks instead: group-by, table sorts and ORDER BY ... LIMIT on 1M-row tables, kwker.frame against pandas, Polars, pyarrow and DuckDB (each one you have installed). Every result is checked against the library's own first. --rows changes the table size.

With --models, it times torchvision image classifiers in float32: KwkCNN against PyTorch eager, torch.compile and OpenVINO, each output checked against eager PyTorch. Name models with --models resnet18,resnet50, and leave out the slow torch.compile step with --no-compile. Add --int8 --images photos/ to also time int8: KwkCNN int8 against OpenVINO int8, both calibrated on the same photos from your folder. The report shows how often each int8 model picks the same class as the float32 model on 100 more of your photos. This uses the pretrained weights, downloaded once.

With --encoders, it times text encoders on padded batches of 1 x 128 and 8 x 128 tokens: KwkEncoder against OpenVINO, ONNX Runtime, torch.compile(backend="kwker") and eager PyTorch, and ONNX Runtime with Kwker's provider added. It uses bert-base-uncased and all-MiniLM-L6-v2 unless you name your own models, for example --encoders BAAI/bge-small-en-v1.5, and --shapes 32x64 changes the batches. --int8 adds KwkEncoder int8 against OpenVINO int8, and ONNX Runtime on an int8 export (made the way the Hub's qint8 files are) with and without Kwker. KwkEncoder needs a CPU with AVX2 or AVX-512; on other CPUs the report says so and times the rest.

With --llm MODEL, it times a Hugging Face language model (an id or a local directory) on every core: KwkDecoder in int4, int8 and bf16 against the model's own eager float32 forward. It reports prompt tokens a second (a 512-token prompt) and generation tokens a second (64 tokens). Add --text with a text file to also get perplexity, scored the way llama.cpp's llama-perplexity scores it. To put llama.cpp in the same table, pass your own build and the model's GGUF files: --llama-bench ./llama-bench --llama-perplexity ./llama-perplexity --gguf model-Q4_0.gguf,model-Q8_0.gguf.

3. Run your own program both ways

ShellOn your machine.
python -m kwker audit -- train.py --epochs 1
python -m kwker audit --output preds.npy --rtol 1e-4 -- predict.py
python -m kwker audit --mode bf16 --rtol 1e-2 -- serve_bench.py
python -m kwker audit --mode numpy -- analysis.py         # a NumPy program, no PyTorch needed
python -m kwker audit --mode scipy -- build_graph.py      # a SciPy program, with kwker.scipy.install()

audit runs your program, unchanged, three times without Kwker and three times with it, alternating, and keeps each side's fastest run. Then it checks that the results match: the standard output, and every file you name with --output. The report shows both times, the speedup and the CPU time saved. Each round runs both sides back to back, so each round gives one paired speedup; the report shows their range under the speedup, and from --repeat 6 a 90% interval for the median instead. A wide range means the machine was busy: run again with a higher --repeat.

--mode What Kwker switches on Results
ops (default) kwker.torch_ops.install(): Kwker runs PyTorch's CPU kernels for sorting, selection, indexing, random numbers, pooling and the other covered operations identical to torch, bit for bit
bf16 ops, plus float32 matmul precision "medium": linears, attention and convolutions in bf16 on AMX within bf16 precision: set --rtol
int8 ops, plus kwker.cpu_backend.set_int8(True): dynamic int8 linears and convolutions change slightly: check them (step 4)
numpy kwker.numpy_ops.install(): Kwker runs NumPy's sort, argsort, partition, unique, searchsorted and set operations (NumPy drop-in) NumPy's, where NumPy defines them

Each call a mode takes over carries one of three promises - exact, equivalent where the library leaves a choice open, or approximate at a precision you asked for (rule 16). kwker.contract checks them on this machine against the library's own functions:

PythonRuns on your machine.
import kwker.contract, kwker.numpy_ops
print(kwker.contract.run(kwker.numpy_ops.routes(), seconds=1.0)["ok"])
Output
True

The report ends with what Kwker ran in your program. It lists each PyTorch kernel, NumPy function and SciPy call it took over, the calls it left to the host library, and the time spent on each side. It also shows how each language model's generate() calls ran. To see that part alone, run the program with KWKER_REPORT=1; the report is printed when the program ends:

ShellOn your machine.
KWKER_REPORT=1 python -m kwker.torch_run train.py --epochs 1

Add --json for a machine-readable report, and --timeout to limit each run.

For a program in C, C++, Rust or another language, build it twice, once without Kwker and once with it. Pass the plain build to --baseline and the Kwker build after --:

ShellOn your machine.
cc -O2 app.c -o app_std && cc -O2 -DUSE_KWKER app.c -lkwker -o app_kwker
python -m kwker audit --baseline "./app_std input.bin" --output result.bin -- ./app_kwker input.bin

Both commands run as they are, alternating, with the same timing and the same output checks.

4. Check accuracy before you use a faster precision

The ops mode never changes a result. Before you switch on bf16, int8 or int4 for a real model, measure the difference on that model with your own inputs:

PythonNeeds torch: runs on your machine.
import torch
import kwker.cpu_backend as scb

model = torch.nn.Sequential(torch.nn.Linear(256, 512), torch.nn.GELU(), torch.nn.Linear(512, 10)).eval()
x = torch.randn(32, 256)
rep = scb.check_accuracy(model, x, mode="int8")   # float32 reference vs int8, settings restored afterwards
print(rep.ok, rep.summary())

check_accuracy reports the relative error, the largest absolute error, cosine similarity and top-1 agreement for every output. See Lower precision with an accuracy check for the limits and how to change them.

What makes a comparison fair

Our published results list the machine, software versions, commands and raw tables for every figure: kwker.io/benchmarks.

Notes