Evaluate Kwker on your machine
You don't have to take our numbers on trust. Three commands show what Kwker does on your hardware and for your code. None of them uploads anything.
| Step | Command | What you learn | Takes |
|---|---|---|---|
| 1. Check the machine | python -m kwker doctor |
your CPU's features, the engine Kwker picked, warnings | seconds |
| 2. Benchmark your libraries | python -m kwker.bench --quick |
Kwker against NumPy, PyTorch, pyarrow and Polars on this CPU | about 10 s |
| 3. Run your own program | python -m kwker audit -- train.py |
end-to-end time, CPU time, and whether the results match | 6 runs of your program |
1. Check the machine
python -m kwker doctor
python -m kwker doctor --json > machine.json # the same report as JSON
doctor prints the Kwker version, the CPU and its instruction sets, the engine Kwker runs (AVX-512, AVX2, NEON, SVE or
portable) and the state of the PyTorch extensions. It also warns about anything that slows Kwker down, such as a torch
release the wheel was not built for.
2. Benchmark the libraries you use
python -m kwker.bench --quick # 100K keys, about 10 seconds
python -m kwker.bench # 1K, 100K and 1M keys, 4 key types, 3 input shapes
python -m kwker.bench --md report.md --json report.json
python -m kwker.bench --native # also x86-simd-sort and VQSort, built on this machine
python -m kwker.bench --suite --native # 43 input patterns x 6 key types, as on kwker.io/benchmarks
python -m kwker.bench --frames # tables: group-by, sorts, ORDER BY ... LIMIT
python -m kwker.bench --models --threads 4 # image classifiers: KwkCNN, PyTorch, OpenVINO
python -m kwker.bench --encoders --threads 4 # text encoders (BERT, MiniLM): KwkEncoder, OpenVINO, PyTorch
python -m kwker.bench --llm HuggingFaceTB/SmolLM2-135M --text wiki.test.raw # a language model: KwkDecoder, transformers
The benchmark times Kwker against every library you have installed: NumPy, PyTorch, pyarrow and Polars. They all get
the same input on one thread (--threads changes that), and every result is checked against NumPy's. The report names
the CPU, the engine and every library version, so you can share it as it is.
With --native, it also times Intel's x86-simd-sort and Google Highway's VQSort, the two libraries
kwker.io/benchmarks compares against. It downloads both at the commits behind the
published numbers and builds them once for your CPU, which takes about a minute. It needs g++ or clang++ and does
not run on Windows yet.
With --suite, it sorts the 43 input patterns behind the sort results on kwker.io/benchmarks
instead: random keys, skewed and duplicate-heavy keys, presorted and reversed runs, timestamps, string-like keys and
more, each as six key types (uint32, int32, float32, uint64, int64, float64) at 1K, 100K and 1M keys. The
report has one table per size: a row per pattern, a column per key type, and Kwker's speed-up over the fastest other
library in each cell. --families uniform,mod-8 picks patterns, --types and --sizes narrow it further, and
--ops sort,argsort adds argsort. The full suite takes a few minutes; add --native to include x86-simd-sort and VQSort.
With --frames, it times the data-frame cases of kwker.io/benchmarks instead: group-by,
table sorts and ORDER BY ... LIMIT on 1M-row tables, kwker.frame against pandas, Polars, pyarrow and DuckDB (each
one you have installed). Every result is checked against the library's own first. --rows changes the table size.
With --models, it times torchvision image classifiers in float32: KwkCNN against PyTorch eager, torch.compile and
OpenVINO, each output checked against eager PyTorch. Name models with --models resnet18,resnet50, and leave out the
slow torch.compile step with --no-compile. Add --int8 --images photos/ to also time int8: KwkCNN int8 against
OpenVINO int8, both calibrated on the same photos from your folder. The report shows how often each int8 model picks
the same class as the float32 model on 100 more of your photos. This uses the pretrained weights, downloaded once.
With --encoders, it times text encoders on padded batches of 1 x 128 and 8 x 128 tokens: KwkEncoder against OpenVINO,
ONNX Runtime, torch.compile(backend="kwker") and eager PyTorch, and ONNX Runtime with Kwker's provider added. It uses
bert-base-uncased and all-MiniLM-L6-v2 unless you name your own models, for example --encoders BAAI/bge-small-en-v1.5,
and --shapes 32x64 changes the batches. --int8 adds KwkEncoder int8 against OpenVINO int8, and ONNX Runtime on an
int8 export (made the way the Hub's qint8 files are) with and without Kwker. KwkEncoder needs a CPU with AVX2 or AVX-512; on other CPUs the report says so
and times the rest.
With --llm MODEL, it times a Hugging Face language model (an id or a local directory) on every core: KwkDecoder in
int4, int8 and bf16 against the model's own eager float32 forward. It reports prompt tokens a second (a 512-token prompt)
and generation tokens a second (64 tokens). Add --text with a text file to also get perplexity, scored the way
llama.cpp's llama-perplexity scores it. To put llama.cpp in the same table, pass your own build and the model's GGUF
files: --llama-bench ./llama-bench --llama-perplexity ./llama-perplexity --gguf model-Q4_0.gguf,model-Q8_0.gguf.
3. Run your own program both ways
python -m kwker audit -- train.py --epochs 1
python -m kwker audit --output preds.npy --rtol 1e-4 -- predict.py
python -m kwker audit --mode bf16 --rtol 1e-2 -- serve_bench.py
python -m kwker audit --mode numpy -- analysis.py # a NumPy program, no PyTorch needed
python -m kwker audit --mode scipy -- build_graph.py # a SciPy program, with kwker.scipy.install()
audit runs your program, unchanged, three times without Kwker and three times with it, alternating, and keeps each
side's fastest run. Then it checks that the results match: the standard output, and every file you name with
--output. The report shows both times, the speedup and the CPU time saved. Each round runs both sides back to back,
so each round gives one paired speedup; the report shows their range under the speedup, and from --repeat 6 a 90%
interval for the median instead. A wide range means the machine was busy: run again with a higher --repeat.
--mode |
What Kwker switches on | Results |
|---|---|---|
ops (default) |
kwker.torch_: Kwker runs PyTorch's CPU kernels for sorting, selection, indexing, random numbers, pooling and the other covered operations |
identical to torch, bit for bit |
bf16 |
ops, plus float32 matmul precision "medium": linears, attention and convolutions in bf16 on AMX |
within bf16 precision: set --rtol |
int8 |
ops, plus kwker.cpu_: dynamic int8 linears and convolutions |
change slightly: check them (step 4) |
numpy |
kwker.numpy_: Kwker runs NumPy's sort, argsort, partition, unique, searchsorted and set operations (NumPy drop-in) |
NumPy's, where NumPy defines them |
Each call a mode takes over carries one of three promises - exact, equivalent where the library leaves a choice open,
or approximate at a precision you asked for (rule 16). kwker.contract checks them on this
machine against the library's own functions:
import kwker.contract, kwker.numpy_ops
print(kwker.contract.run(kwker.numpy_ops.routes(), seconds=1.0)["ok"])
True
The report ends with what Kwker ran in your program. It lists each PyTorch kernel, NumPy function and SciPy call it took
over, the calls it left to the host library, and the time spent on each side. It also shows how each language model's
generate() calls ran. To see that part alone, run the program with KWKER_REPORT=1; the report is printed when the
program ends:
KWKER_REPORT=1 python -m kwker.torch_run train.py --epochs 1
Add --json for a machine-readable report, and --timeout to limit each run.
For a program in C, C++, Rust or another language, build it twice, once without Kwker and once with it. Pass the
plain build to --baseline and the Kwker build after --:
cc -O2 app.c -o app_std && cc -O2 -DUSE_KWKER app.c -lkwker -o app_kwker
python -m kwker audit --baseline "./app_std input.bin" --output result.bin -- ./app_kwker input.bin
Both commands run as they are, alternating, with the same timing and the same output checks.
4. Check accuracy before you use a faster precision
The ops mode never changes a result. Before you switch on bf16, int8 or int4 for a real model, measure the
difference on that model with your own inputs:
import torch
import kwker.cpu_backend as scb
model = torch.nn.Sequential(torch.nn.Linear(256, 512), torch.nn.GELU(), torch.nn.Linear(512, 10)).eval()
x = torch.randn(32, 256)
rep = scb.check_accuracy(model, x, mode="int8") # float32 reference vs int8, settings restored afterwards
print(rep.ok, rep.summary())
check_accuracy reports the relative error, the largest absolute error, cosine similarity and top-1 agreement for
every output. See Lower precision with an accuracy check for the limits and how to change them.
What makes a comparison fair
- Same machine, same inputs, interleaved runs.
kwker.benchandauditdo this for you. Cloud machines vary by 10-30% between runs, so compare runs made together, not on different days. - The fastest alternative, not a slow one. Compare against the library you would use in production. Our own published comparisons use x86-simd-sort, VQSort, llama.cpp, OpenVINO, Polars and DuckDB.
- Threads pinned. Single-thread numbers are the default. For multi-threaded code, give both sides the same cores.
- Results checked. A faster wrong answer is not a result. Keep
audit's output checks on.
Our published results list the machine, software versions, commands and raw tables for every figure: kwker.io/benchmarks.
Notes
- NumPy's own sort already uses x86-simd-sort on AVX-512 and AVX2 CPUs, and Highway's VQSort on ARM.
- x86-simd-sort's argsort is not stable: equal keys come back in any order, which is less work than Kwker's stable argsort.
--nativekeeps its build in~/.cache/kwker/bench.CXXpicks another compiler. Without internet access, setKWKER_BENCH_SRC_HIGHWAYandKWKER_BENCH_SRC_X86_SIMD_SORTto checkouts of the two repositories.auditcompares.npy,.npzand.ptarrays within--rtoland--atol, and other files byte for byte.--repeatsets the number of runs per side.
Related
- Add Kwker to your project: one change per workload.
- Accelerate PyTorch: precision modes, accuracy checks and the model runners.
- Troubleshooting: what to do when
doctorreports a warning.