Measured against what you would use otherwise
Every result compares Kwker with the strongest tool for the same job, on the same machine and the same cores. Ratios are the other tool's time divided by Kwker's: above 1 means Kwker is faster. Losses are published too. Measure it on your own machine.
Language models on the CPU
KwkDecoder against llama.cpp, the reference for CPU inference: a 512-token prompt, then generation one token at a time. Tokens per second, higher is better. Perplexity is the usual accuracy measure for language models: how well the model predicts real text it has not seen (Wikipedia articles). Lower is better. The small figure beside it is the change from the original model at full precision, which is what storing the weights in fewer bits costs.
| Model, 4 cores | Runtime | Prompt, tokens/s | Generation, tokens/s | Perplexity, lower is better |
|---|---|---|---|---|
| SmolLM2-135M | KwkDecoder, int4 | 3,074 | 285 | 20.04+8.4% |
| KwkDecoder, int8 | 2,397 | 208 | 18.51+0.12% | |
| llama.cpp, Q4_0 | 1,039 | 158 | 23.12+25.1% | |
| llama.cpp, Q4_K_M | 773 | 134 | 19.60+6.1% | |
| KwkDecoder, bf16 | 1,911 | 120 | 18.480.00% | |
| llama.cpp, Q8_0 | 802 | 115 | 18.53+0.25% | |
| llama.cpp, BF16 | 826 | 74.1 | 18.48-0.01% | |
| Qwen2.5-0.5B | KwkDecoder, int4 | 1,106 | 87.1 | 15.38+5.1% |
| KwkDecoder, int8 | 850 | 66.6 | 14.65+0.12% | |
| llama.cpp, Q4_0 | 421 | 50.2 | 16.32+11.6% | |
| llama.cpp, Q4_K_M | 286 | 43.7 | 15.17+3.7% | |
| KwkDecoder, bf16 | 625 | 36.5 | 14.630.00% | |
| llama.cpp, Q8_0 | 309 | 36.6 | 14.65+0.16% | |
| llama.cpp, BF16 | 312 | 23.1 | 14.63+0.03% | |
| SmolLM2-1.7B | KwkDecoder, int4 | 276 | 31.4 | 9.77+6.0% |
| KwkDecoder, int8 | 209 | 19.9 | 9.22+0.08% | |
| llama.cpp, Q4_K_M | 115 | 18.3 | 9.99+8.4% | |
| llama.cpp, Q4_0 | 109 | 17.7 | 10.44+13.3% | |
| llama.cpp, Q8_0 | 81.5 | 11.8 | 9.24+0.30% | |
| Granite 3.1 1B-A400M MoE | KwkDecoder, int4 | 827 | 92.9 | 8.85+2.4% |
| KwkDecoder, int8 | 686 | 66.9 | 8.64+0.01% | |
| llama.cpp, Q4_K_M | 247 | 64.3 | 9.11+5.4% | |
| llama.cpp, Q4_0 | 231 | 64.0 | 9.46+9.5% | |
| llama.cpp, Q8_0 | 191 | 44.3 | 8.65+0.17% |
Intel Xeon with AVX-512 VNNI, 4 cores for both. SmolLM2-135M, Qwen2.5-0.5B and SmolLM2-1.7B: one
session on 2026-10-04, the two runtimes' runs interleaved; Granite: an earlier October run (its perplexity measured on 2026-10-04 with today's formats, like every row). Each cell is the median
of repeated runs (KwkDecoder 5, llama.cpp 3); rows run fastest first by generation speed. Against llama.cpp's faster 4-bit file, KwkDecoder int4 generates
1.8× and reads prompts 2.7× faster (geometric means over the three models; the
mixture-of-experts Granite 3.1 1B-A400M - 8 of 32 experts per token -: 1.4× and 3.4×). On SmolLM2-1.7B
(1.7×) int4 keeps its higher-precision tensors at 6 bits from a billion parameters: 5.0 bits per weight
against Q4_K_M's 4.9. The bf16 rows quantize nothing on either side (llama.cpp's BF16 file of the same checkpoint;
KwkDecoder precision="bf16"): KwkDecoder generates 1.6× faster on both models and reads
prompts 2.3× (SmolLM2-135M) and 2.0× (Qwen2.5-0.5B) faster. Perplexity: the wikitext-2 test text,
scored the same way on both sides (llama.cpp's own method: 20 chunks of 512 tokens); the change is measured
against the original model run in 32-bit floats, which reads 18.48 (SmolLM2-135M), 14.63 (Qwen2.5-0.5B), 9.21 (SmolLM2-1.7B) and 8.64 (Granite).
Language models guide · Raw results (text)
Text embeddings
KwkEncoder against OpenVINO's bf16 path and against PyTorch compiled with Kwker's backend. Milliseconds per batch, lower is better.
| Model, batch × tokens | OpenVINO bf16 | KwkEncoder bf16 | torch.compile(backend="kwker") bf16 | vs OpenVINO bf16 |
|---|---|---|---|---|
| BERT-base, 1 × 128 | 17.6 | 10.7 | 18.6 | 1.64× |
| BERT-base, 8 × 128 | 114 | 65.8 | 102 | 1.74× |
| BERT-base, 1 × 512 | 93.1 | 44.4 | 85.2 | 2.10× |
| MiniLM-L6, 1 × 128 | 3.52 | 1.88 | 4.74 | 1.88× |
| MiniLM-L6, 32 × 64 | 33.3 | 26.8 | 37.7 | 1.24× |
Intel Xeon with AMX (bf16), 4 cores, OpenVINO 2026.4, October 2026. Batched rows are sentences of varying length padded to the batch length; KwkEncoder skips the padding. Output error against float32 eager is the same as OpenVINO's (5×10-3). Raw results (text) · KwkEncoder guide
Image models
KwkCNN int8 against OpenVINO int8 (NNCF quantization, oneDNN AMX-INT8 kernels): one image at a time, milliseconds per image, lower is better.
| Model | OpenVINO int8 | KwkCNN int8 | Speed-up |
|---|---|---|---|
| ResNet-18 | 1.5 | 1.0 | 1.39× |
| ResNet-50 | 3.9 | 3.2 | 1.22× |
| ResNeXt-50 32×4d | 4.5 | 3.7 | 1.21× |
| RegNetX-1.6GF | 3.7 | 2.5 | 1.52× |
| DenseNet-121 | 5.8 | 2.5 | 2.35× |
| Inception v3 | 3.7 | 2.9 | 1.24× |
| GoogLeNet | 3.0 | 1.7 | 1.79× |
| MobileNetV2 | 1.7 | 0.8 | 2.10× |
Intel Xeon with AMX, 4 cores, 224 × 224 inputs, October 2026. Both sides calibrate int8 on the same Imagenette images; top-1 accuracy is measured on Imagenette separately. Raw results (text) · KwkCNN guide
DataFrames and analytics
The same table handed to each library and to Kwker's DataFrame functions, 1 million rows. Speed-up of Kwker over each library, every result checked against the library's own.
Group-by, 4 threads each
| Query | pandas 3.0 | Polars 1.44 | pyarrow 25 | DuckDB 1.5 |
|---|---|---|---|---|
| 1 key, 5 aggregations | 4.23× | 2.11× | 2.11× | 5.54× |
| 1 key, 5 aggregations + median | 4.25× | 1.21× | – | 5.40× |
| 2 keys, 3 aggregations | 7.40× | 1.69× | 1.49× | 4.02× |
Sorting a table, 5 columns, 1 thread
| Sort key | pandas | Polars (stable) | pyarrow |
|---|---|---|---|
| int64 | 2.30× | 1.59× | 2.88× |
| float64 with NaN | 2.04× | 1.55× | 2.77× |
| string, 1000 distinct values | 3.32× | 7.83× | 2.47× |
| 2 columns (int, float) | 3.07× | 5.84× | 2.27× |
ORDER BY a, b LIMIT k, 1 thread
| k | DuckDB | Kwker | Speed-up |
|---|---|---|---|
| 10 | 4.27 ms | 2.82 ms | 1.52× |
| 1,000 | 7.39 ms | 3.04 ms | 2.43× |
| 100,000 | 228 ms | 10.5 ms | 21.7× |
Intel Xeon (Cascade Lake, AVX-512), production build, October 2026. The same query against Polars' bottom_k: 11.2× / 9.1× / 4.6× at k = 10 / 1,000 / 100,000; pandas' nsmallest: at least 3.2× at every k. Group-by: one int64 key of 1,000 values with sum, mean, min, max and count of a float column (plus the exact median), or two keys (an int64 of 50 values and a string of 100) with sum, mean and count; results follow each library's own null and NaN rules. DataFrames guide · Raw results (text)
Against each language's own sort
What switching to Kwker means inside an existing service: one million keys, the built-in sort of the language against Kwker's binding, same process, best of several runs.
| Language | Keys | Built-in sort | Built-in | Kwker | Speed-up |
|---|---|---|---|---|---|
| Java 25 | int | Arrays.sort | 9.03 ms | 2.97 ms | 3.0× |
| Java 25 | long | Arrays.sort | 16.78 ms | 5.58 ms | 3.0× |
| Java 25 | float | Arrays.sort | 11.74 ms | 3.20 ms | 3.7× |
| Java 25 | double | Arrays.sort | 19.19 ms | 6.13 ms | 3.1× |
| JavaScript (Deno) | Float32Array | TypedArray.sort | 78.29 ms | 2.71 ms | 28.9× |
| JavaScript (Deno) | Float64Array | TypedArray.sort | 80.68 ms | 5.03 ms | 16.0× |
| JavaScript (Deno) | BigInt64Array | TypedArray.sort | 19.46 ms | 6.61 ms | 2.9× |
AVX-512 engine, October 2026. Raw results (text) · Every language, with examples
Kwker Core
The engine against the fastest sorting libraries
Kwker Core is the sorting, selection and search engine under everything above. It is compared with the fastest public library for each case: Intel's x86-simd-sort, Google's VQSort (Highway), djbsort and Rust's standard sort.
Sorting, one call at a time
Each cell sorts 32,768 keys as separate calls of n keys. Six key types (u32, i32, f32, u64, i64, f64), seven input patterns, 16 sizes: 672 cells per engine. Geometric mean speed-up over the fastest competitor in each cell.
| Engine | 2 | 3 | 4 | 6 | 8 | 10 | 12 | 17 | 24 | 35 | 49 | 70 | 100 | 200 | 400 | 900 | All |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AVX-512 | 3.62× | 2.97× | 2.24× | 2.85× | 2.04× | 2.54× | 2.47× | 2.53× | 2.48× | 3.23× | 2.81× | 4.03× | 3.02× | 2.70× | 3.06× | 3.59× | 2.84× |
| AVX2 | 3.69× | 3.46× | 2.48× | 3.36× | 2.02× | 2.70× | 2.54× | 2.59× | 2.44× | 3.28× | 2.65× | 3.61× | 2.86× | 3.26× | 3.67× | 4.44× | 3.01× |
Where Kwker Core is not faster
| Engine | Cell | Ratio | Faster library |
|---|---|---|---|
| AVX-512 | u32, 8 keys, values 0 to 19 (srr-d20) | 0.92× | VQSort |
| AVX-512 | u64, 200 keys, 95% zeros and 5% random (srr-p5) | 0.995× | VQSort |
| AVX2 | f64, 400 keys, values 0 to 19 (srr-d20) | 0.93× | x86-simd-sort |
AVX-512: 670 of 672 cells faster, 20 below 1.25×, 182 above 4×. AVX2: 671 of 672 faster, 39 below 1.25×, 192 above 4×. Input patterns: uniform random, Zipf, values 0 to 19, 95% zeros with 5% random, 95% sorted with a random tail, sorted, reversed (the srr- patterns are those of sort-research-rs). Every output is checked. Raw results (text) · Sorting guide
Larger arrays
geometric mean speed-up, AVX2 engine
cells where Kwker Core is faster
cells below 1.25× (11 of them below 1×)
of outputs verified against a reference sort
The release regression gate: 6 key types, 40 input patterns (random, skewed, few distinct values, runs, duplicates, ASCII text and more), 1K, 100K and 1M keys, against the faster of x86-simd-sort and VQSort. Single thread, Intel Xeon (Cascade Lake), production build, October 2026. Raw results (TSV)
Method
- Both sides run on the same machine, the same cores and the same thread count. Competitors are current releases (model runtimes, DataFrame libraries) or built from their current sources with the same compiler and flags as Kwker (sorting libraries, including AVX2 builds when the AVX2 engine is measured).
- Runs are interleaved in rounds after a warm-up, with the data prepared before timing starts; the best or the lower median of the rounds counts, as each table says. Where a whole benchmark was repeated, a cell shows the median of the repeats (language models: 5 KwkDecoder runs, 3 llama-bench runs); summary figures are geometric means.
- Every output is checked: sorted arrays against a reference sort, DataFrame results against the library's own, model outputs against float32 eager PyTorch.
- Model timings use the models' real shapes; weights do not change the speed of a dense model, and quality is measured separately (perplexity for language models, top-1 accuracy for image models).
- A held-out set of real data that is never used for tuning (token ids, model weights, file-system metadata) checks that Kwker Core's gains carry over beyond the benchmark patterns.
Every table above has its raw results: machine, date, software versions on both sides, the protocol and the
unrounded numbers, in one text file per section. To time Kwker against the
libraries installed on your own machine, run python -m kwker.bench; it writes a Markdown and JSON report.
python -m kwker.bench --suite --native sorts the 40 input patterns of the regression gate above, plus three
sort-research-rs patterns, as all six key types against x86-simd-sort and VQSort built on your machine.
On your own machine, python -m kwker doctor reports the engine Kwker chose for your processor and checks it; see Troubleshooting.