Kwker
Benchmarks

Measured against what you would use otherwise

Every result compares Kwker with the strongest tool for the same job, on the same machine and the same cores. Ratios are the other tool's time divided by Kwker's: above 1 means Kwker is faster. Losses are published too. Measure it on your own machine.

Language models on the CPU

KwkDecoder against llama.cpp, the reference for CPU inference: a 512-token prompt, then generation one token at a time. Tokens per second, higher is better. Perplexity is the usual accuracy measure for language models: how well the model predicts real text it has not seen (Wikipedia articles). Lower is better. The small figure beside it is the change from the original model at full precision, which is what storing the weights in fewer bits costs.

Model, 4 coresRuntimePrompt, tokens/sGeneration, tokens/sPerplexity, lower is better
SmolLM2-135MKwkDecoder, int43,07428520.04+8.4%
KwkDecoder, int82,39720818.51+0.12%
llama.cpp, Q4_01,03915823.12+25.1%
llama.cpp, Q4_K_M77313419.60+6.1%
KwkDecoder, bf161,91112018.480.00%
llama.cpp, Q8_080211518.53+0.25%
llama.cpp, BF1682674.118.48-0.01%
Qwen2.5-0.5BKwkDecoder, int41,10687.115.38+5.1%
KwkDecoder, int885066.614.65+0.12%
llama.cpp, Q4_042150.216.32+11.6%
llama.cpp, Q4_K_M28643.715.17+3.7%
KwkDecoder, bf1662536.514.630.00%
llama.cpp, Q8_030936.614.65+0.16%
llama.cpp, BF1631223.114.63+0.03%
SmolLM2-1.7BKwkDecoder, int427631.49.77+6.0%
KwkDecoder, int820919.99.22+0.08%
llama.cpp, Q4_K_M11518.39.99+8.4%
llama.cpp, Q4_010917.710.44+13.3%
llama.cpp, Q8_081.511.89.24+0.30%
Granite 3.1 1B-A400M MoEKwkDecoder, int482792.98.85+2.4%
KwkDecoder, int868666.98.64+0.01%
llama.cpp, Q4_K_M24764.39.11+5.4%
llama.cpp, Q4_023164.09.46+9.5%
llama.cpp, Q8_019144.38.65+0.17%

Intel Xeon with AVX-512 VNNI, 4 cores for both. SmolLM2-135M, Qwen2.5-0.5B and SmolLM2-1.7B: one session on 2026-10-04, the two runtimes' runs interleaved; Granite: an earlier October run (its perplexity measured on 2026-10-04 with today's formats, like every row). Each cell is the median of repeated runs (KwkDecoder 5, llama.cpp 3); rows run fastest first by generation speed. Against llama.cpp's faster 4-bit file, KwkDecoder int4 generates 1.8× and reads prompts 2.7× faster (geometric means over the three models; the mixture-of-experts Granite 3.1 1B-A400M - 8 of 32 experts per token -: 1.4× and 3.4×). On SmolLM2-1.7B (1.7×) int4 keeps its higher-precision tensors at 6 bits from a billion parameters: 5.0 bits per weight against Q4_K_M's 4.9. The bf16 rows quantize nothing on either side (llama.cpp's BF16 file of the same checkpoint; KwkDecoder precision="bf16"): KwkDecoder generates 1.6× faster on both models and reads prompts 2.3× (SmolLM2-135M) and 2.0× (Qwen2.5-0.5B) faster. Perplexity: the wikitext-2 test text, scored the same way on both sides (llama.cpp's own method: 20 chunks of 512 tokens); the change is measured against the original model run in 32-bit floats, which reads 18.48 (SmolLM2-135M), 14.63 (Qwen2.5-0.5B), 9.21 (SmolLM2-1.7B) and 8.64 (Granite). Language models guide · Raw results (text)

Text embeddings

KwkEncoder against OpenVINO's bf16 path and against PyTorch compiled with Kwker's backend. Milliseconds per batch, lower is better.

Model, batch × tokensOpenVINO bf16KwkEncoder bf16torch.compile(backend="kwker") bf16vs OpenVINO bf16
BERT-base, 1 × 12817.610.718.61.64×
BERT-base, 8 × 12811465.81021.74×
BERT-base, 1 × 51293.144.485.22.10×
MiniLM-L6, 1 × 1283.521.884.741.88×
MiniLM-L6, 32 × 6433.326.837.71.24×

Intel Xeon with AMX (bf16), 4 cores, OpenVINO 2026.4, October 2026. Batched rows are sentences of varying length padded to the batch length; KwkEncoder skips the padding. Output error against float32 eager is the same as OpenVINO's (5×10-3). Raw results (text) · KwkEncoder guide

Image models

KwkCNN int8 against OpenVINO int8 (NNCF quantization, oneDNN AMX-INT8 kernels): one image at a time, milliseconds per image, lower is better.

ModelOpenVINO int8KwkCNN int8Speed-up
ResNet-181.51.01.39×
ResNet-503.93.21.22×
ResNeXt-50 32×4d4.53.71.21×
RegNetX-1.6GF3.72.51.52×
DenseNet-1215.82.52.35×
Inception v33.72.91.24×
GoogLeNet3.01.71.79×
MobileNetV21.70.82.10×

Intel Xeon with AMX, 4 cores, 224 × 224 inputs, October 2026. Both sides calibrate int8 on the same Imagenette images; top-1 accuracy is measured on Imagenette separately. Raw results (text) · KwkCNN guide

DataFrames and analytics

The same table handed to each library and to Kwker's DataFrame functions, 1 million rows. Speed-up of Kwker over each library, every result checked against the library's own.

Group-by, 4 threads each

Querypandas 3.0Polars 1.44pyarrow 25DuckDB 1.5
1 key, 5 aggregations4.23×2.11×2.11×5.54×
1 key, 5 aggregations + median4.25×1.21×–5.40×
2 keys, 3 aggregations7.40×1.69×1.49×4.02×

Sorting a table, 5 columns, 1 thread

Sort keypandasPolars (stable)pyarrow
int642.30×1.59×2.88×
float64 with NaN2.04×1.55×2.77×
string, 1000 distinct values3.32×7.83×2.47×
2 columns (int, float)3.07×5.84×2.27×

ORDER BY a, b LIMIT k, 1 thread

kDuckDBKwkerSpeed-up
104.27 ms2.82 ms1.52×
1,0007.39 ms3.04 ms2.43×
100,000228 ms10.5 ms21.7×

Intel Xeon (Cascade Lake, AVX-512), production build, October 2026. The same query against Polars' bottom_k: 11.2× / 9.1× / 4.6× at k = 10 / 1,000 / 100,000; pandas' nsmallest: at least 3.2× at every k. Group-by: one int64 key of 1,000 values with sum, mean, min, max and count of a float column (plus the exact median), or two keys (an int64 of 50 values and a string of 100) with sum, mean and count; results follow each library's own null and NaN rules. DataFrames guide · Raw results (text)

Against each language's own sort

What switching to Kwker means inside an existing service: one million keys, the built-in sort of the language against Kwker's binding, same process, best of several runs.

LanguageKeysBuilt-in sortBuilt-inKwkerSpeed-up
Java 25intArrays.sort9.03 ms2.97 ms3.0×
Java 25longArrays.sort16.78 ms5.58 ms3.0×
Java 25floatArrays.sort11.74 ms3.20 ms3.7×
Java 25doubleArrays.sort19.19 ms6.13 ms3.1×
JavaScript (Deno)Float32ArrayTypedArray.sort78.29 ms2.71 ms28.9×
JavaScript (Deno)Float64ArrayTypedArray.sort80.68 ms5.03 ms16.0×
JavaScript (Deno)BigInt64ArrayTypedArray.sort19.46 ms6.61 ms2.9×

AVX-512 engine, October 2026. Raw results (text) · Every language, with examples

Kwker Core

The engine against the fastest sorting libraries

Kwker Core is the sorting, selection and search engine under everything above. It is compared with the fastest public library for each case: Intel's x86-simd-sort, Google's VQSort (Highway), djbsort and Rust's standard sort.

Sorting, one call at a time

Each cell sorts 32,768 keys as separate calls of n keys. Six key types (u32, i32, f32, u64, i64, f64), seven input patterns, 16 sizes: 672 cells per engine. Geometric mean speed-up over the fastest competitor in each cell.

Keys per call
Engine2346810121724354970100200400900All
AVX-5123.62×2.97×2.24×2.85×2.04×2.54×2.47×2.53×2.48×3.23×2.81×4.03×3.02×2.70×3.06×3.59×2.84×
AVX23.69×3.46×2.48×3.36×2.02×2.70×2.54×2.59×2.44×3.28×2.65×3.61×2.86×3.26×3.67×4.44×3.01×

Where Kwker Core is not faster

EngineCellRatioFaster library
AVX-512u32, 8 keys, values 0 to 19 (srr-d20)0.92×VQSort
AVX-512u64, 200 keys, 95% zeros and 5% random (srr-p5)0.995×VQSort
AVX2f64, 400 keys, values 0 to 19 (srr-d20)0.93×x86-simd-sort

AVX-512: 670 of 672 cells faster, 20 below 1.25×, 182 above 4×. AVX2: 671 of 672 faster, 39 below 1.25×, 192 above 4×. Input patterns: uniform random, Zipf, values 0 to 19, 95% zeros with 5% random, 95% sorted with a random tail, sorted, reversed (the srr- patterns are those of sort-research-rs). Every output is checked. Raw results (text) · Sorting guide

Larger arrays

2.22×

geometric mean speed-up, AVX2 engine

709 / 720

cells where Kwker Core is faster

105

cells below 1.25× (11 of them below 1×)

100%

of outputs verified against a reference sort

The release regression gate: 6 key types, 40 input patterns (random, skewed, few distinct values, runs, duplicates, ASCII text and more), 1K, 100K and 1M keys, against the faster of x86-simd-sort and VQSort. Single thread, Intel Xeon (Cascade Lake), production build, October 2026. Raw results (TSV)

Method

  • Both sides run on the same machine, the same cores and the same thread count. Competitors are current releases (model runtimes, DataFrame libraries) or built from their current sources with the same compiler and flags as Kwker (sorting libraries, including AVX2 builds when the AVX2 engine is measured).
  • Runs are interleaved in rounds after a warm-up, with the data prepared before timing starts; the best or the lower median of the rounds counts, as each table says. Where a whole benchmark was repeated, a cell shows the median of the repeats (language models: 5 KwkDecoder runs, 3 llama-bench runs); summary figures are geometric means.
  • Every output is checked: sorted arrays against a reference sort, DataFrame results against the library's own, model outputs against float32 eager PyTorch.
  • Model timings use the models' real shapes; weights do not change the speed of a dense model, and quality is measured separately (perplexity for language models, top-1 accuracy for image models).
  • A held-out set of real data that is never used for tuning (token ids, model weights, file-system metadata) checks that Kwker Core's gains carry over beyond the benchmark patterns.

Every table above has its raw results: machine, date, software versions on both sides, the protocol and the unrounded numbers, in one text file per section. To time Kwker against the libraries installed on your own machine, run python -m kwker.bench; it writes a Markdown and JSON report. python -m kwker.bench --suite --native sorts the 40 input patterns of the regression gate above, plus three sort-research-rs patterns, as all six key types against x86-simd-sort and VQSort built on your machine.

On your own machine, python -m kwker doctor reports the engine Kwker chose for your processor and checks it; see Troubleshooting.