Kwker benchmark results: raw data behind https://kwker.io/benchmarks/ ===================================================================== Every number on the benchmarks page comes from one of these files. Each file names the machine, the date, the software versions on both sides, the protocol (repetitions, best-of rule, thread count) and how correctness was checked. Losses are listed with the wins. llm.txt Language models: KwkDecoder vs llama.cpp (prompt and generation speed, long context, perplexity on wikitext-2) embeddings.txt Text embeddings: KwkEncoder vs OpenVINO bf16 / int8 and torch.compile(backend="kwker") vision.txt Image models: KwkCNN vs OpenVINO int8 / fp32, PyTorch eager and Inductor; int8 accuracy dataframes.txt Group-by, table sorts and ORDER BY ... LIMIT vs pandas, Polars, pyarrow and DuckDB languages.txt Each language's built-in sort (Java, JavaScript) vs Kwker's binding ../grid672_2026-10-03.txt Kwker Core, one call at a time: 672 cells per engine (2 to 900 keys per call) vs x86-simd-sort, VQSort, djbsort and Rust's sort, AVX-512 and AVX2 core-regression-avx2.tsv Kwker Core, larger arrays: the release regression gate, 720 cells (1K / 100K / 1M keys, 6 key types x 40 input patterns) vs x86-simd-sort and VQSort, AVX2 engine How to read the ratios: speed-up = the other tool's time / Kwker's time; above 1 means Kwker is faster. One number per cell: where a benchmark was repeated, the page shows the median of the runs (language models: 5 KwkDecoder runs, 3 llama-bench runs) and the file lists every run; summary figures on the site are geometric means of the cases they cover. Software environment (every file): Ubuntu 24.04 LTS on cloud VMs (Intel Xeon; the CPU model is named in each file). Kwker's C++ engines and every sorting library built from source (x86-simd-sort, VQSort, djbsort) use GCC 13.3.0 with the same flags; llama.cpp is built in the same environment. Python 3.11 with PyTorch 2.14.0 (CPU build), torchvision 0.29.0, transformers 5.17.0, OpenVINO 2026.4.0, pandas 3.0.6, Polars 1.44.2, pyarrow 25.0.1, DuckDB 1.5.6. Kwker production builds (files before 2026-10-08 print the development numbering 0.26.0; the first public release is 0.1.0). Re-measured on 2026-10-08 (Cascade Lake class VM, no AMX): dataframes.txt, languages.txt, core-regression-avx2.tsv and the SmolLM2-135M / Qwen2.5-0.5B rows of llm.txt. Kept from earlier October runs: embeddings.txt and vision.txt (measured on an AMX machine - the 2026-10-08 machine has no AMX), llm.txt's SmolLM2-1.7B and Granite rows (their model files did not fit that machine's disk). The 2026-10-08 host read every runtime slower than the earlier October sessions (sorting libraries about 1.2x, memory-bound table sorts and language models 1.3-2x); each table compares columns of one session only. Run-to-run variation: these are cloud VMs. Single timings vary by 10-30% between runs of the same code; every table is therefore either interleaved (all columns timed in alternating rounds in one process) or best-of several runs, and compares columns within one run. Re-run on your own machine before relying on a number: `python -m kwker.bench` times Kwker against the libraries installed there and writes a Markdown / JSON report, and `python -m kwker.bench --native` also downloads x86-simd-sort (commit fa944ef, 2026-06-25) and Highway's VQSort (commit 32ae733, 2026-09-25) - the versions behind the Kwker Core files above - builds them for that CPU with its C++ compiler and times them directly. `python -m kwker.bench --frames` runs the cases of dataframes.txt (group-by, table sorts, ORDER BY ... LIMIT) against the pandas, Polars, pyarrow and DuckDB installed there, each result checked against the library's own first, and `python -m kwker.bench --models` the float32 rows of vision.txt (KwkCNN vs PyTorch eager, Inductor and OpenVINO), and with `--int8 --images DIR` the int8 rows (KwkCNN int8 vs OpenVINO int8 via NNCF, both calibrated on the same 16 photos from DIR, top-1 agreement with float32 on 100 more). `python -m kwker.bench --encoders --int8 --shapes 1x128,8x128` the rows of embeddings.txt (KwkEncoder bf16 / int8 vs OpenVINO bf16 / int8, torch.compile(backend="kwker"); AMX CPUs). `python -m kwker.bench --llm --text wiki.test.raw` the speed and perplexity columns of llm.txt for one model (KwkDecoder int4 / int8 / bf16 vs transformers eager float32; with --llama-bench / --llama-perplexity / --gguf, the reader's own llama.cpp build in the same table - the protocol above: pp512, tg64, perplexity -c 512 --chunks 20).