Kwker benchmark results: language models on the CPU (KwkDecoder vs llama.cpp) ============================================================================== Published on https://kwker.io/benchmarks/#llm. Each cell on the page is the median of the runs listed here. Machine Intel Xeon (Cascade Lake class, AVX-512 + VNNI, no AMX), cloud VM, 4 cores used by both sides Date October 2026. SmolLM2-135M and Qwen2.5-0.5B (int4, int8 and bf16 against Q4_0, Q4_K_M, Q8_0 and BF16) re-measured on 2026-10-08 in one session, KwkDecoder and llama.cpp runs interleaved per model; the host read both runtimes 25-50% slower than on 2026-10-04 (shared cloud VM), so compare within a model's table only. SmolLM2-1.7B: the 2026-10-04 session (its files do not fit this machine's disk now); Granite: an earlier October run. llama.cpp commit 19e28a2 (2026-09-29), CMake Release build with GGML_NATIVE (AVX-512 VNNI), OpenMP, 4 threads measured with llama-bench (prompt: -p 512, generation: -n 64) Kwker KwkDecoder (kwker.decode), int4 or int8 weights, or bf16 (weights="bf16": nothing quantized - the checkpoint's bf16 weights, float32 activations and sums), fp16 KV cache, 4 threads bf16 rows KwkDecoder bf16 / llama.cpp BF16 (SmolLM2-135M, Qwen2.5-0.5B): the same session, protocol and interleaving; llama.cpp's BF16 GGUF converted from the same snapshot (convert_hf_to_gguf.py --outtype bf16), nothing quantized on either side Models SmolLM2-135M, Qwen2.5-0.5B and SmolLM2-1.7B (Hugging Face checkpoints; llama.cpp GGUF files converted from the same snapshot with convert_hf_to_gguf.py, bf16, then llama-quantize. For the 2026-10-04 session SmolLM2-1.7B's Q4_0 / Q4_K_M files were quantized from its Q8_0 file to fit the disk - a format's speed does not depend on that; the perplexity table below uses files converted from bf16) Protocol a 512-token prompt (prefill), then one token at a time (generation). llama.cpp: 3 llama-bench runs per format, each the mean of llama-bench's 5 repetitions. KwkDecoder: 5 runs per format (kwker-py/tests/fd_bench.py): prompt = the fastest of 3 prefills in a run, generation = 64 greedy tokens in KwkDecoder's native loop after a 64-token prompt, the median of 3 loops' per-token means (llama-bench's tg64 is a native 64-token loop too, from depth 0). Per model: llama.cpp in rounds 1-3, KwkDecoder in rounds 1-5, alternating, all in one session. Every cell below is the median of its runs; all runs follow each table. Speed-ups quoted on the site are geometric means over the three models. Earlier versions of this file had KwkDecoder rows from 2026-10-03 next to llama.cpp rows from earlier in October; llama.cpp read 4-10% slower on this machine by then, so those speed-ups leaned in llama.cpp's favour and these do not. Formats KwkDecoder int4 = 4-bit weights in groups of 32 inputs; each group's scale and offset are 8-bit codes under an fp16 pair per 256 inputs (4.625 bits per weight since 2026-10-03, commit a8fb577; before that an fp16 scale and offset per group, 5.0 bits). The default mix keeps v_proj, lm_head and down_proj on llama.cpp Q4_K_M's "more bits" layers at 8 bits (8.25 bits per weight with their scales since 2026-10-03; 8.75 before, with per-column sums - the rows below were measured with 8.75). With both 2026-10-03 formats SmolLM2-1.7B int4 reads 1,172 MB a token (5.48 bits per weight; tests/pack_bytes.py) and int8 1,765 MB (8.25). From a billion parameters those tensors are 6-bit by default since 2026-10-04 (6.25 bits with their scales): SmolLM2-1.7B int4 reads 1,071 MB a token (5.0 bits per weight; Q4_K_M 1,056 MB) - the SmolLM2-1.7B rows below. KwkDecoder int8 = 8-bit weights, a scale per 64 inputs. Average bits per weight on SmolLM2-1.7B: KwkDecoder int4 about 5.6 (23.5% of the weights at 8 bits, ~1.20 GB read per token; 5.9 and 1.26 GB with the 5.0-bit format); llama.cpp files: Q4_0 4.63 (991 MB), Q4_K_M 4.94 (1,056 MB), Q8_0 8.51 (1,820 MB). Generation is bound by the bytes read per token, so KwkDecoder int4 reads about 13% more per token than Q4_K_M on this model. The 2026-10-04 rows use the 4.625-bit format (the Granite rows: the 5.0-bit one). Same-session A/Bs of the two formats (alternating processes, medians of 5 runs; 3 on SmolLM2-1.7B): generation 3.87 -> 3.69 ms per token on SmolLM2-135M (1.05x), 12.01 -> 11.69 on Qwen2.5-0.5B (1.03x), 36.22 -> 34.58 on SmolLM2-1.7B (1.05x); prompt 512 within noise. Wikitext-2 perplexity (20 chunks) vs float32: +8.29 / +4.91 / +4.77% (was +7.83 / +5.58 / +4.70%). (Corrected 2026-10-03: earlier versions of this file said groups of 64 and 5.2 bits per weight - the old int4 format's figures.) SmolLM2-135M (135M parameters, hidden size 576) prompt, 512 tokens generation, per token KwkDecoder int4 0.18 s (2,774 tokens/s) 5.4 ms (185.2 tokens/s) KwkDecoder int8 0.30 s (1,734 tokens/s) 7.4 ms (135.2 tokens/s) llama.cpp Q4_0 0.73 s ( 701 tokens/s) 9.5 ms (104.8 tokens/s) llama.cpp Q4_K_M 1.18 s ( 432 tokens/s) 11.2 ms ( 89.7 tokens/s) KwkDecoder bf16 0.34 s (1,518 tokens/s) 11.6 ms ( 86.3 tokens/s) llama.cpp Q8_0 1.01 s ( 509 tokens/s) 14.3 ms ( 70.0 tokens/s) llama.cpp BF16 0.85 s ( 600 tokens/s) 20.5 ms ( 48.9 tokens/s) runs (prompt tokens/s | generation tokens/s): KwkDecoder int4 1,574 / 2,052 / 2,791 / 2,774 / 2,777 | 184.1 / 186.3 / 188.2 / 181.3 / 185.2 KwkDecoder int8 1,740 / 1,697 / 1,734 / 1,831 / 1,662 | 106.7 / 135.2 / 150.2 / 115.3 / 148.9 llama.cpp Q4_0 714 / 692 / 701 | 109.3 / 104.8 / 100.6 llama.cpp Q4_K_M 416 / 483 / 432 | 76.9 / 91.6 / 89.7 KwkDecoder bf16 1,217 / 1,576 / 1,159 / 1,626 / 1,518 | 86.3 / 93.8 / 78.7 / 87.1 / 82.0 llama.cpp Q8_0 509 / 535 / 419 | 73.1 / 70.0 / 67.7 llama.cpp BF16 624 / 600 / 528 | 50.1 / 45.8 / 48.9 Qwen2.5-0.5B (494M parameters, hidden size 896, 151,936-token vocabulary) prompt, 512 tokens generation, per token KwkDecoder int4 0.60 s ( 853 tokens/s) 15.6 ms ( 64.3 tokens/s) KwkDecoder int8 0.90 s ( 572 tokens/s) 21.5 ms ( 46.6 tokens/s) llama.cpp Q4_0 1.77 s ( 289 tokens/s) 24.9 ms ( 40.2 tokens/s) llama.cpp Q4_K_M 2.99 s ( 171 tokens/s) 31.7 ms ( 31.5 tokens/s) llama.cpp Q8_0 2.70 s ( 189 tokens/s) 32.4 ms ( 30.9 tokens/s) KwkDecoder bf16 1.07 s ( 477 tokens/s) 35.8 ms ( 27.9 tokens/s) llama.cpp BF16 2.17 s ( 236 tokens/s) 53.1 ms ( 18.8 tokens/s) runs (prompt tokens/s | generation tokens/s): KwkDecoder int4 722 / 853 / 754 / 916 / 929 | 62.0 / 61.6 / 66.3 / 67.5 / 64.3 KwkDecoder int8 523 / 572 / 516 / 605 / 608 | 46.5 / 45.5 / 46.6 / 51.8 / 50.1 llama.cpp Q4_0 291 / 270 / 289 | 43.4 / 39.0 / 40.2 llama.cpp Q4_K_M 171 / 164 / 197 | 31.1 / 31.5 / 35.2 llama.cpp Q8_0 189 / 169 / 191 | 28.0 / 30.9 / 31.4 KwkDecoder bf16 424 / 487 / 445 / 523 / 477 | 25.8 / 26.9 / 28.6 / 27.9 / 28.1 llama.cpp BF16 230 / 237 / 236 | 19.3 / 18.2 / 18.8 SmolLM2-1.7B (1.71B parameters, 24 layers, hidden size 2048; int4 with the 6-bit mix) prompt, 512 tokens generation, per token KwkDecoder int4 1.86 s ( 276 tokens/s) 31.8 ms ( 31.4 tokens/s) KwkDecoder int8 2.45 s ( 209 tokens/s) 50.3 ms ( 19.9 tokens/s) llama.cpp Q4_K_M 4.45 s ( 115 tokens/s) 54.6 ms ( 18.3 tokens/s) llama.cpp Q4_0 4.69 s ( 109 tokens/s) 56.5 ms ( 17.7 tokens/s) llama.cpp Q8_0 6.28 s ( 81.5 tokens/s) 84.5 ms ( 11.8 tokens/s) runs (prompt tokens/s | generation tokens/s): KwkDecoder int4 271 / 270 / 277 / 277 / 276 | 32.2 / 31.4 / 29.8 / 31.1 / 32.1 KwkDecoder int8 209 / 211 / 204 / 202 / 209 | 19.3 / 20.2 / 20.2 / 19.2 / 19.9 llama.cpp Q4_K_M 111 / 116 / 115 | 18.2 / 18.3 / 19.5 llama.cpp Q4_0 110 / 109 / 109 | 19.0 / 17.3 / 17.7 llama.cpp Q8_0 82.9 / 81.5 / 79.0 | 12.4 / 11.8 / 11.8 Granite 3.1 1B-A400M (ibm-granite/granite-3.1-1b-a400m-base; a mixture of experts: 1.33B parameters, about 400M active per token - 32 experts of intermediate size 512 per layer, 8 chosen per token -, hidden size 1024, 24 layers, 49,152-token vocabulary; October 2026, not in the three-model means below). KwkDecoder: the checkpoint in bfloat16 (fd_bench DTYPE=bf16), int4 and int8 in separate processes; llama.cpp: GGUFs converted from the same checkpoint (ss llama --gguf, bfloat16 source) prompt, 512 tokens generation, per token KwkDecoder int4 0.62 s ( 827 tokens/s) 10.8 ms ( 92.9 tokens/s) 3.35x / 1.44x the faster llama.cpp file (Q4_K_M) KwkDecoder int8 0.75 s ( 686 tokens/s) 15.0 ms ( 66.9 tokens/s) 3.60x / 1.51x llama.cpp Q8_0 llama.cpp Q4_K_M 2.08 s ( 247 tokens/s) 15.5 ms ( 64.3 tokens/s) llama.cpp Q4_0 2.22 s ( 231 tokens/s) 15.6 ms ( 64.0 tokens/s) llama.cpp Q8_0 2.68 s ( 191 tokens/s) 22.6 ms ( 44.3 tokens/s) runs (prompt tokens/s | generation tokens/s): KwkDecoder int4 808 / 845 / 913 / 827 / 813 | 96.2 / 91.9 / 90.9 / 93.6 / 92.9 KwkDecoder int8 686 / 692 / 599 / 730 / 686 | 69.7 / 69.3 / 62.2 / 66.9 / 66.5 llama.cpp Q4_K_M 241.6 / 246.7 / 247.6 | 65.08 / 64.34 / 60.12 llama.cpp Q4_0 223.5 / 231.7 / 230.8 | 62.23 / 64.61 / 63.95 llama.cpp Q8_0 187.7 / 190.8 / 192.0 | 44.48 / 44.15 / 44.27 Greedy text against Hugging Face's bfloat16 model, 4 prompts x 24 tokens: KwkDecoder int8 matched all 96 tokens; int4 matched 8 / 24 / 24 / 2 tokens before the first difference ("red, yellow, and blue" for "red, blue, and yellow"). Where a step's time goes (int4, 10.2 ms in a profiled run): experts' gate / up 3.7 ms, down 2.7 ms, q / k / v 1.2 ms, output layer 1.4 ms, o 0.6 ms, attention 0.3 ms - about 285 MB of weights a token, 8.1 ms at this machine's ~35 GB/s. Speed-ups (median against median; 4-bit against the faster of llama.cpp's Q4_0 and Q4_K_M for each model and column; SmolLM2-135M and Qwen2.5-0.5B: 2026-10-08, one session; SmolLM2-1.7B: 2026-10-04) SmolLM2-135M Qwen2.5-0.5B SmolLM2-1.7B geometric mean generation, int4 1.77x 1.60x 1.72x 1.69x prompt, int4 3.96x 2.95x 2.40x 3.04x generation, int8 vs Q8_0 1.93x 1.51x 1.69x 1.70x prompt, int8 vs Q8_0 3.41x 3.02x 2.56x 2.98x generation, bf16 vs BF16 1.77x 1.48x - 1.62x (two models) prompt, bf16 vs BF16 2.53x 2.02x - 2.26x (two models) (Earlier versions of this file also had Gemma 3 270M, Gemma 2 2B and GPT-2 rows; their means were over four models - SmolLM2-135M, Gemma 3 270M, Qwen2.5-0.5B, SmolLM2-1.7B: 1.80x / 2.74x / 1.66x / 2.72x in the earlier 2026-10-04 sessions.) On SmolLM2-1.7B both read about the same bytes per token since int4's 6-bit mix (1,071 MB against Q4_K_M's 1,056). On Qwen2.5-0.5B the tied 151K x 896 output layer, kept at 8 bits, is about 37% of the bytes KwkDecoder int4 reads per token. Speculative decoding (October 2026): SmolLM2-1.7B int4, KwkDecoder.generate, the output = greedy decoding's ------------------------------------------------------------------------------------------------------------- 4 cores; 4 open-ended prompts (kwker-py/tests/spec_bench.py), 128 new tokens each, tokens per second including the prompt pass; 3 runs per series, the median per prompt, then the geometric mean over the prompts. Draft model: SmolLM2-135M int4. Prompt lookup = proposals copied from earlier in the text (generate()'s default, lookup=8). plain prompt lookup draft model draft + lookup geometric mean 25.5 37.1 48.0 51.5 speed-up 1.00x 1.47x 1.88x 2.02x (prompt lookup was measured in its own series: plain 25.3 there; the other columns in one series) runs (tokens/s per prompt, run 1 / 2 / 3): Once upon a time, in a s draft k=4 36.3 / 41.2 / 37.4 Once upon a time, in a s draft+lookup 39.1 / 40.5 / 40.2 Once upon a time, in a s plain 25.5 / 26.9 / 25.6 Photosynthesis is the pr draft k=4 58.1 / 60.8 / 62.0 Photosynthesis is the pr draft+lookup 64.3 / 57.0 / 68.0 Photosynthesis is the pr plain 25.1 / 25.3 / 26.0 The history of the Roman draft k=4 45.4 / 50.2 / 46.2 The history of the Roman draft+lookup 51.0 / 50.1 / 50.9 The history of the Roman plain 24.7 / 26.9 / 25.9 def quicksort(a):\n " draft k=4 51.2 / 50.4 / 48.3 def quicksort(a):\n " draft+lookup 54.3 / 49.5 / 53.5 def quicksort(a):\n " plain 24.4 / 25.2 / 25.5 Once upon a time, in a s lookup series lookup 8 23.6 / 26.9 / 28.9 Once upon a time, in a s lookup series plain 24.6 / 24.8 / 25.1 Photosynthesis is the pr lookup series lookup 8 70.2 / 58.5 / 70.6 Photosynthesis is the pr lookup series plain 25.5 / 25.5 / 25.9 The history of the Roman lookup series lookup 8 27.4 / 26.8 / 30.7 The history of the Roman lookup series plain 24.0 / 25.1 / 25.4 def quicksort(a):\n " lookup series lookup 8 32.1 / 36.7 / 36.8 def quicksort(a):\n " lookup series plain 23.9 / 25.8 / 27.6 Long context (October 2026): SmolLM2-135M decoding at a depth of 2048 tokens (llama-bench -p 0 -n 32 -d 2048: tg32 @ d2048, f16 KV cache; KwkDecoder: fd_bench DEC_CTX=2048, the median of 64 steps after a 2048-token prompt, fp16 KV cache) per token llama.cpp Q4_0 13.6 ms (73.6 tokens/s) llama.cpp Q4_K_M 13.9 ms (72.1 tokens/s) KwkDecoder int4 6.1 ms (165 tokens/s) runs (tokens/s): llama.cpp Q4_0 71.3 / 73.6 / 77.1; Q4_K_M 61.3 / 72.8 / 72.1; KwkDecoder int4 175 / 157 / 157 / 165 / 170. (September 2026, older builds of both: llama.cpp Q4_0 20.1 ms, Q4_K_M 21.2 ms; KwkDecoder int4 9.8 ms with the fp16 cache, 12.1 ms with a float32 cache.) Quality: perplexity, llama.cpp's own protocol (the Perplexity column on kwker.io/benchmarks/#llm) Perplexity = how well the model predicts text it has not seen; lower is better. wikitext-2 test set, 20 chunks of 512 tokens, the second half of each chunk scored: llama-perplexity for llama.cpp, the same chunks and scoring for Kwker (kwker-py/tests/ppl_chunks.py CHUNKS=20). One run each: these numbers are deterministic. Every change is against the original model run in 32-bit floats (Hugging Face eager, ppl_chunks' float32 reference); llama.cpp's BF16 file reads within 0.05% of it on every model, so both runtimes score the same thing. Measured 2026-10-04 with the formats of the speed rows (KwkDecoder at commit 501fae6; llama.cpp 19e28a2; files converted from the same snapshots, bf16, then llama-quantize - except SmolLM2-1.7B's llama.cpp rows, measured in October on bf16-converted files with the same llama.cpp commit and protocol: this session's 1.7B files were re-quantized from Q8_0 for disk space, fine for speed only). SmolLM2-135M perplexity vs 32-bit original (18.4849) KwkDecoder bf16 (weights="bf16") 18.4844 0.00% llama.cpp BF16 18.4833 -0.01% KwkDecoder int8 18.5062 +0.12% llama.cpp Q8_0 18.5308 +0.25% llama.cpp Q4_K_M 19.6047 +6.1% KwkDecoder int4 20.0434 +8.4% llama.cpp Q4_0 23.1154 +25.1% Qwen2.5-0.5B perplexity vs 32-bit original (14.6309) KwkDecoder bf16 (weights="bf16") 14.6309 0.00% llama.cpp BF16 14.6350 +0.03% KwkDecoder int8 14.6490 +0.12% llama.cpp Q8_0 14.6549 +0.16% llama.cpp Q4_K_M 15.1686 +3.7% KwkDecoder int4 15.3780 +5.1% llama.cpp Q4_0 16.3250 +11.6% SmolLM2-1.7B perplexity vs 32-bit original (9.2135; llama.cpp's bf16 file 9.2143) KwkDecoder int8 9.2209 +0.08% llama.cpp Q8_0 9.2408 +0.30% KwkDecoder int4 (6-bit mix, the default here) 9.7695 +6.0% llama.cpp Q4_K_M 9.9895 +8.4% llama.cpp Q4_0 10.4429 +13.3% Granite 3.1 1B-A400M (MoE) perplexity vs 32-bit original (8.6395; llama.cpp's BF16 file 8.6435) KwkDecoder int8 8.6405 +0.01% llama.cpp Q8_0 8.6540 +0.17% KwkDecoder int4 (6-bit mix, the default here) 8.8459 +2.4% llama.cpp Q4_K_M 9.1064 +5.4% llama.cpp Q4_0 9.4618 +9.5% 8-bit: KwkDecoder is closer to the original than llama.cpp's Q8_0 on all four models. 4-bit: KwkDecoder int4 beats Q4_K_M on the two models of a billion parameters or more (its higher-precision tensors are 6-bit there) and Q4_0 on all four; on SmolLM2-135M and Qwen2.5-0.5B Q4_K_M is ahead (there it is mostly 5-bit, see the note below). Note on Q4_K_M: on SmolLM2-135M and Qwen2.5-0.5B it is mostly a 5-bit format. Q4_K needs rows that are a multiple of 256 values; their hidden sizes are 576 and 896, so llama-quantize falls back to Q5_0 for most projections. On SmolLM2-1.7B (hidden size 2048) Q4_K_M is a true 4-bit K-quant format. Serving: many requests at once (October 2026) KwkDecoder: kwker.serve.Engine (continuous batching, lookup 0), max_batch = B, B requests submitted together (kwker-py/tests/serve_bench.py BATCHED=1 PROMPTS=wiki: 64-token wikitext-2 windows, 64 new tokens each); llama.cpp: llama-batched-bench -npp 64 -ntg 64 -npl 1,4,8 on the Q4_0 file. S_TG = generated tokens per second summed over the B requests (the generation phase); S_PP = prompt tokens per second. KwkDecoder: the median of 5 runs in one process (REPS=5); llama.cpp: the median of 3 runs. Both measured on 2026-10-03 on the same machine, after the greedy-argmax and serving-loop changes (commit 7b7fdce). requests S_TG KwkDecoder int4 S_TG llama.cpp Q4_0 speed-up S_PP KwkDecoder / llama.cpp SmolLM2-135M 1 247.9 150.3 1.65x 2,455 / 965 4 751.4 339.3 2.21x 2,753 / 1,115 8 946.2 452.5 2.09x 2,773 / 1,202 SmolLM2-360M 1 107.9 70.9 1.52x 984 / 416 4 323.0 184.3 1.75x 1,109 / 481 8 535.7 228.3 2.35x 1,047 / 494 llama.cpp runs (S_TG at 1 / 4 / 8): SmolLM2-135M 149.5, 339.3, 387.8 | 150.3, 300.3, 464.5 | 153.3, 364.3, 452.5; SmolLM2-360M 64.1, 188.1, 235.0 | 74.5, 177.0, 228.3 | 70.9, 184.3, 199.4. (S_PP at 1 / 4 / 8: SmolLM2-135M 788, 1,115, 1,015 | 996, 1,136, 1,202 | 965, 1,097, 1,299; SmolLM2-360M 397, 481, 494 | 416, 480, 478 | 431, 501, 503.) The earlier table (2026-10-03, before those changes; llama.cpp runs of that session): SmolLM2-135M 262 / 631 / 938 vs 146 / 432 / 485, SmolLM2-360M 104 / 290 / 472 vs 75.8 / 186 / 231. Chat turns: the prompt cache (October 2026, commit 9cf586f) KwkDecoder int4 keeps the keys / values of the previous prompt (prompt_cache=True, the default) and reads only the new tokens when the next prompt extends it. kwker-py/tests/chat_bench.py: a 1,024-token system prompt, then each turn sends the whole conversation so far + a 64-token message (wikitext-2) and gets a 32-token greedy answer. Two decoders on the same weights - prompt_cache=True and prompt_cache=False - alternate per turn; both gave the same answers in every turn. Time to the first token = generate(prompt, 1). Each run: the median over turns 2-6 (SmolLM2-135M) or 2-4 (SmolLM2-1.7B); the table: the median of 3 runs, speed-up median against median. 4 threads. first token, cache first token, no cache speed-up SmolLM2-135M 59 ms 646 ms 10.9x SmolLM2-1.7B 471 ms 4,902 ms 10.4x runs (cache | no cache, ms): SmolLM2-135M 63 / 56 / 59 | 655 / 646 / 597; SmolLM2-1.7B 471 / 523 / 451 | 5,361 / 4,882 / 4,902. Per turn, run 1: SmolLM2-135M 62 / 78 / 63 / 60 / 72 vs 519 / 569 / 655 / 665 / 690; SmolLM2-1.7B 471 / 454 / 643 vs 4,989 / 5,361 / 5,575. Turn 1 (nothing to reuse) costs the same both ways: SmolLM2-135M ~470 ms, SmolLM2-1.7B ~4.3 s. kwker.serve.Engine keeps such a cache per request slot (a new request takes the free slot whose last prompt shares the longest prefix with it). AVX2 CPUs (October 2026): KwkDecoder's AVX2 kernels against llama.cpp built for AVX2 only - the processors without AVX-512 (Intel Core 12th generation and later, AMD Zen 2 / 3). Measured on the same Xeon with AVX-512 switched off on both sides: KwkDecoder with KWKER_GEMM_ISA=avx2; llama.cpp built with GGML_NATIVE=OFF, AVX2 + FMA + F16C + BMI2, no AVX-512 / AVX-VNNI (ss llama --isa avx2, the same 19e28a2 source). The memory system is this machine's, so a laptop's numbers differ; the ratios are the point. Same protocol as above (llama-bench -p 512 -n 64; fd_bench PREFILL=512); medians of 3 llama.cpp and 5 KwkDecoder runs, 2026-10-03. prompt 512 generation SmolLM2-135M llama.cpp Q4_0 731 tokens/s 171.9 tokens/s KwkDecoder int4 1,523 tokens/s (2.08x) 257.7 tokens/s (1.50x) llama.cpp Q8_0 642 tokens/s 121.1 tokens/s KwkDecoder int8 1,194 tokens/s (1.86x) 190.8 tokens/s (1.58x) SmolLM2-360M llama.cpp Q4_0 319 tokens/s 76.5 tokens/s KwkDecoder int4 594 tokens/s (1.86x) 108.0 tokens/s (1.41x) KwkDecoder int8 446 tokens/s 78.7 tokens/s runs: llama.cpp AVX2 Q4_0 135M pp512 731.4 / 442.7 / 770.3, tg64 171.9 / 172.4 / 150.9; Q8_0 642.4 / 572.0 / 654.4, 115.6 / 131.8 / 121.1; 360M Q4_0 325.0 / 317.4 / 319.4, 79.5 / 76.5 / 75.6. KwkDecoder AVX2 (prompt | generation) 135M int4 1,523 / 1,208 / 1,660 / 1,536 / 1,428 | 257.1 / 266.7 / 253.2 / 298.5 / 257.7; int8 1,103 / 1,194 / 1,164 / 1,215 / 1,222 | 178.3 / 201.2 / 190.8 / 191.2 / 189.4; 360M int4 602 / 513 / 600 / 565 / 594 | 107.8 / 109.3 / 108.5 / 101.4 / 108.0; int8 446 / 424 / 457 / 457 / 445 | 78.7 / 78.4 / 79.3 / 78.7 / 79.5. AVX2 re-measured 2026-10-04 (KwkDecoder at commit b75a856 - int4 prompt rows without residual activations on plain AVX2 too, and the 6-bit tensors' prompt kernel; KWKER_TORCH_ISA=avx2; the same protocol, rounds alternating): prompt 512 generation SmolLM2-135M llama.cpp Q4_0 766 tokens/s 148.6 tokens/s KwkDecoder int4 1,571 tokens/s (2.05x) 236.2 tokens/s (1.59x) llama.cpp Q8_0 609 tokens/s 108.9 tokens/s KwkDecoder int8 1,187 tokens/s (1.95x) 168.5 tokens/s (1.55x) SmolLM2-1.7B llama.cpp Q4_0 74 tokens/s 18.8 tokens/s llama.cpp Q4_K_M 93 tokens/s 18.3 tokens/s KwkDecoder int4 135 tokens/s (1.45x) 26.0 tokens/s (1.38x) runs (prompt | generation): 135M llama.cpp Q4_0 766 / 785 / 763 | 139.6 / 148.6 / 153.7; Q8_0 628 / 609 / 605 | 101.1 / 108.9 / 116.9; KwkDecoder int4 1,501 / 1,521 / 1,686 / 1,580 / 1,571 | 236.0 / 236.2 / 262.3 / 248.8 / 235.6; int8 1,170 / 1,134 / 1,211 / 1,198 / 1,187 | 168.5 / 166.5 / 170.7 / 174.9 / 165.7; 1.7B llama.cpp Q4_0 75 / 73 / 74 | 21.2 / 18.8 / 18.4; Q4_K_M 96 / 92 / 93 | 19.0 / 18.3 / 17.9; KwkDecoder int4 137 / 135 / 133 / 135 / 135 | 26.5 / 26.2 / 21.7 / 26.0 / 24.8. (1.7B speed-ups against llama.cpp's faster 4-bit file per column.) On this machine llama.cpp's AVX2 build is about as fast as its AVX-512 build (135M Q4_0 tg64 172 vs 161), and KwkDecoder's AVX2 generation matches its AVX-512 one (135M int4 ~258 vs ~272): decoding is bound by memory, not by the instruction set. The AVX2 kernels give the AVX-512 kernels' packs byte for byte and, with KWKER_A2_I8X=0, their logits bit for bit (tested); by default the int8 sums use a faster form that rounds differently in the last bits. Hugging Face generate() in bfloat16, unchanged code (October 2026): the model loaded in bf16 (Hugging Face's default for bf16 checkpoints), generate() in eager mode, torch's own kernels against kwker.torch_ops.install() (Kwker's bf16 linears - a GEMV at 1-8 rows, K-major panels packed on the fly in 12-row tiles past them - and bf16 decode attention; nothing quantized on either side, float32 sums both). tests/eager_b16_gen.py: a 128-token prompt (TTFT: generate with one new token), then 32 tokens (16 for 1.7B); both sides alternating in one process, median of 4 rounds (2 for 1.7B), 4 cores, this Xeon (AVX-512, no AMX). Greedy tokens identical on both sides for all three models. prompt pass, torch with install() decode, torch with install() SmolLM2-135M 344.0 ms 147.6 ms (2.33x) 35.56 ms/token 27.48 (1.29x) Qwen2.5-0.5B 983.3 ms 347.7 ms (2.83x) 78.56 ms/token 45.99 (1.71x) SmolLM2-1.7B 4,323.3 ms 1,190.7 ms (3.63x) 193.23 ms/token 129.94 (1.49x) torch.compile (October 2026, re-recorded 2026-10-04 after backend="kwker" began running decode steps without Inductor): Hugging Face generate() with a static KV cache, the model's forward compiled with torch.compile(backend="inductor") (stock Inductor) against torch.compile(backend="kwker"); nothing quantized on either side (float32 models from bf16 checkpoints: Kwker reads each weight's exact bf16 copy; bf16 models: the weights as stored; float32 sums both). backend="kwker" picks its eager inner for these decode graphs by itself: the graph runs as a list of native operator calls (kwker::GraphExec) with the KV cache written in place and fused RMSNorm / rotary / SiLU x up operators. kwker-py/tests/compile_variants.py: 32 new tokens after a 32-token prompt, ms per generated token (prefill included once), the best of 3 generates per process; 3 processes per row, interleaved with the other rows; the table: the median of the 3. 4 cores, this Xeon (AVX-512, no AMX). Greedy tokens identical to Inductor's in every run. First call (compile included, Inductor's caches warm): Inductor 14-17 s, backend="kwker" 16-19 s; with empty caches stock Inductor takes about 80 s for SmolLM2-135M, backend="kwker" the same ~18 s. dtype Inductor backend="kwker" SmolLM2-135M float32 26.02 ms/token 16.01 ms/token (1.63x) SmolLM2-135M bf16 26.09 ms/token 17.76 ms/token (1.47x) Qwen2.5-0.5B float32 66.60 ms/token 32.86 ms/token (2.03x) Qwen2.5-0.5B bf16 73.03 ms/token 37.59 ms/token (1.94x) runs (Inductor | kwker, ms/token): SmolLM2-135M float32: 28.90 / 26.02 / 25.87 | 15.38 / 17.97 / 16.01 SmolLM2-135M bf16: 26.09 / 29.65 / 25.70 | 17.76 / 17.80 / 16.77 Qwen2.5-0.5B float32: 66.60 / 66.06 / 70.95 | 32.86 / 32.45 / 37.50 Qwen2.5-0.5B bf16: 69.46 / 76.94 / 73.03 | 37.59 / 37.92 / 35.57 Speculative decoding (October 2026): SmolLM2-1.7B int4, KwkDecoder.generate, the output = greedy decoding's ------------------------------------------------------------------------------------------------------------- 4 cores; 4 open-ended prompts (kwker-py/tests/spec_bench.py), 128 new tokens each, tokens per second including the prompt pass; 3 runs per series, the median per prompt, then the geometric mean over the prompts. Draft model: SmolLM2-135M int4. Prompt lookup = proposals copied from earlier in the text (generate()'s default, lookup=8). plain prompt lookup draft model draft + lookup geometric mean 25.5 37.1 48.0 51.5 speed-up 1.00x 1.47x 1.88x 2.02x (prompt lookup was measured in its own series: plain 25.3 there; the other columns in one series) runs (tokens/s per prompt, run 1 / 2 / 3): Once upon a time, in a s draft k=4 36.3 / 41.2 / 37.4 Once upon a time, in a s draft+lookup 39.1 / 40.5 / 40.2 Once upon a time, in a s plain 25.5 / 26.9 / 25.6 Photosynthesis is the pr draft k=4 58.1 / 60.8 / 62.0 Photosynthesis is the pr draft+lookup 64.3 / 57.0 / 68.0 Photosynthesis is the pr plain 25.1 / 25.3 / 26.0 The history of the Roman draft k=4 45.4 / 50.2 / 46.2 The history of the Roman draft+lookup 51.0 / 50.1 / 50.9 The history of the Roman plain 24.7 / 26.9 / 25.9 def quicksort(a):\n " draft k=4 51.2 / 50.4 / 48.3 def quicksort(a):\n " draft+lookup 54.3 / 49.5 / 53.5 def quicksort(a):\n " plain 24.4 / 25.2 / 25.5 Once upon a time, in a s lookup series lookup 8 23.6 / 26.9 / 28.9 Once upon a time, in a s lookup series plain 24.6 / 24.8 / 25.1 Photosynthesis is the pr lookup series lookup 8 70.2 / 58.5 / 70.6 Photosynthesis is the pr lookup series plain 25.5 / 25.5 / 25.9 The history of the Roman lookup series lookup 8 27.4 / 26.8 / 30.7 The history of the Roman lookup series plain 24.0 / 25.1 / 25.4 def quicksort(a):\n " lookup series lookup 8 32.1 / 36.7 / 36.8 def quicksort(a):\n " lookup series plain 23.9 / 25.8 / 27.6 Long context (October 2026): SmolLM2-135M decoding at a depth of 2048 tokens (llama-bench -p 0 -n 32 -d 2048: tg32 @ d2048, f16 KV cache; KwkDecoder: fd_bench DEC_CTX=2048, the median of 64 steps after a 2048-token prompt, fp16 KV cache) per token llama.cpp Q4_0 13.6 ms (73.6 tokens/s) llama.cpp Q4_K_M 13.9 ms (72.1 tokens/s) KwkDecoder int4 6.1 ms (165 tokens/s) runs (tokens/s): llama.cpp Q4_0 71.3 / 73.6 / 77.1; Q4_K_M 61.3 / 72.8 / 72.1; KwkDecoder int4 175 / 157 / 157 / 165 / 170. (September 2026, older builds of both: llama.cpp Q4_0 20.1 ms, Q4_K_M 21.2 ms; KwkDecoder int4 9.8 ms with the fp16 cache, 12.1 ms with a float32 cache.) Quality: perplexity, llama.cpp's own protocol (the Perplexity column on kwker.io/benchmarks/#llm) Perplexity = how well the model predicts text it has not seen; lower is better. wikitext-2 test set, 20 chunks of 512 tokens, the second half of each chunk scored: llama-perplexity for llama.cpp, the same chunks and scoring for Kwker (kwker-py/tests/ppl_chunks.py CHUNKS=20). One run each: these numbers are deterministic. Every change is against the original model run in 32-bit floats (Hugging Face eager, ppl_chunks' float32 reference); llama.cpp's BF16 file reads within 0.05% of it on every model, so both runtimes score the same thing. Measured 2026-10-04 with the formats of the speed rows (KwkDecoder at commit 501fae6; llama.cpp 19e28a2; files converted from the same snapshots, bf16, then llama-quantize - except SmolLM2-1.7B's llama.cpp rows, measured in October on bf16-converted files with the same llama.cpp commit and protocol: this session's 1.7B files were re-quantized from Q8_0 for disk space, fine for speed only). SmolLM2-135M perplexity vs 32-bit original (18.4849) KwkDecoder bf16 (weights="bf16") 18.4844 0.00% llama.cpp BF16 18.4833 -0.01% KwkDecoder int8 18.5062 +0.12% llama.cpp Q8_0 18.5308 +0.25% llama.cpp Q4_K_M 19.6047 +6.1% KwkDecoder int4 20.0434 +8.4% llama.cpp Q4_0 23.1154 +25.1% Qwen2.5-0.5B perplexity vs 32-bit original (14.6309) KwkDecoder bf16 (weights="bf16") 14.6309 0.00% llama.cpp BF16 14.6350 +0.03% KwkDecoder int8 14.6490 +0.12% llama.cpp Q8_0 14.6549 +0.16% llama.cpp Q4_K_M 15.1686 +3.7% KwkDecoder int4 15.3780 +5.1% llama.cpp Q4_0 16.3250 +11.6% SmolLM2-1.7B perplexity vs 32-bit original (9.2135; llama.cpp's bf16 file 9.2143) KwkDecoder int8 9.2209 +0.08% llama.cpp Q8_0 9.2408 +0.30% KwkDecoder int4 (6-bit mix, the default here) 9.7695 +6.0% llama.cpp Q4_K_M 9.9895 +8.4% llama.cpp Q4_0 10.4429 +13.3% Granite 3.1 1B-A400M (MoE) perplexity vs 32-bit original (8.6395; llama.cpp's BF16 file 8.6435) KwkDecoder int8 8.6405 +0.01% llama.cpp Q8_0 8.6540 +0.17% KwkDecoder int4 (6-bit mix, the default here) 8.8459 +2.4% llama.cpp Q4_K_M 9.1064 +5.4% llama.cpp Q4_0 9.4618 +9.5% 8-bit: KwkDecoder is closer to the original than llama.cpp's Q8_0 on all four models. 4-bit: KwkDecoder int4 beats Q4_K_M on the two models of a billion parameters or more (its higher-precision tensors are 6-bit there) and Q4_0 on all four; on SmolLM2-135M and Qwen2.5-0.5B Q4_K_M is ahead (there it is mostly 5-bit, see the note below). Note on Q4_K_M: on SmolLM2-135M and Qwen2.5-0.5B it is mostly a 5-bit format. Q4_K needs rows that are a multiple of 256 values; their hidden sizes are 576 and 896, so llama-quantize falls back to Q5_0 for most projections. On SmolLM2-1.7B (hidden size 2048) Q4_K_M is a true 4-bit K-quant format. Serving: many requests at once (October 2026) KwkDecoder: kwker.serve.Engine (continuous batching, lookup 0), max_batch = B, B requests submitted together (kwker-py/tests/serve_bench.py BATCHED=1 PROMPTS=wiki: 64-token wikitext-2 windows, 64 new tokens each); llama.cpp: llama-batched-bench -npp 64 -ntg 64 -npl 1,4,8 on the Q4_0 file. S_TG = generated tokens per second summed over the B requests (the generation phase); S_PP = prompt tokens per second. KwkDecoder: the median of 5 runs in one process (REPS=5); llama.cpp: the median of 3 runs. Both measured on 2026-10-03 on the same machine, after the greedy-argmax and serving-loop changes (commit 7b7fdce). requests S_TG KwkDecoder int4 S_TG llama.cpp Q4_0 speed-up S_PP KwkDecoder / llama.cpp SmolLM2-135M 1 247.9 150.3 1.65x 2,455 / 965 4 751.4 339.3 2.21x 2,753 / 1,115 8 946.2 452.5 2.09x 2,773 / 1,202 SmolLM2-360M 1 107.9 70.9 1.52x 984 / 416 4 323.0 184.3 1.75x 1,109 / 481 8 535.7 228.3 2.35x 1,047 / 494 llama.cpp runs (S_TG at 1 / 4 / 8): SmolLM2-135M 149.5, 339.3, 387.8 | 150.3, 300.3, 464.5 | 153.3, 364.3, 452.5; SmolLM2-360M 64.1, 188.1, 235.0 | 74.5, 177.0, 228.3 | 70.9, 184.3, 199.4. (S_PP at 1 / 4 / 8: SmolLM2-135M 788, 1,115, 1,015 | 996, 1,136, 1,202 | 965, 1,097, 1,299; SmolLM2-360M 397, 481, 494 | 416, 480, 478 | 431, 501, 503.) The earlier table (2026-10-03, before those changes; llama.cpp runs of that session): SmolLM2-135M 262 / 631 / 938 vs 146 / 432 / 485, SmolLM2-360M 104 / 290 / 472 vs 75.8 / 186 / 231. Chat turns: the prompt cache (October 2026, commit 9cf586f) KwkDecoder int4 keeps the keys / values of the previous prompt (prompt_cache=True, the default) and reads only the new tokens when the next prompt extends it. kwker-py/tests/chat_bench.py: a 1,024-token system prompt, then each turn sends the whole conversation so far + a 64-token message (wikitext-2) and gets a 32-token greedy answer. Two decoders on the same weights - prompt_cache=True and prompt_cache=False - alternate per turn; both gave the same answers in every turn. Time to the first token = generate(prompt, 1). Each run: the median over turns 2-6 (SmolLM2-135M) or 2-4 (SmolLM2-1.7B); the table: the median of 3 runs, speed-up median against median. 4 threads. first token, cache first token, no cache speed-up SmolLM2-135M 59 ms 646 ms 10.9x SmolLM2-1.7B 471 ms 4,902 ms 10.4x runs (cache | no cache, ms): SmolLM2-135M 63 / 56 / 59 | 655 / 646 / 597; SmolLM2-1.7B 471 / 523 / 451 | 5,361 / 4,882 / 4,902. Per turn, run 1: SmolLM2-135M 62 / 78 / 63 / 60 / 72 vs 519 / 569 / 655 / 665 / 690; SmolLM2-1.7B 471 / 454 / 643 vs 4,989 / 5,361 / 5,575. Turn 1 (nothing to reuse) costs the same both ways: SmolLM2-135M ~470 ms, SmolLM2-1.7B ~4.3 s. kwker.serve.Engine keeps such a cache per request slot (a new request takes the free slot whose last prompt shares the longest prefix with it). AVX2 CPUs (October 2026): KwkDecoder's AVX2 kernels against llama.cpp built for AVX2 only - the processors without AVX-512 (Intel Core 12th generation and later, AMD Zen 2 / 3). Measured on the same Xeon with AVX-512 switched off on both sides: KwkDecoder with KWKER_GEMM_ISA=avx2; llama.cpp built with GGML_NATIVE=OFF, AVX2 + FMA + F16C + BMI2, no AVX-512 / AVX-VNNI (ss llama --isa avx2, the same 19e28a2 source). The memory system is this machine's, so a laptop's numbers differ; the ratios are the point. Same protocol as above (llama-bench -p 512 -n 64; fd_bench PREFILL=512); medians of 3 llama.cpp and 5 KwkDecoder runs, 2026-10-03. prompt 512 generation SmolLM2-135M llama.cpp Q4_0 731 tokens/s 171.9 tokens/s KwkDecoder int4 1,523 tokens/s (2.08x) 257.7 tokens/s (1.50x) llama.cpp Q8_0 642 tokens/s 121.1 tokens/s KwkDecoder int8 1,194 tokens/s (1.86x) 190.8 tokens/s (1.58x) SmolLM2-360M llama.cpp Q4_0 319 tokens/s 76.5 tokens/s KwkDecoder int4 594 tokens/s (1.86x) 108.0 tokens/s (1.41x) KwkDecoder int8 446 tokens/s 78.7 tokens/s runs: llama.cpp AVX2 Q4_0 135M pp512 731.4 / 442.7 / 770.3, tg64 171.9 / 172.4 / 150.9; Q8_0 642.4 / 572.0 / 654.4, 115.6 / 131.8 / 121.1; 360M Q4_0 325.0 / 317.4 / 319.4, 79.5 / 76.5 / 75.6. KwkDecoder AVX2 (prompt | generation) 135M int4 1,523 / 1,208 / 1,660 / 1,536 / 1,428 | 257.1 / 266.7 / 253.2 / 298.5 / 257.7; int8 1,103 / 1,194 / 1,164 / 1,215 / 1,222 | 178.3 / 201.2 / 190.8 / 191.2 / 189.4; 360M int4 602 / 513 / 600 / 565 / 594 | 107.8 / 109.3 / 108.5 / 101.4 / 108.0; int8 446 / 424 / 457 / 457 / 445 | 78.7 / 78.4 / 79.3 / 78.7 / 79.5. AVX2 re-measured 2026-10-04 (KwkDecoder at commit b75a856 - int4 prompt rows without residual activations on plain AVX2 too, and the 6-bit tensors' prompt kernel; KWKER_TORCH_ISA=avx2; the same protocol, rounds alternating): prompt 512 generation SmolLM2-135M llama.cpp Q4_0 766 tokens/s 148.6 tokens/s KwkDecoder int4 1,571 tokens/s (2.05x) 236.2 tokens/s (1.59x) llama.cpp Q8_0 609 tokens/s 108.9 tokens/s KwkDecoder int8 1,187 tokens/s (1.95x) 168.5 tokens/s (1.55x) SmolLM2-1.7B llama.cpp Q4_0 74 tokens/s 18.8 tokens/s llama.cpp Q4_K_M 93 tokens/s 18.3 tokens/s KwkDecoder int4 135 tokens/s (1.45x) 26.0 tokens/s (1.38x) runs (prompt | generation): 135M llama.cpp Q4_0 766 / 785 / 763 | 139.6 / 148.6 / 153.7; Q8_0 628 / 609 / 605 | 101.1 / 108.9 / 116.9; KwkDecoder int4 1,501 / 1,521 / 1,686 / 1,580 / 1,571 | 236.0 / 236.2 / 262.3 / 248.8 / 235.6; int8 1,170 / 1,134 / 1,211 / 1,198 / 1,187 | 168.5 / 166.5 / 170.7 / 174.9 / 165.7; 1.7B llama.cpp Q4_0 75 / 73 / 74 | 21.2 / 18.8 / 18.4; Q4_K_M 96 / 92 / 93 | 19.0 / 18.3 / 17.9; KwkDecoder int4 137 / 135 / 133 / 135 / 135 | 26.5 / 26.2 / 21.7 / 26.0 / 24.8. (1.7B speed-ups against llama.cpp's faster 4-bit file per column.) On this machine llama.cpp's AVX2 build is about as fast as its AVX-512 build (135M Q4_0 tg64 172 vs 161), and KwkDecoder's AVX2 generation matches its AVX-512 one (135M int4 ~258 vs ~272): decoding is bound by memory, not by the instruction set. The AVX2 kernels give the AVX-512 kernels' packs byte for byte and, with KWKER_A2_I8X=0, their logits bit for bit (tested); by default the int8 sums use a faster form that rounds differently in the last bits. Hugging Face generate() in bfloat16, unchanged code (October 2026): the model loaded in bf16 (Hugging Face's default for bf16 checkpoints), generate() in eager mode, torch's own kernels against kwker.torch_ops.install() (Kwker's bf16 linears - a GEMV at 1-8 rows, K-major panels packed on the fly in 12-row tiles past them - and bf16 decode attention; nothing quantized on either side, float32 sums both). tests/eager_b16_gen.py: a 128-token prompt (TTFT: generate with one new token), then 32 tokens (16 for 1.7B); both sides alternating in one process, median of 4 rounds (2 for 1.7B), 4 cores, this Xeon (AVX-512, no AMX). Greedy tokens identical on both sides for all three models. prompt pass, torch with install() decode, torch with install() SmolLM2-135M 344.0 ms 147.6 ms (2.33x) 35.56 ms/token 27.48 (1.29x) Qwen2.5-0.5B 983.3 ms 347.7 ms (2.83x) 78.56 ms/token 45.99 (1.71x) SmolLM2-1.7B 4,323.3 ms 1,190.7 ms (3.63x) 193.23 ms/token 129.94 (1.49x) torch.compile (October 2026, re-recorded 2026-10-04): Hugging Face generate() with a static KV cache, the model's forward compiled with torch.compile(backend="inductor") (stock Inductor) against torch.compile(backend="kwker"); nothing quantized on either side (float32 models from bf16 checkpoints: Kwker reads each weight's exact bf16 copy; bf16 models: the weights as stored; float32 sums both; the fused glue operators stay off under Inductor's Python wrapper - its own fused kernels are faster there). kwker-py/tests/compile_variants.py: 32 new tokens after a 32-token prompt, ms per generated token (prefill included once), the best of 3 generates per process; 3 processes per row, interleaved with the other rows; the table: the median of the 3. 4 cores, this Xeon (AVX-512, no AMX). Greedy tokens identical to Inductor's in every run. dtype Inductor backend="kwker" SmolLM2-135M float32 27.13 ms/token 16.60 ms/token (1.63x) SmolLM2-135M bf16 26.94 ms/token 17.81 ms/token (1.51x) Qwen2.5-0.5B float32 69.30 ms/token 36.60 ms/token (1.89x) Qwen2.5-0.5B bf16 70.82 ms/token 38.61 ms/token (1.83x) runs (Inductor | kwker, ms/token): SmolLM2-135M float32: 27.94 / 25.03 / 27.13 | 16.60 / 15.86 / 17.19 SmolLM2-135M bf16: 27.51 / 26.50 / 26.94 | 17.81 / 17.60 / 18.26 Qwen2.5-0.5B float32: 69.30 / 73.01 / 66.19 | 37.75 / 35.64 / 36.60 Qwen2.5-0.5B bf16: 70.82 / 68.93 / 71.44 | 38.34 / 38.61 / 38.82 Speculative decoding (October 2026): SmolLM2-1.7B at 8 bits, SmolLM2-135M as the draft model, greedy, 64 tokens after each of the three spec_bench prompts, 4 cores, both runtimes alternating in one session. llama.cpp (19e28a2): Q8_0 target, llama-speculative-simple --spec-type draft-simple --spec-draft-n-max 3 (the fastest of 3 / 5 / 8 / 16 in a sweep), the faster of a Q4_0 and a Q8_0 draft per run, decode rate as it prints it; plain = llama-bench tg64. KwkDecoder: int8 target, int4 draft, generate(draft=, draft_k=3), rates including the prompt pass. Per prompt the median of the runs, then the geometric mean. plain speculative llama.cpp Q8_0 12.6 t/s 18.6 t/s (1.48x) KwkDecoder int8 19.8 t/s 42.8 t/s (2.16x) runs: llama.cpp plain 12.50 / 12.81 / 12.56; speculative per prompt 17.7 / 17.7 / 18.6; 21.0 / 21.6 / 22.7; 16.4 / 16.8 / 17.5 KwkDecoder plain per prompt 19.3 / 19.6 / 19.2 / 22.4 / 19.1; 19.5 / 19.5 / 20.8 / 19.8 / 20.0; 20.9 / 20.3 / 20.0 / 20.3 / 20.5 speculative per prompt 42.3 / 41.0 / 40.8 / 47.7 / 41.2; 42.7 / 44.7 / 41.9 / 40.6 / 40.9; 44.9 / 45.3 / 48.1 / 45.4 / 47.2