Serve language models on the CPU servers you have
Chat, extraction, summarization and classification with small and mid-size models, without GPUs. KwkDecoder runs a whole Llama-family model as one native call, with 8-bit or 4-bit weights.
Against llama.cpp on the same cores
Tokens per second on 4 CPU cores: a 512-token prompt, then generation one token at a time. Both sides use 4-bit weights. Perplexity shows what that costs in accuracy: how well the model predicts real text it has not seen, lower is better, with the change from the original full-precision model beside it.
| Model | Prompt, KwkDecoder int4 | Prompt, llama.cpp 4-bit | Generation, KwkDecoder int4 | Generation, llama.cpp 4-bit | Perplexity, KwkDecoder int4 | Perplexity, llama.cpp 4-bit |
|---|---|---|---|---|---|---|
| SmolLM2-135M | 3,074 | 1,039Q4_0 | 285 | 158Q4_0 | 20.04+8.4% | 23.12+25.1% |
| Qwen2.5-0.5B | 1,106 | 421Q4_0 | 87.1 | 50.2Q4_0 | 15.38+5.1% | 16.32+11.6% |
| SmolLM2-1.7B | 276 | 115Q4_K_M | 31.4 | 18.3Q4_K_M | 9.77+6.0% | 9.99+8.4% |
| Granite 3.1 1B-A400M MoE | 827 | 247Q4_K_M | 92.9 | 64.3Q4_K_M | 8.85+2.4% | 9.11+5.4% |
Each cell is the median of repeated runs. llama.cpp's column shows its faster 4-bit file for each model (Q4_0 or Q4_K_M). Intel Xeon with AVX-512 VNNI; SmolLM2-135M, Qwen2.5 and SmolLM2-1.7B measured in one session with both runtimes interleaved, October 2026 (Granite: an earlier run). With 8-bit weights, KwkDecoder generates 1.8× faster than llama.cpp's Q8_0 (geometric mean of the three models). All formats · Run it yourself
Nothing quantized
By default (precision="preserve") KwkDecoder reads a bfloat16 checkpoint's own weights and keeps every
activation and sum in float32: the model's own numbers. Against llama.cpp's BF16 file of the same checkpoint,
tokens per second on 4 CPU cores, and perplexity (lower is better) against the original model's.
| Model | Prompt, KwkDecoder bf16 | Prompt, llama.cpp BF16 | Generation, KwkDecoder bf16 | Generation, llama.cpp BF16 | Perplexity, KwkDecoder bf16 | Perplexity, llama.cpp BF16 |
|---|---|---|---|---|---|---|
| SmolLM2-135M | 1,911 | 826 | 120 | 74.1 | 18.480.00% | 18.48-0.01% |
| Qwen2.5-0.5B | 625 | 312 | 36.5 | 23.1 | 14.630.00% | 14.63+0.03% |
Each cell is the median of repeated runs (KwkDecoder 5, llama.cpp 3), both runtimes interleaved in one
session, October 2026. Logits match the model's float32 forward to about one part in a million, and greedy text is
Hugging Face's. kwker.decode.install(model) runs Hugging Face generate this way:
SmolLM2-135M from 26 tokens per second (float32 eager) to 110, the same tokens.
All formats · The guide
A 1.7-billion-parameter model on four CPU cores
SmolLM2-1.7B with 4-bit weights on 4 cores of a cloud Xeon, no GPU. It generates text several times faster than people read it, and reads a page-long prompt in a few seconds.
tokens per second while generating, KwkDecoder int4 (llama.cpp Q4_K_M: 18.3)
to read a 512-token prompt, KwkDecoder int4 (llama.cpp Q4_K_M: 4.4 s)
perplexity with 8-bit weights, KwkDecoder int8 (llama.cpp Q8_0: +0.30%)
perplexity with 4-bit weights, KwkDecoder int4 (llama.cpp Q4_K_M: +8.4%)
Medians of repeated runs, October 2026. Perplexity (how well the model predicts unseen Wikipedia text; lower is better): wikitext-2, both runtimes scored the same way, each change against the original model run in 32-bit floats. From a billion parameters int4 keeps its higher-precision tensors at 6 bits: 5.0 bits per weight on this model, about Q4_K_M's bytes. Raw results (text) · KwkDecoder guide
Several tokens per pass, the same text
Generation reads every weight once per step, so checking several proposed tokens costs little more than checking one. KwkDecoder proposes tokens copied from earlier in the text (on by default) or from a small model of the same family, checks them all in one pass, and keeps the ones the model agrees with. The output is exactly the model's own greedy text.
| SmolLM2-1.7B, KwkDecoder int4, 4 cores | Tokens per second | Speed-up |
|---|---|---|
| Proposals from SmolLM2-135M, plus copied ones | 51.5 | 2.02× |
| Proposals copied from the text (default) | 37.1 | 1.47× |
| One token per pass | 25.5 | 1.00× |
Geometric means over four prompts, 128 new tokens each, prompt reading included; medians of 3 runs, October 2026. Copying pays most on text that quotes its input (summaries, edits, code): SmolLM2-360M ran 2.4× faster on such prompts. Raw results (text) · Speculative decoding in the guide
Against llama.cpp's speculative decoding
llama.cpp checks a small model's proposals too. Both runtimes ran SmolLM2-1.7B at 8 bits with SmolLM2-135M as the draft model, greedy, so each text is its own model's greedy output.
| SmolLM2-1.7B, 4 cores | Plain decoding | With the 135M draft | Speed-up |
|---|---|---|---|
| KwkDecoder int8 | 19.8 | 42.8 | 2.2× |
| llama.cpp Q8_0 | 12.6 | 18.6 | 1.5× |
Tokens per second, greedy, 64 tokens after each of three short prompts: per prompt the median of
repeated runs (KwkDecoder 5, llama.cpp 3), then the geometric mean over the prompts. llama.cpp:
llama-speculative-simple with up to 3 drafted tokens, its fastest setting here, and the faster of a Q4_0
and a Q8_0 draft in each run; its plain rate is llama-bench's. KwkDecoder: an int4 draft, 3 drafted
tokens, its rates including the prompt. October 2026. · Draft models in the guide
Many requests at once
Generating a token reads every weight once, whether the step serves one request or eight.
kwker.serve puts the requests in flight into one step and admits new ones as slots free up, and each
request gets the tokens it would get alone. python -m kwker.serve --model <id> serves it as an
OpenAI-compatible API.
| 4 cores | Requests at once | KwkDecoder int4 | llama.cpp Q4_0 | Speed-up |
|---|---|---|---|---|
| SmolLM2-135M | 1 | 248 | 150 | 1.65× |
| 8 | 946 | 452 | 2.09× | |
| SmolLM2-360M | 1 | 108 | 70.9 | 1.52× |
| 8 | 536 | 228 | 2.35× |
Generated tokens per second, summed over the requests: 64-token prompts, then 64 new tokens each, as llama.cpp's llama-batched-bench measures it (S_TG; llama.cpp batches the same requests). Medians of 5 runs (KwkDecoder) and 3 runs (llama.cpp), both measured the same day, October 2026. Serving guide · Raw results (text) · Serving guide
The next chat turn starts at once
Each turn of a chat sends the whole conversation again. KwkDecoder keeps what it computed for the
previous prompt and reads only what is new: the last answer and the new message. The tokens are exactly those of a
fresh start. kwker.serve does the same for every request slot, so a shared system prompt is read once.
| KwkDecoder int4, 4 cores | First token, cache | First token, no cache | Speed-up |
|---|---|---|---|
| SmolLM2-135M | 59 ms | 646 ms | 10.9× |
| SmolLM2-1.7B | 471 ms | 4,902 ms | 10.4× |
Time to the first token of chat turns 2 and later: a 1,024-token system prompt, 64-token messages and 32-token answers (wikitext). Medians over turns 2-6 (SmolLM2-135M) and 2-4 (SmolLM2-1.7B), then the median of 3 runs, October 2026; both setups gave the same answers. Guide · Raw results (text) · Prompt cache in the guide
Laptops and desktops too
Most laptops and desktops have no AVX-512: Intel Core 12th generation and later, AMD Zen 2 and 3. There KwkDecoder runs its AVX2 kernels, with the same packed weights and results that match the AVX-512 kernels to within float rounding.
| AVX2 only, 4 cores | Prompt reading | Speed-up | Generation | Speed-up |
|---|---|---|---|---|
| SmolLM2-135M, KwkDecoder int4 | 1,571 | 2.05× | 236.2 | 1.59× |
| SmolLM2-135M, llama.cpp Q4_0 | 766 | 148.6 | ||
| SmolLM2-135M, KwkDecoder int8 | 1,187 | 1.95× | 168.5 | 1.55× |
| SmolLM2-135M, llama.cpp Q8_0 | 609 | 108.9 | ||
| SmolLM2-360M, KwkDecoder int4 | 594 | 1.86× | 108.0 | 1.41× |
| SmolLM2-360M, llama.cpp Q4_0 | 319 | 76.5 | ||
| SmolLM2-1.7B, KwkDecoder int4 | 135 | 1.45× | 26.0 | 1.38× |
| SmolLM2-1.7B, llama.cpp 4-bit | 93Q4_K_M | 18.8Q4_0 |
Tokens per second: a 512-token prompt, then generation. Both sides restricted to AVX2 on the same machine (KwkDecoder's AVX2 kernels; llama.cpp built for AVX2 only), so the memory system is the test machine's and a laptop's own numbers differ. Medians of 5 runs (KwkDecoder) and 3 runs (llama.cpp), October 2026: SmolLM2-135M and SmolLM2-1.7B on 2026-10-04 (after int4 prompt rows stopped carrying residual activations on plain AVX2 too), SmolLM2-360M the day before; for SmolLM2-1.7B llama.cpp's faster 4-bit file per column. Raw results (text) · Supported processors
Your Hugging Face code, unchanged
Already running a bfloat16 model with Hugging Face on a CPU without AMX? One line,
kwker.torch_ops.install(), puts its linears and each new token's attention on Kwker's bf16 kernels. The model and the
generate() call stay as they are, and nothing is quantized on either side.
| Model, bf16, 4 cores | Prompt, torch bf16 | Prompt, Kwker bf16 | Speed-up | Generation, torch bf16 | Generation, Kwker bf16 | Speed-up |
|---|---|---|---|---|---|---|
| SmolLM2-135M | 372 | 867 | 2.3× | 28.1 | 36.4 | 1.3× |
| Qwen2.5-0.5B | 130 | 368 | 2.8× | 12.7 | 21.7 | 1.7× |
| SmolLM2-1.7B | 29.6 | 107 | 3.6× | 5.2 | 7.7 | 1.5× |
Tokens per second: a 128-token prompt, then greedy generation with Hugging Face generate() in
eager mode, both sides alternating in one process (median of the rounds), October 2026. The greedy tokens were the same
on both sides. Kwker adds in another order than torch, so outputs can differ by bf16 rounding. · The CPU backend in the guide
torch.compile, up to 2× Inductor
Compiling your model? Pass backend="kwker" instead of the default. It reads float32 weights of bf16
checkpoints as their exact bf16 values, runs linears that share an input as one call and runs each decode step as a list
of native operator calls, with the KV cache written in place. Nothing is quantized, and the tokens match Inductor's.
| Model, 4 cores | Weights | Inductor, tokens/s | backend="kwker", tokens/s | Speed-up |
|---|---|---|---|---|
| SmolLM2-135M | float32 | 38.4 | 62.5 | 1.6× |
| SmolLM2-135M | bf16 | 38.3 | 56.3 | 1.5× |
| Qwen2.5-0.5B | float32 | 15.0 | 30.4 | 2.0× |
| Qwen2.5-0.5B | bf16 | 13.7 | 26.6 | 1.9× |
Hugging Face generate() with a static KV cache: 32 new tokens after a 32-token prompt (prefill
included), the median of 3 runs per row, October 2026. The first compile of SmolLM2-135M takes about 18 s, with or
without caches (Inductor: about 80 s with empty caches, about 15 s from its cache). · torch.compile in the guide
How it works
- One call per step: the whole model runs in native code, with no Python and no framework overhead between layers.
- Compact weights: 8-bit weights with one scale per 64 values, or 4-bit weights fitted to your own
text (
calib=, optionally with GPTQ). - Built for decoding: a preallocated KV cache, fused attention for the new token, and batching of several sequences into one step.
- Hugging Face's own API: after
install(model),model.generate()keeps sampling, streamers, stopping criteria and batches, and every step runs on Kwker. - Any Hugging Face model:
KwkDecodercovers Llama, Mistral, Qwen2, Qwen3, SmolLM, Phi-3 and Phi-4-mini, Granite, Gemma 2, Gemma 3, the mixture-of-experts Mixtral, Qwen3-MoE and Granite MoE, and GPT-2, Pythia, Phi-2, OPT, StableLM, StarCoder2 and OLMo;Decodercompiles any causal language model.
import kwker
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from kwker.decode import install
name = "HuggingFaceTB/SmolLM2-135M"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, dtype=torch.float32).eval()
install(model, precision="int4") # 4-bit weights; "balanced" for 8-bit, the default keeps the model's own
ids = tok("Sorting algorithms are", return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=24, do_sample=False) # Hugging Face's own API
print(tok.decode(out[0]))Sorting algorithms are used to sort a list of items. The following example shows how to sort a list of numbers using the sort
Quality and requirements
- 8-bit weights stay within a quarter of a percent of the original model's perplexity, closer than llama.cpp's Q8_0 on every model measured (SmolLM2-1.7B: +0.08% against +0.30%; Qwen2.5-0.5B: +0.12% against +0.16%; SmolLM2-135M: +0.12% against +0.25%).
- 4-bit weights trade quality for speed. On SmolLM2-1.7B, KwkDecoder int4 costs +6.0% perplexity against +8.4% for llama.cpp's Q4_K_M, at about the same bytes per token. On the two smallest models Q4_K_M is ahead (SmolLM2-135M +8.4% against +6.1%, Qwen2.5-0.5B +5.1% against +3.7%: there it spends about 5.5 bits per weight), and llama.cpp's 4-bit Q4_0 is well behind (+25.1% and +11.6%). Calibrated 4-bit beats llama.cpp's best 4-bit file on Qwen2.5-0.5B (+2.33% against +2.70%). KwkDecoder's int4 keeps about a quarter of the weights at higher precision: 6 bits from a billion parameters, 8 bits below. On SmolLM2-1.7B that is 5.0 bits per weight on average (4.625 for the 4-bit weights with their scales), Q4_K_M 4.9 and Q4_0 4.6. Measure on your own text before you deploy a format.
- Processor: AVX-512 with VNNI (Intel Xeon Cascade Lake and later, AMD Zen 4 and later) or AVX2 (Intel Core 12th generation and later, AMD Zen 2 and 3); AMX speeds up prompt reading where present. Linux x86-64 with PyTorch.
Perplexity on wikitext-2 with llama.cpp's own protocol (20 chunks of 512 tokens). Details and every format: limitations, guide.
Free to start
Free for companies under 100 people and about $1.3M revenue (PolyForm Small Business). Every product line is in every tier.