Kwker
Language models

Serve language models on the CPU servers you have

Chat, extraction, summarization and classification with small and mid-size models, without GPUs. KwkDecoder runs a whole Llama-family model as one native call, with 8-bit or 4-bit weights.

Against llama.cpp on the same cores

Tokens per second on 4 CPU cores: a 512-token prompt, then generation one token at a time. Both sides use 4-bit weights. Perplexity shows what that costs in accuracy: how well the model predicts real text it has not seen, lower is better, with the change from the original full-precision model beside it.

ModelPrompt, KwkDecoder int4Prompt, llama.cpp 4-bitGeneration, KwkDecoder int4Generation, llama.cpp 4-bitPerplexity, KwkDecoder int4Perplexity, llama.cpp 4-bit
SmolLM2-135M3,0741,039Q4_0285158Q4_020.04+8.4%23.12+25.1%
Qwen2.5-0.5B1,106421Q4_087.150.2Q4_015.38+5.1%16.32+11.6%
SmolLM2-1.7B276115Q4_K_M31.418.3Q4_K_M9.77+6.0%9.99+8.4%
Granite 3.1 1B-A400M MoE827247Q4_K_M92.964.3Q4_K_M8.85+2.4%9.11+5.4%

Each cell is the median of repeated runs. llama.cpp's column shows its faster 4-bit file for each model (Q4_0 or Q4_K_M). Intel Xeon with AVX-512 VNNI; SmolLM2-135M, Qwen2.5 and SmolLM2-1.7B measured in one session with both runtimes interleaved, October 2026 (Granite: an earlier run). With 8-bit weights, KwkDecoder generates 1.8× faster than llama.cpp's Q8_0 (geometric mean of the three models). All formats · Run it yourself

Nothing quantized

By default (precision="preserve") KwkDecoder reads a bfloat16 checkpoint's own weights and keeps every activation and sum in float32: the model's own numbers. Against llama.cpp's BF16 file of the same checkpoint, tokens per second on 4 CPU cores, and perplexity (lower is better) against the original model's.

ModelPrompt, KwkDecoder bf16Prompt, llama.cpp BF16Generation, KwkDecoder bf16Generation, llama.cpp BF16Perplexity, KwkDecoder bf16Perplexity, llama.cpp BF16
SmolLM2-135M1,91182612074.118.480.00%18.48-0.01%
Qwen2.5-0.5B62531236.523.114.630.00%14.63+0.03%

Each cell is the median of repeated runs (KwkDecoder 5, llama.cpp 3), both runtimes interleaved in one session, October 2026. Logits match the model's float32 forward to about one part in a million, and greedy text is Hugging Face's. kwker.decode.install(model) runs Hugging Face generate this way: SmolLM2-135M from 26 tokens per second (float32 eager) to 110, the same tokens. All formats · The guide

A 1.7-billion-parameter model on four CPU cores

SmolLM2-1.7B with 4-bit weights on 4 cores of a cloud Xeon, no GPU. It generates text several times faster than people read it, and reads a page-long prompt in a few seconds.

31.4

tokens per second while generating, KwkDecoder int4 (llama.cpp Q4_K_M: 18.3)

1.9 s

to read a 512-token prompt, KwkDecoder int4 (llama.cpp Q4_K_M: 4.4 s)

+0.08%

perplexity with 8-bit weights, KwkDecoder int8 (llama.cpp Q8_0: +0.30%)

+6.0%

perplexity with 4-bit weights, KwkDecoder int4 (llama.cpp Q4_K_M: +8.4%)

Medians of repeated runs, October 2026. Perplexity (how well the model predicts unseen Wikipedia text; lower is better): wikitext-2, both runtimes scored the same way, each change against the original model run in 32-bit floats. From a billion parameters int4 keeps its higher-precision tensors at 6 bits: 5.0 bits per weight on this model, about Q4_K_M's bytes. Raw results (text) · KwkDecoder guide

Several tokens per pass, the same text

Generation reads every weight once per step, so checking several proposed tokens costs little more than checking one. KwkDecoder proposes tokens copied from earlier in the text (on by default) or from a small model of the same family, checks them all in one pass, and keeps the ones the model agrees with. The output is exactly the model's own greedy text.

SmolLM2-1.7B, KwkDecoder int4, 4 coresTokens per secondSpeed-up
Proposals from SmolLM2-135M, plus copied ones51.52.02×
Proposals copied from the text (default)37.11.47×
One token per pass25.51.00×

Geometric means over four prompts, 128 new tokens each, prompt reading included; medians of 3 runs, October 2026. Copying pays most on text that quotes its input (summaries, edits, code): SmolLM2-360M ran 2.4× faster on such prompts. Raw results (text) · Speculative decoding in the guide

Against llama.cpp's speculative decoding

llama.cpp checks a small model's proposals too. Both runtimes ran SmolLM2-1.7B at 8 bits with SmolLM2-135M as the draft model, greedy, so each text is its own model's greedy output.

SmolLM2-1.7B, 4 coresPlain decodingWith the 135M draftSpeed-up
KwkDecoder int819.842.82.2×
llama.cpp Q8_012.618.61.5×

Tokens per second, greedy, 64 tokens after each of three short prompts: per prompt the median of repeated runs (KwkDecoder 5, llama.cpp 3), then the geometric mean over the prompts. llama.cpp: llama-speculative-simple with up to 3 drafted tokens, its fastest setting here, and the faster of a Q4_0 and a Q8_0 draft in each run; its plain rate is llama-bench's. KwkDecoder: an int4 draft, 3 drafted tokens, its rates including the prompt. October 2026. · Draft models in the guide

Many requests at once

Generating a token reads every weight once, whether the step serves one request or eight. kwker.serve puts the requests in flight into one step and admits new ones as slots free up, and each request gets the tokens it would get alone. python -m kwker.serve --model <id> serves it as an OpenAI-compatible API.

4 coresRequests at onceKwkDecoder int4llama.cpp Q4_0Speed-up
SmolLM2-135M12481501.65×
89464522.09×
SmolLM2-360M110870.91.52×
85362282.35×

Generated tokens per second, summed over the requests: 64-token prompts, then 64 new tokens each, as llama.cpp's llama-batched-bench measures it (S_TG; llama.cpp batches the same requests). Medians of 5 runs (KwkDecoder) and 3 runs (llama.cpp), both measured the same day, October 2026. Serving guide · Raw results (text) · Serving guide

The next chat turn starts at once

Each turn of a chat sends the whole conversation again. KwkDecoder keeps what it computed for the previous prompt and reads only what is new: the last answer and the new message. The tokens are exactly those of a fresh start. kwker.serve does the same for every request slot, so a shared system prompt is read once.

KwkDecoder int4, 4 coresFirst token, cacheFirst token, no cacheSpeed-up
SmolLM2-135M59 ms646 ms10.9×
SmolLM2-1.7B471 ms4,902 ms10.4×

Time to the first token of chat turns 2 and later: a 1,024-token system prompt, 64-token messages and 32-token answers (wikitext). Medians over turns 2-6 (SmolLM2-135M) and 2-4 (SmolLM2-1.7B), then the median of 3 runs, October 2026; both setups gave the same answers. Guide · Raw results (text) · Prompt cache in the guide

Laptops and desktops too

Most laptops and desktops have no AVX-512: Intel Core 12th generation and later, AMD Zen 2 and 3. There KwkDecoder runs its AVX2 kernels, with the same packed weights and results that match the AVX-512 kernels to within float rounding.

AVX2 only, 4 coresPrompt readingSpeed-upGenerationSpeed-up
SmolLM2-135M, KwkDecoder int41,5712.05×236.21.59×
SmolLM2-135M, llama.cpp Q4_0766148.6
SmolLM2-135M, KwkDecoder int81,1871.95×168.51.55×
SmolLM2-135M, llama.cpp Q8_0609108.9
SmolLM2-360M, KwkDecoder int45941.86×108.01.41×
SmolLM2-360M, llama.cpp Q4_031976.5
SmolLM2-1.7B, KwkDecoder int41351.45×26.01.38×
SmolLM2-1.7B, llama.cpp 4-bit93Q4_K_M18.8Q4_0

Tokens per second: a 512-token prompt, then generation. Both sides restricted to AVX2 on the same machine (KwkDecoder's AVX2 kernels; llama.cpp built for AVX2 only), so the memory system is the test machine's and a laptop's own numbers differ. Medians of 5 runs (KwkDecoder) and 3 runs (llama.cpp), October 2026: SmolLM2-135M and SmolLM2-1.7B on 2026-10-04 (after int4 prompt rows stopped carrying residual activations on plain AVX2 too), SmolLM2-360M the day before; for SmolLM2-1.7B llama.cpp's faster 4-bit file per column. Raw results (text) · Supported processors

Your Hugging Face code, unchanged

Already running a bfloat16 model with Hugging Face on a CPU without AMX? One line, kwker.torch_ops.install(), puts its linears and each new token's attention on Kwker's bf16 kernels. The model and the generate() call stay as they are, and nothing is quantized on either side.

Model, bf16, 4 coresPrompt, torch bf16Prompt, Kwker bf16Speed-upGeneration, torch bf16Generation, Kwker bf16Speed-up
SmolLM2-135M3728672.3×28.136.41.3×
Qwen2.5-0.5B1303682.8×12.721.71.7×
SmolLM2-1.7B29.61073.6×5.27.71.5×

Tokens per second: a 128-token prompt, then greedy generation with Hugging Face generate() in eager mode, both sides alternating in one process (median of the rounds), October 2026. The greedy tokens were the same on both sides. Kwker adds in another order than torch, so outputs can differ by bf16 rounding. · The CPU backend in the guide

torch.compile, up to 2× Inductor

Compiling your model? Pass backend="kwker" instead of the default. It reads float32 weights of bf16 checkpoints as their exact bf16 values, runs linears that share an input as one call and runs each decode step as a list of native operator calls, with the KV cache written in place. Nothing is quantized, and the tokens match Inductor's.

Model, 4 coresWeightsInductor, tokens/sbackend="kwker", tokens/sSpeed-up
SmolLM2-135Mfloat3238.462.51.6×
SmolLM2-135Mbf1638.356.31.5×
Qwen2.5-0.5Bfloat3215.030.42.0×
Qwen2.5-0.5Bbf1613.726.61.9×

Hugging Face generate() with a static KV cache: 32 new tokens after a 32-token prompt (prefill included), the median of 3 runs per row, October 2026. The first compile of SmolLM2-135M takes about 18 s, with or without caches (Inductor: about 80 s with empty caches, about 15 s from its cache). · torch.compile in the guide

How it works

  • One call per step: the whole model runs in native code, with no Python and no framework overhead between layers.
  • Compact weights: 8-bit weights with one scale per 64 values, or 4-bit weights fitted to your own text (calib=, optionally with GPTQ).
  • Built for decoding: a preallocated KV cache, fused attention for the new token, and batching of several sequences into one step.
  • Hugging Face's own API: after install(model), model.generate() keeps sampling, streamers, stopping criteria and batches, and every step runs on Kwker.
  • Any Hugging Face model: KwkDecoder covers Llama, Mistral, Qwen2, Qwen3, SmolLM, Phi-3 and Phi-4-mini, Granite, Gemma 2, Gemma 3, the mixture-of-experts Mixtral, Qwen3-MoE and Granite MoE, and GPT-2, Pythia, Phi-2, OPT, StableLM, StarCoder2 and OLMo; Decoder compiles any causal language model.
PythonRuns on your machine.
import kwker
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from kwker.decode import install

name = "HuggingFaceTB/SmolLM2-135M"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, dtype=torch.float32).eval()

install(model, precision="int4")  # 4-bit weights; "balanced" for 8-bit, the default keeps the model's own

ids = tok("Sorting algorithms are", return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=24, do_sample=False)   # Hugging Face's own API
print(tok.decode(out[0]))
Output
Sorting algorithms are used to sort a list of items.

The following example shows how to sort a list of numbers using the sort
The KwkDecoder guide, with runnable examples

Quality and requirements

  • 8-bit weights stay within a quarter of a percent of the original model's perplexity, closer than llama.cpp's Q8_0 on every model measured (SmolLM2-1.7B: +0.08% against +0.30%; Qwen2.5-0.5B: +0.12% against +0.16%; SmolLM2-135M: +0.12% against +0.25%).
  • 4-bit weights trade quality for speed. On SmolLM2-1.7B, KwkDecoder int4 costs +6.0% perplexity against +8.4% for llama.cpp's Q4_K_M, at about the same bytes per token. On the two smallest models Q4_K_M is ahead (SmolLM2-135M +8.4% against +6.1%, Qwen2.5-0.5B +5.1% against +3.7%: there it spends about 5.5 bits per weight), and llama.cpp's 4-bit Q4_0 is well behind (+25.1% and +11.6%). Calibrated 4-bit beats llama.cpp's best 4-bit file on Qwen2.5-0.5B (+2.33% against +2.70%). KwkDecoder's int4 keeps about a quarter of the weights at higher precision: 6 bits from a billion parameters, 8 bits below. On SmolLM2-1.7B that is 5.0 bits per weight on average (4.625 for the 4-bit weights with their scales), Q4_K_M 4.9 and Q4_0 4.6. Measure on your own text before you deploy a format.
  • Processor: AVX-512 with VNNI (Intel Xeon Cascade Lake and later, AMD Zen 4 and later) or AVX2 (Intel Core 12th generation and later, AMD Zen 2 and 3); AMX speeds up prompt reading where present. Linux x86-64 with PyTorch.

Perplexity on wikitext-2 with llama.cpp's own protocol (20 chunks of 512 tokens). Details and every format: limitations, guide.

Free to start

Free for companies under 100 people and about $1.3M revenue (PolyForm Small Business). Every product line is in every tier.