Generate text with a Hugging Face model
KwkDecoder runs a Hugging Face language model's whole decoding step as one native call. Against llama.cpp on the
same CPU, its 4-bit models generated 1.75x faster and read prompts 2.65x faster (geometric means over SmolLM2-135M,
Qwen2.5-0.5B and SmolLM2-1.7B), and its bfloat16 mode, which quantizes nothing, was 1.6x faster than llama.cpp's BF16
files. The examples below use a tiny random model so they run anywhere; with a real model you call
from_pretrained(...) instead.
Prerequisites
- An x86-64 CPU with AVX-512 VNNI, or AVX2 (Intel Core 12th generation and later, AMD Zen 2 and later).
pip install kwker transformers(Installation).
Supported families:
- Llama-family decoders: Llama, Mistral and Ministral, Qwen2, Qwen3, SmolLM2 and SmolLM3, Phi-3 and Phi-4-mini, Granite, Helium, Gemma 2 and Gemma 3.
- Mixture-of-experts models: Mixtral, Qwen3-MoE and Granite MoE.
- GPT-2 and other LayerNorm models: GPT-NeoX and Pythia, Phi-1.5 and Phi-2, OPT, StableLM, StarCoder2, StarCoder and SantaCoder, OLMo.
Step 1: load a model and check it
runner_unsupported(model, max_cache_len) returns None when the model can run, or the reason it can't.
decoder_available() says whether this CPU runs KwkDecoder at all.
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.cpu_backend as scb
from kwker.decode import KwkDecoder, install, runner_unsupported
model = LlamaForCausalLM(LlamaConfig(vocab_size=1000, hidden_size=128, intermediate_size=256, num_hidden_layers=2,
num_attention_heads=2, num_key_value_heads=1)).eval()
# a real model: AutoModelForCausalLM.from_pretrained("HuggingFaceTB/SmolLM2-360M-Instruct")
ok = scb.decoder_available() and runner_unsupported(model, 64) is None
print(ok)
True
Step 2: generate
Build the decoder once with the longest context you need (prompt plus new tokens), then call generate:
dec = KwkDecoder(model, max_cache_len=64) # the model's own precision: nothing rounded
out = dec.generate(torch.tensor([[1, 5, 7, 9]]), max_new_tokens=8) # greedy: prompt + 8 tokens
print(out.shape)
torch.Size([1, 12])
precision="balanced" (int8 weights) or precision="int4" makes the model smaller and faster with a small change in its
output; see
int8 and int4 weights.
Step 3: sample
generate() is greedy by default. For sampling, pass Hugging Face's options, seeded through generator=:
g = torch.Generator().manual_seed(0)
out = dec.generate(torch.tensor([[1, 5, 7, 9]]), max_new_tokens=8, do_sample=True, temperature=0.8, top_p=0.9,
generator=g)
print(out.shape)
torch.Size([1, 12])
Also: top_k, min_p, repetition_penalty, and min_new_tokens=n (no end-of-sequence token in the first n new
tokens).
Step 4: keep model.generate()
If your application already calls Hugging Face's generate, install(model) keeps its whole interface (greedy or
sampling, stopping criteria, streamers, prompt_lookup_num_tokens) and runs every step on a KwkDecoder:
install(model, max_cache_len=256) # the model's own precision; precision="balanced" for int8
torch.manual_seed(0)
out = model.generate(torch.tensor([[1, 5, 7, 9]]), max_new_tokens=8, min_new_tokens=8, do_sample=True,
top_p=0.9, pad_token_id=0)
print(out.shape[1] - 4, model._kwker_fallback)
8 None
Calls it cannot take, such as beam search, run on the model's own forward, and model._kwker_fallback says why.
uninstall(model) restores the original generate.
Other models
For a Hugging Face causal language model outside the families above, Decoder exports and compiles the decode step
once: tens of seconds the first time, then cached on disk. Without either, kwker.torch_ops.install() still speeds up
a plain model.generate(...) (Speed up inference).
Next steps
- int8 and int4 weights: smaller models, with a perplexity check.
- Faster generation: speculative decoding, prompt lookup and batches.
- Chat and serving: multi-turn chat and an OpenAI-compatible server.