Quickstart: language models
Run a Hugging Face language model faster on the CPU. You build a KwkDecoder from the model once, then generate. You
can also keep calling model.generate(...), or serve an OpenAI-compatible API. About five minutes.
KwkDecoder runs Llama, Mistral, Qwen, SmolLM, Phi, Granite, Gemma, Mixtral, GPT-2 and similar models
(the full list), on AVX-512 VNNI CPUs and on AVX2 CPUs (Intel Core
12th generation or newer, AMD Zen 2 or newer).
Install
pip install kwker transformers
python -m kwker doctor # your CPU, the engine Kwker picked, any warnings
Generate
Give a model's name, then chat or generate:
import kwker
llm = kwker.Decoder.from_pretrained("HuggingFaceTB/SmolLM2-360M-Instruct")
print(llm.chat("What is the capital of France?"))
for piece in llm.chat("Name three rivers.", stream=True):
print(piece, end="")
from_pretrained downloads the model and runs it at the checkpoint's own precision, so nothing is rounded.
precision="balanced" stores the weights in int8: about twice as fast, about +0.1% perplexity. llm.load_report()
says what runs. A model KwkDecoder does not cover runs on transformers with Kwker's PyTorch kernels.
With a model you already loaded
Build the decoder from the model object, then generate. precision= works the same way.
import kwker
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from kwker.decode import KwkDecoder
name = "HuggingFaceTB/SmolLM2-360M-Instruct"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, torch_dtype=torch.float32).eval()
dec = KwkDecoder(model, max_cache_len=512, precision="balanced") # int8 weights
ids = tok("The capital of France is", return_tensors="pt").input_ids
out = dec.generate(ids, max_new_tokens=20)
print(tok.decode(out[0]))
The same code on a small random model, so you can run it without a download:
import kwker
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.cpu_backend as scb
from kwker.decode import KwkDecoder, runner_unsupported
model = LlamaForCausalLM(LlamaConfig(vocab_size=1000, hidden_size=128, intermediate_size=256, num_hidden_layers=2,
num_attention_heads=2, num_key_value_heads=1)).eval()
if scb.decoder_available() and runner_unsupported(model, 64) is None:
dec = KwkDecoder(model, max_cache_len=64, precision="int8")
out = dec.generate(torch.tensor([[1, 5, 7, 9]]), max_new_tokens=8)
print(out.shape[0])
runner_unsupported(model, max_cache_len) returns None when the model can run, or the reason it can't.
Keep model.generate()
install(model) keeps Hugging Face's generate and all its options (sampling, stopping criteria, streamers), and runs
every step on a KwkDecoder:
from kwker.decode import install
install(model, max_cache_len=1024, precision="balanced") # int8; the default keeps the model's precision
out = model.generate(ids, max_new_tokens=50, do_sample=True, top_p=0.9)
Serve many requests
kwker.serve batches requests into one decoder and serves an OpenAI-compatible API on port 8000:
python -m kwker.serve --model HuggingFaceTB/SmolLM2-360M-Instruct --max-batch 8
Any OpenAI client works against it with base_url="http://127.0.0.1:8000/v1".
Check quality and measure speed
int8 and int4 weights change the model's output slightly. Measure perplexity and speed on your own text before you deploy a weight format:
python -m kwker.bench --llm HuggingFaceTB/SmolLM2-135M --text wiki.test.raw
The report gives prompt and generation speed next to the original model's, and the perplexity of each weight format.
Next steps
- Generate text with a Hugging Face model: every option, speculative decoding, continuous batching, the prompt cache.
- Evaluate on your machine: comparing with llama.cpp in the same report.