Kwker

Quickstart: language models

Run a Hugging Face language model faster on the CPU. You build a KwkDecoder from the model once, then generate. You can also keep calling model.generate(...), or serve an OpenAI-compatible API. About five minutes.

KwkDecoder runs Llama, Mistral, Qwen, SmolLM, Phi, Granite, Gemma, Mixtral, GPT-2 and similar models (the full list), on AVX-512 VNNI CPUs and on AVX2 CPUs (Intel Core 12th generation or newer, AMD Zen 2 or newer).

Install

ShellOn your machine.
pip install kwker transformers
python -m kwker doctor            # your CPU, the engine Kwker picked, any warnings

Generate

Give a model's name, then chat or generate:

PythonRuns on your machine.
import kwker

llm = kwker.Decoder.from_pretrained("HuggingFaceTB/SmolLM2-360M-Instruct")
print(llm.chat("What is the capital of France?"))
for piece in llm.chat("Name three rivers.", stream=True):
    print(piece, end="")

from_pretrained downloads the model and runs it at the checkpoint's own precision, so nothing is rounded. precision="balanced" stores the weights in int8: about twice as fast, about +0.1% perplexity. llm.load_report() says what runs. A model KwkDecoder does not cover runs on transformers with Kwker's PyTorch kernels.

With a model you already loaded

Build the decoder from the model object, then generate. precision= works the same way.

PythonRuns on your machine.
import kwker
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from kwker.decode import KwkDecoder

name = "HuggingFaceTB/SmolLM2-360M-Instruct"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, torch_dtype=torch.float32).eval()

dec = KwkDecoder(model, max_cache_len=512, precision="balanced")   # int8 weights
ids = tok("The capital of France is", return_tensors="pt").input_ids
out = dec.generate(ids, max_new_tokens=20)
print(tok.decode(out[0]))

The same code on a small random model, so you can run it without a download:

PythonNeeds torch, transformers: runs on your machine.
import kwker
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.cpu_backend as scb
from kwker.decode import KwkDecoder, runner_unsupported

model = LlamaForCausalLM(LlamaConfig(vocab_size=1000, hidden_size=128, intermediate_size=256, num_hidden_layers=2,
                                     num_attention_heads=2, num_key_value_heads=1)).eval()
if scb.decoder_available() and runner_unsupported(model, 64) is None:
    dec = KwkDecoder(model, max_cache_len=64, precision="int8")
    out = dec.generate(torch.tensor([[1, 5, 7, 9]]), max_new_tokens=8)
    print(out.shape[0])

runner_unsupported(model, max_cache_len) returns None when the model can run, or the reason it can't.

Keep model.generate()

install(model) keeps Hugging Face's generate and all its options (sampling, stopping criteria, streamers), and runs every step on a KwkDecoder:

PythonRuns on your machine.
from kwker.decode import install

install(model, max_cache_len=1024, precision="balanced")   # int8; the default keeps the model's precision
out = model.generate(ids, max_new_tokens=50, do_sample=True, top_p=0.9)

Serve many requests

kwker.serve batches requests into one decoder and serves an OpenAI-compatible API on port 8000:

ShellOn your machine.
python -m kwker.serve --model HuggingFaceTB/SmolLM2-360M-Instruct --max-batch 8

Any OpenAI client works against it with base_url="http://127.0.0.1:8000/v1".

Check quality and measure speed

int8 and int4 weights change the model's output slightly. Measure perplexity and speed on your own text before you deploy a weight format:

ShellOn your machine.
python -m kwker.bench --llm HuggingFaceTB/SmolLM2-135M --text wiki.test.raw

The report gives prompt and generation speed next to the original model's, and the perplexity of each weight format.

Next steps