Kwker

Generate text with a Hugging Face model

KwkDecoder runs a Hugging Face language model's whole decoding step as one native call. Against llama.cpp on the same CPU, its 4-bit models generated 1.75x faster and read prompts 2.65x faster (geometric means over SmolLM2-135M, Qwen2.5-0.5B and SmolLM2-1.7B), and its bfloat16 mode, which quantizes nothing, was 1.6x faster than llama.cpp's BF16 files. The examples below use a tiny random model so they run anywhere; with a real model you call from_pretrained(...) instead.

Prerequisites

Supported families:

Step 1: load a model and check it

runner_unsupported(model, max_cache_len) returns None when the model can run, or the reason it can't. decoder_available() says whether this CPU runs KwkDecoder at all.

PythonNeeds torch, transformers: runs on your machine.
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.cpu_backend as scb
from kwker.decode import KwkDecoder, install, runner_unsupported

model = LlamaForCausalLM(LlamaConfig(vocab_size=1000, hidden_size=128, intermediate_size=256, num_hidden_layers=2,
                                     num_attention_heads=2, num_key_value_heads=1)).eval()
# a real model: AutoModelForCausalLM.from_pretrained("HuggingFaceTB/SmolLM2-360M-Instruct")
ok = scb.decoder_available() and runner_unsupported(model, 64) is None
print(ok)
Output
True

Step 2: generate

Build the decoder once with the longest context you need (prompt plus new tokens), then call generate:

PythonRuns on your machine.
dec = KwkDecoder(model, max_cache_len=64)                            # the model's own precision: nothing rounded
out = dec.generate(torch.tensor([[1, 5, 7, 9]]), max_new_tokens=8)   # greedy: prompt + 8 tokens
print(out.shape)
Output
torch.Size([1, 12])

precision="balanced" (int8 weights) or precision="int4" makes the model smaller and faster with a small change in its output; see int8 and int4 weights.

Step 3: sample

generate() is greedy by default. For sampling, pass Hugging Face's options, seeded through generator=:

PythonRuns on your machine.
g = torch.Generator().manual_seed(0)
out = dec.generate(torch.tensor([[1, 5, 7, 9]]), max_new_tokens=8, do_sample=True, temperature=0.8, top_p=0.9,
                   generator=g)
print(out.shape)
Output
torch.Size([1, 12])

Also: top_k, min_p, repetition_penalty, and min_new_tokens=n (no end-of-sequence token in the first n new tokens).

Step 4: keep model.generate()

If your application already calls Hugging Face's generate, install(model) keeps its whole interface (greedy or sampling, stopping criteria, streamers, prompt_lookup_num_tokens) and runs every step on a KwkDecoder:

PythonRuns on your machine.
install(model, max_cache_len=256)              # the model's own precision; precision="balanced" for int8
torch.manual_seed(0)
out = model.generate(torch.tensor([[1, 5, 7, 9]]), max_new_tokens=8, min_new_tokens=8, do_sample=True,
                     top_p=0.9, pad_token_id=0)
print(out.shape[1] - 4, model._kwker_fallback)
Output
8 None

Calls it cannot take, such as beam search, run on the model's own forward, and model._kwker_fallback says why. uninstall(model) restores the original generate.

Other models

For a Hugging Face causal language model outside the families above, Decoder exports and compiles the decode step once: tens of seconds the first time, then cached on disk. Without either, kwker.torch_ops.install() still speeds up a plain model.generate(...) (Speed up inference).

Next steps