Compile with torch.compile
torch.compile(model, backend="kwker") compiles a model for the CPU with Kwker's kernels inside. It is never slower
than plain Inductor on the graphs it tunes, and it is 1.08-6.4x faster than stock Inductor on the models measured.
No import is needed: the backend is found through the package's entry point.
import torch
model = ... # any model
fast = torch.compile(model, backend="kwker")
out = fast(inputs) # the first call compiles
What it changes
The backend replaces what it can with Kwker operators, then compiles with Inductor:
- sorting ops, median pooling, k-nearest-neighbours (
topk(cdist(...))), top-k of a product, k-winners-take-all - linear layers with fused activations, and conv / batch-norm / residual / ReLU chains
For inference graphs it times Inductor, eager and its own version once and keeps the fastest. Training graphs are tuned the same way at their first call. Graphs on CUDA, XPU or MPS go to Inductor unchanged.
Language models get more. Linear layers read fewer weight bytes (bfloat16 weights as stored, a bfloat16 copy of
bfloat16-valued float32 weights), linear layers that share an input run as one call, and the small operations between
them are fused. Decode steps that write a static KV cache skip Inductor and run as a list of native calls: as fast or
faster, with a much shorter first compile. Hugging Face generate() compiled this way decodes 1.47-2.03x faster than
with stock Inductor, with the same tokens.
Options
mode= and options= work as they do with Inductor. Three options are Kwker's own:
fast = torch.compile(model, backend="kwker", options={"inner": "eager"}) # skip Inductor: the quickest first compile
fast = torch.compile(model, backend="kwker", options={"inner": "inductor"}) # always compile with Inductor
fast = torch.compile(model, backend="kwker", options={"glue": False}) # keep the model's own norm / rotary ops
With {"inner": "eager"} the first compile of a small language model takes about 18 seconds instead of over a
minute, and its decode step is within a few percent of the Inductor version.
Skip guard checks in a decode loop
Once a decode loop has compiled every shape it uses, you can skip Dynamo's guard checks on each call:
with torch.compiler.set_stance("default", skip_guard_eval_unsafe=True):
out = model.generate(input_ids, max_new_tokens=64)
Only do this after a warm-up call: with the guards off, a changed input shape or dtype is no longer detected.
Notes
- Compiled linear layers can round differently in the last bits, so compare outputs with a tolerance.
- Models with sliding-window layers (Gemma 2, Gemma 3) compile once for decoding. Hugging Face's static cache recompiles the decode step at every token for them, so the backend stores those layers like full-length layers when the cache holds at most 4x the window. The model attends the same tokens.
- With
options={"inner": "eager"}, Gemma's RMSNorm and GELU product are fused too. KWKER_COMPILE_DEBUG=1prints what the backend rewrote; the other switches are in Runtime controls.
Next steps
- A faster model, step by step: compile a model and time it against eager.
- Lower precision with an accuracy check: bfloat16 and int8 inside the compiled model.
- Generate text with a Hugging Face model: KwkDecoder, faster still for supported models.