Kwker

Compile with torch.compile

torch.compile(model, backend="kwker") compiles a model for the CPU with Kwker's kernels inside. It is never slower than plain Inductor on the graphs it tunes, and it is 1.08-6.4x faster than stock Inductor on the models measured. No import is needed: the backend is found through the package's entry point.

PythonRuns on your machine.
import torch

model = ...                                        # any model
fast = torch.compile(model, backend="kwker")
out = fast(inputs)                                 # the first call compiles

What it changes

The backend replaces what it can with Kwker operators, then compiles with Inductor:

For inference graphs it times Inductor, eager and its own version once and keeps the fastest. Training graphs are tuned the same way at their first call. Graphs on CUDA, XPU or MPS go to Inductor unchanged.

Language models get more. Linear layers read fewer weight bytes (bfloat16 weights as stored, a bfloat16 copy of bfloat16-valued float32 weights), linear layers that share an input run as one call, and the small operations between them are fused. Decode steps that write a static KV cache skip Inductor and run as a list of native calls: as fast or faster, with a much shorter first compile. Hugging Face generate() compiled this way decodes 1.47-2.03x faster than with stock Inductor, with the same tokens.

Options

mode= and options= work as they do with Inductor. Three options are Kwker's own:

PythonRuns on your machine.
fast = torch.compile(model, backend="kwker", options={"inner": "eager"})     # skip Inductor: the quickest first compile
fast = torch.compile(model, backend="kwker", options={"inner": "inductor"})  # always compile with Inductor
fast = torch.compile(model, backend="kwker", options={"glue": False})        # keep the model's own norm / rotary ops

With {"inner": "eager"} the first compile of a small language model takes about 18 seconds instead of over a minute, and its decode step is within a few percent of the Inductor version.

Skip guard checks in a decode loop

Once a decode loop has compiled every shape it uses, you can skip Dynamo's guard checks on each call:

PythonRuns on your machine.
with torch.compiler.set_stance("default", skip_guard_eval_unsafe=True):
    out = model.generate(input_ids, max_new_tokens=64)

Only do this after a warm-up call: with the guards off, a changed input shape or dtype is no longer detected.

Notes

Next steps