Kwker

Quickstart: PyTorch

Make a PyTorch model faster on the CPU without editing it. You switch Kwker on in one line, check that the model still gives the same results, then measure the difference on your machine. About five minutes.

Install

You need Linux on x86-64 and the Kwker wheel built for your torch release (Installation):

ShellOn your machine.
pip install kwker
python -m kwker doctor            # your CPU, the engine Kwker picked, any warnings

Switch it on

install() makes torch's CPU sort, top-k, unique, quantile, median and similar operations run on Kwker, everywhere in the process: your code and every library you call.

PythonNeeds torch: runs on your machine.
import torch
import kwker.torch_ops

kwker.torch_ops.install()
scores = torch.randn(100_000)
best = torch.topk(scores, 5)
print(best.values.shape, best.indices.dtype)
Output
torch.Size([5]) torch.int64

The results are torch's own, bit for bit. kwker.torch_ops.uninstall() switches it off again.

Or compile the model

torch.compile with the kwker backend also runs linear layers, attention and convolutions on Kwker. Where Inductor's own code is faster for a graph, it keeps that. Nothing to import:

PythonRuns on your machine.
fast = torch.compile(model, backend="kwker")
out = fast(inputs)

Check the results

Compare your model's output with and without Kwker before you rely on it. For the drop-in kernels the answer is exact:

PythonNeeds torch: runs on your machine.
import torch
import kwker.torch_ops

x = torch.randn(50_000)
plain = torch.sort(x).values
with kwker.torch_ops.accelerated():        # Kwker inside the block only
    fast = torch.sort(x).values
print(torch.equal(plain, fast))
Output
True

Lower-precision modes (bfloat16, int8) change results slightly. kwker.cpu_backend.check_accuracy(model, x, mode="int8") measures how much on your model before you turn one on; see Lower precision with an accuracy check.

Measure it on your machine

ShellOn your machine.
python -m kwker.bench --models --threads 4       # image models: Kwker vs PyTorch eager, Inductor and OpenVINO

To time your own model, run it a few times with and without install() (or the compiled version) on the same inputs.

Next steps