Quickstart: PyTorch
Make a PyTorch model faster on the CPU without editing it. You switch Kwker on in one line, check that the model still gives the same results, then measure the difference on your machine. About five minutes.
Install
You need Linux on x86-64 and the Kwker wheel built for your torch release (Installation):
pip install kwker
python -m kwker doctor # your CPU, the engine Kwker picked, any warnings
Switch it on
install() makes torch's CPU sort, top-k, unique, quantile, median and similar operations run on Kwker, everywhere in
the process: your code and every library you call.
import torch
import kwker.torch_ops
kwker.torch_ops.install()
scores = torch.randn(100_000)
best = torch.topk(scores, 5)
print(best.values.shape, best.indices.dtype)
torch.Size([5]) torch.int64
The results are torch's own, bit for bit. kwker.torch_ops.uninstall() switches it off again.
Or compile the model
torch.compile with the kwker backend also runs linear layers, attention and convolutions on Kwker. Where
Inductor's own code is faster for a graph, it keeps that. Nothing to import:
fast = torch.compile(model, backend="kwker")
out = fast(inputs)
Check the results
Compare your model's output with and without Kwker before you rely on it. For the drop-in kernels the answer is exact:
import torch
import kwker.torch_ops
x = torch.randn(50_000)
plain = torch.sort(x).values
with kwker.torch_ops.accelerated(): # Kwker inside the block only
fast = torch.sort(x).values
print(torch.equal(plain, fast))
True
Lower-precision modes (bfloat16, int8) change results slightly. kwker.cpu_backend.check_accuracy(model, x, mode="int8") measures how much on your model before you turn one on; see Lower precision with an accuracy check.
Measure it on your machine
python -m kwker.bench --models --threads 4 # image models: Kwker vs PyTorch eager, Inductor and OpenVINO
To time your own model, run it a few times with and without install() (or the compiled version) on the same inputs.
Next steps
- Accelerate PyTorch: inference, torch.compile, lower precision, training and the operators, one page each.
- Quickstart: language models, embeddings and image models: runners that execute a whole model in one native call.
- Tutorial: speed up a PyTorch model: the same steps on a model that scores and ranks items.