Kwker

Accelerate PyTorch

Kwker makes PyTorch faster on the CPU at four levels. Each is one change to your code; pick the one that fits how much you want to change.

Level What you write What runs on Kwker Measured speed-up
Drop-in kernels kwker.torch_ops.install() once torch's own sort, topk, unique, quantile, optimizer steps, ... for every caller each call about 9x torch's own (median of 97 real ML inputs, one thread)
torch.compile torch.compile(model, backend="kwker") linear layers, attention and convolutions, plus the sorting ops language-model decoding 1.5-2.0x Inductor; float32 CNNs 1.3-2.3x eager (batch 1)
Whole-model runners KwkCNN(model), KwkEncoder(model), KwkDecoder(model, ...) the entire forward pass as one native call KwkCNN 2.8x eager in float32; KwkDecoder bf16 4.2x float32 eager, the same tokens
Operators torch.ops.kwker.sort(x) and friends that call the same 9x per call

All of it runs on the CPU, on x86-64 Linux for now (Installation has the wheel for your torch release). Tensors on other devices stay with PyTorch. The figures are from 4-core Intel Xeon servers; the full tables are at kwker.io/benchmarks.

Examples

Speed up a model you already have

Train on the CPU

Run whole models natively

Build with Kwker's operators

Next steps