Accelerate PyTorch
Kwker makes PyTorch faster on the CPU at four levels. Each is one change to your code; pick the one that fits how much you want to change.
| Level | What you write | What runs on Kwker | Measured speed-up |
|---|---|---|---|
| Drop-in kernels | kwker.torch_ once |
torch's own sort, topk, unique, quantile, optimizer steps, ... for every caller |
each call about 9x torch's own (median of 97 real ML inputs, one thread) |
| torch.compile | torch.compile(model, backend="kwker") |
linear layers, attention and convolutions, plus the sorting ops | language-model decoding 1.5-2.0x Inductor; float32 CNNs 1.3-2.3x eager (batch 1) |
| Whole-model runners | KwkCNN(model), KwkEncoder(model), KwkDecoder(model, ...) |
the entire forward pass as one native call | KwkCNN 2.8x eager in float32; KwkDecoder bf16 4.2x float32 eager, the same tokens |
| Operators | torch.ops.kwker.sort(x) and friends |
that call | the same 9x per call |
All of it runs on the CPU, on x86-64 Linux for now (Installation has the wheel for your torch release). Tensors on other devices stay with PyTorch. The figures are from 4-core Intel Xeon servers; the full tables are at kwker.io/benchmarks.
Examples
Speed up a model you already have
- Speed up inference with one line:
install(), the same results, measured withkwker audit. - Compile with torch.compile:
backend="kwker", never slower than Inductor on the graphs it tunes. - Lower precision with an accuracy check: bfloat16 and int8 on AMX and VNNI CPUs.
- A faster model, step by step: drop-in kernels, then torch.compile, timed on one model.
Train on the CPU
- Train faster on the CPU: fused Adam, AdamW and SGD steps, the same parameters bit for bit.
Run whole models natively
- Generate text with a Hugging Face model: KwkDecoder, ahead of llama.cpp.
- Sentence embeddings: KwkEncoder for BERT-family models.
- Classify images: KwkCNN in float32 or calibrated int8.
Build with Kwker's operators
- Sorting operators and layers:
torch.ops.kwker, median pooling, k-winners-take-all, torch.export. - Accelerate JAX: the same calls under
jitandvmap.
Next steps
- Quickstart: PyTorch: five minutes from install to a checked speed-up.
- Runtime controls: the switches for single features.