Add Kwker to your project
Find your workload below. Each one is a single change to code you already have, and the results stay the same: exactly, or within a precision you pick. Install once:
pip install kwker
python -m kwker doctor # your CPU, the engine Kwker picked, warnings
| Your workload | The change | Full guide |
|---|---|---|
| Any PyTorch model on the CPU | kwker.torch_, or torch.compile(model, backend="kwker") |
PyTorch |
| Text generation with an LLM | KwkDecoder(model, max_ |
KwkDecoder |
| Embeddings, rerankers, text classifiers | KwkEncoder(model) |
KwkEncoder |
| Image classifiers | KwkCNN(model, example_ |
KwkCNN |
| pandas, Polars, pyarrow, DuckDB | kwker.frame.group_, .sort(...), .top_ |
DataFrames |
| Arrays, in Python and 19 more languages | kwker.sort(a), kwker.top_ |
Quickstart: Kwker Core |
| An existing NumPy program | kwker.numpy_: np.sort, np.argsort, np.unique and others run Kwker |
NumPy drop-in |
The PyTorch features ship in the Linux x86-64 wheel built for your torch release (Installation). To see what Kwker changes for your own program before you commit to it, read Evaluate on your machine.
PyTorch models
Two ways in, and neither touches your model code. install() makes torch's CPU sort, top-k, unique, quantile and
related operations run on Kwker, everywhere in the process:
import torch
import kwker.torch_ops
kwker.torch_ops.install() # torch.sort / topk / unique / quantile / ... now run on Kwker
x = torch.randn(1_000_000)
print(torch.topk(x, 3).values.shape) # same results as torch, bit for bit
torch.compile with the kwker backend also speeds up linear layers, attention and convolutions. Where Inductor's
own code is faster for a graph, it keeps that:
fast = torch.compile(model, backend="kwker")
Text generation
KwkDecoder runs these models on their own weights (the default, precision="preserve") or with int8 or int4 weights
(precision="balanced" or "int4"):
- Llama-family decoders: Llama, Mistral and Ministral, Qwen2, Qwen3, SmolLM2 and SmolLM3, Phi-3 and Phi-4-mini, Granite, Helium, Gemma 2 and Gemma 3.
- Mixture-of-experts models: Mixtral, Qwen3-MoE and Granite MoE.
- The GPT-2 family and other LayerNorm models: GPT-2, GPT-NeoX and Pythia, Phi-1.5 and Phi-2, OPT, StableLM, StarCoder2, StarCoder and SantaCoder, OLMo.
It runs on AVX-512 CPUs with VNNI (Intel Xeon Cascade Lake or newer, AMD Zen 4 or newer) and on AVX2 CPUs (Intel Core 12th generation or newer, AMD Zen 2 and 3). Load the model, build the decoder once, then generate:
import kwker
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.cpu_backend as scb
from kwker.decode import KwkDecoder, runner_unsupported
model = LlamaForCausalLM(LlamaConfig(vocab_size=1000, hidden_size=128, intermediate_size=256, num_hidden_layers=2,
num_attention_heads=2, num_key_value_heads=1)).eval()
if scb.decoder_available() and runner_unsupported(model, 64) is None:
dec = KwkDecoder(model, max_cache_len=64, precision="int8") # or "int4"; the default keeps the model's precision
out = dec.generate(torch.tensor([[1, 5, 7, 9]]), max_new_tokens=8)
print(out.shape)
For a real checkpoint, load it with from_pretrained(name, torch_dtype=torch.float32).
Embeddings
KwkEncoder runs BERT, RoBERTa, XLM-R, DistilBERT, MPNet, ModernBERT, NomicBERT, EuroBERT and the sentence embedders and rerankers
built on them, on x86 CPUs with AVX2 or AVX-512 (fastest with Intel AMX). It skips padding tokens instead of computing them:
import kwker
import torch
from transformers import BertConfig, BertModel
from kwker.encode import KwkEncoder, runner_unsupported
model = BertModel(BertConfig(vocab_size=1000, hidden_size=128, num_hidden_layers=2, num_attention_heads=2,
intermediate_size=256)).eval()
if runner_unsupported(model) is None:
enc = KwkEncoder(model)
ids = torch.randint(0, 1000, (4, 32))
hidden = enc(ids, attention_mask=torch.ones_like(ids)) # [4, 32, 128] float32
Image models
KwkCNN runs ResNet, MobileNet, EfficientNet, RegNet, ResNeXt, ConvNeXt and similar classifiers as one native call, in
float32 or calibrated int8:
import torch, torchvision
from kwker.cnn import KwkCNN
model = torchvision.models.resnet18(weights=None).eval()
fast = KwkCNN(model, example_shape=(1, 3, 64, 64)) # float32: matches eager to rounding
x = torch.randn(2, 3, 64, 64)
with torch.no_grad():
print(float((fast(x) - model(x)).abs().max()) < 1e-3) # True
For int8, add int8=True, calib=<images like your real inputs>, then check top-1 accuracy on your own data.
DataFrames and SQL
kwker.frame takes a pandas, Polars or pyarrow table and returns the same type:
import pandas as pd
import kwker.frame as sf
orders = pd.DataFrame({"customer": ["b", "a", "b", "a", "c"], "amount": [10.0, 2.0, 5.0, 4.0, 1.0]})
print(sf.group_by(orders, "customer", {"orders": "size", "revenue": ("amount", "sum")}))
print(sf.top_k(orders, "amount", 2, descending=True)) # ORDER BY amount DESC LIMIT 2
DuckDB relations work through kwker.duck; see DataFrames, Arrow and DuckDB.
Arrays in any language
The engine under everything above sorts, selects and searches arrays directly. In Python it works on NumPy arrays; the same calls exist in Rust, C, C++, JavaScript, Go, Java, C# and 12 more languages (All languages).
import numpy as np
import kwker
latency_ms = np.array([120.0, 85.0, 430.0, 95.0, 610.0, 150.0, 520.0])
slowest, at = kwker.top_k(latency_ms, 3, descending=True)
print(slowest, at)
kwker.sort(latency_ms)
print(latency_ms)
[610. 520. 430.] [4 6 2]
[ 85. 95. 120. 150. 430. 520. 610.]
Next steps
- Evaluate on your machine: measure the difference on your hardware and your program.
- Accelerate PyTorch: precision modes, accuracy checks and every runner option.
- DataFrames, Arrow and DuckDB: sorting, top-k and group-by rules per library.
- Quickstarts: one per workflow; Kwker Core has the array API in eight languages.