Kwker

Add Kwker to your project

Find your workload below. Each one is a single change to code you already have, and the results stay the same: exactly, or within a precision you pick. Install once:

ShellOn your machine.
pip install kwker
python -m kwker doctor            # your CPU, the engine Kwker picked, warnings
Your workload The change Full guide
Any PyTorch model on the CPU kwker.torch_ops.install(), or torch.compile(model, backend="kwker") PyTorch
Text generation with an LLM KwkDecoder(model, max_cache_len=...) KwkDecoder
Embeddings, rerankers, text classifiers KwkEncoder(model) KwkEncoder
Image classifiers KwkCNN(model, example_shape=...) KwkCNN
pandas, Polars, pyarrow, DuckDB kwker.frame.group_by(...), .sort(...), .top_k(...) DataFrames
Arrays, in Python and 19 more languages kwker.sort(a), kwker.top_k(a, k) Quickstart: Kwker Core
An existing NumPy program kwker.numpy_ops.install(): np.sort, np.argsort, np.unique and others run Kwker NumPy drop-in

The PyTorch features ship in the Linux x86-64 wheel built for your torch release (Installation). To see what Kwker changes for your own program before you commit to it, read Evaluate on your machine.

PyTorch models

Two ways in, and neither touches your model code. install() makes torch's CPU sort, top-k, unique, quantile and related operations run on Kwker, everywhere in the process:

PythonNeeds torch: runs on your machine.
import torch
import kwker.torch_ops

kwker.torch_ops.install()                 # torch.sort / topk / unique / quantile / ... now run on Kwker
x = torch.randn(1_000_000)
print(torch.topk(x, 3).values.shape)      # same results as torch, bit for bit

torch.compile with the kwker backend also speeds up linear layers, attention and convolutions. Where Inductor's own code is faster for a graph, it keeps that:

PythonRuns on your machine.
fast = torch.compile(model, backend="kwker")

Text generation

KwkDecoder runs these models on their own weights (the default, precision="preserve") or with int8 or int4 weights (precision="balanced" or "int4"):

It runs on AVX-512 CPUs with VNNI (Intel Xeon Cascade Lake or newer, AMD Zen 4 or newer) and on AVX2 CPUs (Intel Core 12th generation or newer, AMD Zen 2 and 3). Load the model, build the decoder once, then generate:

PythonNeeds torch, transformers: runs on your machine.
import kwker
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.cpu_backend as scb
from kwker.decode import KwkDecoder, runner_unsupported

model = LlamaForCausalLM(LlamaConfig(vocab_size=1000, hidden_size=128, intermediate_size=256, num_hidden_layers=2,
                                     num_attention_heads=2, num_key_value_heads=1)).eval()
if scb.decoder_available() and runner_unsupported(model, 64) is None:
    dec = KwkDecoder(model, max_cache_len=64, precision="int8")   # or "int4"; the default keeps the model's precision
    out = dec.generate(torch.tensor([[1, 5, 7, 9]]), max_new_tokens=8)
    print(out.shape)

For a real checkpoint, load it with from_pretrained(name, torch_dtype=torch.float32).

Embeddings

KwkEncoder runs BERT, RoBERTa, XLM-R, DistilBERT, MPNet, ModernBERT, NomicBERT, EuroBERT and the sentence embedders and rerankers built on them, on x86 CPUs with AVX2 or AVX-512 (fastest with Intel AMX). It skips padding tokens instead of computing them:

PythonNeeds torch, transformers: runs on your machine.
import kwker
import torch
from transformers import BertConfig, BertModel
from kwker.encode import KwkEncoder, runner_unsupported

model = BertModel(BertConfig(vocab_size=1000, hidden_size=128, num_hidden_layers=2, num_attention_heads=2,
                             intermediate_size=256)).eval()
if runner_unsupported(model) is None:
    enc = KwkEncoder(model)
    ids = torch.randint(0, 1000, (4, 32))
    hidden = enc(ids, attention_mask=torch.ones_like(ids))   # [4, 32, 128] float32

Image models

KwkCNN runs ResNet, MobileNet, EfficientNet, RegNet, ResNeXt, ConvNeXt and similar classifiers as one native call, in float32 or calibrated int8:

PythonNeeds torch, torchvision: runs on your machine.
import torch, torchvision
from kwker.cnn import KwkCNN

model = torchvision.models.resnet18(weights=None).eval()
fast = KwkCNN(model, example_shape=(1, 3, 64, 64))      # float32: matches eager to rounding
x = torch.randn(2, 3, 64, 64)
with torch.no_grad():
    print(float((fast(x) - model(x)).abs().max()) < 1e-3)   # True

For int8, add int8=True, calib=<images like your real inputs>, then check top-1 accuracy on your own data.

DataFrames and SQL

kwker.frame takes a pandas, Polars or pyarrow table and returns the same type:

PythonNeeds pandas: runs on your machine.
import pandas as pd
import kwker.frame as sf

orders = pd.DataFrame({"customer": ["b", "a", "b", "a", "c"], "amount": [10.0, 2.0, 5.0, 4.0, 1.0]})
print(sf.group_by(orders, "customer", {"orders": "size", "revenue": ("amount", "sum")}))
print(sf.top_k(orders, "amount", 2, descending=True))   # ORDER BY amount DESC LIMIT 2

DuckDB relations work through kwker.duck; see DataFrames, Arrow and DuckDB.

Arrays in any language

The engine under everything above sorts, selects and searches arrays directly. In Python it works on NumPy arrays; the same calls exist in Rust, C, C++, JavaScript, Go, Java, C# and 12 more languages (All languages).

import numpy as np
import kwker

latency_ms = np.array([120.0, 85.0, 430.0, 95.0, 610.0, 150.0, 520.0])
slowest, at = kwker.top_k(latency_ms, 3, descending=True)
print(slowest, at)
kwker.sort(latency_ms)
print(latency_ms)
[610. 520. 430.] [4 6 2]
[ 85.  95. 120. 150. 430. 520. 610.]

Next steps