Kwker

Quickstart: embeddings

Encode text faster on the CPU. KwkEncoder runs a BERT-family model's whole forward pass in one native call and skips padding tokens instead of computing them. About five minutes.

It runs BERT, RoBERTa, XLM-R, CamemBERT, DistilBERT, MPNet, ModernBERT, NomicBERT, EuroBERT and the sentence embedders, rerankers and classifiers built on them, on x86 CPUs with AVX2 or AVX-512 (Intel since 2013, AMD since 2015). It is fastest with Intel AMX (Xeon Sapphire Rapids or newer).

Install

ShellOn your machine.
pip install kwker transformers
python -m kwker doctor            # shows this CPU's features (AVX2, AVX-512, AMX)

Encode

Give a model's name, then encode texts:

PythonRuns on your machine.
import kwker

enc = kwker.Encoder.from_pretrained("BAAI/bge-small-en-v1.5")
vectors = enc.encode(["Kwker runs on the CPU.", "Embeddings, faster."])
print(vectors.shape)                               # (2, 384)

encode returns one float32 vector per text, pooled and normalized the way the model's sentence-transformers configuration says (CLS or mean pooling, unit length or not).

With a model you already loaded

Wrap the model once, then call it like the model: token IDs and an attention mask in, hidden states out.

PythonRuns on your machine.
import kwker
import torch
from transformers import AutoModel, AutoTokenizer
from kwker.encode import KwkEncoder

name = "sentence-transformers/all-MiniLM-L6-v2"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModel.from_pretrained(name).eval()
enc = KwkEncoder(model)

batch = tok(["Kwker runs on the CPU.", "Embeddings, faster."], padding=True, return_tensors="pt")
hidden = enc(batch.input_ids, attention_mask=batch.attention_mask)
mask = batch.attention_mask.unsqueeze(-1)
emb = (hidden * mask).sum(1) / mask.sum(1)        # mean pooling, as the model card does it
print(emb.shape)                                   # [2, 384]

The same on a small random model, so you can run it without a download:

PythonNeeds torch, transformers: runs on your machine.
import kwker
import torch
from transformers import BertConfig, BertModel
from kwker.encode import KwkEncoder, runner_unsupported

model = BertModel(BertConfig(vocab_size=1000, hidden_size=128, num_hidden_layers=2, num_attention_heads=2,
                             intermediate_size=256)).eval()
why = runner_unsupported(model)
if why is None:
    enc = KwkEncoder(model)
    ids = torch.randint(0, 1000, (4, 32))
    print(enc(ids, attention_mask=torch.ones_like(ids)).shape)
else:
    print("KwkEncoder unavailable:", why)

runner_unsupported(model) returns None when the model can run here, or the reason it can't (for example, an unsupported model type).

Check the results

Compare with the model's own output on a few real sentences:

PythonRuns on your machine.
with torch.no_grad():
    ref = model(batch.input_ids, attention_mask=batch.attention_mask).last_hidden_state
print(torch.nn.functional.cosine_similarity(hidden.flatten(1), ref.flatten(1)))   # close to 1.0

KwkEncoder(model, int8=True, calib=<example batches>) uses int8 weights and activations: faster, with a small change in the output. Check it the same way.

Measure it on your machine

ShellOn your machine.
python -m kwker.bench --encoders --threads 4     # KwkEncoder vs PyTorch and OpenVINO, batch 1 and 8

Next steps