Quickstart: embeddings
Encode text faster on the CPU. KwkEncoder runs a BERT-family model's whole forward pass in one native call and skips
padding tokens instead of computing them. About five minutes.
It runs BERT, RoBERTa, XLM-R, CamemBERT, DistilBERT, MPNet, ModernBERT, NomicBERT, EuroBERT and the sentence embedders, rerankers and classifiers built on them, on x86 CPUs with AVX2 or AVX-512 (Intel since 2013, AMD since 2015). It is fastest with Intel AMX (Xeon Sapphire Rapids or newer).
Install
pip install kwker transformers
python -m kwker doctor # shows this CPU's features (AVX2, AVX-512, AMX)
Encode
Give a model's name, then encode texts:
import kwker
enc = kwker.Encoder.from_pretrained("BAAI/bge-small-en-v1.5")
vectors = enc.encode(["Kwker runs on the CPU.", "Embeddings, faster."])
print(vectors.shape) # (2, 384)
encode returns one float32 vector per text, pooled and normalized the way the model's sentence-transformers
configuration says (CLS or mean pooling, unit length or not).
With a model you already loaded
Wrap the model once, then call it like the model: token IDs and an attention mask in, hidden states out.
import kwker
import torch
from transformers import AutoModel, AutoTokenizer
from kwker.encode import KwkEncoder
name = "sentence-transformers/all-MiniLM-L6-v2"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModel.from_pretrained(name).eval()
enc = KwkEncoder(model)
batch = tok(["Kwker runs on the CPU.", "Embeddings, faster."], padding=True, return_tensors="pt")
hidden = enc(batch.input_ids, attention_mask=batch.attention_mask)
mask = batch.attention_mask.unsqueeze(-1)
emb = (hidden * mask).sum(1) / mask.sum(1) # mean pooling, as the model card does it
print(emb.shape) # [2, 384]
The same on a small random model, so you can run it without a download:
import kwker
import torch
from transformers import BertConfig, BertModel
from kwker.encode import KwkEncoder, runner_unsupported
model = BertModel(BertConfig(vocab_size=1000, hidden_size=128, num_hidden_layers=2, num_attention_heads=2,
intermediate_size=256)).eval()
why = runner_unsupported(model)
if why is None:
enc = KwkEncoder(model)
ids = torch.randint(0, 1000, (4, 32))
print(enc(ids, attention_mask=torch.ones_like(ids)).shape)
else:
print("KwkEncoder unavailable:", why)
runner_unsupported(model) returns None when the model can run here, or the reason it can't (for example, an
unsupported model type).
Check the results
Compare with the model's own output on a few real sentences:
with torch.no_grad():
ref = model(batch.input_ids, attention_mask=batch.attention_mask).last_hidden_state
print(torch.nn.functional.cosine_similarity(hidden.flatten(1), ref.flatten(1))) # close to 1.0
KwkEncoder(model, int8=True, calib=<example batches>) uses int8 weights and activations: faster, with a small change
in the output. Check it the same way.
Measure it on your machine
python -m kwker.bench --encoders --threads 4 # KwkEncoder vs PyTorch and OpenVINO, batch 1 and 8
Next steps
- Sentence embeddings: the supported models and options.
- Evaluate on your machine: what each benchmark measures.