Sentence embeddings
KwkEncoder runs a BERT-family encoder (BERT, RoBERTa, XLM-R, CamemBERT, DistilBERT, MPNet, ModernBERT, NomicBERT, EuroBERT and
the sentence embedders and classifiers built on them) in one native call. Padding tokens are skipped, not computed. On a 4-core Xeon
without AMX, all-MiniLM-L6-v2 on a padded batch of 8 x 128 tokens took 44 ms in the default mode and 17 ms in int8,
against 84 ms in PyTorch eager and 60 / 24 ms in OpenVINO float32 / int8.
Prerequisites
- An x86-64 CPU with AVX2 or AVX-512; int8 needs AMX, AVX-512 VNNI or AVX-VNNI (Intel Core 12th generation and later).
pip install kwker transformers(Installation).
Step 1: build the encoder
A tiny random model stands in for a real one here; with a real model you call AutoModel.from_pretrained(...):
import torch
from transformers import BertConfig, BertModel
from kwker.encode import KwkEncoder, runner_unsupported
model = BertModel(BertConfig(vocab_size=1000, hidden_size=128, num_hidden_layers=2, num_attention_heads=2,
intermediate_size=256)).eval()
# a real model: AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
print(runner_unsupported(model))
enc = KwkEncoder(model)
None
Step 2: embed a batch
enc(ids, attention_mask=...) returns the last hidden state, like the model's own forward. Sentence embedders such as
all-MiniLM-L6-v2 average it over the real tokens:
ids = torch.randint(0, 1000, (4, 32))
mask = torch.ones_like(ids)
mask[2, 20:] = 0 # a shorter sentence, padded
hidden = enc(ids, attention_mask=mask) # [4, 32, 128]
emb = (hidden * mask[..., None]).sum(1) / mask.sum(1, keepdim=True)
emb = torch.nn.functional.normalize(emb, dim=-1)
print(emb.shape)
torch.Size([4, 128])
Step 3: check against the original model
with torch.no_grad():
ref = model(ids, attention_mask=mask).last_hidden_state
cos = torch.nn.functional.cosine_similarity(hidden[mask.bool()], ref[mask.bool()], dim=-1)
print(bool(cos.min() > 0.999))
True
Step 4: int8
int8=True with typical inputs as calib= (a tokenizer batch, a list of them, or token ids) runs int8 weights and
activations. It changes the embeddings slightly, so compare it the same way on your own sentences:
enc8 = KwkEncoder(model, int8=True, calib={"input_ids": ids, "attention_mask": mask}) # or a list of batches
python -m kwker.bench --encoders --int8 compares KwkEncoder with OpenVINO and PyTorch on your machine. ModernBERT,
NomicBERT and EuroBERT models run in bf16 or float32 only; int8 raises ValueError for them.
Next steps
- ONNX Runtime: the same encoder speed for an ONNX file, inside ONNX Runtime.
- Lower precision with an accuracy check: bfloat16 and int8 for other models.
- Evaluate on your machine: what each benchmark measures.