Kwker

Sentence embeddings

KwkEncoder runs a BERT-family encoder (BERT, RoBERTa, XLM-R, CamemBERT, DistilBERT, MPNet, ModernBERT, NomicBERT, EuroBERT and the sentence embedders and classifiers built on them) in one native call. Padding tokens are skipped, not computed. On a 4-core Xeon without AMX, all-MiniLM-L6-v2 on a padded batch of 8 x 128 tokens took 44 ms in the default mode and 17 ms in int8, against 84 ms in PyTorch eager and 60 / 24 ms in OpenVINO float32 / int8.

Prerequisites

Step 1: build the encoder

A tiny random model stands in for a real one here; with a real model you call AutoModel.from_pretrained(...):

PythonNeeds torch, transformers: runs on your machine.
import torch
from transformers import BertConfig, BertModel
from kwker.encode import KwkEncoder, runner_unsupported

model = BertModel(BertConfig(vocab_size=1000, hidden_size=128, num_hidden_layers=2, num_attention_heads=2,
                             intermediate_size=256)).eval()
# a real model: AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
print(runner_unsupported(model))
enc = KwkEncoder(model)
Output
None

Step 2: embed a batch

enc(ids, attention_mask=...) returns the last hidden state, like the model's own forward. Sentence embedders such as all-MiniLM-L6-v2 average it over the real tokens:

PythonRuns on your machine.
ids = torch.randint(0, 1000, (4, 32))
mask = torch.ones_like(ids)
mask[2, 20:] = 0                                            # a shorter sentence, padded
hidden = enc(ids, attention_mask=mask)                      # [4, 32, 128]
emb = (hidden * mask[..., None]).sum(1) / mask.sum(1, keepdim=True)
emb = torch.nn.functional.normalize(emb, dim=-1)
print(emb.shape)
Output
torch.Size([4, 128])

Step 3: check against the original model

PythonRuns on your machine.
with torch.no_grad():
    ref = model(ids, attention_mask=mask).last_hidden_state
cos = torch.nn.functional.cosine_similarity(hidden[mask.bool()], ref[mask.bool()], dim=-1)
print(bool(cos.min() > 0.999))
Output
True

Step 4: int8

int8=True with typical inputs as calib= (a tokenizer batch, a list of them, or token ids) runs int8 weights and activations. It changes the embeddings slightly, so compare it the same way on your own sentences:

PythonRuns on your machine.
enc8 = KwkEncoder(model, int8=True, calib={"input_ids": ids, "attention_mask": mask})   # or a list of batches

python -m kwker.bench --encoders --int8 compares KwkEncoder with OpenVINO and PyTorch on your machine. ModernBERT, NomicBERT and EuroBERT models run in bf16 or float32 only; int8 raises ValueError for them.

Next steps