Kwker benchmark results: text embeddings (KwkEncoder vs OpenVINO and torch.compile) ==================================================================================== Published on https://kwker.io/benchmarks/#enc. Raw numbers as recorded. Machine Intel Xeon with AMX (bf16 / int8), Emerald Rapids class cloud VM, 4 cores used by every column Date October 2026 OpenVINO 2026.4, bf16 inference precision hint (int8: NNCF post-training quantization, transformer preset, 8 calibration batches of the same shape) PyTorch torch.compile(backend="kwker") in its "medium" precision mode (bf16 matrix products, float32 elsewhere) Kwker KwkEncoder (kwker.encode): bf16 operands on AMX tiles, float32 accumulation, LayerNorm / softmax / GELU in float32; padded batches skip the padding (unpad=True, the default) Protocol interleaved rounds of every column, best round per column, milliseconds per batch. Batched rows are sentences of L/2..L tokens padded to L. Output error against float32 eager PyTorch, relative to the largest output: OpenVINO bf16 5e-3, KwkEncoder 5e-3 (MiniLM: 1.7e-3 both). Random weights, bf16 OpenVINO bf16 torch.compile(backend="kwker") bf16 KwkEncoder bf16 vs OpenVINO bf16 BERT-base 1 x 128 17.6 ms 18.6 ms 10.7 ms 1.64x BERT-base 8 x 128 114 ms 102 ms 65.8 ms 1.74x BERT-base 1 x 512 93.1 ms 85.2 ms 44.4 ms 2.10x MiniLM-L6 1 x 128 3.52 ms 4.74 ms 1.88 ms 1.88x MiniLM-L6 32 x 64 33.3 ms 37.7 ms 26.8 ms 1.24x Without skipping the padding (unpad=False), same run: BERT-base 8 x 128 79.7 ms (61.1 ms with), MiniLM 32 x 64 33.4 ms (23.4 ms with). Real checkpoints (bert-base-uncased, all-MiniLM-L6-v2), bf16 and int8; error in parentheses OpenVINO bf16 OpenVINO int8 KwkEncoder bf16 KwkEncoder int8 int8 vs OpenVINO int8 BERT-base 1 x 128 18.3 ms (1.1e-2) 12.1 ms (7.9e-2) 10.9 ms (8.7e-3) 9.25 ms (1.4e-1) 1.31x BERT-base 8 x 128 96.3 ms 72.7 ms (1.3e-1) 57.6 ms 42.2 ms (2.7e-1) 1.72x MiniLM-L6 1 x 128 3.73 ms 3.20 ms (5.0e-2) 2.03 ms 1.49 ms (3.0e-2) 2.15x MiniLM-L6 32 x 64 28.0 ms 30.9 ms (1.1e-1) 20.7 ms 15.9 ms (5.3e-2) 1.94x KwkEncoder bf16 alone is faster than OpenVINO int8 in every row. Uncalibrated int8 has a larger error than OpenVINO's on BERT-base (outlier channels in its LayerNorm and GELU outputs); check accuracy on your own data before using int8. This VM's AMX throughput dips for a few milliseconds at random times (another tenant on the core): single timings vary, best-of rounds are stable.