Kwker benchmark results: image models (KwkCNN vs OpenVINO, PyTorch eager and Inductor) ======================================================================================= Published on https://kwker.io/benchmarks/#cnn. Raw numbers as recorded; milliseconds per image, lower is better. Machine Intel Xeon with AMX (bf16 / int8), Emerald Rapids class cloud VM, 4 cores used by every column Date 2026-10-02 Models torchvision classifiers, pretrained weights, 224 x 224 inputs, channels-last, batch 1 unless noted OpenVINO 2026.4; float32 pinned to f32 (INFERENCE_PRECISION_HINT); int8 = NNCF post-training quantization, oneDNN AMX-INT8 kernels PyTorch eager, Inductor (torch.compile), Inductor with freezing Kwker KwkCNN (kwker.cnn): the whole network as one native call; int8 = calibrated u8 activations / s8 weights Protocol every column timed round-robin in one process: a 10 ms pause, one warm call, then the best of 3 back-to-back calls; 10 rounds, best round per column. int8 on both sides is calibrated on the same Imagenette training images; the input is a real Imagenette image. Headline table: int8 on AMX-INT8 tiles for both (first AMX round of KwkCNN) OpenVINO int8 KwkCNN int8 (VNNI, earlier run) KwkCNN int8 (AMX) KwkCNN vs OpenVINO int8 resnet18 1.5 1.8 1.0 1.39x resnet50 3.9 4.0 3.2 1.22x resnext50_32x4d 4.5 5.3 3.7 1.21x regnet_x_1_6gf 3.7 2.6 2.5 1.52x densenet121 5.8 3.6 2.5 2.35x inception_v3 3.7 3.3 2.9 1.24x googlenet 3.0 1.7 1.7 1.79x mobilenet_v2 1.7 1.1 0.8 2.10x (OpenVINO's own times moved between runs on this VM, e.g. regnet_x 4.4 -> 3.7 ms: compare within a row.) Second AMX round, same protocol, including batch 8 (ms per batch) batch OpenVINO int8 KwkCNN int8 (VNNI) KwkCNN int8 (AMX) vs OpenVINO int8 resnet18 1 1.5 1.69 0.9 1.60x resnet50 1 3.6 4.10 2.9 1.25x resnext50_32x4d 1 4.9 4.85 3.9 1.28x resnet50 8 16.8 33.9 16.5 1.02x resnext50_32x4d 8 24.3 35.3 23.8 1.02x All 14 models, every engine (earlier run the same day: KwkCNN int8 still on VNNI kernels, no AMX) eager Inductor Inductor+freezing OpenVINO fp32 OpenVINO int8 KwkCNN fp32 KwkCNN int8 int8 vs OpenVINO int8 mobilenet_v2 9.1 7.9 7.5 2.8 1.7 2.9 1.1 1.65x mobilenet_v3_large 12.8 9.3 9.0 3.5 2.9 2.7 1.2 2.43x efficientnet_b0 15.1 10.3 9.6 4.6 4.4 3.5 1.5 2.86x regnet_y_800mf 14.2 11.0 10.1 5.9 2.8 5.9 1.7 1.65x resnet18 11.6 9.9 9.7 7.1 1.5 6.2 1.8 0.82x resnet50 31.3 27.4 27.4 17.8 4.8 16.4 4.0 1.18x convnext_tiny 35.0 35.2 31.4 30.8 15.7 * 23.4 6.3 2.51x * densenet121 37.6 30.2 30.6 17.6 6.9 12.7 3.6 1.94x googlenet 17.1 11.6 11.4 7.6 3.0 6.7 1.7 1.74x squeezenet1_1 4.8 3.6 4.1 2.7 1.4 2.4 0.6 2.21x shufflenet_v2_x1_0 8.6 6.0 5.9 3.0 2.4 1.3 0.5 5.05x resnext50_32x4d 36.7 33.5 32.0 23.9 4.8 21.4 5.3 0.92x regnet_x_1_6gf 20.7 16.3 16.7 9.5 4.4 11.2 2.6 1.70x inception_v3 31.6 27.5 28.3 16.6 4.5 13.7 3.3 1.36x * OpenVINO int8's convnext_tiny output was wrong in this run (error 1.0 of the largest logit); every other int8 column agrees with float32 on top-1. Ratios are computed from unrounded times. The two losses in this table (resnet18, resnext50) came from OpenVINO's AMX-INT8 kernels; KwkCNN's own AMX-INT8 kernels (tables above) turned both into wins. Accuracy (int8 top-1 on 200 Imagenette validation images, 1000-way, pretrained; float32 -> KwkCNN int8): mobilenet_v2 62.5 -> 66.0%, mobilenet_v3_large 72.0 -> 73.5%, efficientnet_b0 79.0 -> 78.0%, resnet18 67.5 -> 69.0%, densenet121 69.5 -> 73.5%, googlenet 64.0 -> 64.0%, inception_v3 74.0 -> 77.5%, squeezenet1_1 48.5 -> 49.0%, resnext50 85.5 -> 88.5%, regnet_x_1_6gf 75.0 -> 83.5%. With 200 images the sampling error is about +-3 points. KwkCNN float32 equals eager PyTorch (largest difference 2.6e-6).