Kwker
Image models

Classify and tag images without a GPU

Content moderation, product tagging, document routing and quality inspection run on convolutional networks. KwkCNN runs the whole network as one native call, in float32 or in int8 calibrated on your own images.

Against OpenVINO int8 on the same cores

Milliseconds per image, one image at a time, 4 cores, lower is better. Both sides quantize to int8 with the same calibration images; OpenVINO uses its AMX-INT8 kernels.

ModelOpenVINO int8KwkCNN int8Speed-up
ResNet-181.51.01.39×
ResNet-503.93.21.22×
ResNeXt-50 32×4d4.53.71.21×
RegNetX-1.6GF3.72.51.52×
DenseNet-1215.82.52.35×
Inception v33.72.91.24×
GoogLeNet3.01.71.79×
MobileNetV21.70.82.10×

Intel Xeon with AMX, 224 × 224 inputs, October 2026. Method · Run it yourself

How it works

  • Standard architectures: ResNet, ResNeXt, RegNet, MobileNetV2 and V3, EfficientNet, ConvNeXt and models built from the same parts, straight from torchvision.
  • Calibrated int8: weight and activation scales fitted on a few dozen of your images; AMX-INT8 tiles where the processor has them, AVX-512 VNNI elsewhere.
  • Exact float32: without int8, outputs match PyTorch to rounding (within 3×10-6).
  • Any batch size: build once for an input size, then call with one image or a batch.
PythonRuns on your machine.
import torch, torchvision
from kwker.cnn import KwkCNN

model = torchvision.models.resnet50(weights="DEFAULT").eval()
calib = torch.randn(32, 3, 224, 224)   # 32 of your own images, preprocessed
fast = KwkCNN(model, example_shape=(1, 3, 224, 224),
              int8=True, calib=calib)

images = torch.randn(4, 3, 224, 224)
with torch.no_grad():
    logits = fast(images)              # one image or a batch
print(logits.shape)
Output
torch.Size([4, 1000])
The KwkCNN guide, with runnable examples

Quality and requirements

  • int8 accuracy stays close to float32: top-1 on 200 Imagenette images moved by at most 2 points (ResNet-18 67.5 → 68.0%, MobileNetV2 62.5 → 64.5%, EfficientNet-B0 79.0 → 77.5%). Check top-1 on your own data after calibrating.
  • Processor: AVX-512 for float32; AVX-512 VNNI for int8 (Intel Xeon Cascade Lake and later, AMD Zen 4 and later), with AMX tiles on Sapphire Rapids and later. Linux x86-64 with PyTorch.
  • Other networks: runner_unsupported(model) says why a model cannot run; torch.compile(model, backend="kwker") covers the rest.

Guide, limitations.

Free to start

Free for companies under 100 people and about $1.3M revenue (PolyForm Small Business). Every product line is in every tier.