Image models
Classify and tag images without a GPU
Content moderation, product tagging, document routing and quality inspection run on convolutional networks. KwkCNN runs the whole network as one native call, in float32 or in int8 calibrated on your own images.
Against OpenVINO int8 on the same cores
Milliseconds per image, one image at a time, 4 cores, lower is better. Both sides quantize to int8 with the same calibration images; OpenVINO uses its AMX-INT8 kernels.
| Model | OpenVINO int8 | KwkCNN int8 | Speed-up |
|---|---|---|---|
| ResNet-18 | 1.5 | 1.0 | 1.39× |
| ResNet-50 | 3.9 | 3.2 | 1.22× |
| ResNeXt-50 32×4d | 4.5 | 3.7 | 1.21× |
| RegNetX-1.6GF | 3.7 | 2.5 | 1.52× |
| DenseNet-121 | 5.8 | 2.5 | 2.35× |
| Inception v3 | 3.7 | 2.9 | 1.24× |
| GoogLeNet | 3.0 | 1.7 | 1.79× |
| MobileNetV2 | 1.7 | 0.8 | 2.10× |
Intel Xeon with AMX, 224 × 224 inputs, October 2026. Method · Run it yourself
How it works
- Standard architectures: ResNet, ResNeXt, RegNet, MobileNetV2 and V3, EfficientNet, ConvNeXt and models built from the same parts, straight from torchvision.
- Calibrated int8: weight and activation scales fitted on a few dozen of your images; AMX-INT8 tiles where the processor has them, AVX-512 VNNI elsewhere.
- Exact float32: without int8, outputs match PyTorch to rounding (within 3×10-6).
- Any batch size: build once for an input size, then call with one image or a batch.
import torch, torchvision
from kwker.cnn import KwkCNN
model = torchvision.models.resnet50(weights="DEFAULT").eval()
calib = torch.randn(32, 3, 224, 224) # 32 of your own images, preprocessed
fast = KwkCNN(model, example_shape=(1, 3, 224, 224),
int8=True, calib=calib)
images = torch.randn(4, 3, 224, 224)
with torch.no_grad():
logits = fast(images) # one image or a batch
print(logits.shape)Output
torch.Size([4, 1000])
Quality and requirements
- int8 accuracy stays close to float32: top-1 on 200 Imagenette images moved by at most 2 points (ResNet-18 67.5 → 68.0%, MobileNetV2 62.5 → 64.5%, EfficientNet-B0 79.0 → 77.5%). Check top-1 on your own data after calibrating.
- Processor: AVX-512 for float32; AVX-512 VNNI for int8 (Intel Xeon Cascade Lake and later, AMD Zen 4 and later), with AMX tiles on Sapphire Rapids and later. Linux x86-64 with PyTorch.
- Other networks:
runner_unsupported(model)says why a model cannot run;torch.compile(model, backend="kwker")covers the rest.
Free to start
Free for companies under 100 people and about $1.3M revenue (PolyForm Small Business). Every product line is in every tier.