Kwker

Quickstart: image models

Run an image classifier faster on the CPU. KwkCNN runs the whole model in one native call, in float32 or in int8. About five minutes.

It runs ResNet, MobileNetV2 and V3, EfficientNet, RegNet, ResNeXt, ConvNeXt and models built from the same layers, on AVX-512 CPUs. int8 is fastest on CPUs with VNNI or AMX.

Install

ShellOn your machine.
pip install kwker torchvision
python -m kwker doctor            # your CPU, the engine Kwker picked, any warnings

Run a model

Build the runner once for your input size (any batch size works), then call it like the model:

PythonNeeds torch, torchvision: runs on your machine.
import torch, torchvision
from kwker.cnn import KwkCNN

model = torchvision.models.resnet18(weights=None).eval()     # weights="DEFAULT" for the pretrained ones
fast = KwkCNN(model, example_shape=(1, 3, 224, 224))
x = torch.randn(2, 3, 224, 224)
with torch.no_grad():
    print(float((fast(x) - model(x)).abs().max()) < 1e-3)
Output
True

In float32 the outputs match the model's to rounding.

Use int8

int8 is faster and changes the outputs slightly. Give it calibration images that look like your real inputs; about 16 photos are enough:

PythonRuns on your machine.
calib = torch.stack([preprocess(img) for img in my_16_photos])     # shape [16, 3, 224, 224]
fast8 = KwkCNN(model, example_shape=(1, 3, 224, 224), int8=True, calib=calib)

Check the results on your photos

Before you use int8, check how often it picks the same class as float32 on your own photos:

ShellOn your machine.
python -m kwker.bench --models resnet50 --int8 --images photos/

The report shows top-1 agreement with float32 for Kwker int8 and OpenVINO int8, both calibrated on the same photos, and the speed of each.

Measure it on your machine

ShellOn your machine.
python -m kwker.bench --models --threads 4       # KwkCNN vs PyTorch eager, Inductor and OpenVINO

Next steps