Quickstart: image models
Run an image classifier faster on the CPU. KwkCNN runs the whole model in one native call, in float32 or in int8.
About five minutes.
It runs ResNet, MobileNetV2 and V3, EfficientNet, RegNet, ResNeXt, ConvNeXt and models built from the same layers, on AVX-512 CPUs. int8 is fastest on CPUs with VNNI or AMX.
Install
pip install kwker torchvision
python -m kwker doctor # your CPU, the engine Kwker picked, any warnings
Run a model
Build the runner once for your input size (any batch size works), then call it like the model:
import torch, torchvision
from kwker.cnn import KwkCNN
model = torchvision.models.resnet18(weights=None).eval() # weights="DEFAULT" for the pretrained ones
fast = KwkCNN(model, example_shape=(1, 3, 224, 224))
x = torch.randn(2, 3, 224, 224)
with torch.no_grad():
print(float((fast(x) - model(x)).abs().max()) < 1e-3)
True
In float32 the outputs match the model's to rounding.
Use int8
int8 is faster and changes the outputs slightly. Give it calibration images that look like your real inputs; about 16 photos are enough:
calib = torch.stack([preprocess(img) for img in my_16_photos]) # shape [16, 3, 224, 224]
fast8 = KwkCNN(model, example_shape=(1, 3, 224, 224), int8=True, calib=calib)
Check the results on your photos
Before you use int8, check how often it picks the same class as float32 on your own photos:
python -m kwker.bench --models resnet50 --int8 --images photos/
The report shows top-1 agreement with float32 for Kwker int8 and OpenVINO int8, both calibrated on the same photos, and the speed of each.
Measure it on your machine
python -m kwker.bench --models --threads 4 # KwkCNN vs PyTorch eager, Inductor and OpenVINO
Next steps
- Classify images: the supported models, int8 and the comparison with OpenVINO.
- Evaluate on your machine: what each benchmark measures.