Classify images
KwkCNN runs a whole torchvision classifier in one native call: no graph capture and no compile step, built in
seconds. In float32 it ran 2.8x faster than PyTorch eager across 14 torchvision models (batch 1, 4 cores); in int8 it
was 1.25-5.05x faster than OpenVINO int8 on every one of them.
Prerequisites
- An x86-64 CPU with AVX-512 for float32; int8 is fastest with VNNI or AMX.
pip install kwker torchvision(Installation).
Supported: ResNet, MobileNetV2 and V3, EfficientNet, RegNet, ResNeXt, ConvNeXt, and models built from the same layers.
Step 1: check the model
runner_unsupported returns None when the model can run, or the reason it can't:
import torch, torchvision
from kwker.cnn import KwkCNN, runner_unsupported
model = torchvision.models.resnet18(weights=None).eval() # weights="DEFAULT" for the pretrained ones
print(runner_unsupported(model))
None
Step 2: run it in float32
Build the runner once for your input size; any batch size works afterwards:
fast = KwkCNN(model, example_shape=(1, 3, 64, 64))
x = torch.randn(2, 3, 64, 64)
with torch.no_grad():
print(float((fast(x) - model(x)).abs().max()) < 1e-3)
True
In float32 the outputs match the model's to rounding.
Step 3: int8, calibrated on your photos
int8 is faster and changes the outputs slightly. Pass calibration images that look like your real inputs; about 16 photos are enough. Random tensors stand in for them here:
calib = torch.randn(8, 3, 64, 64) # in practice: 16 preprocessed photos
fast8 = KwkCNN(model, example_shape=(1, 3, 64, 64), int8=True, calib=calib)
with torch.no_grad():
print(fast8(x).shape)
torch.Size([2, 1000])
Step 4: check int8 on your own photos
Before you use int8, check how often it picks the same class as float32 on photos like yours:
python -m kwker.bench --models resnet50 --int8 --images photos/
The report shows the top-1 agreement with float32 for Kwker int8 and OpenVINO int8, both calibrated on the same
photos, and the speed of each. python -m kwker.bench --models --threads 4 compares float32 KwkCNN with PyTorch eager,
Inductor and OpenVINO.
Notes
- The image preprocessing in torchvision transforms (color adjustments,
ToDtype+Normalize, PIL resizing) also runs on Kwker afterkwker.torch_ops.install(), bit for bit the same. - A model KwkCNN does not run still gets faster with torch.compile.
Next steps
- Lower precision with an accuracy check: int8 for other models.
- Evaluate on your machine: what each benchmark measures.