Kwker

Classify images

KwkCNN runs a whole torchvision classifier in one native call: no graph capture and no compile step, built in seconds. In float32 it ran 2.8x faster than PyTorch eager across 14 torchvision models (batch 1, 4 cores); in int8 it was 1.25-5.05x faster than OpenVINO int8 on every one of them.

Prerequisites

Supported: ResNet, MobileNetV2 and V3, EfficientNet, RegNet, ResNeXt, ConvNeXt, and models built from the same layers.

Step 1: check the model

runner_unsupported returns None when the model can run, or the reason it can't:

PythonNeeds torch, torchvision: runs on your machine.
import torch, torchvision
from kwker.cnn import KwkCNN, runner_unsupported

model = torchvision.models.resnet18(weights=None).eval()     # weights="DEFAULT" for the pretrained ones
print(runner_unsupported(model))
Output
None

Step 2: run it in float32

Build the runner once for your input size; any batch size works afterwards:

PythonRuns on your machine.
fast = KwkCNN(model, example_shape=(1, 3, 64, 64))
x = torch.randn(2, 3, 64, 64)
with torch.no_grad():
    print(float((fast(x) - model(x)).abs().max()) < 1e-3)
Output
True

In float32 the outputs match the model's to rounding.

Step 3: int8, calibrated on your photos

int8 is faster and changes the outputs slightly. Pass calibration images that look like your real inputs; about 16 photos are enough. Random tensors stand in for them here:

PythonRuns on your machine.
calib = torch.randn(8, 3, 64, 64)                            # in practice: 16 preprocessed photos
fast8 = KwkCNN(model, example_shape=(1, 3, 64, 64), int8=True, calib=calib)
with torch.no_grad():
    print(fast8(x).shape)
Output
torch.Size([2, 1000])

Step 4: check int8 on your own photos

Before you use int8, check how often it picks the same class as float32 on photos like yours:

ShellOn your machine.
python -m kwker.bench --models resnet50 --int8 --images photos/

The report shows the top-1 agreement with float32 for Kwker int8 and OpenVINO int8, both calibrated on the same photos, and the speed of each. python -m kwker.bench --models --threads 4 compares float32 KwkCNN with PyTorch eager, Inductor and OpenVINO.

Notes

Next steps