Deploy and operate
Kwker is a library, not a service: install the package and import it. At run time it opens no network connections, sends no telemetry and checks no license. Each process picks the engine for the CPU it runs on, so one build serves every machine.
Containers
Any image with glibc 2.34 or newer works with the Linux wheels: the official python images (Debian 12), Ubuntu
22.04 and later, and RHEL 9 and later.
FROM python:3.12-slim
RUN pip install --no-cache-dir kwker numpy
# the engine depends on the CPU the container runs on, so check it when the container starts
CMD ["sh", "-c", "python -m kwker doctor && exec python app.py"]
- Alpine and other musl-based images have no prebuilt wheel. Use a glibc image.
- For the PyTorch extensions, pin torch and the matching wheel together (
kwker==0.1.0+torch2.14withtorch>=2.14,<2.15).doctorwarns when they do not match. threads=0and the default thread count follow the CPUs the container may use: its CPU affinity and its cgroup CPU quota. A container limited to 2 CPUs gets 2 threads, not one per CPU of the host.
Machines with different CPUs
Each process picks its engine when Kwker loads: AVX-512, AVX2, NEON, SVE or portable. Every engine returns the same results (rule 11); only the speed differs. A fleet of mixed machines needs one build and no settings.
To make every machine run the same engine - for example to compare timings across a fleet - cap it:
KWKER_ISA=avx2 python app.py # at most AVX2, even on AVX-512 machines
Offline installs
Download the wheels on a machine with network access, copy them over, and install without an index:
pip download kwker numpy -d wheels/ # on a connected machine (same platform and Python)
pip install --no-index --find-links wheels/ kwker numpy # on the offline machine
Kwker itself never connects to the network. The one exception is python -m kwker.bench --native, which downloads
the sources of the libraries it compares against; the command-line reference lists
its options.
Memory limits
Kwker keeps the working memory its calls allocate for reuse, up to 1 GiB by default. Under a tight container memory limit, keep less, or cap what one thread may allocate:
import kwker
kwker.set_scratch_policy(cache_bytes=64 << 20) # keep at most 64 MiB of buffers between calls
kwker.set_scratch_limit(256 << 20) # this thread's calls allocate at most 256 MiB
print(sorted(kwker.scratch_use()))
kwker.set_scratch_limit(None)
['cached', 'in_use', 'peak']
kwker.release_scratch() returns the kept buffers to the system, for example before a process goes idle. The
Runtime controls page lists every memory setting.
Files Kwker writes
| Path | What | Change it with |
|---|---|---|
~/.cache/kwker |
everything below, plus the PyTorch self-check and torch.compile tuning choices |
KWKER_ |
~/.cache/kwker/decode |
compiled Decoder packages |
KWKER_ ("" turns it off) |
~/.cache/kwker/packs |
KwkDecoder(cache_ packed weights |
KWKER_ |
~/.cache/kwker/bench |
kwker.bench --native sources and builds |
KWKER_ |
The core library writes nothing. On a read-only file system, point these variables at a writable directory.
python -m kwker cache shows what is there; python -m kwker cache --clear deletes it (it is rebuilt when needed).
Monitoring and bug reports
python -m kwker doctor --jsonprints the machine report as JSON, for a startup log or a health check.python -m kwker support-bundlewrites a local archive to attach to a bug report. Nothing is uploaded.- Troubleshooting explains each warning
doctorprints.
Notes
- Upgrades follow the compatibility policy.
- Security reports: kwker.io/security.
Related
- Installation: requirements, and checking a download's checksums, signature and SBOM.
- Runtime controls: engines, threads, memory and every environment variable.
- Command-line reference:
doctor,support-bundleand the other commands.