Kwker

Runtime controls

Kwker configures itself when it loads. It reads the CPU's features and cache sizes, picks the engine, and sets its thresholds from them. There is nothing to tune. This page covers the switches that exist anyway: for production fallbacks, debugging, resource limits and A/B timings. None of them changes a result; they only change speed and resource use.

What runs here

PythonRuns on your machine.
import kwker
caps = kwker.capabilities()
print(caps["isa"], caps["engines_usable"], caps["l2"], caps["cpus"])
r, report = kwker.observe(kwker.sorted, list(range(1000, 0, -1)))   # which engine stages ran
print(report["isa"], [s["name"] for s in report["stages"]][:3])

capabilities() returns the version, the engine in use, the engines built and usable, the CPU features, the cache sizes, the CPU count and the build. observe(fn, ...) runs one call with the engines' path report switched on. The report lists the stages that ran (with counts and cycles) and the scratch peak. Stage reports exist on the x86 engines; elsewhere traced is False.

From a shell, run python -m kwker doctor, or kwker-doctor (the C / C++ package's bin/). Add --json for scripts. To collect everything for a support request, run python -m kwker support-bundle. It writes a .tar.gz holding the doctor report, CPU details, package versions, the KWKER_* and threading environment variables, and a self-check on generated data. It holds none of your data.

Fallback switches

Switch Scope Effect
KWKER_DISABLE=1 (environment) the process every install() and backend="kwker" does nothing: the program runs on PyTorch, NumPy and transformers alone
KWKER_ISA=avx2 / sse42 / portable (environment) the process, from start-up caps the engine
kwker.set_isa("portable"), set_isa(None) the process, from now on caps the engine at run time; None restores the best
kwker.set_algorithms({"adaptive"}) the calling thread permits only some algorithm classes (below)
kwker.set_scratch_limit(0) the calling thread no scratch allocations at all
kwker.torch_ops.uninstall() the process torch's own kernels again
KWKER_SELF_CHECK=force next torch_ops.install() re-runs the torch override self-check
KWKER_AMX=0 the process no AMX kernels in the CPU backend

If you ever suspect Kwker in a production problem, KWKER_ISA=portable is the first check. It runs the scalar engine, which produces the same results. If the problem stays, the cause is not in an engine's SIMD code.

The same cap from Rust, C, C++ and Go, with a check that the results do not change:

RustAdd kwker to Cargo.toml, then cargo run.
use kwker::Isa;

fn main() {
    kwker::set_isa(Some(Isa::Portable));   // the scalar engine from now on
    println!("{:?}", kwker::isa());
    let mut a = [3, 1, 2];
    kwker::sort(&mut a);                     // the same result on every engine
    println!("{a:?}");
    kwker::set_isa(None);                    // back to the best engine
}
Output
Portable
[1, 2, 3]

Algorithm classes

Besides the comparison core, which always runs, Kwker uses three optional classes of algorithm:

set_algorithms restricts them for the calling thread. A typical use is a latency-critical service that wants predictable timings on adversarial inputs, or isolating a slowdown. Results, including stable orders, are identical under every setting.

PythonRuns on your machine.
import numpy as np, kwker
prev = kwker.set_algorithms({"adaptive"})       # comparison core + presortedness checks only
a = np.random.default_rng(0).integers(0, 4, 100_000)
kwker.sort(a)                                    # same result, without the counting path
kwker.set_algorithms(prev)

Threads

Calls are single-threaded unless you pass threads=. Kwker never starts threads behind your back.

Call Effect
sort(a, threads=4) that call on four threads
set_default_threads(n) the count used by threads=0 (default 1; 0 = every CPU the process may use)
set_max_threads(n) the process-wide budget shared by all concurrent parallel calls (default: the CPUs available)
set_worker_cpus([0, 2, 4, 6]) pins the worker threads (Linux)
parallel_threads() the threads working now, and the most since the last reset

A parallel call made from inside another one's task runs on its caller's thread, so nesting never multiplies thread counts. Worker threads are kept between calls and exit after one second idle. Under PyTorch, Kwker uses PyTorch's own worker threads.

Memory

Call Effect
set_scratch_limit(n) caps the scratch this thread's in-place calls may allocate; 0 = allocation-free
set_scratch_policy(cache_bytes=, mapped=, huge_pages=) how Kwker keeps its own buffers between calls (default: up to 1 GiB kept for reuse, no huge pages)
scratch_use() bytes in use, cached, and the peak
release_scratch() returns the cached buffers to the system
Plan(...) keeps one caller's buffers between repeated calls
PythonRuns on your machine.
import numpy as np, kwker
kwker.set_scratch_limit(0)                       # allocation-free from here on (this thread)
a = np.random.default_rng(1).random(200_000); kwker.sort(a)
kwker.set_scratch_limit(None)
print(kwker.scratch_use())

In Rust, scratch_bound() returns the most scratch memory an operation can use.

Environment variables

Variable Default Meaning
KWKER_REPORT off 1 prints what Kwker ran when the program ends: calls and time per PyTorch kernel and per NumPy, SciPy and data-frame call, generate() paths (kwker.report)
KWKER_DISABLE off 1 turns every integration off: install() calls and torch.compile(backend="kwker") change nothing; names (scipy,numpy; also torch, cpu_backend, decode, encode, compile) turn off only those
KWKER_ISA best available cap the engine: avx512, avx2, sse42, neon, portable
KWKER_HUGEPAGE off huge-page advice for large scratch mappings
KWKER_NO_FAST off Python: use the ctypes path instead of the compiled fast path
KWKER_API_KEY none the key python -m kwker.serve requires on every request (as --api-key)
KWKER_MEMORY_CHECK on 0 skips Decoder / Encoder.from_pretrained's check that the weight files fit in the memory left
KWKER_SELF_CHECK on 0 skips the torch override self-check, force re-runs it
KWKER_TORCH_ISA best available cap the torch extensions' own kernels: avx512 (or 2), avx2 (1), portable (0)
KWKER_AMX on 0 turns off the AMX kernels of the CPU backend
KWKER_CACHE_DIR ~/.cache/kwker Kwker's one cache directory: self-checks, packed weights, compiled decoders, tuning choices (python -m kwker cache lists and clears it)
KWKER_CACHE_GB 32 the cache directory's size limit: after Kwker writes an entry, the least recently used entries are removed until it fits
KWKER_DECODE_CACHE ~/.cache/kwker/decode compiled Decoder packages ("" turns the cache off)

python -m kwker doctor lists every KWKER_* variable that is set in the environment.

PyTorch switches

Environment variables that turn single PyTorch features off or change their limits:

Variable Effect
KWKER_EAGER_B16=0 eager bfloat16 linear layers on torch's kernels
KWKER_EAGER_F32B=0 no bfloat16 copy of float32 weights in eager code
KWKER_EAGER_GLUE=0 Hugging Face's own norm, rotary and SiLU code in eager mode
KWKER_B16_CACHE_MB the memory limit of the bfloat16 weight copies (at most half of the machine's memory)
KWKER_HF_SLIDING=0 compiled: Hugging Face's own sliding-window cache layer; a number sets the 4x limit
KWKER_COMPILE_DEBUG=1 print what torch.compile rewrote
KWKER_COMPILE_F32B=0, KWKER_COMPILE_B16=0 compiled: no bfloat16 copy, no bfloat16 linear layers
KWKER_COMPILE_GROUP=0 compiled: no grouped linear layers
KWKER_COMPILE_ASSERTS=1 compiled: keep Inductor's per-call shape checks
KWKER_COMPILE_GLUE=1 or 0 compiled: fuse the glue ops or not, as options={"glue": ...}
KWKER_COMPILE_INNER=inductor, KWKER_COMPILE_AUTO=0 compiled: always Inductor, as options={"inner": "inductor"}
KWKER_COMPILE_GEXEC=0 eager inner: run the graph through Python, not the native call list
KWKER_COMPILE_CPP=1 Inductor's C++ wrapper, as options={"cpp_wrapper": True}: about 12% faster steps, a first compile more than twice as long
KWKER_LOAD_REPORT=0 install(model) prints no line when a language model runs at a lower precision than its checkpoint
KWKER_MIX_BITS=6 or 8 the default mix_bits of int4 language models
KWKER_PACK_CACHE_GB the size limit of the packed-weight cache (16 GB)