Runtime controls
Kwker configures itself when it loads. It reads the CPU's features and cache sizes, picks the engine, and sets its thresholds from them. There is nothing to tune. This page covers the switches that exist anyway: for production fallbacks, debugging, resource limits and A/B timings. None of them changes a result; they only change speed and resource use.
What runs here
import kwker
caps = kwker.capabilities()
print(caps["isa"], caps["engines_usable"], caps["l2"], caps["cpus"])
r, report = kwker.observe(kwker.sorted, list(range(1000, 0, -1))) # which engine stages ran
print(report["isa"], [s["name"] for s in report["stages"]][:3])
capabilities() returns the version, the engine in use, the engines built and usable, the CPU features, the cache
sizes, the CPU count and the build. observe(fn, ...) runs one call with the engines' path report switched on. The
report lists the stages that ran (with counts and cycles) and the scratch peak. Stage reports exist on the x86 engines;
elsewhere traced is False.
From a shell, run python -m kwker doctor, or kwker-doctor (the C / C++ package's bin/). Add --json for
scripts. To collect everything for a support request,
run python -m kwker support-bundle. It writes a .tar.gz holding the doctor report, CPU details, package
versions, the KWKER_* and threading environment variables, and a self-check on generated data. It holds
none of your data.
Fallback switches
| Switch | Scope | Effect |
|---|---|---|
KWKER_ (environment) |
the process | every install() and backend="kwker" does nothing: the program runs on PyTorch, NumPy and transformers alone |
KWKER_ / sse42 / portable (environment) |
the process, from start-up | caps the engine |
kwker.set_, set_ |
the process, from now on | caps the engine at run time; None restores the best |
kwker.set_ |
the calling thread | permits only some algorithm classes (below) |
kwker.set_ |
the calling thread | no scratch allocations at all |
kwker.torch_ |
the process | torch's own kernels again |
KWKER_ |
next torch_ |
re-runs the torch override self-check |
KWKER_ |
the process | no AMX kernels in the CPU backend |
If you ever suspect Kwker in a production problem, KWKER_ISA=portable is the first check. It runs the scalar
engine, which produces the same results. If the problem stays, the cause is not in an engine's SIMD code.
The same cap from Rust, C, C++ and Go, with a check that the results do not change:
use kwker::Isa;
fn main() {
kwker::set_isa(Some(Isa::Portable)); // the scalar engine from now on
println!("{:?}", kwker::isa());
let mut a = [3, 1, 2];
kwker::sort(&mut a); // the same result on every engine
println!("{a:?}");
kwker::set_isa(None); // back to the best engine
}
Portable [1, 2, 3]
#include <stdio.h>
#include <kwker.h>
int main(void) {
kwker_set_isa("portable"); /* the scalar engine from now on */
printf("%s\n", kwker_isa());
int32_t a[] = {3, 1, 2};
kwker_i32_sort(a, 3); /* the same result on every engine */
printf("%d %d %d\n", a[0], a[1], a[2]);
kwker_set_isa(NULL); /* back to the best engine */
return 0;
}
portable 1 2 3
#include <iostream>
#include <vector>
#include <kwker.hpp>
int main() {
kwker::set_isa("portable"); // the scalar engine from now on
std::cout << kwker::isa() << '\n';
std::vector<int> a{3, 1, 2};
kwker::sort(a); // the same result on every engine
std::cout << a[0] << ' ' << a[1] << ' ' << a[2] << '\n';
kwker::set_isa(); // back to the best engine
}
portable 1 2 3
package main
import (
"fmt"
"kwker.io/go/kwker"
)
func main() {
isa, _ := kwker.SetISA("portable") // the scalar engine from now on
fmt.Println(isa)
a := []int32{3, 1, 2}
kwker.Sort(a) // the same result on every engine
fmt.Println(a)
kwker.SetISA("") // back to the best engine
}
portable [1 2 3]
import io.kwker.Kwker;
import java.util.Arrays;
public class Example {
public static void main(String[] args) {
System.out.println(Kwker.setIsa("portable")); // the scalar engine from now on
int[] a = {3, 1, 2};
Kwker.sort(a); // the same result on every engine
System.out.println(Arrays.toString(a));
Kwker.setIsa(null); // back to the best engine
}
}
portable [1, 2, 3]
using Kwker;
Console.WriteLine(Sorter.SetIsa("portable")); // the scalar engine from now on
var a = new[] { 3, 1, 2 };
Sorter.Sort(a); // the same result on every engine
Console.WriteLine(string.Join(" ", a));
Sorter.SetIsa(null); // back to the best engine
portable 1 2 3
Algorithm classes
Besides the comparison core, which always runs, Kwker uses three optional classes of algorithm:
adaptive: finishes early on inputs that are already sorted, reversed, all equal or nearly sorted, and sorts inputs made of runs of equal keys (timestamps, repeated readings) as one entry per run.counting: handles few distinct values, dominant or heavy values, and narrow key ranges.radix: radix and value-bucket passes for large inputs.
set_algorithms restricts them for the calling thread. A typical use is a latency-critical service that wants
predictable timings on adversarial inputs, or isolating a slowdown. Results, including stable orders, are identical
under every setting.
import numpy as np, kwker
prev = kwker.set_algorithms({"adaptive"}) # comparison core + presortedness checks only
a = np.random.default_rng(0).integers(0, 4, 100_000)
kwker.sort(a) # same result, without the counting path
kwker.set_algorithms(prev)
Threads
Calls are single-threaded unless you pass threads=. Kwker never starts threads behind your back.
| Call | Effect |
|---|---|
sort(a, threads=4) |
that call on four threads |
set_ |
the count used by threads=0 (default 1; 0 = every CPU the process may use) |
set_ |
the process-wide budget shared by all concurrent parallel calls (default: the CPUs available) |
set_ |
pins the worker threads (Linux) |
parallel_ |
the threads working now, and the most since the last reset |
A parallel call made from inside another one's task runs on its caller's thread, so nesting never multiplies thread counts. Worker threads are kept between calls and exit after one second idle. Under PyTorch, Kwker uses PyTorch's own worker threads.
Memory
| Call | Effect |
|---|---|
set_ |
caps the scratch this thread's in-place calls may allocate; 0 = allocation-free |
set_ |
how Kwker keeps its own buffers between calls (default: up to 1 GiB kept for reuse, no huge pages) |
scratch_ |
bytes in use, cached, and the peak |
release_ |
returns the cached buffers to the system |
Plan(...) |
keeps one caller's buffers between repeated calls |
import numpy as np, kwker
kwker.set_scratch_limit(0) # allocation-free from here on (this thread)
a = np.random.default_rng(1).random(200_000); kwker.sort(a)
kwker.set_scratch_limit(None)
print(kwker.scratch_use())
In Rust, scratch_bound() returns the most scratch memory an operation can use.
Environment variables
| Variable | Default | Meaning |
|---|---|---|
KWKER_ |
off | 1 prints what Kwker ran when the program ends: calls and time per PyTorch kernel and per NumPy, SciPy and data-frame call, generate() paths (kwker.report) |
KWKER_ |
off | 1 turns every integration off: install() calls and torch.compile(backend="kwker") change nothing; names (scipy,numpy; also torch, cpu_, decode, encode, compile) turn off only those |
KWKER_ |
best available | cap the engine: avx512, avx2, sse42, neon, portable |
KWKER_ |
off | huge-page advice for large scratch mappings |
KWKER_ |
off | Python: use the ctypes path instead of the compiled fast path |
KWKER_ |
none | the key python -m kwker.serve requires on every request (as --api-key) |
KWKER_ |
on | 0 skips Decoder / Encoder.from_'s check that the weight files fit in the memory left |
KWKER_ |
on | 0 skips the torch override self-check, force re-runs it |
KWKER_ |
best available | cap the torch extensions' own kernels: avx512 (or 2), avx2 (1), portable (0) |
KWKER_ |
on | 0 turns off the AMX kernels of the CPU backend |
KWKER_ |
~/.cache/kwker |
Kwker's one cache directory: self-checks, packed weights, compiled decoders, tuning choices (python -m kwker cache lists and clears it) |
KWKER_ |
32 |
the cache directory's size limit: after Kwker writes an entry, the least recently used entries are removed until it fits |
KWKER_ |
~/.cache/kwker/decode |
compiled Decoder packages ("" turns the cache off) |
python -m kwker doctor lists every KWKER_* variable that is set in the environment.
PyTorch switches
Environment variables that turn single PyTorch features off or change their limits:
| Variable | Effect |
|---|---|
KWKER_ |
eager bfloat16 linear layers on torch's kernels |
KWKER_ |
no bfloat16 copy of float32 weights in eager code |
KWKER_ |
Hugging Face's own norm, rotary and SiLU code in eager mode |
KWKER_ |
the memory limit of the bfloat16 weight copies (at most half of the machine's memory) |
KWKER_ |
compiled: Hugging Face's own sliding-window cache layer; a number sets the 4x limit |
KWKER_ |
print what torch.compile rewrote |
KWKER_, KWKER_ |
compiled: no bfloat16 copy, no bfloat16 linear layers |
KWKER_ |
compiled: no grouped linear layers |
KWKER_ |
compiled: keep Inductor's per-call shape checks |
KWKER_ or 0 |
compiled: fuse the glue ops or not, as options={"glue": ...} |
KWKER_, KWKER_ |
compiled: always Inductor, as options={"inner": "inductor"} |
KWKER_ |
eager inner: run the graph through Python, not the native call list |
KWKER_ |
Inductor's C++ wrapper, as options={"cpp_: about 12% faster steps, a first compile more than twice as long |
KWKER_ |
install(model) prints no line when a language model runs at a lower precision than its checkpoint |
KWKER_ or 8 |
the default mix_ of int4 language models |
KWKER_ |
the size limit of the packed-weight cache (16 GB) |
Related
- Core concepts: engines, threads and memory explained.
- Large data: several cores and a memory limit for one call.
- Troubleshooting: what each doctor warning means.