Kwker

Troubleshooting

Start with python -m kwker doctor, or bin/kwker-doctor from the C / C++ package. It prints the engine in use, the CPU, the installed frameworks, any set KWKER_* variables, a self-check result and, under "warnings", every known reason Kwker may be running below its best. The native kwker-doctor also checks that sorting gives correct results on every engine your CPU runs (--json for scripts; exit status 1 on a wrong result). The entries below follow those warnings, then the questions that come up most often.

Doctor warnings

"the engine is capped at avx2 / portable ... although this CPU runs avx512" Something chose a lower engine: KWKER_ISA in the environment, or a kwker.set_isa(...) call. Unset the variable, or call set_isa(None).

"this x86-64 CPU has no AVX2" Kwker is using its SSE4.2 engine, or its portable engine on CPUs without SSE4.2. This happens on pre-2013 CPUs, on low-power cores without AVX2, and on virtual machines that hide CPU features from the guest (some hypervisor CPU models, emulators, -cpu qemu64). Results are correct but slower. On a VM, ask for a CPU model that passes AVX2 / AVX-512 through ("host-passthrough").

"the compiled PyTorch operators were built for torch X, not Y" The PyTorch extensions are tied to one torch minor version. With another torch release, Kwker keeps its Python-registered operators, and the drop-in kernels and the CPU backend are off. Install the Kwker wheel built for your torch (its version ends in +torch2.M), or the torch version the wheel names.

"these PyTorch overrides differ from this torch's own kernels and stay off" The install-time self-check found a kernel whose output did not match this torch's bit for bit. That kernel stays off and torch's own runs instead, so nothing is wrong in your results. Please report it with a support bundle (below). KWKER_SELF_CHECK=force re-runs the check after an upgrade.

"the CPU has AMX but the CPU backend's AMX kernels are off" Linux grants AMX state per process, so a kernel older than 5.16 or a sandbox that blocks arch_prctl keeps AMX off. Other causes are KWKER_AMX=0 or mismatched torch extensions. The int8 / int4 decode kernels still run on AVX-512.

"this process may run on N of M CPUs" A CPU affinity mask or a container CPU limit is in force. Multithreaded calls then use the CPUs that are available, which is usually what you want. Check taskset -p $$ or your container's CPU settings if it is not.

"the compiled per-call fast path (_fast) is not loaded" Calls go through ctypes instead, which adds 5-10 us each. This only matters for many small calls. The fast path needs CPython 3.11 or newer, or 3.13t (the free-threaded build).

"free-threaded Python with the GIL enabled" A module without free-threading support was imported, or PYTHON_GIL=1 is set. Kwker itself supports running without the GIL.

Common questions

Kwker is slower than NumPy on my tiny arrays. Below about 100 elements, Python call overhead dominates the sort itself. Check that the fast path is loaded (doctor), and batch rows into one call with sort(m, axis=1) or sort_segments instead of looping in Python.

argsort of a 2-D array returns a flat permutation. axis=None (the default) sorts the flattened array. Pass axis=-1 for numpy.argsort's per-row behaviour.

My argsort indices differ from NumPy's. NumPy's default argsort (kind="quicksort") is not stable, so equal keys come back in any order. Kwker's argsort is always stable. Compare against np.argsort(a, kind="stable").

NaNs come out in a different place than in pandas / Polars. Kwker's NumPy functions put NaNs last unless nans_first=True, in both directions. kwker.frame follows each dataframe library's own rules instead; see DataFrames.

The first torch.compile(..., backend="kwker") call is slow. For inference graphs the backend times several candidates once (Inductor, eager, the rewritten graph) and caches the choice on disk, so later runs reuse it. Decoder packages are compiled once per architecture and also cached.

A mode changed my model's outputs. "high" is about float32. "medium", int8 and int4 change results. Measure on your model before relying on a mode: kwker.cpu_backend.check_accuracy(model, *inputs, mode=...) (see PyTorch).

Memory stays high after a large sort. Kwker keeps up to 1 GiB of released scratch buffers so repeated large calls do not page-fault. Call kwker.release_scratch(), or set_scratch_policy(cache_bytes=0) to keep nothing.

Many threads are busy although I passed threads=4. Kwker's own threads stay within the budget you set (set_max_threads). If PyTorch, NumPy's BLAS or OpenMP code in the same process also run thread pools, cap those too, for example with OMP_NUM_THREADS or torch.set_num_threads. Doctor lists the threading variables that are set.

Ruling Kwker in or out

  1. Run the failing program with KWKER_DISABLE=1: every install() call and torch.compile(backend="kwker") then does nothing, without code changes. If the problem stays, it is not in Kwker's integrations.
  2. Run it with KWKER_ISA=portable. Every engine gives identical results, so if the problem stays, it is not in SIMD code.
  3. For PyTorch, compare with kwker.torch_ops.uninstall(), or run python -m kwker audit -- your_script.py. Audit runs the script with Kwker and with KWKER_DISABLE=1, and reports wall time, CPU time and whether the outputs agree.
  4. kwker.set_algorithms({"adaptive"}) and set_scratch_limit(0) switch off the optional paths (Runtime controls).

Reporting a problem

ShellOn your machine.
python -m kwker support-bundle          # writes kwker-support-<time>.tar.gz

The bundle holds:

It contains none of your data, and you can read every file in it before sending. Use --no-probe to skip the sorts, and --all-packages to list every installed package instead of the relevant ones.

Send the bundle, with a few lines on what you ran and what happened, through the support form or to mail@kwker.io.