Kwker

Command-line reference

Installing the Python package puts the kwker command on your PATH. Its file commands sort, count and rank text files, CSV / TSV tables and NumPy arrays - faster than sort, uniq and the database tools, with the same results:

Command What it does
kwker sort Sorts text files (GNU sort's flags, the same bytes out), CSV / TSV tables by column and NumPy arrays.
kwker top Writes the first N lines or records of the sorted input, as `sort
kwker count Counts the distinct lines, fields or column values, most frequent first.
kwker unique Writes the distinct lines, fields or column values once each, in sorted order.
kwker rank Adds a rank column to a table: SQL's RANK, DENSE_RANK, ROW_NUMBER or pandas' average rank.

kwker sort

Sorts the lines of text files - with GNU sort's flags and byte order (LC_ALL=C sort), byte for byte the same output - or the records of a CSV / TSV table by named columns (--by), or a NumPy array file (.npy). Inputs larger than the memory budget are sorted in parts on disk and merged.

ShellOn your machine.
kwker sort access.log -k2,2 -k6,6n -o sorted.log
kwker sort sales.csv --by amount:desc,date -o sorted.csv
kwker sort values.npy -o sorted.npy
kwker sort huge.txt -S 4G -T /scratch -o huge.sorted
Option What it does
-k, --key=KEYDEF Sort by a key; KEYDEF is F[.C][OPTS][,F[.C][OPTS]] (OPTS: b n r).
-t, --field-separator=SEP Use SEP instead of the blank-to-non-blank transition.
-n, --numeric-sort Compare by numeric value (GNU's -n).
-r, --reverse Reverse the result of comparisons.
-b, --ignore-leading-blanks Ignore leading blanks in keys.
-s, --stable No last-resort comparison: equal keys keep their input order.
-u, --unique Output only the first of an equal run.
-c, -C, --check Check for sorted input (-C: quietly).
-m, --merge Merge files that are already sorted; do not sort.
--by=COLS Table mode (CSV / TSV): sort by named columns, NAME[:int|float|str][:desc],... - numbers as numbers, text in byte order, empty fields last, ties in input order (DuckDB's ORDER BY).
--format=csv|tsv|npy The input format (default: from the file name; else lines / CSV with --by).
--argsort .npy arrays: write the sorting positions (int64, stable) instead of the values.
--nulls=first|last Where empty fields go in table mode (default last).
--explain Print the columns' detected types (table mode) on standard error.
-o, --output=FILE Write the result to FILE (it may be an input).
-z, --zero-terminated Lines end with NUL, not newline.
-S, --buffer-size=SIZE Memory budget (GNU units; default: 3/4 of the available memory) - larger inputs are sorted in runs on disk, then merged.
-T, --temporary-directory=DIR Where those runs go (default $TMPDIR, else /tmp).
--parallel=N Threads (default: the available cores, at most 8).
--stats Time per phase on standard error.

kwker top

Writes the first N lines (or records) of the sorted input - what sort | head -n N writes. It takes the options of kwker sort, the number first.

ShellOn your machine.
kwker top 100 events.csv --by latency_ms:desc
kwker top 10 -t, -k4,4nr data.csv
kwker top 5 -r scores.npy

kwker count

Counts how often each distinct line - or one field of each line - occurs and writes them most frequent first, as cut | sort | uniq -c | sort -rn does. With --by it counts the distinct values of named table columns (SQL's GROUP BY with count).

ShellOn your machine.
kwker count -d' ' -f2 access.log --top 20
kwker count orders.csv --by country
Option What it does
-d, --delimiter=SEP The field separator for -f (default TAB, as for cut).
-f, --field=N Count field N (1-based) instead of the whole line; a line without SEP counts whole, as cut prints it.
--top=N Only the N most frequent values.
--by=NAME[,NAME] Table mode: count the distinct values of these columns of a CSV / TSV table with a header (types and options as for kwker sort --by; equal counts in --by order); writes a table: the columns, then count.
-o, --output=FILE Write the result to FILE.
-z, --zero-terminated Lines end with NUL, not newline.
--parallel=N Threads (default: the available cores, at most 8).
--stats Time per phase on standard error.

kwker unique

Writes the distinct lines - or fields, or the values of named table columns - once each in sorted order, as sort | uniq does; -c adds the counts. On a .npy file it writes the distinct values as np.unique gives them.

ShellOn your machine.
kwker unique -c words.txt
kwker unique users.csv --by email
kwker unique ids.npy -o distinct.npy
Option What it does
-d, --delimiter=SEP The field separator for -f (default TAB, as for cut).
-f, --field=N Field N (1-based) instead of the whole line; a line without SEP counts whole, as cut prints it.
-c, --count Each value after its count (uniq -c).
--by=NAME[,NAME] Table mode: the distinct values of these columns of a CSV / TSV table with a header, in --by order (types and options as for kwker sort --by); writes a table (with -c a count column).
-o, --output=FILE Write the result to FILE.
-z, --zero-terminated Lines end with NUL, not newline.
--parallel=N Threads (default: the available cores, at most 8).
--stats Time per phase on standard error.

kwker rank

Writes a table's records in file order, each with its rank by the --by columns appended as a last column - SQL's RANK, DENSE_RANK and ROW_NUMBER, or pandas' average rank.

ShellOn your machine.
kwker rank scores.csv --by score:desc --method dense --as place
Option What it does
--by=COLUMNS The columns to rank by, as for kwker sort --by: NAME[:int|float|str][:asc|desc], comma-separated.
--method=M Min (default; SQL RANK: 1 + the records ranked before), max, dense (SQL DENSE_RANK), first (SQL ROW_NUMBER: ties in file order) or average (pandas' default: the mean of min and max).
--as=NAME The rank column's name (default: rank).
--nulls=first|last Empty fields rank first or last (default: last, as DuckDB).
-t, --field-separator=SEP The field separator (default: a tab for .tsv files, else a comma).
--format=csv|tsv The table format (default: from the file name; else CSV).
-o, --output=FILE Write the result to FILE.
--explain Print the columns' detected types on standard error.
--parallel=N Threads (default: the available cores, at most 8).
--stats Time per phase on standard error.

Python package commands

The package's own commands run as kwker doctor, kwker bench, kwker serve and so on, or with python -m as below - the same commands either way:

Command What it does
python -m kwker doctor Prints what Kwker sees on this machine: the CPU and its features, the engine in use, threads, the operating system and the framework versions, with a warning for anything that limits speed or compatibility.
python -m kwker support-bundle Writes a diagnostic archive for a bug report: the doctor report, the CPU description, the relevant environment variables, framework versions and which engine a few test sorts took.
python -m kwker.bench Times Kwker against the libraries installed on your machine and writes a report you can share.
python -m kwker audit Runs your own program with and without Kwker and reports the wall time, the CPU time, the speed-up and whether the results agree.
python -m kwker.serve Serves a language model over an OpenAI-compatible HTTP API (/v1/completions and /v1/chat/completions), with requests batched together on KwkDecoder.
python -m kwker.torch_run Runs a PyTorch program with Kwker as its CPU sorting kernels, without changing its code.

python -m kwker doctor Page

Prints what Kwker sees on this machine: the CPU and its features, the engine in use, threads, the operating system and the framework versions, with a warning for anything that limits speed or compatibility.

ShellOn your machine.
python -m kwker doctor
python -m kwker doctor --json > machine.json
Option What it does
--json The report as JSON.
--no-torch Do not import torch (skip the PyTorch extensions' checks).

python -m kwker support-bundle Page

Writes a diagnostic archive for a bug report: the doctor report, the CPU description, the relevant environment variables, framework versions and which engine a few test sorts took. It stays on your machine: nothing is uploaded and none of your data is read. Look inside before you share it.

ShellOn your machine.
python -m kwker support-bundle
python -m kwker support-bundle -o report.tar.gz
Option What it does
-o FILE, --output FILE Archive path (default ./kwker-support-<time>.tar.gz).
--all-packages List every installed Python package, not only the integrated frameworks.
--no-probe Skip the probe sorts of generated keys.

python -m kwker.bench Page

Times Kwker against the libraries installed on your machine and writes a report you can share. Without options it times sort, argsort, top-k and smallest-k on 1K, 100K and 1M values of four types. The other suites - input patterns, data frames, image models, text encoders and language models - are chosen with one option each.

ShellOn your machine.
python -m kwker.bench --quick
python -m kwker.bench --md report.md --json report.json
python -m kwker.bench --suite --native
python -m kwker.bench --frames
python -m kwker.bench --llm HuggingFaceTB/SmolLM2-135M
Option What it does
--quick 100K keys only.
--sizes N,.. Comma-separated key counts.
--types TYPE,.. Comma-separated NumPy dtypes (default int32,int64,float32,float64; --suite: uint32,int32,float32,uint64,int64,float64).
--ops OP,.. Sort, argsort, top-k (with indices), smallest-k (values only); default all four (--suite: sort; sort and argsort there).
--reps N Timed runs per cell; the median counts. Default: 5.
--threads N Let the libraries use their default thread pools (the value is recorded).
--native Also time x86-simd-sort and VQSort, downloaded at pinned commits and built here with the C++ compiler.
--suite The input-family suite instead: 40 key patterns (random, skewed, duplicate-heavy, presorted, string-like, ...) and 3 sort-research-rs patterns, each as the six key types.
--frames The data-frame suite instead: kwker.frame against pandas, Polars, pyarrow and DuckDB.
--models [NAME,..] The image-model suite instead: KwkCNN against PyTorch eager, Inductor and OpenVINO (float32); optionally a comma-separated list of torchvision model names.
--encoders [NAME,..] The text-encoder suite instead: KwkEncoder against OpenVINO, torch.compile(backend="kwker") and eager PyTorch; optionally comma-separated Hugging Face ids or directories.
--llm MODEL The language-model suite instead: KwkDecoder against transformers eager (and your llama.cpp build) on a Hugging Face model id or local directory.
--md FILE Write the Markdown report here.
--json FILE Write the raw results here.

With --suite:

Option What it does
--families FAMILIES Comma-separated pattern names (default all 43).

With --frames:

Option What it does
--rows N Rows per table. Default: 1,000,000.

With --models:

Option What it does
--no-compile Leave out torch.compile (its compiles take a while).
--int8 Also KwkCNN int8 against OpenVINO int8 (NNCF), pretrained weights, calibrated on --images; --encoders: also KwkEncoder int8 against OpenVINO int8, calibrated on the same batches.

With --encoders:

Option What it does
--shapes BxL,.. Comma-separated BATCHxLENGTH batch shapes. Default: 1x128,8x128.

With --models --int8:

Option What it does
--images DIR A folder of photos (searched recursively; 116 used: 16 to calibrate, 100 to check top-1 agreement with float32).

With --llm:

Option What it does
--llm-modes MODE,.. KwkDecoder weight modes (int4, int8, bf16). Default: int4,int8,bf16.
--text FILE A text file for perplexity (llama.cpp's protocol; wikitext-2's test split is the usual one).
--chunks N Perplexity chunks of 512 tokens. Default: 20.
--bos Put the BOS token first in each chunk, as llama.cpp does when the vocabulary asks for one.
--llama-bench PATH Path of your llama.cpp llama-bench (with --gguf).
--llama-perplexity PATH Path of your llama.cpp llama-perplexity (with --gguf and --text).
--gguf FILE,.. Comma-separated GGUF files of the same model for llama.cpp.

python -m kwker audit Page

Runs your own program with and without Kwker and reports the wall time, the CPU time, the speed-up and whether the results agree. The program runs unchanged; give it after --. The exit status is 1 when the results differ.

ShellOn your machine.
python -m kwker audit -- train.py --epochs 1
python -m kwker audit --output preds.npy -- predict.py
python -m kwker audit --mode numpy -- analysis.py
Option What it does
--mode {ops,bf16,int8,numpy,scipy} ops: Kwker as PyTorch's sorting kernels (results identical); numpy: the NumPy drop-in functions; scipy: kwker.scipy's sparse conversions, rank statistics and window filters (results identical); bf16 / int8: the CPU backend's faster precision modes (results change a little). Default: ops.
--repeat N Runs of each side, alternating; the fastest counts. Default: 3.
--output FILE A file the program writes, compared between the runs (repeatable).
--rtol X Relative tolerance when --output arrays are compared. Default: 1e-05.
--atol X Absolute tolerance when --output arrays are compared. Default: 1e-08.
--timeout SECONDS Seconds per run.
--json The report as JSON.
--baseline CMD A command for the plain side, in one quoted string (for example "./app_std input.bin"): both sides then run as commands, not Python - two builds of a C, C++, Rust or other program.
program The program (or give it after --).

python -m kwker.serve Page

Serves a language model over an OpenAI-compatible HTTP API (/v1/completions and /v1/chat/completions), with requests batched together on KwkDecoder.

ShellOn your machine.
python -m kwker.serve --model HuggingFaceTB/SmolLM2-360M-Instruct --port 8000
Option What it does
--model MODEL A Hugging Face model id or local path, of a model family KwkDecoder supports (/v1/completions, /v1/chat/completions).
--embeddings EMBEDDINGS A Hugging Face embedding model id or local path (/v1/embeddings: pooled and normalized as its sentence-transformers configuration says).
--precision {preserve,balanced,float32,bf16,int8,int4} preserve (the default: nothing lost - a model runs on its own bf16 or float32 weights), balanced (int8 weights: about +0.1% perplexity), or float32 / bf16 / int8 / int4 by name.
--max-batch N Requests decoded together (cache slots). Default: 8.
--max-cache-len N Prompt + generated tokens per request. Default: 4096.
--max-tokens N max_tokens for requests that do not set it. Default: 256.
--cache-dir DIR Keep the packed weights on disk (KwkDecoder cache_dir).
--host ADDRESS The address to listen on. Default: 127.0.0.1.
--port PORT The port to listen on. Default: 8000.
--api-key API_KEY Require this key on every request, sent as OpenAI clients send it (Authorization: Bearer <key>); the KWKER_API_KEY environment variable sets it too. Without a key, the server answers every request.

python -m kwker.torch_run Page

Runs a PyTorch program with Kwker as PyTorch's CPU sorting kernels, without changing its code. It calls kwker.torch_ops.install(), then runs the program as __main__ with its own arguments. Results equal PyTorch's.

ShellOn your machine.
python -m kwker.torch_run train.py --epochs 3
python -m kwker.torch_run -m mypackage.eval

It has no options of its own: everything after torch_run is the program and its arguments. Set KWKER_RUN_QUIET=1 to drop the one-line note it prints on standard error.

Applies to Kwker 0.1 · Python
Last updated
Was this page helpful?
Kwker 0.1.x: the engines each platform chooses from at run time (details)
PlatformEngines
Linux x86-64AVX-512, AVX2, SSE4.2, portable
Linux ARM64SVE / SVE2 (64-bit keys), NEON, portable
Windows x64AVX-512, AVX2, SSE4.2, portable
Windows ARM64NEON, portable
macOS ARM64NEON, portable
macOS x86-64AVX2, SSE4.2, portable
Other CPUs (RISC-V, POWER, x86 without SSE4.2, ...)portable