Command-line reference
Installing the Python package puts the kwker command on your PATH. Its file commands sort, count and rank text files, CSV / TSV tables and NumPy arrays - faster than sort, uniq and the database tools, with the same results:
| Command | What it does |
|---|---|
kwker sort |
Sorts text files (GNU sort's flags, the same bytes out), CSV / TSV tables by column and NumPy arrays. |
kwker top |
Writes the first N lines or records of the sorted input, as `sort |
kwker count |
Counts the distinct lines, fields or column values, most frequent first. |
kwker unique |
Writes the distinct lines, fields or column values once each, in sorted order. |
kwker rank |
Adds a rank column to a table: SQL's RANK, DENSE_RANK, ROW_NUMBER or pandas' average rank. |
kwker sort
Sorts the lines of text files - with GNU sort's flags and byte order (LC_ALL=C sort), byte for byte the same output - or the records of a CSV / TSV table by named columns (--by), or a NumPy array file (.npy). Inputs larger than the memory budget are sorted in parts on disk and merged.
kwker sort access.log -k2,2 -k6,6n -o sorted.log
kwker sort sales.csv --by amount:desc,date -o sorted.csv
kwker sort values.npy -o sorted.npy
kwker sort huge.txt -S 4G -T /scratch -o huge.sorted
| Option | What it does |
|---|---|
-k, --key=KEYDEF |
Sort by a key; KEYDEF is F[.C][OPTS][,F[.C][OPTS]] (OPTS: b n r). |
-t, --field-separator=SEP |
Use SEP instead of the blank-to-non-blank transition. |
-n, --numeric-sort |
Compare by numeric value (GNU's -n). |
-r, --reverse |
Reverse the result of comparisons. |
-b, --ignore-leading-blanks |
Ignore leading blanks in keys. |
-s, --stable |
No last-resort comparison: equal keys keep their input order. |
-u, --unique |
Output only the first of an equal run. |
-c, -C, --check |
Check for sorted input (-C: quietly). |
-m, --merge |
Merge files that are already sorted; do not sort. |
--by=COLS |
Table mode (CSV / TSV): sort by named columns, NAME[:int|float|str][:desc],... - numbers as numbers, text in byte order, empty fields last, ties in input order (DuckDB's ORDER BY). |
--format=csv|tsv|npy |
The input format (default: from the file name; else lines / CSV with --by). |
--argsort |
.npy arrays: write the sorting positions (int64, stable) instead of the values. |
--nulls=first|last |
Where empty fields go in table mode (default last). |
--explain |
Print the columns' detected types (table mode) on standard error. |
-o, --output=FILE |
Write the result to FILE (it may be an input). |
-z, --zero-terminated |
Lines end with NUL, not newline. |
-S, --buffer-size=SIZE |
Memory budget (GNU units; default: 3/4 of the available memory) - larger inputs are sorted in runs on disk, then merged. |
-T, --temporary-directory=DIR |
Where those runs go (default $TMPDIR, else /tmp). |
--parallel=N |
Threads (default: the available cores, at most 8). |
--stats |
Time per phase on standard error. |
kwker top
Writes the first N lines (or records) of the sorted input - what sort | head -n N writes. It takes the options of kwker sort, the number first.
kwker top 100 events.csv --by latency_ms:desc
kwker top 10 -t, -k4,4nr data.csv
kwker top 5 -r scores.npy
kwker count
Counts how often each distinct line - or one field of each line - occurs and writes them most frequent first, as cut | sort | uniq -c | sort -rn does. With --by it counts the distinct values of named table columns (SQL's GROUP BY with count).
kwker count -d' ' -f2 access.log --top 20
kwker count orders.csv --by country
| Option | What it does |
|---|---|
-d, --delimiter=SEP |
The field separator for -f (default TAB, as for cut). |
-f, --field=N |
Count field N (1-based) instead of the whole line; a line without SEP counts whole, as cut prints it. |
--top=N |
Only the N most frequent values. |
--by=NAME[,NAME] |
Table mode: count the distinct values of these columns of a CSV / TSV table with a header (types and options as for kwker sort --by; equal counts in --by order); writes a table: the columns, then count. |
-o, --output=FILE |
Write the result to FILE. |
-z, --zero-terminated |
Lines end with NUL, not newline. |
--parallel=N |
Threads (default: the available cores, at most 8). |
--stats |
Time per phase on standard error. |
kwker unique
Writes the distinct lines - or fields, or the values of named table columns - once each in sorted order, as sort | uniq does; -c adds the counts. On a .npy file it writes the distinct values as np.unique gives them.
kwker unique -c words.txt
kwker unique users.csv --by email
kwker unique ids.npy -o distinct.npy
| Option | What it does |
|---|---|
-d, --delimiter=SEP |
The field separator for -f (default TAB, as for cut). |
-f, --field=N |
Field N (1-based) instead of the whole line; a line without SEP counts whole, as cut prints it. |
-c, --count |
Each value after its count (uniq -c). |
--by=NAME[,NAME] |
Table mode: the distinct values of these columns of a CSV / TSV table with a header, in --by order (types and options as for kwker sort --by); writes a table (with -c a count column). |
-o, --output=FILE |
Write the result to FILE. |
-z, --zero-terminated |
Lines end with NUL, not newline. |
--parallel=N |
Threads (default: the available cores, at most 8). |
--stats |
Time per phase on standard error. |
kwker rank
Writes a table's records in file order, each with its rank by the --by columns appended as a last column - SQL's RANK, DENSE_RANK and ROW_NUMBER, or pandas' average rank.
kwker rank scores.csv --by score:desc --method dense --as place
| Option | What it does |
|---|---|
--by=COLUMNS |
The columns to rank by, as for kwker sort --by: NAME[:int|float|str][:asc|desc], comma-separated. |
--method=M |
Min (default; SQL RANK: 1 + the records ranked before), max, dense (SQL DENSE_RANK), first (SQL ROW_NUMBER: ties in file order) or average (pandas' default: the mean of min and max). |
--as=NAME |
The rank column's name (default: rank). |
--nulls=first|last |
Empty fields rank first or last (default: last, as DuckDB). |
-t, --field-separator=SEP |
The field separator (default: a tab for .tsv files, else a comma). |
--format=csv|tsv |
The table format (default: from the file name; else CSV). |
-o, --output=FILE |
Write the result to FILE. |
--explain |
Print the columns' detected types on standard error. |
--parallel=N |
Threads (default: the available cores, at most 8). |
--stats |
Time per phase on standard error. |
Python package commands
The package's own commands run as kwker doctor, kwker bench, kwker serve and so on, or with python -m as below - the same commands either way:
| Command | What it does |
|---|---|
python -m kwker doctor |
Prints what Kwker sees on this machine: the CPU and its features, the engine in use, threads, the operating system and the framework versions, with a warning for anything that limits speed or compatibility. |
python -m kwker support-bundle |
Writes a diagnostic archive for a bug report: the doctor report, the CPU description, the relevant environment variables, framework versions and which engine a few test sorts took. |
python -m kwker.bench |
Times Kwker against the libraries installed on your machine and writes a report you can share. |
python -m kwker audit |
Runs your own program with and without Kwker and reports the wall time, the CPU time, the speed-up and whether the results agree. |
python -m kwker.serve |
Serves a language model over an OpenAI-compatible HTTP API (/v1/completions and /v1/chat/completions), with requests batched together on KwkDecoder. |
python -m kwker.torch_ |
Runs a PyTorch program with Kwker as its CPU sorting kernels, without changing its code. |
python -m kwker doctor Page
Prints what Kwker sees on this machine: the CPU and its features, the engine in use, threads, the operating system and the framework versions, with a warning for anything that limits speed or compatibility.
python -m kwker doctor
python -m kwker doctor --json > machine.json
| Option | What it does |
|---|---|
--json |
The report as JSON. |
--no-torch |
Do not import torch (skip the PyTorch extensions' checks). |
python -m kwker support-bundle Page
Writes a diagnostic archive for a bug report: the doctor report, the CPU description, the relevant environment variables, framework versions and which engine a few test sorts took. It stays on your machine: nothing is uploaded and none of your data is read. Look inside before you share it.
python -m kwker support-bundle
python -m kwker support-bundle -o report.tar.gz
| Option | What it does |
|---|---|
-o FILE, --output FILE |
Archive path (default ./kwker-support-<time>.tar.gz). |
--all-packages |
List every installed Python package, not only the integrated frameworks. |
--no-probe |
Skip the probe sorts of generated keys. |
python -m kwker.bench Page
Times Kwker against the libraries installed on your machine and writes a report you can share. Without options it times sort, argsort, top-k and smallest-k on 1K, 100K and 1M values of four types. The other suites - input patterns, data frames, image models, text encoders and language models - are chosen with one option each.
python -m kwker.bench --quick
python -m kwker.bench --md report.md --json report.json
python -m kwker.bench --suite --native
python -m kwker.bench --frames
python -m kwker.bench --llm HuggingFaceTB/SmolLM2-135M
| Option | What it does |
|---|---|
--quick |
100K keys only. |
--sizes N,.. |
Comma-separated key counts. |
--types TYPE,.. |
Comma-separated NumPy dtypes (default int32,int64,float32,float64; --suite: uint32,int32,float32,uint64,int64,float64). |
--ops OP,.. |
Sort, argsort, top-k (with indices), smallest-k (values only); default all four (--suite: sort; sort and argsort there). |
--reps N |
Timed runs per cell; the median counts. Default: 5. |
--threads N |
Let the libraries use their default thread pools (the value is recorded). |
--native |
Also time x86-simd-sort and VQSort, downloaded at pinned commits and built here with the C++ compiler. |
--suite |
The input-family suite instead: 40 key patterns (random, skewed, duplicate-heavy, presorted, string-like, ...) and 3 sort-research-rs patterns, each as the six key types. |
--frames |
The data-frame suite instead: kwker.frame against pandas, Polars, pyarrow and DuckDB. |
--models [NAME,..] |
The image-model suite instead: KwkCNN against PyTorch eager, Inductor and OpenVINO (float32); optionally a comma-separated list of torchvision model names. |
--encoders [NAME,..] |
The text-encoder suite instead: KwkEncoder against OpenVINO, torch.compile(backend="kwker") and eager PyTorch; optionally comma-separated Hugging Face ids or directories. |
--llm MODEL |
The language-model suite instead: KwkDecoder against transformers eager (and your llama.cpp build) on a Hugging Face model id or local directory. |
--md FILE |
Write the Markdown report here. |
--json FILE |
Write the raw results here. |
With --suite:
| Option | What it does |
|---|---|
--families FAMILIES |
Comma-separated pattern names (default all 43). |
With --frames:
| Option | What it does |
|---|---|
--rows N |
Rows per table. Default: 1,000,000. |
With --models:
| Option | What it does |
|---|---|
--no-compile |
Leave out torch.compile (its compiles take a while). |
--int8 |
Also KwkCNN int8 against OpenVINO int8 (NNCF), pretrained weights, calibrated on --images; --encoders: also KwkEncoder int8 against OpenVINO int8, calibrated on the same batches. |
With --encoders:
| Option | What it does |
|---|---|
--shapes BxL,.. |
Comma-separated BATCHxLENGTH batch shapes. Default: 1x128,8x128. |
With --models --int8:
| Option | What it does |
|---|---|
--images DIR |
A folder of photos (searched recursively; 116 used: 16 to calibrate, 100 to check top-1 agreement with float32). |
With --llm:
| Option | What it does |
|---|---|
--llm-modes MODE,.. |
KwkDecoder weight modes (int4, int8, bf16). Default: int4,int8,bf16. |
--text FILE |
A text file for perplexity (llama.cpp's protocol; wikitext-2's test split is the usual one). |
--chunks N |
Perplexity chunks of 512 tokens. Default: 20. |
--bos |
Put the BOS token first in each chunk, as llama.cpp does when the vocabulary asks for one. |
--llama-bench PATH |
Path of your llama.cpp llama-bench (with --gguf). |
--llama-perplexity PATH |
Path of your llama.cpp llama-perplexity (with --gguf and --text). |
--gguf FILE,.. |
Comma-separated GGUF files of the same model for llama.cpp. |
python -m kwker audit Page
Runs your own program with and without Kwker and reports the wall time, the CPU time, the speed-up and whether the results agree. The program runs unchanged; give it after --. The exit status is 1 when the results differ.
python -m kwker audit -- train.py --epochs 1
python -m kwker audit --output preds.npy -- predict.py
python -m kwker audit --mode numpy -- analysis.py
| Option | What it does |
|---|---|
--mode {ops,bf16,int8,numpy,scipy} |
ops: Kwker as PyTorch's sorting kernels (results identical); numpy: the NumPy drop-in functions; scipy: kwker.scipy's sparse conversions, rank statistics and window filters (results identical); bf16 / int8: the CPU backend's faster precision modes (results change a little). Default: ops. |
--repeat N |
Runs of each side, alternating; the fastest counts. Default: 3. |
--output FILE |
A file the program writes, compared between the runs (repeatable). |
--rtol X |
Relative tolerance when --output arrays are compared. Default: 1e-05. |
--atol X |
Absolute tolerance when --output arrays are compared. Default: 1e-08. |
--timeout SECONDS |
Seconds per run. |
--json |
The report as JSON. |
--baseline CMD |
A command for the plain side, in one quoted string (for example "./app_std input.bin"): both sides then run as commands, not Python - two builds of a C, C++, Rust or other program. |
program |
The program (or give it after --). |
python -m kwker.serve Page
Serves a language model over an OpenAI-compatible HTTP API (/v1/completions and /v1/chat/completions), with requests batched together on KwkDecoder.
python -m kwker.serve --model HuggingFaceTB/SmolLM2-360M-Instruct --port 8000
| Option | What it does |
|---|---|
--model MODEL |
A Hugging Face model id or local path, of a model family KwkDecoder supports (/v1/completions, /v1/chat/completions). |
--embeddings EMBEDDINGS |
A Hugging Face embedding model id or local path (/v1/embeddings: pooled and normalized as its sentence-transformers configuration says). |
--precision {preserve,balanced,float32,bf16,int8,int4} |
preserve (the default: nothing lost - a model runs on its own bf16 or float32 weights), balanced (int8 weights: about +0.1% perplexity), or float32 / bf16 / int8 / int4 by name. |
--max-batch N |
Requests decoded together (cache slots). Default: 8. |
--max-cache-len N |
Prompt + generated tokens per request. Default: 4096. |
--max-tokens N |
max_tokens for requests that do not set it. Default: 256. |
--cache-dir DIR |
Keep the packed weights on disk (KwkDecoder cache_dir). |
--host ADDRESS |
The address to listen on. Default: 127.0.0.1. |
--port PORT |
The port to listen on. Default: 8000. |
--api-key API_ |
Require this key on every request, sent as OpenAI clients send it (Authorization: Bearer <key>); the KWKER_API_KEY environment variable sets it too. Without a key, the server answers every request. |
python -m kwker.torch_run Page
Runs a PyTorch program with Kwker as PyTorch's CPU sorting kernels, without changing its code. It calls kwker.torch_ops.install(), then runs the program as __main__ with its own arguments. Results equal PyTorch's.
python -m kwker.torch_run train.py --epochs 3
python -m kwker.torch_run -m mypackage.eval
It has no options of its own: everything after torch_run is the program and its arguments. Set KWKER_RUN_QUIET=1 to drop the one-line note it prints on standard error.
Related
- Evaluate on your machine: the commands in the order you use them.
- Troubleshooting: what each doctor warning means.
- Runtime controls: the environment variables.