Python API reference
Every public function and class of the Python package (Kwker 0.1.0). For the concepts behind the parameters - orders, NaN placement, stability, threads - see Core concepts; for worked examples, the operation guides.
- NumPy and Python lists -
kwker - NumPy drop-in functions -
kwker.numpy_ops - DataFrames (pandas, Polars, pyarrow) -
kwker.frame - DuckDB -
kwker.duck - SciPy drop-in -
kwker.scipy - PyTorch operators and drop-in kernels -
kwker.torch_ops - PyTorch CPU backend -
kwker.cpu_backend - KwkCNN -
kwker.cnn - KwkEncoder -
kwker.encode - Hugging Face models by name: kwker.Decoder, kwker.Encoder -
kwker.hub - ONNX Runtime provider -
kwker.onnx - KwkDecoder and Decoder -
kwker.decode - Serving: continuous batching and an HTTP API -
kwker.serve - Constrained decoding: JSON output -
kwker.constrain - JAX -
kwker.jax_ops - Diagnostics: doctor and support bundle -
kwker.doctor - Diagnostics: audit -
kwker.audit - Diagnostics: what Kwker ran -
kwker.report - Diagnostics: the replacement contract -
kwker.contract - Benchmark -
kwker.bench
NumPy and Python lists
import kwker
algorithms() Page
Return the algorithm families this thread may use (see set_algorithms), as a frozenset of names.
argpartition(a, kth, axis=-1, descending=False, nans_first=False) Page
Return positions that partition an array around its kth value, like numpy.argpartition.
Arguments
a: A NumPy array. It is not changed.kth: The rank to partition at (a negative value counts from the end).axis: The axis to work along (the last by default).descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
An int64 array of a's shape. Along the axis, position kth holds the position of the kth value; the positions of smaller values come before it and of larger values after it.
Example
import numpy as np
import kwker
print(kwker.argpartition(np.array([7, 1, 9, 4, 3]), 2))
[1 4 3 0 2]
Notes
- Unlike NumPy, the result is deterministic: equal values are ordered by position.
Errors
TypeError: kth must be one integerValueError: a must have at least one dimensionValueError: kth ... out of range for a lane of ...
argselect(a, k, descending=False, nans_first=False) Page
Return the positions of the k smallest values (the k largest with descending=True), in no particular order.
Arguments
a: A NumPy array. It is not changed.k: How many positions.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
A uint64 NumPy array of k positions into the flattened array.
Example
import numpy as np
import kwker
print(np.sort(kwker.argselect(np.array([7, 1, 9, 4, 3]), 2)))
[1 4]
Notes
- Among equal values, the earlier positions are chosen.
Remarks: The positions of the stable order's first k keys, in any order (rule 6). Keys follow the key order: -0.0 before +0.0, every NaN in one block, last unless NaNs-first is asked for.
Examples: Top-k and selection: Only the positions
argsort(a, descending=False, nans_first=False, axis=None, threads=1, progress=None) Page
Return the positions that sort an array, like numpy.argsort. Equal values keep their input order.
Arguments
a: A NumPy array. It is not changed.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.axis: None treats the array as one flat list; an integer works on each 1-D slice along that axis.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).progress: Optional callback progress(done, total) for long calls; return True from it to cancel (Cancelled is raised).
Returns
A uint64 NumPy array of positions: one per value for axis=None, else an array of a's shape holding the positions within each slice.
Example
import numpy as np
import kwker
print(kwker.argsort(np.array([30, 10, 20])))
[1 2 0]
Notes
- threads and progress apply to a flat argsort (axis=None).
Remarks: Stable (rule 6): equal keys keep their input order, so the same input always gives the same positions. Keys follow the key order: -0.0 before +0.0, every NaN in one block, last unless NaNs-first is asked for. Threads change only the speed: the result follows the same rules on any thread count (rule 10).
Examples: Coming from another library: NumPy and PyTorch, Languages: The same calls in every language, Large data: Use several cores, Order and ranking: The order of an array: argsort
argsort_int4_packed(data, n=None, signed=False, descending=False) Page
Return the positions that sort packed 4-bit values (see sort_int4_packed). Equal values keep their order.
Arguments
data: A uint8 array; value i is in byte i // 2, the low 4 bits first.n: How many 4-bit values (default: 2 * data.size).signed: True reads the values as -8..7, False (the default) as 0..15.descending: True for largest first; the default is smallest first.
Returns
An int64 array of n positions.
Example
import numpy as np
import kwker
print(kwker.argsort_int4_packed(np.array([0x31, 0x02], dtype=np.uint8)))
[3 0 2 1]
Examples: 4-bit packed values: The largest values and their positions
argsort_masked(a, mask, descending=False, nans_first=False, bitmap=False) Page
argsort over only the positions a mask selects.
Equivalent to np.flatnonzero(mask)[np.argsort(a[mask], kind="stable")], without the copies.
Arguments
a: A NumPy array. It is not changed.mask: A boolean (or uint8) array of a's size, or an Arrow-style validity bitmap with bitmap=True.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.bitmap: True when mask is a bitmap.
Returns
A uint64 array of the selected positions, in sorted order of their values.
Example
import numpy as np
import kwker
print(kwker.argsort_masked(np.array([5, 9, 1, 7]), np.array([True, False, True, True])))
[2 0 3]
argsort_strings(strings, collation='bytes') Page
Return the positions that sort a list or array of strings. Equal strings keep their input order.
Arguments
strings: A sequence of str or bytes, or a NumPy string array (flattened).collation: The order: "bytes" (as Python compares strings, the default), "caseless" (ASCII letters without case), "natural" (numbers by value: "file2" before "file10"), "natural_caseless", or a collating sequence from collation_table (or any 256 byte weights).
Returns
A uint64 array of positions.
Example
import numpy as np
import kwker
print(kwker.argsort_strings(["file10", "File2", "file1"], collation="natural_caseless"))
[2 1 0]
Errors
ValueError: collation must be one of ... or 256 byte weights
Examples: Strings: The order of strings
arrow_argsort(arr, order='ascending', null_placement='at_end', by_codes=False) Page
Return the positions that sort an Arrow array, like pyarrow.compute.array_sort_indices. The data is read in place.
Arguments
arr: A pyarrow Array or ChunkedArray of almost any type: numbers, booleans, dates and times, decimals, strings, binaries, dictionaries, lists, maps, structs, run-end encoded arrays and unions.order: "ascending" (the default) or "descending".null_placement: "at_end" (the default) or "at_start".by_codes: True sorts a dictionary array by its codes instead of its values.
Returns
A uint64 NumPy array of positions. Equal values keep their order.
Example
import numpy as np
import kwker
import pyarrow as pa
print(kwker.arrow_argsort(pa.array(["b", None, "a", "c"])))
[2 0 3 1]
Notes
- Matches pyarrow: -0.0 equals +0.0, NaN sorts after the numbers and before the nulls, strings compare as bytes.
- Lists and structs sort element by element; unions by member, then value, as DuckDB does.
Examples: DataFrames, Arrow and DuckDB: Arrow arrays
arrow_dense_ranks(arr, threads=1) Page
Give every string of an Arrow array its rank among the distinct strings (dense ranks), plus the distinct count.
Use the ranks as compact integer keys for sorting or grouping strings.
Arguments
arr: A pyarrow string, binary (or view) or fixed-size binary array.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).
Returns
(ranks, distinct): a uint32 rank per row (rows ordered by their bytes; null rows get 0) and the number of distinct values.
Example
import numpy as np
import kwker
import pyarrow as pa
print(kwker.arrow_dense_ranks(pa.array(["b", "a", "b", "c"])))
(array([1, 0, 1, 2], dtype=uint32), 3)
arrow_top_k(arr, k, order='ascending', null_placement='at_end', by_codes=False) Page
Return the first k positions of arrow_argsort's order without sorting every row, like ORDER BY ... LIMIT k.
Arguments
arr: A pyarrow Array or ChunkedArray (the types arrow_argsort takes).k: How many positions.order: "ascending" (the default) or "descending".null_placement: "at_end" (the default) or "at_start".by_codes: True orders a dictionary array by its codes.
Returns
A uint64 array of up to k positions, in order.
Example
import numpy as np
import kwker
import pyarrow as pa
print(kwker.arrow_top_k(pa.array([5, 1, 4, None]), 2))
[1 2]
Examples: DataFrames, Arrow and DuckDB: Arrow arrays
asof_indices(left_on, right_on, left_by=None, right_by=None, direction='backward') Page
As-of join, like pandas merge_asof: for each left row, the index of the matching right row by time.
Arguments
left_on: The left rows' times: an integer or datetime64 array.right_on: The right rows' times. Neither side needs to be sorted.left_by: Optional left keys: only rows with equal keys match.right_by: The right keys, when left_by is given.direction: "backward" (the latest right time at or before the left time, the default), "forward" (the earliest at or after it) or "nearest" (the closer of the two).
Returns
An int64 array with one right-row index per left row; -1 where nothing matches.
Example
import numpy as np
import kwker
trades = np.array([5, 12, 20])
quotes = np.array([1, 10, 15])
print(kwker.asof_indices(trades, quotes))
[0 1 2]
Notes
- Among equal right times, "backward" takes the last in input order and "forward" the first; a tie in "nearest" takes the backward match.
Errors
ValueError: direction must be 'backward', 'forward' or 'nearest'ValueError: give both left_by and right_by, or neitherTypeError: ... must be integers or datetimes, not ...ValueError: ... past the int64 rangeValueError: by keys of other lengths than the times
Examples: Statistics and data helpers: Point-in-time lookups (as-of joins)
average_precision_score(y_true, y_score, *, pos_label=1) Page
Average precision of binary predictions, like sklearn.metrics.average_precision_score.
Arguments
y_true: The labels: one per prediction; equal to pos_label for a positive.y_score: The predicted scores, one per label (higher means more likely positive).pos_label: The label value that marks a positive.
Returns
A float: the area under the precision-recall curve, as scikit-learn computes it.
Example
import numpy as np
import kwker
print(kwker.average_precision_score(np.array([0, 1, 1, 0]), np.array([0.1, 0.8, 0.4, 0.35])))
1.0
Notes
- Without any positive the result is 0.0 (scikit-learn 1.9). A NaN score raises ValueError. No sample weights.
Examples: Statistics and data helpers: Ranking metrics
bucket_counts(x, boundaries, right=False, descending=False, nans_first=False) Page
Count how many values of x fall into each bucket between sorted boundaries: a histogram with your own edges.
Arguments
x: The values: an array of any shape.boundaries: The bucket edges: a sorted 1-D array.right: As in bucketize: False puts a value equal to an edge in the lower bucket, True in the upper one.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
An int64 array of len(boundaries) + 1 counts.
Example
import numpy as np
import kwker
print(kwker.bucket_counts(np.array([1, 5, 10, 15]), np.array([5, 10])))
[2 1 1]
Errors
TypeError: values of dtype ... do not convert exactly to ...
Remarks: -0.0 and +0.0 compare equal here, as in NumPy (rule 14).
Examples: Searching sorted data: How many in each bucket? bucket_counts
bucketize(x, boundaries, right=False, descending=False, nans_first=False) Page
Return the bucket of every value of x between sorted boundaries, like torch.bucketize.
Arguments
x: The values: an array of any shape.boundaries: The bucket edges: a sorted 1-D array.right: False (the default): bucket i holds boundaries[i - 1] < x <= boundaries[i]. True: boundaries[i - 1] <= x < boundaries[i].descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
An int64 array of x's shape: bucket numbers from 0 to len(boundaries).
Example
import numpy as np
import kwker
print(kwker.bucketize(np.array([1, 5, 10, 15]), np.array([5, 10])))
[0 0 1 2]
Remarks: -0.0 and +0.0 compare equal here, as in NumPy (rule 14).
Examples: Searching sorted data: Which bucket? bucketize
class Cancelled(...) Page
Raised when a progress callback cancels an operation.
capabilities() Page
Describe what Kwker can do on this machine, as a dict.
Returns
A dict with version, isa (the engine in use), engines_built, engines_usable, features, sve_vector_bits, cache sizes (l1d / l2 / l3 in bytes, or None), cpus, default_threads, scratch_limit and build.
Example
import numpy as np
import kwker
print(sorted(kwker.capabilities())[:4])
['build', 'cpus', 'default_threads', 'engines_built']
Examples: Runtime controls: What runs here
cdf_distance(a, b, kind='wasserstein') Page
Distance between the distributions of two samples, as SciPy's two-sample tests compute it.
Arguments
a: The first sample: a 1-D array of numbers.b: The second sample.kind: "wasserstein" (scipy.stats.wasserstein_distance, the default), "ks" (the Kolmogorov-Smirnov statistic of scipy.stats.ks_2samp) or "energy" (scipy.stats.energy_distance).
Returns
A float. NaN when either sample is empty or holds a NaN.
Example
import numpy as np
import kwker
print(kwker.cdf_distance(np.array([1.0, 2.0, 3.0]), np.array([2.0, 3.0, 4.0])))
1.0
Errors
ValueError: kind must be 'ks', 'wasserstein' or 'energy'
Examples: Statistics and data helpers: Compare two samples
collation_table(name) Page
Return a built-in collating sequence for argsort_strings: 256 byte weights.
Arguments
name: "ebcdic037" (order text as IBM EBCDIC code page 037 does) or "from_ebcdic037" (order EBCDIC 037 bytes as Latin-1).
Returns
bytes of length 256: weights[b] is byte b's weight.
Example
import numpy as np
import kwker
w = kwker.collation_table("ebcdic037")
print(kwker.sort_strings([b"a1", b"A1", b"1a"], collation=w))
[b'a1', b'A1', b'1a']
Errors
ValueError: name must be one of ...
coo_coalesce(rows, cols, vals, shape, reduce='sum') Page
Sort sparse COO entries row by row and combine duplicate (row, col) entries, like torch's coalesce().
Arguments
rows: The row index of every entry.cols: The column index of every entry.vals: The value of every entry.shape: The matrix shape (rows, columns).reduce: How duplicates combine: "sum" (the default), "prod", "min", "max" or "mean".
Returns
(rows, cols, vals): int64 indices in row-major order, one entry per distinct position.
Example
import numpy as np
import kwker
print(kwker.coo_coalesce(np.array([1, 0, 1]), np.array([0, 2, 0]), np.array([1.0, 2.0, 3.0]), (2, 3)))
(array([0, 1]), array([2, 0]), array([2., 4.]))
Errors
ValueError: an index is negative or outside the shape (or the shape needs > 64 key bits)
Examples: Sparse matrices: Combine duplicates another way
coo_to_csc(rows, cols, vals, shape, reduce='sum') Page
Convert sparse COO entries to CSC, combining duplicates.
Arguments
rows: The row index of every entry.cols: The column index of every entry.vals: The value of every entry.shape: The matrix shape (rows, columns).reduce: How duplicates combine: "sum" (the default), "prod", "min", "max" or "mean".
Returns
(indptr, indices, data): indptr has columns + 1 entries; row indices ascend within each column.
Example
import numpy as np
import kwker
print(kwker.coo_to_csc(np.array([1, 0, 1]), np.array([0, 2, 0]), np.array([1.0, 2.0, 3.0]), (2, 3)))
(array([0, 1, 1, 2]), array([1, 0]), array([4., 2.]))
coo_to_csr(rows, cols, vals, shape, reduce='sum') Page
Convert sparse COO entries to CSR, combining duplicates.
Arguments
rows: The row index of every entry.cols: The column index of every entry.vals: The value of every entry.shape: The matrix shape (rows, columns).reduce: How duplicates combine: "sum" (the default), "prod", "min", "max" or "mean".
Returns
(indptr, indices, data): indptr has rows + 1 entries; column indices ascend within each row.
Example
import numpy as np
import kwker
print(kwker.coo_to_csr(np.array([1, 0, 1]), np.array([0, 2, 0]), np.array([1.0, 2.0, 3.0]), (2, 3)))
(array([0, 1, 2]), array([2, 0]), array([2., 4.]))
Examples: Sparse matrices: Build a CSR matrix from entries
csc_to_csr(colptr, indices, data, shape, threads=1) Page
Convert a CSC matrix to CSR, without sorting.
Arguments
colptr: The CSC column pointers (columns + 1 entries).indices: The row index of every stored value.data: The stored values.shape: The matrix shape (rows, columns).threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).
Returns
(indptr, column indices, data): column indices ascend within each row; duplicates are kept.
Example
import numpy as np
import kwker
print(kwker.csc_to_csr(np.array([0, 1, 2, 3]), np.array([1, 1, 0]), np.array([2.0, 3.0, 1.0]), (2, 3)))
(array([0, 1, 3]), array([2, 0, 1]), array([1., 2., 3.]))
csr_to_csc(indptr, indices, data, shape, threads=1) Page
Convert a CSR matrix to CSC, like scipy's tocsc(), without sorting.
Arguments
indptr: The CSR row pointers (rows + 1 entries).indices: The column index of every stored value.data: The stored values.shape: The matrix shape (rows, columns).threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).
Returns
(colptr, row indices, data): row indices ascend within each column; duplicates are kept. int32 / uint32 inputs give int32 outputs, others int64.
Example
import numpy as np
import kwker
print(kwker.csr_to_csc(np.array([0, 1, 3]), np.array([2, 0, 1]), np.array([1.0, 2.0, 3.0]), (2, 3)))
(array([0, 1, 2, 3]), array([1, 1, 0]), array([2., 3., 1.]))
Examples: Sparse matrices: Convert between CSR and CSC
group_codes(columns, descending=False, nans_first=False, threads=1) Page
Number the groups of equal rows across one or more key columns, the first step of a group-by.
Arguments
columns: A list of 1-D arrays of the same length (the key columns, the first most significant).descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).
Returns
(codes, first, sizes): each row's group number (uint32; groups numbered in sorted key order), and each group's first row and size (uint64).
Example
import numpy as np
import kwker
codes, first, sizes = kwker.group_codes([np.array([3, 1, 3, 2])])
print(codes, first, sizes)
[2 0 2 1] [1 3 0] [1 1 2]
Notes
- Floats group by value, as pandas and Polars do: -0.0 and +0.0 form one group, and all NaN values one group.
Examples: Groups, merges and sets: Group several key columns: group_codes
group_reduce(codes, groups, values, op, valid=None, skip_nan=False, threads=1, q=0.5, interpolation='linear') Page
Compute one aggregate per group (sum, min, max, count, median, ...) from group codes.
Arguments
codes: Each row's group, as group_codes returns it (0xFFFFFFFF: the row belongs to no group).groups: The number of groups.values: One value per row.op: "sum", "min", "max", "count", "first_row", "last_row", "median" or "quantile".valid: Optional boolean mask of the rows that have a value.skip_nan: True treats NaN as a missing value.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).q: The quantile for op="quantile", from 0 to 1.interpolation: For op="quantile": "linear" (the default), "lower", "higher", "midpoint" or "nearest".
Returns
(result, counts): one result per group, and for sum / min / max / median / quantile the number of values each group had (None for the other ops). A group with count 0 has no meaningful result.
Example
import numpy as np
import kwker
codes = np.array([0, 1, 0, 1], dtype=np.uint32)
result, counts = kwker.group_reduce(codes, 2, np.array([1.0, 10.0, 2.0, 20.0]), "sum")
print(result, counts)
[ 3. 30.] [2 2]
Notes
- Sums are float64 for floats and int64 / uint64 for integers (integers wrap around on overflow).
- first_row / last_row return row numbers (uint64; 2**64 - 1 for an empty group).
- op="quantile" with skip_nan=True matches pandas groupby().quantile exactly.
Errors
ValueError: op must be one of ...TypeError: unsupported values dtype ...ValueError: ... values for ... codesValueError: ... validity flags for ... codesValueError: q ... must be in [0, 1] and interpolation one of ...
group_stats(codes, groups, values, valid=None, skip_nan=False, threads=1) Page
Compute each group's sum, count, minimum and maximum in one pass.
Arguments
codes: Each row's group, as group_codes returns it.groups: The number of groups.values: One value per row.valid: Optional boolean mask of the rows that have a value.skip_nan: True treats NaN as a missing value.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).
Returns
(sums, counts, mins, maxs), one entry per group. A group without values has sum 0 and no meaningful min / max.
Example
import numpy as np
import kwker
codes = np.array([0, 1, 0, 1], dtype=np.uint32)
print(kwker.group_stats(codes, 2, np.array([1.0, 10.0, 2.0, 20.0])))
(array([ 3., 30.]), array([2, 2], dtype=uint64), array([ 1., 10.]), array([ 2., 20.]))
Errors
TypeError: unsupported values dtype ...ValueError: ... values for ... codesValueError: ... validity flags for ... codes
histogram(x, bins=10, range=None) Page
Count values in equal-width bins, like np.histogram for float32 / float64 data.
Arguments
x: The values: a float array.bins: The number of bins.range: (low, high) of the bins; by default the data's minimum and maximum.
Returns
(counts, edges): int64 counts per bin and the bins + 1 edges, as np.histogram returns them.
Example
import numpy as np
import kwker
counts, edges = kwker.histogram(np.array([0.5, 1.5, 1.7, 3.0]), bins=3)
print(counts, edges)
[1 2 1] [0.5 1.33333333 2.16666667 3. ]
Notes
- The last bin includes its upper edge. Values outside the range and NaN values are not counted.
Errors
ValueError: bins must be positiveValueError: range must be finite with lo <= hi
Examples: Statistics and data helpers: Histograms
intersect1d(ar1, ar2, return_indices=False) Page
Return the sorted distinct values found in both arrays, like numpy.intersect1d.
Arguments
ar1: The first array (any shape and order; flattened).ar2: The second array.return_indices: True also returns the position of each common value's first occurrence in ar1 and in ar2.
Returns
The common values, or (values, positions in ar1, positions in ar2) with return_indices=True.
Example
import numpy as np
import kwker
print(kwker.intersect1d([3, 1, 2, 3], [3, 4, 1]))
[1 3]
Examples: Groups, merges and sets: Set operations
intersection_indices_sorted(a, b, multiset=False, descending=False, nans_first=False) Page
Return where the common values of two sorted 1-D arrays sit in each array.
Arguments
a: A sorted 1-D array.b: A sorted 1-D array of the same dtype.multiset: False (the default): each common value's first position in a and in b. True: one pair of positions per matching copy.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
(positions in a, positions in b): two int64 arrays of the same length.
Example
import numpy as np
import kwker
print(kwker.intersection_indices_sorted(np.array([1, 2, 4, 6]), np.array([2, 3, 6])))
(array([1, 3]), array([0, 2]))
Errors
TypeError: dtypes differ (... vs ...)
isa() Page
Return the engine in use: "avx512", "avx2", "sse42", "neon" or "portable".
Remarks: The engine changes only the speed, never a result (rule 11).
Examples: Core concepts: Engines, Kwker documentation: Check your installation, Quickstart: Kwker Core: Check what runs on your machine
isin(element, test_elements, *, invert=False) Page
Test which values of element appear in test_elements, like numpy.isin.
Arguments
element: The values to test: an array of any shape.test_elements: The values to look for: an array of any shape and order.invert: True returns the opposite: True where the value is not in test_elements.
Returns
A boolean NumPy array of element's shape.
Example
import numpy as np
import kwker
print(kwker.isin(np.array([1, 2, 3, 4]), [2, 4]))
[False True False True]
Notes
- -0.0 and +0.0 are equal; NaN matches nothing, as in NumPy.
Examples: Groups, merges and sets: Membership tests
kway_merge(runs, descending=False, nans_first=False) Page
Merge several sorted arrays into one sorted array.
Arguments
runs: A list of sorted 1-D arrays of one dtype.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
A new sorted 1-D array with all their values. Equal values keep their order: by run, then by position.
Example
import numpy as np
import kwker
print(kwker.kway_merge([np.array([1, 4, 9]), np.array([2, 3, 10]), np.array([5])]))
[ 1 2 3 4 5 9 10]
Examples: Groups, merges and sets: Merge sorted lists: kway_merge
kway_merge_kv(key_runs, value_runs, descending=False, nans_first=False) Page
Merge several sorted key arrays into one, moving each key's value along with it.
Arguments
key_runs: A list of sorted 1-D key arrays of one dtype.value_runs: One value array per key run, of matching length (rows along the first axis).descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
(keys, values): the merged keys and their values. Equal keys keep their order: by run, then by position.
Example
import numpy as np
import kwker
keys, values = kwker.kway_merge_kv([np.array([1, 4]), np.array([2, 3])], [np.array([10, 40]), np.array([20, 30])])
print(keys, values)
[1 2 3 4] [10 20 30 40]
Errors
ValueError: one value run per key run, a row per keyTypeError: the value runs must share one dtype and row shapeValueError: empty value rows
lex_select(columns, k, descending=False, nans_first=False) Page
Return the row at position k of a multi-column order, without sorting.
Arguments
columns: A list of 1-D arrays of the same length, the most important first.k: The position, from 0 to the number of rows - 1.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
The row position (an int).
Example
import numpy as np
import kwker
city = np.array([2, 1, 2, 1])
age = np.array([30, 40, 20, 35])
print(kwker.lex_select([city, age], 2))
2
Errors
ValueError: k = ... out of range for ... rows
lex_top_k(columns, k, descending=False, nans_first=False) Page
Return the first k rows of a multi-column order, like SQL ORDER BY a, b LIMIT k, without sorting every row.
Arguments
columns: A list of 1-D arrays of the same length, the most important first.k: How many rows.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
A uint64 array of up to k row positions, in order.
Example
import numpy as np
import kwker
city = np.array([2, 1, 2, 1])
age = np.array([30, 40, 20, 35])
print(kwker.lex_top_k([city, age], 2))
[3 1]
Errors
ValueError: k must be >= 0
Examples: Order and ranking: Sort by several columns, Tutorial: rank a leaderboard: Step 4: only the top three, Tutorial: rank a leaderboard: Step 5: a million players
lexsort(columns, descending=False, nans_first=False, threads=1) Page
Return the row order for sorting by several columns, like SQL ORDER BY a, b, ...
The first column decides; ties are broken by the next column, and so on. (numpy.lexsort takes the columns in the reverse order.)
Arguments
columns: A list of 1-D arrays of the same length, the most important first.descending: One flag for all columns, or a list with one per column.nans_first: One flag for all columns, or a list with one per column.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).
Returns
A uint64 array of row positions. Rows equal in every column keep their input order.
Example
import numpy as np
import kwker
city = np.array([2, 1, 2, 1])
age = np.array([30, 40, 20, 35])
print(kwker.lexsort([city, age]))
[3 1 2 0]
Examples: Order and ranking: Sort by several columns, Tutorial: rank a leaderboard: Step 3: the sorted table, Tutorial: rank a leaderboard: Step 5: a million players
max_threads() Page
Return the thread budget in force (see set_max_threads).
Remarks: Threads change only the speed: the result follows the same rules on any thread count (rule 10).
min_p_filter(logits, p=0.05, fill=-inf) Page
Min-p sampling filter for LLM logits: drop tokens whose probability is under p times the top token's.
Arguments
logits: float32 logits, one row per sequence (the last axis): a NumPy array or a torch tensor.p: The fraction of the top token's probability a token needs to stay.fill: The value written over the filtered logits.
Returns
The filtered logits, of the input's type and shape.
Example
import numpy as np
import kwker
print(kwker.min_p_filter(np.array([[2.0, 1.0, -2.0]], dtype=np.float32), p=0.1))
[[ 2. 1. -inf]]
Errors
ValueError: p must be in (0, 1]
Examples: Statistics and data helpers: Sampling filters for language models
observe(fn, *args, **kwargs) Page
Run a function and report which internal stages Kwker ran, with their counts and CPU cycles.
Arguments
fn: The function to call.*args: Its positional arguments.**kwargs: Its keyword arguments.
Returns
(result, report): fn's result and a dict with isa, traced, scratch_peak (bytes) and stages (each stage's id, name, description, count and cycles).
Example
import numpy as np
import kwker
result, report = kwker.observe(kwker.sorted, np.arange(1000)[::-1])
print(report["isa"], report["traced"] in (True, False))
avx512 True
Notes
- Stages are traced on the x86 engines only (traced is False elsewhere).
- Calls from other threads during the observation are counted too; a second concurrent observe raises RuntimeError.
Errors
RuntimeError: another observation runs
Examples: Runtime controls: What runs here
parallel_threads(reset=False) Page
Return (threads running parallel work now, the most at once since the last reset).
Arguments
reset: True restarts the peak after this reading.
partial_sort(a, k, descending=False, nans_first=False) Page
Sort only the front of an array: the k smallest values, in order, at positions 0 to k - 1.
Arguments
a: A NumPy array, changed in place (as one flat array).k: How many values to put in order at the front.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
None. The values after position k - 1 are in no particular order.
Example
import numpy as np
import kwker
a = np.array([7, 1, 9, 4, 3])
kwker.partial_sort(a, 2)
print(a[:2])
[1 3]
Remarks: The first k positions hold exactly what a full sort puts there; the rest hold the other keys in any order (rule 8). Keys follow the key order: -0.0 before +0.0, every NaN in one block, last unless NaNs-first is asked for.
Examples: Top-k and selection: Sort only the beginning
partial_sort_kv(keys, values, k, descending=False, nans_first=False) Page
partial_sort for keys with values: the k smallest keys in order at the front, each with its value.
Arguments
keys: A NumPy array, changed in place.values: An array of the same length, moved with its keys.k: How many keys to put in order at the front.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
None.
Example
import numpy as np
import kwker
keys = np.array([7, 1, 9, 4])
values = np.array([70, 10, 90, 40])
kwker.partial_sort_kv(keys, values, 2)
print(keys[:2], values[:2])
[1 4] [10 40]
Errors
ValueError: k must be >= 0
Remarks: The first k positions hold exactly what a full sort puts there; the rest hold the other keys in any order (rule 8). Unless the stable form is asked for, pairs with equal keys may come out in any order; the stable form keeps their input order (rule 6).
Examples: Sorting keys with values: Only the first k pairs
partition_indices(a, splitters, base=0, descending=False, nans_first=False) Page
Group the positions of an array by part (by splitters), without moving the data.
Arguments
a: A NumPy array (flattened). It is not changed.splitters: A Splitters, as splitters() returns it.base: Added to every position, as when the splitters were made.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
(positions, offsets): part j's positions are positions[offsets[j]:offsets[j + 1]], ascending. Send those elements to worker j.
Example
import numpy as np
import kwker
a = np.array([5, 1, 4, 2, 3, 6])
print(kwker.partition_indices(a, kwker.splitters_exact(a, 2)))
(array([1, 3, 4, 0, 2, 5], dtype=uint64), array([0, 3, 6], dtype=uint64))
partition_splitters(a, splitters, base=0, descending=False, nans_first=False) Page
Reorder a 1-D array in place so that each part (by splitters) is contiguous; return the part boundaries.
Arguments
a: A 1-D C-contiguous NumPy array, changed in place.splitters: A Splitters, as splitters() returns it.base: Added to every position, as when the splitters were made.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
The parts + 1 offsets (uint64): part j is a[off[j]:off[j + 1]], in no particular order inside.
Example
import numpy as np
import kwker
a = np.array([5, 1, 4, 2, 3, 6])
off = kwker.partition_splitters(a, kwker.splitters_exact(a, 2))
print(a, off)
[2 1 3 5 4 6] [0 3 6]
Errors
TypeError: a writable C-contiguous 1-D numpy.ndarray is required
percent_rank(a, descending=False, nans_first=False, axis=None) Page
Return each value's relative rank between 0 and 1, like SQL PERCENT_RANK: (rank - 1) / (n - 1).
Arguments
a: A NumPy array. It is not changed.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.axis: None treats the array as one flat list; an integer works on each 1-D slice along that axis.
Returns
A float64 array of a's shape. Equal values share the lowest rank of their group.
Example
import numpy as np
import kwker
print(kwker.percent_rank(np.array([30, 10, 30, 20])))
[0.66666667 0. 0.66666667 0.33333333]
Remarks: Equal values are one key here, as in NumPy: -0.0 and +0.0 are equal, and so are all NaNs (rule 13).
Examples: Order and ranking: Ranks
permute_in_place(a, perm) Page
Reorder an array in place by a permutation: afterwards a[j] is the old a[perm[j]].
Use it to apply an argsort to rows, records or a structured array without making a copy.
Arguments
a: A C-contiguous writable array, reordered along its first axis.perm: A permutation of 0 .. len(a) - 1.
Returns
a.
Example
import numpy as np
import kwker
a = np.array([10, 20, 30])
print(kwker.permute_in_place(a, np.array([2, 0, 1])))
[30 10 20]
Notes
- A perm that is not a permutation raises ValueError and moves nothing.
Errors
TypeError: a C-contiguous writable numpy.ndarray of at least one dimension is requiredValueError: ... indices for ... records
Examples: Order and ranking: Reorder records in place
class Plan(descending=False, nans_first=False, threads=1, scratch_limit=<unset>, algorithms=<unset>) Page
Settings and working memory reused across many calls: the order, the thread count and a memory cap, chosen once.
Repeated calls through a plan reuse its buffers, so argsort and top_k allocate nothing once the plan has grown. Use a plan from one thread at a time; close it (or use a with block) to free its memory early.
Arguments
descending: True orders the largest first in every call.nans_first: True puts NaN values first in every call.threads: Threads for sort and sort_kv: 1 (the default), 0 for the default count, or a count.scratch_limit: A memory cap applied during each call (default: the thread's own setting).algorithms: Algorithm families applied during each call, as set_algorithms takes them (default: the thread's).
Example
import numpy as np
import kwker
with kwker.Plan(descending=True) as p:
a = np.array([3, 1, 2])
p.sort(a)
print(a, p.argsort(np.array([5, 9, 1])))
[3 2 1] [1 0 2]
Plan.__enter__(self)
Plan.argsort(self, a, out=None)
argsort(a) of the flattened array in the plan's order, reusing the plan's buffers.
Arguments
a: A NumPy array. It is not changed.out: Optional uint64 array of a.size to write the positions into.
Returns
A uint64 array of positions (out, when given).
Plan.close(self)
Free the plan's working memory now (otherwise it is freed when the plan is garbage-collected). Do not use the plan afterwards.
Plan.partial_sort(self, a, k)
Sort only the front of an array in the plan's order: the first k values, in order (as partial_sort).
Arguments
a: A NumPy array, changed in place (as one flat array).k: How many values to put in order at the front.
Returns
None. The values after position k - 1 are in no particular order.
Plan.partial_sort_kv(self, keys, values, k)
The first k keys in order at the front, each with its value, in the plan's order (as partial_sort_kv).
Arguments
keys: A NumPy array, changed in place.values: An array of the same length, moved with its keys.k: How many keys to put in order at the front.
Returns
None.
Plan.select(self, a, k)
Put the value of rank k at position k in the plan's order, without sorting the rest (as select).
Arguments
a: A NumPy array, changed in place (as one flat array).k: The position, from 0 to a.size - 1.
Returns
None. The answer is a.flat[k].
Plan.select_kv(self, keys, values, k)
The key of rank k and its value end up at position k, in the plan's order (as select_kv).
Arguments
keys: A NumPy array, changed in place.values: An array of the same length, moved with its keys.k: The position, from 0 to len(keys) - 1.
Returns
None. Smaller keys (with their values) come before position k, larger ones after it.
Plan.sort(self, a)
Sort an array in place in the plan's order, on the plan's threads.
Arguments
a: A NumPy array, sorted in place (as one flat array).
Returns
None. a holds the sorted values.
Plan.sort_kv(self, keys, values, stable=False)
Sort keys in place and move values along with them, in the plan's order and thread count (as sort_kv).
Arguments
keys: A NumPy array of numbers, sorted in place.values: An array of the same length, reordered in place to follow its keys.stable: True keeps equal keys in their input order.
Returns
None.
Plan.top_k(self, a, k, sorted=True, out=None)
top_k(a, k, sorted) in the plan's order, reusing the plan's buffers.
Arguments
a: A NumPy array. It is not changed.k: How many values.sorted: True (the default) returns them in order.out: Optional (values, positions) arrays of k elements (a's dtype and uint64) to write the result into.
Returns
(values, positions), as top_k returns them.
Plan.tune(self, sample, max_threads=0, reps=3)
Pick the fastest thread count for inputs like sample, keep it in the plan and return it.
Arguments
sample: An input like the ones the plan will sort (it is copied, not changed).max_threads: The largest count to try; 0 means every CPU this process may run on.reps: Timed runs per count; the best is kept.
Returns
The chosen thread count. A larger count must be at least 3% faster to be chosen. Results never depend on it.
Plan.tune_algorithms(self, sample, reps=3)
Pick the fastest algorithm families for inputs like sample, keep them in the plan and return them.
Arguments
sample: An input like the ones the plan will sort (it is copied, not changed).reps: Timed runs per choice; the best is kept.
Returns
The chosen families, a frozenset of names. A narrower set must be at least 3% faster to be chosen. Results never depend on it.
Plan.workspace_bytes(self)
Return the bytes of working memory the plan holds now.
rank(a, method='average', descending=False, nans_first=False, axis=None) Page
Return the rank of every value (1 for the smallest), like scipy.stats.rankdata.
Arguments
a: A NumPy array. It is not changed.method: How equal values are ranked: "average" (they share the mean of their ranks, the default), "min" (SQL RANK), "max", "dense" (SQL DENSE_RANK) or "ordinal" (each its own rank, in input order).descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.axis: None treats the array as one flat list; an integer works on each 1-D slice along that axis.
Returns
An array of a's shape: float64 for "average", int64 for the other methods.
Example
import numpy as np
import kwker
print(kwker.rank(np.array([30, 10, 30, 20])))
print(kwker.rank(np.array([30, 10, 30, 20]), method="dense"))
[3.5 1. 3.5 2. ] [3 1 3 2]
Notes
- NaN values are ranked last (first with nans_first); -0.0 and +0.0 are equal, as in SciPy.
Errors
ValueError: unknown method ...
Remarks: Ordinal ranks give equal keys increasing ranks in input order (rule 6). Equal values are one key here, as in NumPy: -0.0 and +0.0 are equal, and so are all NaNs (rule 13).
Examples: Order and ranking: Ranks, Tutorial: rank a leaderboard: Step 2: ranks with ties, Tutorial: rank a leaderboard: Step 5: a million players
reduce_by_key(keys, values=None, op='sum', descending=False, nans_first=False, threads=1) Page
Group-by in one call: the distinct keys in sorted order and one aggregate per key.
Arguments
keys: The group keys: a 1-D array.values: The values to aggregate, one per key (not needed for op="count").op: "sum" (the default), "min", "max", "first", "last", "mean" or "count".descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).
Returns
(keys, results): the distinct keys, sorted, and one result per key.
Example
import numpy as np
import kwker
k, r = kwker.reduce_by_key(np.array([2, 1, 2]), np.array([10, 5, 7]), "sum")
print(k, r)
[1 2] [ 5 17]
Notes
- Integer sums wrap around on overflow; float sums add in input order. mean returns float64, count uint64.
- first / last follow the input order.
Errors
TypeError: unsupported key dtype ...TypeError: unsupported value dtype ... (int32 / int64 / uint32 / uint64 / float32 / float64)ValueError: keys and values differ in lengthValueError: op must be sum, min, max, first, last, mean or count (not ...)
Examples: Groups, merges and sets: Totals per key: reduce_by_key
release_scratch() Page
Give the kept buffer memory back to the system, including this thread's argsort and top-k buffers.
roc_auc_score(y_true, y_score, *, pos_label=1) Page
Area under the ROC curve of binary predictions, like sklearn.metrics.roc_auc_score.
Arguments
y_true: The labels: one per prediction; equal to pos_label for a positive.y_score: The predicted scores, one per label (higher means more likely positive).pos_label: The label value that marks a positive.
Returns
A float: the probability that a positive scores above a negative (ties count half).
Example
import numpy as np
import kwker
print(kwker.roc_auc_score(np.array([0, 1, 1, 0]), np.array([0.1, 0.8, 0.4, 0.35])))
1.0
Notes
- With only one class present the result is NaN, with a warning (scikit-learn 1.9). A NaN score raises ValueError.
Examples: Statistics and data helpers: Ranking metrics
rolling_mad(a, window) Page
Rolling median absolute deviation (MAD), the spread measure of Hampel outlier filters.
Value i is the median of |x - m| over the window ending at i, where m is that window's median.
Arguments
a: A 1-D array of numbers.window: The window length.
Returns
A float64 array of len(a). Multiply by 1.4826 to compare with a standard deviation.
Example
import numpy as np
import kwker
print(kwker.rolling_mad(np.array([1.0, 5.0, 2.0, 8.0, 3.0]), 3))
[nan nan 1. 3. 1.]
Notes
- The first window - 1 values are NaN, as is every window that holds a NaN. Even windows use the mean of the two middle values, as np.median does.
Errors
ValueError: window ... must be at least 1
Examples: Top-k and selection: Medians and quantiles over a sliding window
rolling_median(a, window) Page
Rolling median over a sliding window, like pandas Series.rolling(window).median().
Arguments
a: A 1-D array of numbers.window: The window length.
Returns
A float64 array of len(a): value i is the median of a[i + 1 - window : i + 1].
Example
import numpy as np
import kwker
print(kwker.rolling_median(np.array([1.0, 5.0, 2.0, 8.0, 3.0]), 3))
[nan nan 2. 5. 3.]
Notes
- The first window - 1 values are NaN, as is every window that holds a NaN. Even windows take the mean of the two middle values.
Examples: Top-k and selection: Medians and quantiles over a sliding window
rolling_quantile(a, window, q=0.5, interpolation='linear') Page
Rolling quantile over a sliding window, like pandas Series.rolling(window).quantile(q).
Arguments
a: A 1-D array of numbers.window: The window length.q: The quantile, from 0 to 1 (0.5 is the median).interpolation: "linear" (the default, pandas' result exactly), "lower", "higher", "midpoint" or "nearest".
Returns
A float64 array of len(a): value i is the quantile of a[i + 1 - window : i + 1].
Example
import numpy as np
import kwker
print(kwker.rolling_quantile(np.array([1.0, 5.0, 2.0, 8.0, 3.0]), 3, 0.5))
[nan nan 2. 5. 3.]
Notes
- The first window - 1 values are NaN, as is every window that holds a NaN (pandas' default).
Errors
ValueError: interpolation must be one of ...ValueError: window ... must be at least 1 and q ... in [0, 1]
Examples: Top-k and selection: Medians and quantiles over a sliding window
row_unique_count(a) Page
Count the distinct values in each row of a 2-D array.
Arguments
a: A 2-D NumPy array.
Returns
A uint64 array with one count per row.
Example
import numpy as np
import kwker
print(kwker.row_unique_count(np.array([[1, 2, 2], [5, 5, 5]])))
[2 1]
Notes
- -0.0 and +0.0 count as one value, and all NaN values as one.
Errors
ValueError: a 2-D arrayTypeError: unsupported dtype ...
Examples: Statistics and data helpers: Sizes of row-wise sets
sample(a, m, base=0, seed=0) Page
Take a random sample of m keys, one from each of m equal slices of the data.
Arguments
a: A NumPy array (flattened).m: The sample size (at most a.size is taken).base: Added to every position (to tell workers apart).seed: The random seed; the same seed gives the same sample.
Returns
A Splitters with the sampled keys and their positions.
Example
import numpy as np
import kwker
s = kwker.sample(np.arange(100), 4, seed=1)
print(len(s.keys))
4
sample_sorted(a, m, base=0) Page
Take m evenly spaced keys from the sorted data (regular sampling, as in parallel sorting by regular sampling).
Arguments
a: A NumPy array (sorted first; flattened).m: The sample size.base: Added to every position.
Returns
A Splitters with the keys at ranks (j + 1) n / (m + 1).
Example
import numpy as np
import kwker
print(kwker.sample_sorted(np.array([9, 1, 5, 3, 7]), 2).keys)
[1 3]
scratch_limit() Page
Return this thread's scratch memory cap in bytes, or None when there is none (see set_scratch_limit).
Remarks: The scratch limit changes only the speed and memory use, never a result (rule 12).
Examples: Large data: Limit the extra memory
scratch_policy() Page
Return the buffer policy (see set_scratch_policy) as a dict: cache_bytes, mapped, huge_pages.
scratch_use(reset_peak=False) Page
Report the library's buffer memory now, as a dict: in_use, cached and peak (bytes).
Arguments
reset_peak: True restarts the peak measurement after this reading.
Returns
{"in_use": bytes held by running calls, "cached": bytes kept for later calls, "peak": the most in use at once}.
Example
import numpy as np
import kwker
print(sorted(kwker.scratch_use()))
['cached', 'in_use', 'peak']
Examples: Runtime controls: Memory, Deploy and operate: Memory limits
searchsorted(a, v, side='left', descending=False, nans_first=False) Page
Find where values would be inserted into a sorted array to keep it sorted, like numpy.searchsorted.
Arguments
a: A sorted 1-D NumPy array (in the order given by descending / nans_first; not checked).v: The values to look up: a scalar or an array of any shape.side: "left" (the default) returns the position before equal values, "right" the position after them.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
An int64 array of v's shape, or a Python int when v is a scalar.
Example
import numpy as np
import kwker
a = np.array([10, 20, 30, 40])
print(kwker.searchsorted(a, [25, 10, 99]))
[2 0 4]
Errors
ValueError: side must be 'left' or 'right', not ...TypeError: values of dtype ... do not convert exactly to ...
Remarks: -0.0 and +0.0 compare equal here, as in NumPy (rule 14).
Examples: Searching sorted data: Where does a value go? searchsorted
select(a, k, descending=False, nans_first=False) Page
Put the value of rank k at position k, without sorting the rest, like numpy.partition.
Afterwards a[k] holds the value a full sort would put there, the values before it are no larger and the values after it no smaller. Use it for a median or a percentile.
Arguments
a: A NumPy array, changed in place (as one flat array).k: The position, from 0 to a.size - 1.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
None. The answer is a.flat[k].
Example
import numpy as np
import kwker
a = np.array([7, 1, 9, 4, 3])
kwker.select(a, 2)
print(a[2])
4
Remarks: Position k holds the key a full sort puts there; the keys before it are ordered before or equal to it, the keys after it after or equal (rule 7). Keys follow the key order: -0.0 before +0.0, every NaN in one block, last unless NaNs-first is asked for.
Examples: Performance: Ask only for what you need, Quickstart: Kwker Core: The median, without a full sort, Top-k and selection: The median and other positions
select_kv(keys, values, k, descending=False, nans_first=False) Page
select for keys with values: the key of rank k and its value end up at position k.
Arguments
keys: A NumPy array, changed in place.values: An array of the same length, moved with its keys.k: The position, from 0 to len(keys) - 1.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
None. Smaller keys (with their values) come before position k, larger ones after it, in no particular order.
Example
import numpy as np
import kwker
keys = np.array([7, 1, 9, 4])
values = np.array([70, 10, 90, 40])
kwker.select_kv(keys, values, 1)
print(keys[1], values[1])
4 40
Errors
ValueError: k = ... out of range for ... keys
Remarks: Position k holds the key a full sort puts there; the keys before it are ordered before or equal to it, the keys after it after or equal (rule 7). Unless the stable form is asked for, pairs with equal keys may come out in any order; the stable form keeps their input order (rule 6).
set_algorithms(classes) Page
Allow only some algorithm families in this thread's later calls, for testing and timing; return the previous set.
The results never change, only the time. The comparison-based core always runs.
Arguments
classes: Names from "adaptive" (shortcuts for sorted, reversed or nearly sorted input), "counting" (shortcuts for few distinct values), "radix" (radix passes), or "all" (the default). An int mask (1 / 2 / 4) works too.
Returns
The previous classes, as a frozenset of names.
Example
import numpy as np
import kwker
previous = kwker.set_algorithms({"adaptive"})
kwker.set_algorithms(previous)
print(sorted(previous))
['adaptive', 'counting', 'radix']
Errors
ValueError: unknown class ... (adaptive, counting, radix, all)
Examples: Runtime controls: Algorithm classes
set_default_threads(n) Page
Set how many threads calls made with threads=0 use (process-wide; 1 at start).
Arguments
n: The thread count; 0 means every CPU this process may run on.
Remarks: Threads change only the speed: the result follows the same rules on any thread count (rule 10).
set_isa(name=None) Page
Switch later calls to a slower engine, for timing or debugging (process-wide); return the engine now in use.
Arguments
name: "avx512", "avx2", "sse42", "neon", "portable", or None for the best this CPU supports.
Returns
The engine in use after the change (a cap above what the CPU runs gives the best it does).
Example
import numpy as np
import kwker
print(kwker.set_isa("portable"))
kwker.set_isa(None)
portable
Notes
- Results never depend on the engine, only the speed.
Errors
ValueError: unknown engine ...
Remarks: The engine changes only the speed, never a result (rule 11).
Examples: Core concepts: Engines
set_max_threads(n) Page
Set the total number of threads all parallel calls may use together (process-wide).
Concurrent multithreaded calls share this budget instead of each starting its own threads.
Arguments
n: The budget; 0 (the default) means every CPU this process may run on.
Remarks: Threads change only the speed: the result follows the same rules on any thread count (rule 10).
set_op_sorted(a, b, op, multiset=False, descending=False, nans_first=False) Page
Intersection, union, difference or symmetric difference of two already-sorted 1-D arrays, in one linear pass.
Arguments
a: A sorted 1-D array (not checked).b: A sorted 1-D array of the same dtype.op: "intersection", "union", "difference" (a minus b) or "symmetric_difference".multiset: False (the default) treats the arrays as sets; True counts repeated values, as C++ std::set_* does.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
A sorted 1-D array.
Example
import numpy as np
import kwker
print(kwker.set_op_sorted(np.array([1, 2, 4, 6]), np.array([2, 3, 6]), "intersection"))
[2 6]
Notes
- Values compare as NumPy compares them: -0.0 equals +0.0, and all NaN values are equal.
Errors
TypeError: dtypes differ (... vs ...)ValueError: unknown op ...
Remarks: Equal values are one key here, as in NumPy: -0.0 and +0.0 are equal, and so are all NaNs (rule 13).
Examples: Groups, merges and sets: Set operations
set_scratch_limit(nbytes) Page
Cap the extra memory this thread's later calls may allocate.
Arguments
nbytes: The cap in bytes; None (the default) means no cap, 0 means no allocation at all.
Returns
None. Results are the same under any cap; some inputs take longer.
Example
import numpy as np
import kwker
kwker.set_scratch_limit(0)
a = np.array([3, 1, 2])
kwker.sort(a)
kwker.set_scratch_limit(None)
print(a)
[1 2 3]
Notes
- Applies to the in-place calls: sort, select, partial_sort and sort_kv.
Remarks: The scratch limit changes only the speed and memory use, never a result (rule 12).
Examples: Runtime controls: Memory, Deploy and operate: Memory limits, Large data: Limit the extra memory
set_scratch_policy(cache_bytes=None, mapped=None, huge_pages=None) Page
Set how the library keeps the memory buffers its calls allocate (process-wide). None leaves a setting unchanged.
Arguments
cache_bytes: How many bytes of released buffers to keep for later calls (default 1 GiB; 0 returns every buffer when its call ends).mapped: True (the default on Linux) takes large buffers from the library's own memory mappings; False from the heap.huge_pages: True asks the system for huge pages on new mappings (default off).
Returns
None.
Example
import numpy as np
import kwker
kwker.set_scratch_policy(cache_bytes=0)
print(kwker.scratch_policy()["cache_bytes"])
kwker.set_scratch_policy(cache_bytes=1 << 30)
0
Examples: Deploy and operate: Memory limits
set_worker_cpus(cpus) Page
Pin the worker threads of later parallel calls to these CPUs (Linux only): worker w runs on cpus[w % len(cpus)].
Arguments
cpus: A list of CPU numbers; None or [] removes the pinning (the default).
setdiff1d(ar1, ar2) Page
Return the sorted distinct values of ar1 that are not in ar2, like numpy.setdiff1d.
Arguments
ar1: The first array (any shape and order; flattened).ar2: The values to remove.
Returns
A sorted 1-D array.
Example
import numpy as np
import kwker
print(kwker.setdiff1d([5, 1, 3, 1], [3]))
[1 5]
Examples: Groups, merges and sets: Set operations
setxor1d(ar1, ar2) Page
Return the sorted distinct values found in exactly one of the two arrays, like numpy.setxor1d.
Arguments
ar1: The first array (any shape and order; flattened).ar2: The second array.
Returns
A sorted 1-D array.
Example
import numpy as np
import kwker
print(kwker.setxor1d([1, 2, 3], [2, 4]))
[1 3 4]
sort(a, descending=False, nans_first=False, axis=None, threads=1, progress=None) Page
Sort an array in place, smallest first (largest first with descending=True).
Arguments
a: A NumPy array of numbers, sorted in place.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.axis: None treats the array as one flat list; an integer works on each 1-D slice along that axis.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).progress: Optional callback progress(done, total) for long calls; return True from it to cancel (Cancelled is raised).
Returns
None. The array itself is sorted.
Example
import numpy as np
import kwker
a = np.array([3, 1, 2])
kwker.sort(a)
print(a)
[1 2 3]
Notes
- threads and progress apply to a flat sort (axis=None).
- Cancelling through progress raises Cancelled; the array then holds its own values in some order.
Errors
TypeError: a writable numpy.ndarray is required
Remarks: Not stable (rule 5): keys that are equal but can be told apart (NaNs with different bits) may change places. Keys follow the key order: -0.0 before +0.0, every NaN in one block, last unless NaNs-first is asked for. Threads change only the speed: the result follows the same rules on any thread count (rule 10).
Examples: PyTorch and JAX (archived): Operators, Core concepts: In place or a copy, Runtime controls: Algorithm classes, Runtime controls: Memory
sort_file(input, output, dtype, memory=None, temp_dir=None, descending=False, nans_first=False, threads=1, progress=None, key=None, record_size=None, key_offset=0, sync=False) Page
Sort a binary file that may be larger than memory, writing the sorted result to output.
Arguments
input: The path of a file of keys stored back to back in native byte order (what numpy's tofile writes).output: The path to write; it may be the same as input.dtype: The key type, or a structured dtype for files of fixed-size records (with key=).memory: The memory to use for buffers, in bytes (default 256 MiB).temp_dir: Where to put the temporary run files (default: the output's directory). They are removed afterwards.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).progress: Optional callback progress(done, total) for long calls; return True from it to cancel (Cancelled is raised).key: For record files: the name of the key field in the structured dtype.record_size: For record files without a structured dtype: the bytes per record.key_offset: With record_size: the byte offset of the key inside each record.sync: True flushes the output to disk (fsync) before returning.
Returns
None.
Example
import numpy as np
import kwker
import os, tempfile
path = os.path.join(tempfile.mkdtemp(), "keys.bin")
np.array([3, 1, 2], dtype=np.int64).tofile(path)
kwker.sort_file(path, path, np.int64)
print(np.fromfile(path, dtype=np.int64))
[1 2 3]
Notes
- Record files sort stably by their key; the record bytes are moved unchanged.
- threads and progress apply to plain key files; cancelling raises Cancelled and leaves the output incomplete.
Errors
TypeError: unsupported dtype ...Cancelled: cancelled by the progress callback: ...OSError: I/O error or a file that is not a whole number of ...-byte keys: ...TypeError: unsupported key dtype ...ValueError: a ...-byte key at offset ... does not fit ...-byte recordsOSError: I/O error or a file that is not a whole number of ...-byte records: ...TypeError: key=... needs a structured dtype with that field, got ...
Examples: Large data: Sort a file larger than memory, Tutorial: sort a file larger than memory: Step 2: sort it with 16 MiB, Tutorial: sort a file larger than memory: Step 4: follow the progress, Tutorial: sort a file larger than memory: Step 5: sort records by one field
sort_int4_packed(data, n=None, signed=False, descending=False) Page
Sort 4-bit values stored two per byte (the int4 layout of GGML, ONNX and quint4x2) in place.
Arguments
data: A uint8 array; value i is in byte i // 2, the low 4 bits first.n: How many 4-bit values (default: 2 * data.size).signed: True reads the values as -8..7, False (the default) as 0..15.descending: True for largest first; the default is smallest first.
Returns
None.
Example
import numpy as np
import kwker
data = np.array([0x31, 0x02], dtype=np.uint8)
kwker.sort_int4_packed(data)
print([hex(b) for b in data])
['0x10', '0x32']
Examples: 4-bit packed values: Sort packed values
sort_kv(keys, values, descending=False, nans_first=False, stable=False, threads=1, progress=None) Page
Sort keys in place and move a second array of values along with them.
Arguments
keys: A NumPy array of numbers, sorted in place.values: An array of the same length, reordered in place to follow its keys. Elements of 1 to 32 bytes.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.stable: True keeps equal keys in their input order.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).progress: Optional callback progress(done, total) for long calls; return True from it to cancel (Cancelled is raised).
Returns
None.
Example
import numpy as np
import kwker
keys = np.array([3, 1, 2])
values = np.array([30.0, 10.0, 20.0])
kwker.sort_kv(keys, values)
print(keys, values)
[1 2 3] [10. 20. 30.]
Notes
- With stable=False (the default) the values of equal keys come out in any order.
- Cancelling through progress raises Cancelled and leaves keys and values unchanged.
Errors
ValueError: keys and values differ in lengthValueError: keys and values must be C-contiguous
Remarks: Unless the stable form is asked for, pairs with equal keys may come out in any order; the stable form keeps their input order (rule 6). Keys follow the key order: -0.0 before +0.0, every NaN in one block, last unless NaNs-first is asked for. Threads change only the speed: the result follows the same rules on any thread count (rule 10).
Examples: Core concepts: In place or a copy, Sorting keys with values: Sort keys and values together, Sorting keys with values: Keep equal keys in order: the stable sort
sort_masked(a, mask, descending=False, nans_first=False, bitmap=False) Page
Sort in place only the values at the positions a mask selects; the other positions are not touched.
Arguments
a: A NumPy array, changed in place.mask: A boolean (or uint8) array of a's size, or an Arrow-style validity bitmap with bitmap=True.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.bitmap: True when mask is a bitmap.
Returns
None.
Example
import numpy as np
import kwker
a = np.array([5, 9, 1, 7])
kwker.sort_masked(a, np.array([True, False, True, True]))
print(a)
[1 9 5 7]
sort_rows(a, descending=False, nans_first=False, progress=None, threads=1) Page
Sort every row of a 2-D array in place.
Arguments
a: A 2-D NumPy array (C order); each row is sorted along the last axis.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.progress: Optional callback progress(done, total) for long calls; return True from it to cancel (Cancelled is raised).threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).
Returns
None.
Example
import numpy as np
import kwker
a = np.array([[3, 1, 2], [9, 7, 8]])
kwker.sort_rows(a)
print(a)
[[1 2 3] [7 8 9]]
Notes
- progress is called about every 64K values; cancelling raises Cancelled and leaves the rows done so far sorted.
- threads cannot be combined with progress.
Errors
ValueError: a 2-D array is required
Remarks: Not stable (rule 5): keys that are equal but can be told apart (NaNs with different bits) may change places. Keys follow the key order: -0.0 before +0.0, every NaN in one block, last unless NaNs-first is asked for. Threads change only the speed: the result follows the same rules on any thread count (rule 10).
sort_segments(a, offsets, descending=False, nans_first=False, progress=None, threads=1) Page
Sort each segment of a 1-D array in place: segment i is a[offsets[i]:offsets[i + 1]].
Arguments
a: A 1-D NumPy array.offsets: The segment boundaries: nondecreasing, the last at most len(a).descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.progress: Optional callback progress(done, total) for long calls; return True from it to cancel (Cancelled is raised).threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).
Returns
None. Values outside the segments are not touched.
Example
import numpy as np
import kwker
a = np.array([3, 1, 2, 9, 7])
kwker.sort_segments(a, [0, 3, 5])
print(a)
[1 2 3 7 9]
Errors
ValueError: a 1-D array is requiredValueError: offsets must be nondecreasing within 0..len(a)
Remarks: Not stable (rule 5): keys that are equal but can be told apart (NaNs with different bits) may change places. Keys follow the key order: -0.0 before +0.0, every NaN in one block, last unless NaNs-first is asked for. Threads change only the speed: the result follows the same rules on any thread count (rule 10).
sort_strings(strings, collation='bytes') Page
Return strings sorted (in argsort_strings' order): a new list for a sequence, a new array for a NumPy array.
Arguments
strings: A sequence of str or bytes, or a NumPy string array.collation: As in argsort_strings.
Returns
The sorted strings.
Example
import numpy as np
import kwker
print(kwker.sort_strings(["file10", "file2", "file1"], collation="natural"))
['file1', 'file2', 'file10']
Examples: Strings: Sort a list of strings, Strings: NumPy string arrays
sorted(a, descending=False, nans_first=False) Page
Return a sorted copy of an array, flattened to one dimension. The input is not changed.
Arguments
a: A NumPy array, or anything NumPy converts (a list, a tuple).descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
A new 1-D NumPy array of a's dtype, sorted.
Example
import numpy as np
import kwker
print(kwker.sorted([3, 1, 2]))
[1 2 3]
Remarks: Not stable (rule 5): keys that are equal but can be told apart (NaNs with different bits) may change places. Keys follow the key order: -0.0 before +0.0, every NaN in one block, last unless NaNs-first is asked for.
Examples: Coming from another library: NumPy and PyTorch, Core concepts: In place or a copy, Core concepts: Special float values, Quickstart: Kwker Core: Largest first, or a sorted copy
split_points(sorted_keys, splitters, base=0, descending=False, nans_first=False) Page
Return where splitters cut an already sorted array: the parts + 1 offsets.
Arguments
sorted_keys: A sorted 1-D array.splitters: A Splitters, as splitters() returns it.base: Added to every position, as when the splitters were made.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
A uint64 array of parts + 1 offsets.
Example
import numpy as np
import kwker
a = np.array([1, 2, 3, 4, 5, 6])
print(kwker.split_points(a, kwker.splitters_exact(a, 3)))
[0 2 4 6]
class Splitters(keys, pos) Page
A set of keys with positions: a sample of the data, or the splitters that cut it into parts.
Positions break ties between equal keys, so a boundary can fall between two equal values. Each position is base + the element's index in the array a call sees (for example base = worker << 40 for distributed data).
It is the tuple (keys, pos) - the keys and their uint64 positions - with the properties keys, pos and parts (the number of parts these splitters make: len(keys) + 1).
Arguments
-
keys -
pos -
Splitters.keys(property) -
Splitters.parts(property) -
Splitters.pos(property)
splitters(samples, parts, descending=False, nans_first=False) Page
Pick parts - 1 splitters from one or more samples, to cut data into parts of similar size.
Arguments
samples: A Splitters, or a list of them (one per worker).parts: How many parts.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
A Splitters with parts - 1 entries. An empty sample gives one part.
Example
import numpy as np
import kwker
s = kwker.sample(np.arange(1000), 64, seed=1)
print(len(kwker.splitters(s, 4).keys))
3
Errors
ValueError: no samplesValueError: parts must be at least 1
splitters_exact(a, parts, base=0, descending=False, nans_first=False) Page
Compute exact splitters that cut the data into parts of equal size, duplicates included.
Arguments
a: A NumPy array (flattened).parts: How many parts.base: Added to every position.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
A Splitters with parts - 1 entries; each part gets floor or ceil(n / parts) keys.
Example
import numpy as np
import kwker
print(kwker.splitters_exact(np.array([5, 1, 4, 2, 3, 6]), 3).keys)
[3 5]
Errors
ValueError: parts must be at least 1
top_k(a, k, sorted=True, descending=False, nans_first=False) Page
Return the k smallest values and their positions (the k largest with descending=True).
Arguments
a: A NumPy array. It is not changed.k: How many values.sorted: True (the default) returns them in order; False in any order, which is faster.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
(values, positions): k values of a's dtype and their uint64 positions in the flattened array.
Example
import numpy as np
import kwker
values, positions = kwker.top_k(np.array([5, 9, 1, 7]), 2, descending=True)
print(values, positions)
[9 7] [1 3]
Notes
- Among equal values, the earlier positions are chosen.
Remarks: The first k keys of the stable order and their positions; equal keys keep their input order (rule 9). Keys follow the key order: -0.0 before +0.0, every NaN in one block, last unless NaNs-first is asked for.
Examples: Coming from another library: NumPy and PyTorch, Kwker documentation: Try it, Add Kwker to your project: Arrays in any language, Languages: The same calls in every language
top_k_by_group(a, groups, k, descending=False, nans_first=False) Page
Top k per group, like SQL ROW_NUMBER() OVER (PARTITION BY group ORDER BY value) <= k.
Arguments
a: The values to rank: a 1-D array.groups: An integer group label per value, in any order.k: How many values per group.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.
Returns
(labels, offsets, positions): the distinct labels in ascending order; group i's best positions into a are positions[offsets[i]:offsets[i + 1]], in order.
Example
import numpy as np
import kwker
a = np.array([5, 9, 1, 7, 3])
g = np.array([0, 1, 0, 1, 0])
print(kwker.top_k_by_group(a, g, 2, descending=True))
(array([0, 1]), array([0, 2, 4], dtype=uint64), array([0, 4, 1, 3], dtype=uint64))
Notes
- For labels that are not integers, pass numpy.unique(labels, return_inverse=True)[1].
Errors
TypeError: integer labels (other types: numpy.unique(labels, return_inverse=True)[1])ValueError: ... keys but ... labels
Examples: Top-k and selection: Top-k per group, Tutorial: top-k recommendations: Step 3: every user at once, Tutorial: top-k recommendations: Step 4: skip what they already bought, Tutorial: top-k recommendations: Step 5: a million candidates
top_k_filter(logits, k, fill=-inf) Page
Top-k sampling filter for LLM logits: keep each row's k largest logits and fill the rest.
Arguments
logits: float32 logits, one row per sequence (the last axis): a NumPy array or a torch tensor.k: One k for every row, or one per row (a list, array or tensor). k <= 0 or k >= the row length keeps the row.fill: The value written over the filtered logits.
Returns
The filtered logits, of the input's type and shape. Logits equal to the k-th largest stay.
Example
import numpy as np
import kwker
print(kwker.top_k_filter(np.array([[2.0, 1.0, 3.0, 0.0]], dtype=np.float32), 2))
[[ 2. -inf 3. -inf]]
Notes
- NaN counts as the largest value, as in torch.topk.
Errors
TypeError: float32 logitsTypeError: integer kValueError: ... k values for ... rows
Examples: Statistics and data helpers: Sampling filters for language models
top_k_int4_packed(data, k, n=None, signed=False, descending=True) Page
Return the positions of the k largest packed 4-bit values (the k smallest with descending=False).
Arguments
data: A uint8 array; value i is in byte i // 2, the low 4 bits first.k: How many positions.n: How many 4-bit values (default: 2 * data.size).signed: True reads the values as -8..7, False (the default) as 0..15.descending: True for largest first; the default is smallest first.
Returns
An int64 array of k positions, in value order (equal values by position).
Example
import numpy as np
import kwker
print(kwker.top_k_int4_packed(np.array([0x31, 0x92], dtype=np.uint8), 2))
[3 1]
Examples: 4-bit packed values: The largest values and their positions
top_k_masked(a, mask, k, sorted=True, descending=False, nans_first=False, bitmap=False) Page
top_k over only the positions a mask selects.
Arguments
a: A NumPy array. It is not changed.mask: A boolean (or uint8) array of a's size: nonzero positions take part. With bitmap=True, an Arrow-style validity bitmap instead (one bit per value, least significant bit first).k: How many values.sorted: True (the default) returns them in order.descending: True for largest first; the default is smallest first.nans_first: True puts NaN values first; the default puts them last.bitmap: True when mask is a bitmap.
Returns
(values, positions) as top_k returns them; fewer than k when fewer positions are selected.
Example
import numpy as np
import kwker
a = np.array([5, 9, 1, 7])
print(kwker.top_k_masked(a, np.array([True, False, True, True]), 2, descending=True))
(array([7, 5]), array([3, 0], dtype=uint64))
Errors
ValueError: a validity bitmap is uint8 with ceil(n / 8) bytesValueError: ... keys but ... mask entries
Examples: Top-k and selection: Top-k with a filter
top_k_segments(a, offsets, k, largest=True, sorted=True) Page
Return the k largest values of every segment of a 1-D array (the k smallest with largest=False).
Segment i is a[offsets[i]:offsets[i + 1]]. Use it for the top items per user, the top terms per document, or the best neighbors of each node in a sparse graph.
Arguments
a: A 1-D NumPy array.offsets: The segment boundaries: nondecreasing, the last at most len(a).k: How many values per segment.largest: True (the default) for the largest values, False for the smallest.sorted: True (the default) returns each segment's values in order.
Returns
(values, positions), each of shape (segments, k); positions are within the segment. A segment shorter than k fills its remaining slots with value 0 and position -1.
Example
import numpy as np
import kwker
values, positions = kwker.top_k_segments(np.array([5, 1, 4, 9, 2]), [0, 3, 5], 2)
print(values)
print(positions)
[[5 4] [9 2]] [[0 2] [0 1]]
Notes
- NaN counts as the largest value, as in torch.topk.
Errors
ValueError: offsets must be nondecreasing within 0..len(a)ValueError: k must be at least 0
Examples: Top-k and selection: Top-k per group
top_n_sigma(logits, n=1.0, fill=-inf) Page
Top-n-sigma sampling filter for LLM logits: keep the tokens within n standard deviations of the best one.
Logits below max - n x std of their row are replaced by fill, so sampling picks only among the strongest tokens.
Arguments
logits: float32 logits, one row per sequence (the last axis): a NumPy array or a torch tensor.n: How many standard deviations below the maximum to keep.fill: The value written over the filtered logits.
Returns
The filtered logits, of the input's type and shape.
Example
import numpy as np
import kwker
print(kwker.top_n_sigma(np.array([[2.0, 1.0, 0.0, -5.0]], dtype=np.float32), n=1.0))
[[ 2. 1. 0. -inf]]
Notes
- A row holding a NaN is left unchanged.
Examples: Statistics and data helpers: Sampling filters for language models
trim_mean(a, proportiontocut, axis=0) Page
Mean after cutting a fraction of the smallest and largest values, like scipy.stats.trim_mean.
Arguments
a: A NumPy array of numbers.proportiontocut: The fraction to cut from each end, from 0 to 0.5.axis: The axis to average along (0 by default); None averages the flattened array.
Returns
The trimmed means: float input gives its own float type, integer input float64.
Example
import numpy as np
import kwker
print(kwker.trim_mean(np.array([1.0, 2.0, 3.0, 4.0, 100.0]), 0.2))
3.0
Notes
- NaN sorts above every number, so a result is NaN only when more NaN values remain than were cut.
Errors
ValueError: axis out of range for a ...-dimensional arrayValueError: Proportion too big.
Examples: Statistics and data helpers: Trimmed means and weighted quantiles
union1d(ar1, ar2) Page
Return the sorted distinct values found in either array, like numpy.union1d.
Arguments
ar1: The first array (any shape and order; flattened).ar2: The second array.
Returns
A sorted 1-D array of the distinct values.
Example
import numpy as np
import kwker
print(kwker.union1d([3, 1, 2], [2, 5]))
[1 2 3 5]
Examples: Groups, merges and sets: Set operations
unique(a, return_index=False, return_inverse=False, return_counts=False, *, threads=1) Page
The sorted distinct values of an array, as numpy.unique.
Arguments
a: A NumPy array or a list of integers, floats or booleans. It is flattened first.return_index: Also return where each distinct value first occurs ina.return_inverse: Also return, for every element ofa, the index of its value in the result.return_counts: Also return how often each distinct value occurs.threads: The threads to use for large arrays.
Returns
The distinct values in ascending order, a 1-D array of a's dtype. With any return_* flag, a tuple of the values and the arrays asked for, in the order index, inverse, counts (all int64; the inverse has a's shape).
Example
import numpy as np
import kwker
values, counts = kwker.unique(np.array([3, 1, 3, 2, 1, 3]), return_counts=True)
print(values, counts)
[1 2 3] [2 1 3]
Notes
- All NaN values count as one value, sorted last, and -0.0 and +0.0 as one, as in numpy.unique.
Errors
TypeError: unsupported dtype ...
Remarks: As numpy.unique: -0.0 and +0.0 are one value, and so are all NaNs, sorted last (rule 15). Threads change only the speed: the result follows the same rules on any thread count (rule 10).
Examples: Groups, merges and sets: Distinct values: unique
version() Page
Return the library version string, for example "0.1.0".
Examples: Kwker documentation: Check your installation, Quickstart: Kwker Core: Check what runs on your machine
weighted_quantile(a, q, weights) Page
Weighted quantiles, like np.quantile(a, q, weights=weights, method="inverted_cdf") in NumPy 2.
Arguments
a: The values (flattened).q: One quantile or a list of them, each from 0 to 1.weights: Non-negative weights, the shape of a.
Returns
The quantile values, matching NumPy's result exactly: a scalar for a single q, else an array.
Example
import numpy as np
import kwker
print(kwker.weighted_quantile(np.array([1.0, 2.0, 3.0]), 0.5, np.array([1.0, 1.0, 4.0])))
3.0
Notes
- A NaN in a makes every result NaN, as in NumPy.
Errors
ValueError: ... weights for ... valuesValueError: no valuesValueError: Weights must be non-negative.ValueError: Quantiles must be in the range [0, 1]
Examples: Statistics and data helpers: Trimmed means and weighted quantiles
NumPy drop-in functions
import kwker.numpy_ops
Kwker as NumPy's sorting functions for a whole program, with no code change: install() replaces numpy.sort,
argsort, partition, argpartition, lexsort, searchsorted, unique, intersect1d, union1d, setdiff1d and setxor1d with
wrappers that run Kwker for the arrays it handles and NumPy's own function for everything else; uninstall() puts
NumPy's back, and with kwker.numpy_ops.accelerated(): does both around a block.
import kwker.numpy_ops as knp
knp.install() # every later np.sort(...) / np.argsort(...) / ... call in the process
with knp.accelerated(): ... # only inside the block
Results are NumPy's wherever NumPy defines them: sorted values, stable permutations (kind="stable" / "mergesort", stable=True), lexsort, searchsorted positions, unique values, first indices, inverse and counts, the set operations, and the k-th element of a partition with smaller keys before it and larger after. NaNs sort last, as in NumPy. Where NumPy leaves the result open, Kwker may give another valid answer: argsort's default (unstable) kind returns the stable permutation, so equal keys keep their index order; partition / argpartition arrange the keys on either side of kth in their own order (both also differ between NumPy versions); -0.0 sorts before +0.0, which NumPy treats as equal keys.
Kwker runs a call when the input is a plain numpy.ndarray (or a list / tuple converted to one) of a supported dtype -
bool, int8-int64, uint8-uint64, float16, float32, float64, datetime64, timedelta64, native byte order - with at least
min_size() elements (KWKER_NUMPY_MIN, default 1024; below it NumPy's call is quicker; lexsort from 16 x min_size() keys
over all its columns), and the arguments are ones
Kwker implements (no structured order=, no sorter=, a single kth, 1-D keys for lexsort and the set operations). Anything
else - masked arrays and other subclasses, object / string / complex dtypes, axis forms not covered - goes to NumPy
unchanged. Only calls through the numpy namespace are replaced: ndarray methods (a.sort(), a.argsort()) and functions
imported by name before install() (from numpy import sort) keep NumPy's code. One thread, as NumPy.
class numpy_ops.accelerated() Page
A with block that installs Kwker's NumPy functions for its duration and restores the previous state after it.
numpy_ops.accelerated.__enter__(self)
numpy_ops.FUNCTIONS Page
numpy_ops.FUNCTIONS = ('sort', 'argsort', 'partition', 'argpartition', 'lexsort', 'searchsorted', 'unique', 'intersect1d', 'union1d', 'setdiff1d', 'setxor1d')
numpy_ops.install(on=True) Page
Make NumPy use Kwker for sort, argsort, partition, argpartition, lexsort, searchsorted, unique and the set operations, with no code changes.
Applies to every later call through the numpy namespace. Arrays smaller than min_size() stay on NumPy.
Arguments
on: True installs, False restores NumPy's own functions.
Returns
The previous state: True when it was installed before the call.
Example
import numpy as np
import kwker
import kwker.numpy_ops
kwker.numpy_ops.install()
print(np.sort(np.array([3, 1, 2])))
kwker.numpy_ops.install(False)
[1 2 3]
Examples: Quickstart: NumPy: Switch it on
numpy_ops.installed() Page
Return True while install() is in effect.
numpy_ops.min_size() Page
Return the smallest array size (in elements) Kwker takes over from NumPy; smaller arrays stay on NumPy.
numpy_ops.routes() Page
The calls install() takes over from NumPy, as kwker.contract routes: NumPy's function, Kwker's, and the promise.
Returns
A list of kwker.contract.Route - every function in FUNCTIONS, argsort / unique also with their other forms. Kwker's side takes arrays of every size (min_size() set to 1 for the call).
Example
import numpy as np
import kwker
import kwker.contract, kwker.numpy_ops
print(kwker.contract.run(kwker.numpy_ops.routes(), seconds=1.0)["ok"]) # True
True
Examples: Evaluate Kwker on your machine: 3. Run your own program both ways
numpy_ops.set_min_size(n) Page
Set the smallest array size Kwker takes over (0: every size); return the previous value.
Arguments
n
numpy_ops.uninstall() Page
Restore NumPy's own functions (the same as install(False)).
DataFrames (pandas, Polars, pyarrow)
import kwker.frame
DataFrame sorting with Kwker: pandas DataFrames, Polars DataFrames and pyarrow Tables keep their own types and ordering rules; the row order comes from Kwker (one argsort, or lexsort over the key columns), then one take of the rows.
import kwker.frame as sf
sf.sort(df, ["region", "price"], descending=[False, True]) # the same type back
sf.top_k(df, "price", 100, descending=True) # ORDER BY price DESC LIMIT 100
sf.argsort(df, "price") # the row order (int64)
Each library's missing-value rules are kept unless nulls_last says otherwise:
-
pandas: NaN, None, NaT and pd.NA are missing; missing rows last by default in both directions (na_position). Categoricals sort by their category order.
-
Polars: nulls first by default (nulls_last=False); NaN is a value above every number (last ascending, first descending). Categoricals and strings sort lexically, Enums by their category order.
-
pyarrow: nulls last by default; NaN after every number in both directions (Arrow's sort_indices).
Supported key columns: integers, floats, booleans, datetimes / durations / dates, strings, binary, categoricals / dictionaries / Enums, and any other Arrow type through pyarrow's dense rank. Stable: equal keys keep their row order (pandas' default sort_values kind is not stable; compare with kind="stable").
frame.argsort(df, by, descending=False, nulls_last=None, threads=1) Page
Return the row order that sorts a DataFrame by one or more columns.
Arguments
df: A pandas DataFrame, Polars DataFrame or pyarrow Table.by: A column name, or a list of names (the first decides first).descending: One flag, or one per column.nulls_last: One flag, or one per column; None uses the library's own default.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).
Returns
An int64 NumPy array of row positions. Rows with equal keys keep their order.
Example
import numpy as np
import kwker
import pandas as pd
import kwker.frame
df = pd.DataFrame({"sales": [3, 5, 1]})
print(kwker.frame.argsort(df, "sales"))
[2 0 1]
frame.group_by(df, by, aggs, dropna=None, threads=1, rules=None) Page
Group a DataFrame by key columns and aggregate, like SQL GROUP BY ... ORDER BY the keys.
Returns one row per distinct key combination, in sorted key order, with the same results the input's own library (pandas, Polars or pyarrow) gives.
Arguments
df: A pandas DataFrame, Polars DataFrame or pyarrow Table.by: A column name, or a list of names.aggs: The aggregations: {output name: (column, op)} or [(column, op), ...] (named "<column>_<op>"). ops: "sum", "mean", "min", "max", "count" (values that are not missing), "size" (rows; {name: "size"}), "first", "last", "median".dropna: Drop the group of missing keys; None uses the library's own default (pandas drops it, Polars and pyarrow keep it).threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).rules: "duckdb" applies DuckDB's rules instead of the input library's (used by kwker.duck).
Returns
A new frame of the same type: the key columns, then one column per aggregation. pandas results have a plain RangeIndex, as groupby(..., as_index=False) gives.
Example
import numpy as np
import kwker
import pandas as pd
import kwker.frame
df = pd.DataFrame({"city": ["b", "a", "b"], "sales": [3, 5, 1]})
print(kwker.frame.group_by(df, "city", {"total": ("sales", "sum")}))
city total 0 a 5 1 b 4
Notes
- Missing values follow each library's rules: pandas skips NaN; Polars and pyarrow skip nulls and treat NaN as a value.
- Integer sums are int64 (uint64 for unsigned); means and medians are float64; decimals up to 18 digits sum exactly.
- Float sums are added per chunk of rows, so the last bits can differ from the library's own.
Errors
ValueError: no key columnsValueError: duplicate output column names in ...ValueError: rules is None or 'duckdb' (pyarrow tables)TypeError: ... of the temporal column ...
frame.routes() Page
The calls kwker.frame stands in for, as kwker.contract routes: each library's own stable sort and group-by, and the kwker.frame call that gives the same rows.
Returns
A list of kwker.contract.Route for the libraries installed here: pandas sort_values(kind="stable") / head(k) / groupby, Polars sort(maintain_order=True), pyarrow sort_indices. Each route's inputs are random frames of 1-3 key columns (integers, floats with NaN / -0.0 / inf, strings, booleans, missing values).
Example
import numpy as np
import kwker
import kwker.contract, kwker.frame
print(kwker.contract.run(kwker.frame.routes(), seconds=1.0)["ok"]) # True
True
frame.sort(df, by, descending=False, nulls_last=None, threads=1) Page
Sort a DataFrame's rows by one or more columns; works with pandas, Polars and pyarrow tables.
Arguments
df: A pandas DataFrame, Polars DataFrame or pyarrow Table.by: A column name, or a list of names (the first decides first).descending: One flag, or one per column.nulls_last: One flag, or one per column; None uses the library's own default.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).
Returns
A new frame of the same type with the rows in order (pandas keeps each row's index label).
Example
import numpy as np
import kwker
import pandas as pd
import kwker.frame
df = pd.DataFrame({"city": ["b", "a", "b"], "sales": [3, 5, 1]})
print(kwker.frame.sort(df, ["city", "sales"]))
city sales 1 a 5 2 b 1 0 b 3
frame.top_k(df, by, k, descending=False, nulls_last=None) Page
Return the first k rows of a sorted DataFrame, like ORDER BY ... LIMIT k, without sorting every row.
Arguments
df: A pandas DataFrame, Polars DataFrame or pyarrow Table.by: A column name, or a list of names.k: How many rows.descending: One flag, or one per column.nulls_last: One flag, or one per column; None uses the library's own default.
Returns
A new frame of the same type with up to k rows, in order.
Example
import numpy as np
import kwker
import pandas as pd
import kwker.frame
df = pd.DataFrame({"name": ["a", "b", "c"], "score": [7, 9, 4]})
print(kwker.frame.top_k(df, "score", 2, descending=True))
name score 1 b 9 0 a 7
Errors
ValueError: k must be >= 0
DuckDB
import kwker.duck
Kwker for DuckDB: ORDER BY, ORDER BY ... LIMIT and GROUP BY ... ORDER BY computed by Kwker over a DuckDB relation's (or a pyarrow Table's) Arrow data with DuckDB's rules - NaN is the largest value, NULLs last unless asked, NULL keys a group - and the result back as a DuckDB relation (arrow=True: the pyarrow Table).
import duckdb, kwker.duck as sd
con = duckdb.connect()
rel = con.sql("SELECT * FROM 'events.parquet'")
sd.sort(rel, ["user", "ts"], descending=[False, True], connection=con)
sd.group_by(rel, ["user"], {"n": "size", "total": ("amount", "sum"), "p50": ("amount", "median")}, connection=con)
DuckDB's own sort is not stable: rows with equal keys may come in any order there; here they keep their input order.
duck.argsort(rel, by, descending=False, nulls_last=True, threads=1) Page
Return the row order of ORDER BY by, with DuckDB's ordering rules.
Arguments
rel: A DuckDB relation.by: A column name, or a list of names.descending: One flag, or one per column.nulls_last: True (the default) puts NULLs last in both directions.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).
Returns
An int64 NumPy array of row positions. Rows with equal keys keep their order.
Notes
- As in DuckDB, NaN sorts above every number.
duck.group_by(rel, by, aggs, threads=1, arrow=False, connection=None) Page
GROUP BY by ORDER BY by for a DuckDB relation, with DuckDB's results.
Arguments
rel: A DuckDB relation.by: A column name, or a list of names.aggs: The aggregations, as kwker.frame.group_by takes them: sum, mean, min, max, count, size, first, last, median.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).arrow: True returns a pyarrow Table.connection: The DuckDB connection for the result.
Returns
A relation (or pyarrow Table): the key columns, then one column per aggregation.
Example
import numpy as np
import kwker
import duckdb
import kwker.duck
rel = duckdb.sql("SELECT * FROM (VALUES ('b', 3), ('a', 5), ('b', 1)) t(city, sales)")
print(kwker.duck.group_by(rel, "city", {"total": ("sales", "sum")}).fetchall())
[('a', 5), ('b', 4)]
Notes
- NULL keys form one group, sorted last; NaN keys form one group above every number. Decimal results are typed as DuckDB types them.
duck.routes() Page
The SQL kwker.duck computes, as kwker.contract routes: DuckDB's own ORDER BY, ORDER BY ... LIMIT and GROUP BY ... ORDER BY against the kwker.duck call.
Returns
A list of kwker.contract.Route (empty without duckdb and pyarrow). Inputs are random tables of 1-3 key columns (integers, doubles with NaN / -0.0 / inf, strings, NULLs) and an integer value column.
Example
import numpy as np
import kwker
import kwker.contract, kwker.duck
print(kwker.contract.run(kwker.duck.routes(), seconds=1.0)["ok"]) # True
True
duck.sort(rel, by, descending=False, nulls_last=True, threads=1, arrow=False, connection=None) Page
Sort a DuckDB relation's rows, like ORDER BY by, with DuckDB's ordering rules.
Arguments
rel: A DuckDB relation (or anything kwker.duck reads as one).by: A column name, or a list of names.descending: One flag, or one per column.nulls_last: True (the default) puts NULLs last in both directions, as DuckDB does.threads: How many threads to use: 1 (the default) uses the calling thread only; 0 uses the default count (set_default_threads).arrow: True returns a pyarrow Table instead of a relation.connection: The DuckDB connection for the result (default: the relation's own).
Returns
A relation (or pyarrow Table) with the rows in order.
Example
import numpy as np
import kwker
import duckdb
import kwker.duck
rel = duckdb.sql("SELECT * FROM (VALUES (3), (1), (2)) t(x)")
print(kwker.duck.sort(rel, "x").fetchall())
[(1,), (2,), (3,)]
duck.top_k(rel, by, k, descending=False, nulls_last=True, arrow=False, connection=None) Page
ORDER BY by LIMIT k for a DuckDB relation, without sorting every row.
Arguments
rel: A DuckDB relation.by: A column name, or a list of names.k: How many rows.descending: One flag, or one per column.nulls_last: True (the default) puts NULLs last.arrow: True returns a pyarrow Table.connection: The DuckDB connection for the result.
Returns
A relation (or pyarrow Table) with up to k rows, in order.
SciPy drop-in
import kwker.scipy (needs scipy)
SciPy, faster, with the same answers: kwker.scipy.install() routes SciPy's sparse-matrix conversions through Kwker.
import kwker.scipy
kwker.scipy.install() # existing SciPy code below runs unchanged
import scipy.sparse as sp
A = sp.coo_array((vals, (rows, cols)), shape=(n, m)).tocsr() # duplicates summed, as before
print(kwker.scipy.report()) # which calls Kwker ran, which went back to SciPy, and why
Every routed call returns SciPy's answer bit for bit: the same arrays, index data types and format flags. A call goes back to SciPy when Kwker cannot promise that - float matrices whose duplicate entries SciPy sums in an order it does not define, value types Kwker's sparse kernels do not take (complex, bool, 8- and 16-bit), other dimensions than 2 - and when SciPy is faster: small matrices (_SMALL). Routed now: coo tocsr() / tocsc() / sum_duplicates(), csr tocsc(), csc tocsr(), csr / csc sort_indices() and sum_duplicates(), scipy.stats.rankdata and the private _rankdata (and so the tests built on them: spearmanr, kruskal, mannwhitneyu, wilcoxon, ...), scipy.ndimage.median_filter / rank_filter / percentile_filter (3 x 3 / 5 x 5 / 7 x 7 windows), scipy.signal.medfilt2d, scipy.stats.wasserstein_distance / energy_distance.
scipy.install(on=True) Page
Route SciPy's sparse conversions, rank statistics, window filters and distances through Kwker, with the same results.
Arguments
on: True installs; False restores SciPy's own functions and methods.
Returns
The previous state: True when Kwker was installed before the call.
Example
import numpy as np
import kwker
import kwker.scipy
kwker.scipy.install()
import scipy.sparse as sp
print(sp.coo_array(([1, 2], ([0, 1], [1, 0])), shape=(2, 2)).tocsr().toarray())
[[0 1] [2 0]]
scipy.installed() Page
Return True while install() is in effect.
scipy.report() -> dict Page
What kwker.scipy ran: per call, the calls Kwker ran and the ones it left to SciPy, with the reasons.
Returns
A dict: "installed", "scipy" (the version), "self_check" ({route: "ok" or the difference found}, None before install() ran it) and "calls": {call: {"kwker": n, "scipy": n, "reasons": {reason: n}, "kwker_s" / "scipy_s": the calls' wall time in seconds on each side}}.
scipy.routes() Page
The SciPy calls install() takes over, as kwker.contract routes (EXACT: bit for bit, index dtypes and flags).
Returns
A list of kwker.contract.Route; inputs are random sparse matrices (duplicates, empty rows, every routed value type).
Example
import numpy as np
import kwker
import kwker.contract, kwker.scipy
print(kwker.contract.run(kwker.scipy.routes(), seconds=1.0)["ok"]) # True
True
scipy.self_check(force=False) -> dict Page
Check every call install() takes over against SciPy's own on a few small inputs, once per build and CPU.
Arguments
force: True runs the probes again instead of reading the cached result.
Returns
A dict: route name -> None when Kwker's result is SciPy's, else the first difference found. install() leaves the calls of a differing route with SciPy.
Notes
- The result is cached in Kwker's cache directory (selfcheck-scipy-*.json); KWKER_SELF_CHECK=0 skips the probes (every route counted as agreeing), KWKER_SELF_CHECK=force runs them again.
scipy.uninstall() Page
Restore SciPy's own methods (the same as install(False)).
scipy.VERSIONS Page
scipy.VERSIONS = ('1.13', '1.14', '1.15', '1.16', '1.17')
PyTorch operators and drop-in kernels
import kwker.torch_ops (needs torch)
Kwker operators for PyTorch CPU tensors (namespace kwker):
torch.ops.kwker.sort(x, dim=-1, descending=False) -> (values, indices) like torch.sort(stable=True)
torch.ops.kwker.sort_values(x, dim=-1, descending=False) -> values (no indices)
torch.ops.kwker.argsort(x, dim=-1, descending=False) -> indices (int64) like torch.argsort(stable=True)
torch.ops.kwker.topk(x, k, dim=-1, largest=True, sorted=True) -> (values, indices) like torch.topk
torch.ops.kwker.kthvalue(x, k, dim=-1, keepdim=False) / median(x, dim=-1, keepdim=False) like torch's
torch.ops.kwker.median_pool2d(x, kernel_size=3, padding="reflect") -> 3 x 3 (or 5 x 5) window medians of the last two
dimensions, stride 1: the median of every window of F.pad(x, (1, 1, 1, 1), mode=padding) - the
unfold + torch.median MedianPool2d in one pass (padding "reflect", "replicate", "zeros", "valid": no padding)
torch.ops.kwker.kwta(x, k) -> k-winners-take-all per sample (dim 0): x * (x >= the k-th largest of its sample),
the thresholds by selection (no sorted top-k), one masking pass - topk(x.flatten(1), k)[0][:, -1:] + the mask
torch.ops.kwker.topk_reduce(x, k, dim=-1, reduction="mean") -> the mean (or "sum") of the k largest keys along
dim (OHEM: the k hardest losses) - topk(x, k, dim).values.mean(dim) without the sorted values or indices: the
k-th key by selection, one summing pass; backward: grad / k at the selected keys (ties: the first by index)
The compiled operators also take torch's out= form (torch.ops.kwker.sort(x, dim, descending, values=v, indices=i), sort_values(..., out=o), argsort(..., out=o), topk(..., values=v, indices=i): outputs resized to the result's shape, written in place when contiguous and dim is the last dimension; sort_values(x, out=x) sorts x in place; not differentiable).
import kwker.torch_ops registers them; the same functions are this module's sort / sort_values /
argsort / topk. COMPILED says which implementation runs: the compiled library kwker/_torch_ops.so
(TORCH_LIBRARY: C++ CPU and Meta kernels over the C library's batched row kernels, one call per block of rows,
blocks spread over torch's intra-op threads - usable from C++ / torch.export without Python) when it is built, else
the Python-registered fallback (_torch_py.py: torch.library.custom_op through ctypes, a Python loop per row for argsort
and topk). KWKER_TORCH_PY=1 forces the fallback. Both carry autograd (sort, sort_values, topk: the gradient
scattered back through the indices), vmap rules and fake kernels (torch.compile).
Order as torch's: NaNs are the largest values (last ascending, first descending / largest), ties keep their index order (the stable order; topk's ties too). dtypes: uint8, int8, int16, int32, int64, float16, float32, float64, bfloat16, and uint16 / uint32 / uint64 where torch has them (the compiled operators also float8_e5m2 / e4m3fn). -0.0 sorts before +0.0.
class torch_ops.accelerated(release=True) Page
A with block that installs Kwker's kernels for its duration and restores the previous state after it.
Arguments
release: Also free the cached scratch memory.
Example
import numpy as np
import kwker
import torch
import kwker.torch_ops
with kwker.torch_ops.accelerated():
print(torch.topk(torch.tensor([5.0, 9.0, 1.0]), 2).values)
tensor([9., 5.])
torch_ops.accelerated.__enter__(self)
torch_ops.argsort(*args: _P.args, **kwargs: _P.kwargs) -> ~_T Page
Documented in the module overview above.
torch_ops.built_for_this_torch() Page
Return True when the compiled PyTorch extensions match this PyTorch's major.minor version.
A mismatch warns and keeps only the Python-registered operators. A development tree without a build stamp returns True.
torch_ops.COMPILED Page
torch_ops.COMPILED = True
torch_ops.install(on=True, release=True) Page
Make PyTorch run Kwker for its sorting-style CPU operations, with no code changes.
Covers torch.sort, argsort, msort, topk, kthvalue, median, nanmedian, unique, searchsorted, bucketize, quantile, nanquantile, sparse coalesce and the weight gradients of nn.Embedding / nn.EmbeddingBag. It applies to every caller: your code, libraries, TorchScript and torch.compile. Results are identical to PyTorch's.
Arguments
on: True installs Kwker's kernels; False restores PyTorch's own.release: With on=False, True also frees the memory Kwker's tensor allocator keeps.
Returns
The previous state: True when Kwker was installed before the call.
Example
import numpy as np
import kwker
import torch
import kwker.torch_ops
kwker.torch_ops.install()
print(torch.sort(torch.tensor([3, 1, 2])).values)
kwker.torch_ops.install(False)
tensor([1, 2, 3])
Notes
- The first install() in a process checks each kernel against PyTorch's own (see self_check); a kernel that differs stays off. The check is cached on disk.
- Cases Kwker does not cover (bool and complex tensors, empty tensors, unusual arguments) run PyTorch's kernel.
- KWKER_DISABLE=1 in the environment makes install() do nothing.
Errors
RuntimeError: needs the compiled operators (kwker/_torch_ops.so)
Examples: Add Kwker to your project: PyTorch models, Quickstart: PyTorch: Switch it on
torch_ops.installed() Page
Return True while install() is in effect.
torch_ops.kthvalue(*args: _P.args, **kwargs: _P.kwargs) -> ~_T Page
Documented in the module overview above.
torch_ops.kwta(*args: _P.args, **kwargs: _P.kwargs) -> ~_T Page
Documented in the module overview above.
torch_ops.median(*args: _P.args, **kwargs: _P.kwargs) -> ~_T Page
Documented in the module overview above.
torch_ops.median_pool2d(*args: _P.args, **kwargs: _P.kwargs) -> ~_T Page
Documented in the module overview above.
torch_ops.nanquantile(x, q, dim=None, keepdim=False, *, interpolation='linear') Page
torch.nanquantile (quantile ignoring NaN values), computed with Kwker's selection; the results are identical to PyTorch's.
Arguments
x: The input tensor.q: The quantile (a float or a 1-D tensor of floats, each from 0 to 1).dim: The dimension to reduce; None reduces the flattened tensor.keepdim: True keeps the reduced dimension with size 1.interpolation: "linear" (the default), "lower", "higher", "midpoint" or "nearest".
Returns
A tensor of quantiles, as torch.nanquantile returns it (NaN where a slice holds only NaN values).
torch_ops.quantile(x, q, dim=None, keepdim=False, *, interpolation='linear') Page
torch.quantile, computed with Kwker's selection; the results are identical to PyTorch's.
Arguments
x: The input tensor.q: The quantile (a float or a 1-D tensor of floats, each from 0 to 1).dim: The dimension to reduce; None reduces the flattened tensor.keepdim: True keeps the reduced dimension with size 1.interpolation: "linear" (the default), "lower", "higher", "midpoint" or "nearest".
Returns
A tensor of quantiles, as torch.quantile returns it.
torch_ops.release_cached_memory() Page
Free the memory blocks Kwker's tensor allocator keeps for reuse (tensors still in use are not affected).
torch_ops.routes() Page
The PyTorch calls install() takes over, as kwker.contract routes: torch's own CPU kernel, Kwker's, and the promise.
Returns
A list of kwker.contract.Route. The host side runs with every Kwker kernel off on the calling thread (torch's own kernels, nested calls included); the Kwker side runs with install() on - routes() installs it.
Example
import numpy as np
import kwker
import kwker.contract, kwker.torch_ops
print(kwker.contract.run(kwker.torch_ops.routes(), seconds=1.0)["ok"]) # True
True
torch_ops.self_check(force=False, verbose=False) Page
Check every Kwker kernel against this PyTorch's own and switch off any that differs; return the results.
install() runs this once per process and caches the result on disk (per PyTorch build, CPU and library build), so it costs nothing after the first run. Each kernel runs a small probe through both implementations; the results must match bit for bit.
Arguments
force: True runs the probes now, ignoring the cache.verbose: True prints each result.
Returns
A dict {kernel name: status}: "ok", "differs" (switched off), "error: ..." (switched off) or "not reached" (the probe never reached Kwker's kernel; left on). Entries named "transforms ..." cover the torchvision transforms.
Notes
- KWKER_SELF_CHECK=0 skips the check; KWKER_SELF_CHECK=force reruns it.
- A kernel that differs (another PyTorch release, for example) stays off with a RuntimeWarning naming it.
torch_ops.sort(*args: _P.args, **kwargs: _P.kwargs) -> ~_T Page
Documented in the module overview above.
torch_ops.sort_values(*args: _P.args, **kwargs: _P.kwargs) -> ~_T Page
Documented in the module overview above.
torch_ops.topk(*args: _P.args, **kwargs: _P.kwargs) -> ~_T Page
Documented in the module overview above.
torch_ops.topk_reduce(x, k, dim=-1, reduction='mean') Page
The mean or sum of the k largest values along a dimension, in one call: topk(x, k, dim).values.mean(dim).
Used for hard-example losses (OHEM), where only the k worst samples count.
Arguments
x: A float tensor.k: How many of the largest values to take.dim: The dimension to reduce (the last by default).reduction: "mean" (the default) or "sum".
Returns
A tensor with dim removed. Gradients flow to the k selected values.
Notes
- NaN counts as the largest value; equal values are taken in index order.
torch_ops.uninstall(release=True) Page
Restore PyTorch's own CPU kernels (the same as install(False)).
Arguments
release: Also free the cached scratch memory.
torch_ops.unique(x, sorted=True, return_inverse=False, return_counts=False, dim=None) Page
The distinct values of a tensor, like torch.unique, with the same results.
Arguments
x: A tensor.sorted: Kept for torch.unique's signature; the values always come back in ascending order.return_inverse: True also returns, for every element of x, the position of its value in the result.return_counts: True also returns how often each distinct value occurs.dim: None (the default) works on the flattened tensor; a dimension falls back to torch.unique.
Returns
The distinct values as a 1-D tensor; with return_inverse and / or return_counts, a tuple (values, inverse, counts) of the parts asked for, as torch.unique returns them.
PyTorch CPU backend
import kwker.cpu_backend (needs torch)
Kwker's CPU execution backend for PyTorch: faster linears, attention and convolutions, used by torch.compile(backend="kwker") and by the precision settings below.
On Intel CPUs with AMX (Sapphire Rapids and later), float32 linears follow torch.set_float32_matmul_precision:
-
"highest" (PyTorch's default): torch's own float32 matrix multiplications, unchanged.
-
"high": Kwker's AMX kernels at about float32 accuracy (each value is split into two bfloat16 parts).
-
"medium": bfloat16 products with float32 sums, as PyTorch itself does for "medium" - faster, slightly less accurate.
KWKER_AMX=0 turns the AMX kernels off.
class cpu_backend.AccuracyError(report) Page
Raised by enable_mode when the model's outputs exceed the accuracy limits; .report holds the AccuracyReport.
Arguments
class cpu_backend.AccuracyReport(mode, compiled, outputs, tol, seconds, active=True) Page
The result of check_accuracy: per output its relative error, largest absolute error, cosine similarity and top-1 agreement, plus ok (all within the limits in tol) and active (whether the mode's kernels actually ran here).
Arguments
modecompiledoutputstolsecondsactive
cpu_backend.AccuracyReport.summary(self) -> str
One line: the mode, relative error, cosine similarity, top-1 agreement, the output count and the verdict.
cpu_backend.amx() -> bool Page
Return True when this CPU has AMX (Intel Xeon Sapphire Rapids and later) and Kwker's AMX kernels run.
cpu_backend.available() -> bool Page
Return True when Kwker's PyTorch CPU kernels are loaded and this CPU can run them (AVX-512 or AMX).
cpu_backend.check_accuracy(model, *args, mode='int8', compile=False, tol=None, timed=False, **kwargs) -> 'AccuracyReport' Page
Measure how much a faster precision mode changes your model's outputs, before you turn it on.
Runs the model once in full float32 and once under mode, and compares every floating-point output. Your settings are restored afterwards.
Arguments
model: The PyTorch model (any callable).*args: The model's example inputs.mode: "int8" (the default), "int4", "medium" (bf16), "high" (bf16x3, close to float32) or "highest".compile: True measures torch.compile(model, backend="kwker") instead of eager execution.tol: Optional limits: a dict of "rel", "cos" and "top1" (default: TOLERANCES[mode]).timed: True also times one call of each after a warm-up (in .seconds).**kwargs: The model's keyword inputs.
Returns
An AccuracyReport: per output the relative error, the largest absolute error, the cosine similarity and (for 2-D and larger outputs) the top-1 agreement, plus ok against the limits.
Example
import numpy as np
import kwker
import torch
import kwker.cpu_backend as cb
model = torch.nn.Linear(64, 8)
r = cb.check_accuracy(model, torch.randn(4, 64), mode="medium")
print(type(r).__name__)
AccuracyReport
Notes
- Runs under torch.no_grad(). With compile=True, torch._dynamo's caches are reset afterwards.
Errors
ValueError: unknown mode ... (one of ...)RuntimeError: the outputs differ in number or shape (... vs ...)
cpu_backend.conv_mode() -> int Page
Return the convolution kernel mode as a number, from torch's float32 convolution precision: 3 ("tf32", run as bf16x3), 1 ("bf16"), 0 (exact float32), or 8 while set_int8 is on.
cpu_backend.decoder_available() -> bool Page
Return True when KwkDecoder and kwker.serve can run on this CPU.
They need AVX-512 VNNI, or AVX2 with FMA and F16C (Intel Core 12th generation and later, AMD Zen 2 and later). Both give the same results.
cpu_backend.enable_mode(mode, model=None, *args, tol=None, compile=False, **kwargs) Page
Turn on a precision mode for later calls; with a model, check its accuracy first.
Arguments
mode: "highest", "high", "medium", "int8" or "int4" (see check_accuracy).model: Optional model to check first; its example inputs follow as *args / **kwargs.*argstol: The accuracy limits for the check (see check_accuracy).compile: True checks the compiled model.**kwargs
Returns
The AccuracyReport of the check, or None without a model.
Notes
- When the outputs exceed the limits, AccuracyError is raised and the previous settings stay.
Errors
cpu_backend.install(on: bool = True) -> bool Page
Use Kwker's matrix-multiply kernels as PyTorch's own for float32 linear layers (called by kwker.torch_ops.install()).
On AMX CPUs this applies while torch's float32 matmul precision is "high" or "medium"; on other CPUs bfloat16 linear layers use Kwker's bf16 kernels. Shapes the kernels do not cover run PyTorch's.
Arguments
on: True installs, False restores PyTorch's kernels.
Returns
True when the kernels are installed (never while KWKER_DISABLE=1 is set).
cpu_backend.matmul_mode() -> int Page
Return the matrix-multiply kernel mode as a number: 8 (int8), 4 (int4), 3 ("high", bf16x3), 1 ("medium", bf16) or 0 ("highest").
cpu_backend.mode() -> str Page
Return the precision mode in force: "int8" or "int4" when set, else "highest", "high" or "medium" (torch's float32 matmul precision).
cpu_backend.MODES Page
cpu_backend.MODES = ('highest', 'high', 'medium', 'int8', 'int4')
cpu_backend.set_int4(on: bool = True) -> bool Page
Opt in to 4-bit weights for float32 linear layers: under a third of bf16's memory, for fast token-by-token decoding.
Errors are larger than int8's (around 5% of the values); check your model's outputs first. Convolutions run in int8 meanwhile. Needs AVX-512 VNNI.
Arguments
on: True turns int4 on, False off.
Returns
The previous state.
cpu_backend.set_int8(on: bool = True) -> bool Page
Opt in to int8 linear layers (weights and activations in 8 bits) for float32 models: faster, slightly less exact.
Errors are around 1% of the values; check your model with check_accuracy first. Applies to eager calls under install() and to torch.compile(backend="kwker") graphs compiled while it is on.
Arguments
on: True turns int8 on, False off.
Returns
The previous state.
cpu_backend.TOLERANCES Page
cpu_backend.TOLERANCES = {'highest': {'rel': 1e-06, 'cos': 0.999999}, 'high': {'rel': 0.0001, 'cos': 0.99999}, 'medium': {'rel': 0.03, 'cos': 0.999}, 'int8': {'rel': 0.08, 'cos': 0.995}, 'int4': {'rel': 0.2, 'cos': 0.98}}
KwkCNN
import kwker.cnn (needs torch)
KwkCNN: run a torchvision-style image model (ResNet, MobileNet, EfficientNet, RegNet, ConvNeXt and similar) as one fast native call on the CPU, in float32 or int8.
cnn = kwker.cnn.KwkCNN(model) # model in eval mode; built once
y = cnn(x) # x: [batch, 3, H, W] float32, the H and W it was built for
runner_unsupported(model) tells you why a model cannot run.
class cnn.KwkCNN(model, example_shape=(1, 3, 224, 224), int8=False, calib=None) Page
A PyTorch image model's inference as one native call: the same outputs as the model (float32), or faster int8.
Build it once from the model; then call it like the model. Any batch size works; the image size is the one it was built for.
Arguments
model: A torchvision-style image model in eval mode, float32.example_shape: The input shape to build for: (batch, 3, H, W); (1, 3, 224, 224) by default.int8: True runs int8 (AVX-512 VNNI, or AMX where present): faster, slightly less exact.calib: With int8, a float32 batch of typical images [N, 3, H, W] for calibration (default: 8 random images).
Example
import numpy as np
import kwker
import torch, torchvision
import kwker.cnn
model = torchvision.models.resnet18().eval()
cnn = kwker.cnn.KwkCNN(model)
print(cnn(torch.randn(1, 3, 224, 224)).shape)
torch.Size([1, 1000])
Notes
- In int8 mode the squeeze-and-excitation blocks and the classifier stay float32.
cnn.KwkCNN.__call__(self, x)
The model's output for x ([B, 3, H, W] float32 at the H, W the runner was built for; any batch B).
Arguments
x: The input tensor.
cnn.runner_unsupported(model, example_shape=(1, 3, 224, 224)) Page
Return None when KwkCNN can run the model, else a short reason why not.
Arguments
model: The PyTorch image model.example_shape: The input shape to plan for: (batch, 3, H, W).
KwkEncoder
import kwker.encode (needs torch)
KwkEncoder: run a BERT-family text encoder (BERT, RoBERTa, XLM-RoBERTa, DistilBERT, MPNet, ModernBERT, MiniLM, E5, BGE, GTE and sentence-transformers models built on them) as one fast native call per batch.
enc = kwker.encode.KwkEncoder(model) # packs the weights once
h = enc(input_ids, attention_mask) # last_hidden_state: [batch, length, hidden] float32
enc.install() # or: the model itself (any task head) runs on it from now on
It computes in bf16 with float32 sums, like torch's "medium" precision. It runs on x86 CPUs with AVX2 or AVX-512. With AMX (Intel Sapphire Rapids or later) both weights and activations are bf16; without AMX the weights are bf16 and the activations float32. precision="float32" keeps the model's float32 weights and float32 attention instead (nothing rounded). runner_unsupported(model) tells you why a model cannot run.
class encode.KwkEncoder(model, unpad: bool = True, int8: bool = False, calib=None, smooth: float = 0.5, precision: str | None = None) Page
A BERT-family encoder's forward pass as one native call: call it with token ids, get the last hidden state.
Call it as enc(input_ids, attention_mask=None, token_type_ids=None); it returns last_hidden_state [batch, length, hidden] (float32). The weights are packed when it is built, so later changes to the model are not seen.
Arguments
model: A Hugging Face BERT-family model (float32, float16 or bfloat16 weights).unpad: True (the default) skips the padding tokens entirely: faster, the same results; padded output rows are zero.int8: True runs the linear layers in int8 (AMX-INT8, AVX-512 VNNI or AVX-VNNI): faster, less exact. Check accuracy on your data.calib: With int8, typical inputs (a tokenizer batch, a list of them, or token ids) to calibrate with (SmoothQuant).smooth: With calib, the SmoothQuant strength from 0 to 1 (0.5 by default).precision: "bf16" (the default: bf16 weights, float32 sums), "float32" (the model's own float32 weights and float32 attention: nothing rounded) or "int8" (the same as int8=True).
Example
import numpy as np
import kwker
import torch
from transformers import BertConfig, BertModel
import kwker.encode
model = BertModel(BertConfig(hidden_size=64, num_hidden_layers=2, num_attention_heads=2, intermediate_size=128)).eval()
if kwker.encode.runner_unsupported(model) is None:
print(kwker.encode.KwkEncoder(model)(torch.tensor([[101, 2023, 102]])).shape)
torch.Size([1, 3, 64])
Notes
- Without AMX the padding tokens are always skipped (unpad=False has no effect there).
- int8 on a CPU without AMX-INT8, AVX-512 VNNI or AVX-VNNI raises ValueError.
- ModernBERT runs in bf16 or float32 only (int8 raises ValueError).
- float16 / bfloat16 weights are read as they are (precision="float32" keeps them exactly); the output is float32, or the model's own dtype through install().
encode.KwkEncoder.__call__(self, input_ids, attention_mask=None, token_type_ids=None)
last_hidden_state [batch, length, hidden] float32 for token ids [batch, length] (or [length]); see the class.
Arguments
input_idsattention_masktoken_type_ids
encode.KwkEncoder.install(self, model=None)
Make the model itself run on this encoder: model(...) calls with input_ids go through KwkEncoder from now on.
Works for any task head on top (classification, token tagging, sentence embeddings). Calls it cannot serve (inputs_embeds, attention outputs, gradients and the like) run the model's original code.
Arguments
model: The model to route (default: the one the encoder was built from).
Returns
The model. uninstall() undoes it. With KWKER_DISABLE=1 set it stays the model's own code.
encode.KwkEncoder.positions(self, input_ids)
Return the position ids the model's embeddings use for input_ids (0 .. L - 1, or RoBERTa's padding-aware positions).
Arguments
input_ids
encode.KwkEncoder.set_timing(self, on: bool = True)
Turn per-phase timing on or off (read it with stats()).
Arguments
on
encode.KwkEncoder.stats(self) -> dict
Return the seconds spent per phase since the last stats() call (with set_timing(True)).
encode.KwkEncoder.uninstall(self, model=None)
Undo install(): the model runs its original code again.
Arguments
model
encode.runner_unsupported(model) -> str | None Page
Return None when KwkEncoder can run the model, else a short reason why not (for example "no AVX2 or AVX-512 on this CPU").
Arguments
model: A Hugging Face encoder model.
Hugging Face models by name: kwker.Decoder, kwker.Encoder
import kwker.hub (needs torch)
Hugging Face models by name: a Hub id in, a ready model out.
import kwker
llm = kwker.Decoder.from_pretrained("HuggingFaceTB/SmolLM2-135M-Instruct")
print(llm.chat("Explain sorting networks in two sentences."))
enc = kwker.Encoder.from_pretrained("BAAI/bge-small-en-v1.5")
vectors = enc.encode(["first text", "second text"]) # float32, pooled and normalized as the model specifies
Decoder runs a causal language model on KwkDecoder at the checkpoint's own precision by default (precision= picks another); architectures KwkDecoder does not cover run on transformers with Kwker's PyTorch kernels. Encoder runs a BERT-family embedding model on KwkEncoder and pools as its sentence-transformers configuration says. Both print nothing and download nothing beyond what transformers.from_pretrained downloads.
class hub.Decoder(model, tokenizer, precision='preserve', max_cache_len=4096) Page
A Hugging Face causal language model ready to generate text: KwkDecoder where it covers the architecture, else transformers with Kwker's PyTorch kernels.
Build it with Decoder.from_pretrained(name). Attributes: model (the transformers model), tokenizer, decoder (the KwkDecoder, or None on the fallback route), precision (what runs) and why (the fallback's reason, or None).
Arguments
modeltokenizerprecisionmax_cache_len
hub.Decoder.chat(self, messages, max_new_tokens=256, stream=False, system=None, schema=None, **sampling)
Answer a chat message through the model's chat template and return the reply.
Arguments
messages: The user's message as a string, or a list of {"role": ..., "content": ...} messages.max_new_tokens: The most tokens in the reply.stream: True returns an iterator of text pieces as they are generated.system: An optional system message put first.schema: A JSON Schema (dict) the reply must match as one JSON value, {} for any JSON value, or None for free text.**sampling: do_sample, temperature, top_k, top_p, ... as transformers' generate() takes them.
Returns
The reply's text, or with stream=True an iterator of its pieces.
Example
import numpy as np
import kwker
import json, kwker
llm = kwker.Decoder.from_pretrained("HuggingFaceTB/SmolLM2-135M-Instruct")
for piece in llm.chat("Name three sorting algorithms.", stream=True):
print(piece, end="")
city = {"type": "object", "properties": {"city": {"type": "string"}, "country": {"type": "string"}},
"required": ["city", "country"], "additionalProperties": False}
print(json.loads(llm.chat("Name a large city.", schema=city)))
Here are three sorting algorithms:
1. **Quick Sort Algorithm**: A divide-and-conquer algorithm that partitions the data into two halves, then recursively sorts the remaining half.
2. **Merge Sort Algorithm**: A divide-and-conquer algorithm that merges two sorted lists into a single sorted list.
3. **Heap Sort Algorithm**: A divide-and-conquer algorithm that uses a heap data structure to sort the data. It uses a heap to keep track of the elements to be sorted and to fill the heap as soon as it is full.{'city': 'New York', 'country': 'United States'}
Notes
- With schema, sampling takes do_sample, temperature, top_k, top_p, min_p, repetition_penalty and seed on KwkDecoder; the schema keywords enforced are listed in the structured output guide.
hub.Decoder.from_pretrained(cls, name, precision='preserve', max_cache_len=4096, revision=None, **model_kwargs)
Load a causal language model by Hub id (or local directory) and make it ready to generate.
Arguments
name: A Hugging Face model id ("HuggingFaceTB/SmolLM2-135M-Instruct") or a local model directory.precision: "preserve" (the default: the checkpoint's own weights, nothing rounded), "balanced" (int8 weights, about +0.1% perplexity), or "float32", "bf16", "int8", "int4" by name.max_cache_len: The longest prompt plus generated text, in tokens.revision: A Hub revision (branch, tag or commit).**model_kwargs: Passed to transformers' AutoModelForCausalLM.from_pretrained.
Returns
A Decoder.
Example
import numpy as np
import kwker
import kwker
llm = kwker.Decoder.from_pretrained("HuggingFaceTB/SmolLM2-135M-Instruct")
print(llm.chat("Say hello."))
Hello! I'm here to help you with any questions or issues you might have. What's on your mind?
Notes
- Raises MemoryError before loading when the weight files are larger than the memory left (KWKER_MEMORY_CHECK=0 skips the check).
hub.Decoder.generate(self, prompt, max_new_tokens=256, stream=False, schema=None, **sampling)
Continue a text (no chat template) and return the new text.
Arguments
prompt: The text to continue, or its token ids.max_new_tokens: The most tokens to generate.stream: True returns an iterator of text pieces as they are generated.schema: A JSON Schema (dict) the text must match as one JSON value, {} for any JSON value, or None for free text.**sampling: do_sample, temperature, top_k, top_p, ... as transformers' generate() takes them.
Returns
The generated text (without the prompt), or with stream=True an iterator of its pieces.
hub.Decoder.load_report(self) -> str
Return one line on what runs: the decoder and its weights, or the fallback and its reason.
class hub.Encoder(model, tokenizer, pooling='mean', normalize=False, max_length=512, int8=False) Page
A Hugging Face embedding model ready to encode texts: KwkEncoder where it covers the architecture, pooling and normalization as the model's sentence-transformers configuration says.
Build it with Encoder.from_pretrained(name). Attributes: model, tokenizer, encoder (the KwkEncoder, or None on the fallback route), pooling ("cls", "mean" or "max"), normalize (bool) and why (the fallback's reason, or None).
Arguments
modeltokenizerpoolingnormalizemax_lengthint8
hub.Encoder.encode(self, texts, batch_size=32)
Embed texts: one vector per text, pooled (and normalized) as the model specifies.
Arguments
texts: A string or a list of strings.batch_size: Texts per forward pass.
Returns
A float32 NumPy array [len(texts), hidden size] (one row for a single string).
hub.Encoder.from_pretrained(cls, name, pooling=None, normalize=None, max_length=None, int8=False, revision=None, **model_kwargs)
Load an embedding model by Hub id (or local directory) and make it ready to encode.
Arguments
name: A Hugging Face model id ("BAAI/bge-small-en-v1.5") or a local model directory.pooling: "cls", "mean" or "max"; None reads the model's sentence-transformers configuration (mean without one).normalize: True scales every vector to length 1; None reads the configuration (a Normalize module).max_length: The longest input in tokens; None uses the configuration's max_seq_length, else 512.int8: True runs the linear layers in int8 (faster, less exact - check retrieval quality on your data).revision: A Hub revision.**model_kwargs: Passed to transformers' AutoModel.from_pretrained.
Returns
An Encoder.
Example
import numpy as np
import kwker
import kwker
enc = kwker.Encoder.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
print(enc.encode(["a sentence"]).shape)
(1, 384)
Notes
- Raises MemoryError before loading when the weight files are larger than the memory left (KWKER_MEMORY_CHECK=0 skips the check).
ONNX Runtime provider
import kwker.onnx (needs onnxruntime)
Kwker as an ONNX Runtime execution provider.
Run an ONNX model you already have through ONNX Runtime with Kwker underneath: Kwker takes the models it runs faster (BERT-family encoders - BERT, RoBERTa, XLM-R and the MiniLM, BGE, E5, GTE embedding and reranking models as optimum, torch.onnx and transformers.js export them: float32, float16, int8 or 4-bit weights) and leaves every other model to ONNX Runtime's own providers, so adding it never breaks a session.
Example:
import onnxruntime as ort
import kwker.onnx
sess = ort.InferenceSession("model.onnx", sess_options=kwker.onnx.session_options())
hidden = sess.run(None, {"input_ids": ids, "attention_mask": mask})[0]
onnx.library_path() -> str Page
Return the path of the Kwker Runtime library that ONNX Runtime loads as a plugin (libkwker_rt).
Returns
str: KWKER_RT_LIB when set, else the library next to this package, else the development build (target/rt).
Errors
FileNotFoundError: ... not found (set KWKER_RT_LIB)
onnx.register(ort=None, name: str = 'kwker') Page
Register Kwker with ONNX Runtime (once per process) and return its devices.
Arguments
ort: The onnxruntime module (default: imported here).name: The provider name ONNX Runtime lists it under.
Returns
list: Kwker's entries of ort.get_ep_devices() - the CPU, or none on a CPU without AVX2.
onnx.session_options(precision: str = 'preserve', threads: int | None = None, so=None, ort=None, spinning: bool = False) Page
Return ONNX Runtime SessionOptions that run models on Kwker where it can.
Arguments
precision: "preserve" (default: the file's own precision - float32 for a float, float16 or 4-bit weight model, int8 for an int8 one such as model_qint8_avx512.onnx), "float32", "bf16" or "int8" (faster, approximate).threads: Threads for Kwker's part (default: all cores).so: SessionOptions to extend (default: new ones).ort: The onnxruntime module (default: imported here).spinning: Let ONNX Runtime's own threads spin while they wait for work (its default). Off by default here: nodes left to ONNX Runtime (pooling, a classifier head) otherwise leave its threads spinning against Kwker's.
Returns
onnxruntime.SessionOptions: pass it as InferenceSession(path, sess_options=...).
Examples: ONNX Runtime: Step 1: open a session with Kwker, ONNX Runtime: Step 2: run it, ONNX Runtime: Options
KwkDecoder and Decoder
import kwker.decode (needs torch)
Fast LLM text generation on the CPU, for Hugging Face models.
KwkDecoder runs a Llama-family model (Llama, Mistral, Qwen, SmolLM, Phi, Gemma, Granite and more) one token at a time as one native call per step. It builds in about a second:
dec = kwker.decode.KwkDecoder(model, max_cache_len=1024) # int4 weights by default
out = dec.generate(input_ids, max_new_tokens=64) # greedy
install(model) puts it under Hugging Face's own model.generate(), so sampling, logits processors, streamers and batches keep working:
kwker.decode.install(model)
out = model.generate(input_ids, max_new_tokens=64, do_sample=True, top_p=0.9)
Decoder covers other architectures by compiling the model ahead of time (once, tens of seconds, then cached).
decode.checkpoint_precision(model) -> str Page
The precision a model's weight matrices are stored in.
Arguments
model: A PyTorch model (its weight matrices are read).
Returns
"bf16", "fp16", "float32 (bf16 values)" (float32 weights that are all exact bf16 numbers, as a model trained in bf16 and saved in float32), "float32" or "mixed".
Example
import numpy as np
import kwker
import torch
import kwker.decode
print(kwker.decode.checkpoint_precision(torch.nn.Linear(4, 4).to(torch.bfloat16))) # bf16
bf16
class decode.Decoder(model, max_cache_len: int, batch_size: int = 1, package_path: str | None = None, cache_dir=None, mix: bool = True, calib=None, calib_ctx: int = 512, gptq: bool = False) Page
A Hugging Face language model compiled ahead of time for fast token-by-token generation (any architecture).
The first Decoder of an architecture compiles (tens of seconds); the compiled code is cached on disk, so later ones load in about a second. Use KwkDecoder for the families it supports: it builds faster and runs faster.
Arguments
model: A Hugging Face causal LM in eval mode.max_cache_len: The longest context: prompt + generated tokens.batch_size: How many sequences per call.package_path: Where to write the compiled package (default: a temporary path).cache_dir: The compile cache (default ~/.cache/kwker/decode; False or "" turns it off).mix: int4 mode: keep the most sensitive tensors in int8.calib: int4 mode: optional token ids to calibrate the int4 weights on.calib_ctx: The chunk length used to run the calibration tokens.gptq: With calib: GPTQ on the int4 tensors.
Notes
- The precision is fixed when it is built: kwker.cpu_backend.set_int4 / set_int8, or torch's float32 matmul precision ("medium": bf16 weights).
decode.Decoder.__call__(self, input_ids: torch.Tensor, cache_position: torch.Tensor) -> torch.Tensor
One forward of input_ids [B, L] at cache positions cache_position (L positions): logits [B, L, vocab]; the keys and values are written into the static cache at those positions.
Arguments
input_idscache_position: logits [B, L, vocab]
decode.Decoder.generate(self, input_ids: torch.Tensor, max_new_tokens: int, eos_token_id: int | None = None, lookup: int | None = None, ngram: int = 3, draft=None, draft_k: int = 4, do_sample: bool = False, temperature: float = 1.0, top_k: int = 0, top_p: float = 1.0, min_p: float = 0.0, repetition_penalty: float = 1.0, generator=None, min_new_tokens: int = 0, streamer=None) -> torch.Tensor
Generate tokens after a prompt: greedy by default, sampling with do_sample=True.
Arguments
input_ids: The prompt token ids [batch, length].max_new_tokens: How many tokens to generate.eos_token_id: Stop at this token.lookup: Prompt-lookup speculative decoding: up to this many draft tokens copied from the context (default 8 for one sequence; 0 turns it off). The tokens are the same as plain greedy decoding's.ngram: The longest context n-gram prompt lookup matches.draft: A smaller decoder of the same tokenizer that proposes tokens for this model to verify (same tokens, faster).draft_k: How many tokens the draft proposes per step.do_sample: True samples instead of greedy decoding.temperature: Sampling temperature.top_k: Sample from the k most likely tokens (0: off).top_p: Sample from the smallest set of tokens whose probability reaches top_p (1.0: off).min_p: Drop tokens below min_p times the top token's probability (0.0: off).repetition_penalty: Penalty for tokens already in the context (1.0: off).generator: A torch.Generator for reproducible sampling.min_new_tokens: Do not stop before this many tokens.streamer: A Hugging Face streamer that receives tokens as they are generated.
Returns
The prompt followed by the new tokens, [batch, length + new tokens].
decode.install(model, decoder=None, max_cache_len: int = 4096, max_batch: int = 1, precision=None, ops: bool = True, draft=None, mode=None, **kw) Page
Run Hugging Face's model.generate() on KwkDecoder, with no other code changes.
generate() keeps its whole interface: greedy or sampling, logits processors, stopping criteria, streamers, batches up to max_batch. Each step of the model runs as one native KwkDecoder call.
Arguments
model: A Hugging Face causal LM.decoder: An existing KwkDecoder (or Decoder) of this model to use; else one is built.max_cache_len: The longest context for the decoder that is built.max_batch: How many sequences one generate() call may hold.precision: "preserve" (the default: nothing lost - a bf16 model decodes on its bf16 weights, a float32 model on its float32 weights), "balanced" (int8 weights: about +0.1% perplexity), or "float32", "bf16", "int8", "int4".ops: True (the default) also installs Kwker's PyTorch kernels (kwker.torch_ops.install()), which speeds up sampling.draft: A smaller model of the same family and tokenizer for speculative decoding (the same tokens or distribution, about 2x faster on large models).mode**kw: More KwkDecoder options.
Returns
The model. uninstall(model) undoes it.
Notes
- Calls the decoder cannot take (beam search, inputs_embeds, use_cache=False, your own past_key_values, more rows or tokens than it was built for) run the model's own code; model._kwker_fallback says why.
- Models outside the KwkDecoder families get a compiled Decoder (one sequence).
- With precision="preserve", mixture-of-experts models keep generating on transformers itself, with Kwker's PyTorch kernels; precision="balanced" runs them on KwkDecoder.
- A line on stderr says when the weights run at a lower precision than the checkpoint's (KWKER_LOAD_REPORT=0: none).
- KWKER_DISABLE=1 in the environment leaves the model untouched.
Errors
ValueError: precision ...; one of ...ValueError: precision ... and mode ... disagreeRuntimeError: several new tokens per sequence in a batch
class decode.KwkDecoder(model, max_cache_len: int, mix=True, int8_groups=None, kv_fp16=None, kv_int8=False, max_batch: int = 1, calib=None, calib_ctx: int = 512, awq: bool = False, gptq: bool = False, cache_dir=None, mix_bits=None, prompt_cache: bool = True, precision=None, weights=None) Page
A Hugging Face language model's generation step as one fast native call; build it once, then generate.
Arguments
model: A Hugging Face causal LM (see runner_unsupported for the supported families).max_cache_len: The longest context: prompt + generated tokens.mix: int4 mode: keep the most sensitive tensors in more bits (True, the default).int8_groups: int8 mode: finer scales (True) or one per channel (False); None: the default.kv_fp16: The attention cache in fp16 or float32 (False). None (the default): fp16, float32 for a float32 checkpoint under precision="preserve".kv_int8: True stores the attention cache in int8 (about half the memory of fp16).max_batch: How many sequences can decode together (prefill, decode_batch, generate_batch).calib: Optional token ids (a 1-D LongTensor) to calibrate int4 weights on: better quality. Use text other than the text you evaluate on.calib_ctx: The chunk length used to run the calibration tokens.awq: With calib: AWQ-style channel scaling before the int4 fit (helps some models, hurts others).gptq: With calib: GPTQ on the int4 tensors.cache_dir: Keep the packed weights on disk and reuse them next time (True: ~/.cache/kwker/packs).mix_bits: int4 mode: 6 or 8 bits for those tensors; None picks 6 from a billion parameters, else 8.prompt_cache: True (the default) reuses the previous prompt's work when the next prompt starts with it (chat turns).precision: "preserve" (the default: nothing lost - a bf16 model runs on its bf16 weights, a float32 model on its float32 weights), "balanced" (int8 weights: about +0.1% perplexity, 1.7x bf16's speed), or "float32", "bf16", "int8", "int4" by name. See resolve_precision.weights
Example
import numpy as np
import kwker
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.cpu_backend, kwker.decode
cfg = LlamaConfig(vocab_size=128, hidden_size=64, intermediate_size=128, num_hidden_layers=2, num_attention_heads=2)
model = LlamaForCausalLM(cfg).eval()
if kwker.cpu_backend.decoder_available():
dec = kwker.decode.KwkDecoder(model, max_cache_len=64)
print(dec.generate(torch.tensor([[1, 2, 3]]), max_new_tokens=4).shape)
torch.Size([1, 7])
Notes
- Greedy tokens and logits are the same on AVX-512 and AVX2 CPUs.
- precision="preserve" raises ValueError for mixture-of-experts models: they run in int8 or int4 only.
- With cache_dir, the cache key covers every weight, so a changed model packs anew; the directory is kept under KWKER_PACK_CACHE_GB (16 GB by default).
Inherits: from decode.Decoder: generate
decode.KwkDecoder.__call__(self, input_ids: torch.Tensor, cache_position: torch.Tensor) -> torch.Tensor
One forward of input_ids [1, L] starting at cache_position[0]: logits [1, L, vocab] (as Decoder's).
Arguments
input_idscache_position
decode.KwkDecoder.decode_batch(self, tokens: torch.Tensor, positions: torch.Tensor, seqs: torch.Tensor) -> torch.Tensor
One decoding step for several sequences at once: token i of sequence seqs[i] at position positions[i].
Arguments
tokenspositionsseqs
Returns
Logits [batch, vocab]: the same as separate steps, at nearly the cost of one.
decode.KwkDecoder.generate_batch(self, prompts, max_new_tokens: int, eos_token_id: int | None = None)
Greedy decoding of up to max_batch prompts together.
Arguments
prompts: A list of 1-D LongTensors of any lengths.max_new_tokens: How many tokens to generate per prompt.eos_token_id: Stop a sequence at this token.
Returns
Each prompt followed by its new tokens: the same tokens generate() gives it alone.
decode.KwkDecoder.keep(self, pos: int, rows) -> None
After tree_step at pos: keep the accepted path's rows (increasing, rows[0] = 0) in the cache at positions pos, pos + 1, ...
Arguments
pos: keep the accepted path's rows (increasing, rows[0] = 0) in the cache at positions pos, pos + 1, ..rows
decode.KwkDecoder.load_report(self) -> str
What this decoder runs: the checkpoint's precision, the one it runs at, the weights' size, and a faster or more exact choice.
Returns
One or two lines of text, for example: kwker: checkpoint bf16 -> bf16 (precision="preserve": lossless), weight matrices 0.27 GB -> 0.27 GB
decode.KwkDecoder.prefill(self, input_ids: torch.Tensor, seq: int = 0) -> torch.Tensor
Read a prompt into the cache of batched sequence seq (0 .. max_batch - 1); return the last token's logits [vocab].
Arguments
input_idsseq
decode.KwkDecoder.tree_step(self, tokens: torch.Tensor, pos: int, parents) -> torch.Tensor
Verify a tree of draft tokens in one step (speculative decoding): token i follows token parents[i].
Arguments
tokens: The draft tokens (up to 64).pos: The cache position of the root token.parents: parents[i] < i is the index of token i's parent; parents[0] = -1.
Returns
Logits [L, vocab]: row i equals what an ordinary step along the root-to-i path would give. Then call keep().
decode.resolve_precision(model, precision='preserve') Page
The weights KwkDecoder runs a model with, for a precision name.
Arguments
model: A Hugging Face causal LM.precision: "preserve" (only lossless choices: bf16 weights for a bf16 checkpoint or float32 weights holding bf16 values, float32 weights for float32 and fp16 checkpoints), "balanced" (int8 weights, about +0.1% perplexity), or "float32", "bf16", "int8", "int4" by name.
Returns
(weights, why): weights is "float32", "bf16", "int8" or "int4"; None when "preserve" has no lossless native choice for this model (mixture-of-experts models), and why says what to pass instead.
Example
import numpy as np
import kwker
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.decode
cfg = LlamaConfig(vocab_size=128, hidden_size=64, intermediate_size=128, num_hidden_layers=2, num_attention_heads=2)
model = LlamaForCausalLM(cfg).to(torch.bfloat16).eval()
print(kwker.decode.resolve_precision(model)) # ('bf16', None)
('bf16', None)
decode.runner_unsupported(model, max_cache_len: int | None = None) -> str | None Page
Return None when KwkDecoder can run the model, else a short reason why not (then use Decoder).
Supported: Llama-family models (Llama, Mistral, Qwen2 / Qwen3, SmolLM2 / SmolLM3, Phi-3 / Phi-4-mini, Granite, Gemma 2 / 3), mixture-of-experts models (Mixtral, Qwen3-MoE, Granite MoE) and the LayerNorm family (GPT-2, GPT-NeoX / Pythia, Phi-1.5 / Phi-2, OPT, StableLM, StarCoder2, OLMo), with float32, bfloat16 or float16 weights.
Arguments
model: A Hugging Face causal language model.max_cache_len: The context length you plan to use (prompt + generated tokens).
decode.swap_ops(ep, mode: int, mix: bool = True) -> dict Page
Replace the linear-layer and attention calls of an exported program with Kwker's operators, in place (used by Decoder).
Arguments
ep: A torch.export ExportedProgram of a language model step.mode: The weight mode: 16 (bf16), 8 (int8) or 4 (int4).mix: int4 mode: keep the most sensitive tensors in int8.
Returns
The number of replacements per operator.
decode.uninstall(model) Page
Undo install(model): generate() runs the model's own code again.
Arguments
model
Serving: continuous batching and an HTTP API
import kwker.serve (needs torch)
Serve many generation requests at once on one KwkDecoder (continuous batching), from Python or as an OpenAI-compatible HTTP server.
from kwker.serve import Engine
eng = Engine(KwkDecoder(model, max_cache_len=2048, max_batch=8)).start() # serves on a background thread
req = eng.submit(input_ids, max_new_tokens=64)
for tok in req: # token ids as they are generated
...
Every step reads the model's weights once for all active requests, so 8 requests cost little more than one. Each request gets exactly the tokens it would get alone.
python -m kwker.serve --model <Hugging Face id> # the HTTP server (/v1/completions, /v1/chat/completions)
python -m kwker.serve --embeddings <embedding model> # /v1/embeddings (both flags: one server for both)
class serve.Engine(decoder, prefill_chunk: int = 256, lookup: int = 8, ngram: int = 3, prompt_cache: bool = True) Page
Continuous batching over a KwkDecoder: many requests share each pass over the weights.
Arguments
decoder: A KwkDecoder built with max_batch > 1 (its cache slots are the batch).prefill_chunk: Prompt tokens read per iteration, so long prompts do not stall the requests already generating.lookup: Prompt-lookup speculative decoding inside the batch: up to this many draft tokens per request (0: off).ngram: The longest context n-gram prompt lookup matches.prompt_cache: True (the default) reuses a free slot's work when a new prompt starts with that slot's last prompt.
Example
import numpy as np
import kwker
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.cpu_backend
from kwker.decode import KwkDecoder
from kwker.serve import Engine
if kwker.cpu_backend.decoder_available():
cfg = LlamaConfig(vocab_size=128, hidden_size=64, intermediate_size=128, num_hidden_layers=2, num_attention_heads=2)
eng = Engine(KwkDecoder(LlamaForCausalLM(cfg).eval(), max_cache_len=64, max_batch=2))
req = eng.submit(torch.tensor([1, 2, 3]), max_new_tokens=4)
eng.run_until_done()
print(len(req.result()))
4
serve.Engine.run_until_done(self)
Serve every queued request to the end on this thread (offline batch use).
serve.Engine.start(self)
Start serving on a background thread; return the engine.
serve.Engine.step(self) -> bool
Run one iteration on this thread: admit queued requests, read up to prefill_chunk prompt tokens, then one decoding step for every generating request. Returns False when there was nothing to do.
serve.Engine.stop(self)
Stop the background thread; requests still in flight stay unfinished.
serve.Engine.submit(self, input_ids, max_new_tokens: int = 128, eos_token_id: int | None = None, do_sample: bool = False, temperature: float = 1.0, top_k: int = 0, top_p: float = 1.0, min_p: float = 0.0, repetition_penalty: float = 1.0, seed: int | None = None, min_new_tokens: int = 0, constraint=None, stop=None) -> kwker.serve.Request
Queue a generation request and return its Request.
Arguments
input_ids: The prompt: a 1-D sequence of token ids.max_new_tokens: How many tokens to generate at most.eos_token_id: The end-of-sequence token id, or a list of them.do_sample: True samples (Hugging Face's sampling rules); False is greedy.temperature: Sampling temperature.top_k: Sample from the k most likely tokens (0: off).top_p: Sample from the smallest set of tokens whose probability reaches top_p (1.0: off).min_p: Drop tokens below min_p times the top token's probability (0.0: off).repetition_penaltyseed: The sampling seed: the same seed gives the same text (None: drawn from torch's generator).min_new_tokensconstraint: What the generated text may be - a kwker.constrain.JsonConstraint for one JSON value (None: free).stop: A function of the generated token ids (a list) that returns True to end the request there, with finish_reason "stop" (None: off) - for example a check for stop strings.
class serve.Request(ids, max_new_tokens, eos, warp, base, min_new_tokens) Page
One generation request of an Engine: iterate it for token ids as they arrive, or call result() for all of them.
finish_reason is "stop" (end-of-sequence token), "length" (max_new_tokens or the cache length) or "error".
Arguments
idsmax_new_tokenseoswarpbasemin_new_tokens
serve.Request.aresult(self)
Wait (asyncio) for the request to finish and return its new token ids.
serve.Request.result(self, timeout=None)
Wait for the request to finish and return its new token ids.
Arguments
-
timeout: Seconds to wait at most (None: no limit). -
serve.Request.stats(property): Return the request's timings: queue_ms (submit to start), ttft_ms (submit to the first token), tokens, and reused (prompt tokens taken from the cache of an earlier request).
Constrained decoding: JSON output
import kwker.constrain (needs torch)
Constrained decoding: generation that can only produce text of a given form - one JSON value, or one JSON value that matches a JSON Schema (the OpenAI API's response_format json_object / json_schema).
A character-level automaton tracks the JSON text generated so far. Each step, the most likely tokens are checked against it in order: greedy decoding takes the first that keeps the text a valid prefix, sampling draws among the valid ones. End-of-sequence is allowed only once the value is complete, and once it is complete, end-of-sequence is chosen.
class constrain.JsonConstraint(tokenizer, eos, top=64, schema=None) Page
Keeps one request's generated text a valid JSON value prefix, ending at a complete value - with a schema, a value that matches the schema.
Arguments
tokenizer: The model's tokenizer (each token's text is decoded once and kept on the tokenizer object).eos: The end-of-sequence token id(s).top: How many of the most likely tokens are checked per step before the rest are scanned too.schema: A JSON Schema (a dict) the value must match, or None for any JSON value.
Notes
- Raises ValueError for a schema it cannot use (a $ref outside the schema, a reference that points nowhere).
constrain.JsonConstraint.advance(self, t)
Move the automaton past a token.
Arguments
t: The token id emitted last.
constrain.JsonConstraint.pick(self, row, sample=None)
The next token: from logits row, the most likely one that keeps the JSON valid (greedy), or with sample(masked_row) a draw among the valid ones of the top candidates.
Arguments
rowsample
constrain.JsonConstraint.processor(self, prompt_len)
A transformers LogitsProcessor that keeps one generate() call (batch 1) on this constraint.
Arguments
prompt_len: The prompt's length in tokens (tokens after it are fed to the constraint).
Returns
The processor, for generate(logits_processor=LogitsProcessorList([...])).
constrain.JsonConstraint.valid(self, row, first=False)
The tokens allowed next, most likely first.
Arguments
row: The next token's logits (1-D).first: True stops at the most likely valid token (greedy decoding needs no more).
Returns
A list of token ids: the valid ones among the top candidates, else the most likely valid one beyond them; end-of-sequence alone once the value is complete and closed.
JAX
import kwker.jax_ops (needs jax)
Kwker's sort, argsort, top_k and rank as JAX functions on the CPU; jit, vmap and differentiation work.
kwker.jax_ops.sort(x, axis=-1, descending=False) # like jnp.sort
kwker.jax_ops.argsort(x, axis=-1, descending=False) # like jnp.argsort(stable=True)
kwker.jax_ops.top_k(x, k) # like jax.lax.top_k
kwker.jax_ops.rank(x, axis=-1, method="average") # like scipy.stats.rankdata
Equal values are ordered as in jnp: -0.0 equals 0.0 and all NaN values are equal. FFI says how the calls run: as XLA custom calls when the compiled handlers are built (one call per array, no copies), else through jax.pure_callback.
jax_ops.argsort(x, axis=-1, descending=False) Page
Return the positions that sort x along an axis, like jnp.argsort(stable=True).
Arguments
x: A JAX array.axis: The axis to sort along (the last by default).descending: True for largest first.
Returns
An int32 array of x's shape (int64 with jax_enable_x64).
jax_ops.FFI Page
jax_ops.FFI = True
jax_ops.rank(x, axis=-1, method='average', descending=False) Page
Return the rank of every value along an axis (1 for the smallest), like scipy.stats.rankdata.
Arguments
x: A JAX array.axis: The axis to rank along (the last by default).method: How equal values are ranked: "average" (the default), "min", "max", "dense" or "ordinal".descending: True ranks the largest value 1.
Returns
An array of x's shape: float for "average", integers for the other methods.
Errors
ValueError: method must be one of ... (got ...)
jax_ops.sort(x, axis=-1, descending=False) Page
Sort along an axis, like jnp.sort (NaN values last; first when descending).
Arguments
x: A JAX array.axis: The axis to sort along (the last by default).descending: True for largest first.
Returns
The sorted array, of x's shape and dtype.
jax_ops.top_k(x, k) Page
Return the k largest values along the last axis and their positions, like jax.lax.top_k (largest first).
Arguments
x: A JAX array.k: How many values.
Returns
(values, positions), each with k entries along the last axis.
Errors
ValueError: k (...) out of range for a last axis of shape ...
Diagnostics: doctor and support bundle
import kwker.doctor
Diagnostics: what Kwker sees and does on this machine, and a support bundle for bug reports.
python -m kwker doctor [--json]
python -m kwker support-bundle [-o FILE] [--all-packages] [--no-probe]
The report covers the version and build, the CPU and its features, the engine in use, Python and thread settings, the framework versions and the PyTorch extensions' state, with warnings for anything that limits speed or compatibility. Nothing is uploaded, and no user data is read.
doctor.doctor(torch=True) Page
Describe Kwker on this machine as a dict, including a list of warnings (format_report() prints it as text).
Arguments
torch: False skips importing PyTorch (the extensions' state is then not checked).
Example
import numpy as np
import kwker
import kwker.doctor
print("warnings" in kwker.doctor.doctor(torch=False))
True
doctor.format_report(r) Page
Return doctor()'s dict as readable text.
Arguments
r
doctor.main(argv=None) Page
The command line (python -m kwker, or the kwker command): doctor [--json] [--no-torch], support-bundle [-o FILE] [--all-packages] [--no-probe], cache [--clear [GROUP]], bench and audit.
Arguments
argv
doctor.support_bundle(path=None, all_packages=False, probe=True) Page
Write a diagnostic archive (tar.gz) to attach to a bug report and return its path.
It holds the doctor report (text and JSON), relevant environment variables, the CPU description, installed framework versions and, unless probe=False, the engine path of a few sorts of generated data. Nothing is uploaded and no user data is read; look inside before sharing it.
Arguments
path: Where to write it (default: a new file in the current directory).all_packages: True lists every installed package, not only the related frameworks.probe: False skips the probe sorts.
Diagnostics: audit
import kwker.audit
Measure your own program with and without Kwker: wall time, CPU time, speed-up, and whether the results agree.
python -m kwker audit [options] -- train.py --epochs 1
python -m kwker audit [options] -- -m mypackage.eval ...
python -m kwker audit --baseline "./app_std input.bin" -- ./app_kwker input.bin
The program runs unchanged in fresh interpreters, alternating plain runs and Kwker runs, --repeat times each; the fastest run of each side counts. The results must agree: the standard output must match, and every --output file must be equal (arrays within --rtol / --atol). With --baseline, both sides are commands run as they are - two builds of a C, C++, Rust or any other program, one without Kwker and one with it.
Modes: ops (the default: kwker.torch_ops.install(), results identical to PyTorch's), numpy (kwker.numpy_ops.install()), bf16 and int8 (faster PyTorch precision modes; results change a little). Nothing is uploaded.
audit.audit(program, mode='ops', repeat=3, outputs=(), rtol=1e-05, atol=1e-08, timeout=None, python=None, baseline=None) Page
Run a program alternately without and with Kwker and return the report as a dict (format_report() prints it).
Arguments
program: The script and its arguments, or ["-m", module, ...]. With baseline: the Kwker build's command.mode: "ops" (the default), "numpy", "scipy", "bf16" or "int8". Ignored with baseline.repeat: Runs per side; the fastest counts.outputs: Files the program writes that must be equal on both sides.rtol: Relative tolerance for array outputs (.npy / .npz / .pt).atol: Absolute tolerance for array outputs.timeout: Seconds per run at most (None: no limit).python: The Python interpreter to run (default: this one).baseline: A command (a list of arguments) for the plain side - for example the program built without Kwker. Both sides then run as given, without Python; the report's mode is "commands".
Errors
ValueError: mode ... (ops, bf16, int8, numpy, scipy)ValueError: no program givenValueError: an empty baseline command
audit.format_report(r) Page
Return audit()'s result as the text report the command prints.
Arguments
r
audit.main(argv=None) Page
The command line: python -m kwker audit [--mode ops|numpy|bf16|int8] [--repeat N] [--output FILE ...] [--rtol R] [--atol A] [--timeout S] [--json] [--baseline CMD] -- program.py [args]. Exits with status 1 when the outputs differ.
Arguments
argv
audit.run_child(mode, argv) Page
The Kwker side of an audit: switch the mode on, then run the program as __main__.
Arguments
modeargv
Diagnostics: what Kwker ran
import kwker.report
What Kwker did in this process: which calls it ran, which fell back to the host library, and why.
Each integration adds its part - the PyTorch kernels (kwker.torch_ops.install()), the NumPy drop-in (kwker.numpy_ops.install()), Hugging Face models under kwker.decode.install(model), and torch.compile(backend="kwker"). KWKER_REPORT=1 prints the report on stderr when the program ends; python -m kwker audit shows it for the program it runs.
report.format_report(r) -> str Page
report()'s result as text, one line per thing Kwker took over.
Arguments
r: The dict report() returns.
Returns
The text, starting with "Kwker report".
report.report() -> dict Page
What Kwker ran in this process so far, integration by integration.
Returns
A dict: "disabled" (KWKER_DISABLE is set), "pytorch" (installed, and per kernel the calls Kwker and PyTorch ran - with KWKER_REPORT or KWKER_REPORT_FILE set before install(), their wall time too), "numpy" (installed; with KWKER_REPORT or KWKER_REPORT_FILE set before install(), per function the calls and wall time on Kwker and on NumPy), "scipy" (kwker.scipy's calls and their wall time: Kwker / SciPy, with reasons), "dataframes" (kwker.frame / kwker.duck calls per library and call, with their wall time), "language_models" (each model under kwker.decode.install: its decoder, precision and how its generate() calls ran, with the reasons for calls that ran on the model's own code) and "compile" (graphs compiled by backend="kwker" and what was rewritten). A part is None when that integration was never imported.
Example
import numpy as np
import kwker
import kwker.report
print(kwker.report.format_report(kwker.report.report()))
Kwker report
PyTorch kernels: installed; Kwker ran 2,162 calls, PyTorch ran 6,124 (cases Kwker does not cover)
sort 895 Kwker - 1,049 PyTorch -
gather 82 Kwker - 780 PyTorch -
bincount 79 Kwker - 551 PyTorch -
searchsorted 76 Kwker - 434 PyTorch -
nonzero 54 Kwker - 352 PyTorch -
masked_select 41 Kwker - 284 PyTorch -
nanmedian 23 Kwker - 281 PyTorch -
topk 69 Kwker - 229 PyTorch -
scatter_add 8 Kwker - 283 PyTorch -
median.dim 124 Kwker - 158 PyTorch -
median 22 Kwker - 252 PyTorch -
mode 48 Kwker - 226 PyTorch -
kthvalue 113 Kwker - 143 PyTorch -
_unique2 107 Kwker - 135 PyTorch -
bucketize 35 Kwker - 187 PyTorch -
index_select 39 Kwker - 155 PyTorch -
index 54 Kwker - 109 PyTorch -
index_add 7 Kwker - 147 PyTorch -
quantile 58 Kwker - 74 PyTorch -
scatter_reduce 1 Kwker - 123 PyTorch -
NumPy drop-in: not installed
SciPy 1.17.1: installed
coo.sum_duplicates 128 Kwker 0 SciPy
coo.tocsc 133 Kwker 79 SciPy (429us / 7.3ms) (float duplicates in long rows (SciPy's summation order is not defined) 16; small matrix (SciPy is faster) 63)
coo.tocsr 119 Kwker 60 SciPy (1.3ms / 7.0ms) (small matrix (SciPy is faster) 49; float duplicates in long rows (SciPy's summation order is not defined) 11)
csc.sort_indices 7 Kwker 155 SciPy (- / 7.1ms) (small matrix (SciPy is faster) 145; repeated indices in long rows (SciPy's order is not defined) 10)
csc.sum_duplicates 26 Kwker 124 SciPy (- / 8.5ms) (small matrix (SciPy is faster) 121; float duplicates in long rows (SciPy's summation order is not defined) 3)
csc.tocsr 73 Kwker 0 SciPy
csr.sort_indices 26 Kwker 144 SciPy (- / 6.6ms) (small matrix (SciPy is faster) 131; repeated indices in long rows (SciPy's order is not defined) 13)
csr.sum_duplicates 37 Kwker 112 SciPy (- / 8.8ms) (small matrix (SciPy is faster) 108; float duplicates in long rows (SciPy's summation order is not defined) 4)
csr.tocsc 59 Kwker 0 SciPy
ndimage.median_filter 106 Kwker 13 SciPy (NaN or -0.0 values 4; cval not exactly a value of the image's type 6; 64-bit integers beyond 2^53 3)
ndimage.percentile_filter 119 Kwker 20 SciPy (64-bit integers beyond 2^53 8; NaN or -0.0 values 6; cval not exactly a value of the image's type 6)
ndimage.rank_filter 90 Kwker 59 SciPy (- / 6.8ms) (cval not exactly a value of the image's type 8; 64-bit integers beyond 2^53 6; NaN or -0.0 values 3; not a 3 x 3 / 5 x 5 / 7 x 7 window on 2-D planes 42)
signal.medfilt2d 116 Kwker 47 SciPy (uint64[2-D] input 4; NaN or -0.0 values 4; not a 3 x 3 / 5 x 5 / 7 x 7 window 22; int8[2-D] input 3; int32[2-D] input 2; uint32[2-D] input 5; uint16[2-D] input 3; int64[2-D] input 3; int16[2-D] input 1)
stats._cdf_distance 108 Kwker 21 SciPy (a zero distance with -0.0 values (its sign follows NumPy's order of the zeros) 16; NaN values 5)
stats._rankdata 234 Kwker 34 SciPy (5.7ms / 5.4ms) (small slices (SciPy is faster) 15)
stats.rankdata 111 Kwker 15 SciPy
kwker.frame / kwker.duck calls: arrow.argsort 273 (27.0ms), polars.argsort 193 (44.7ms), pandas.argsort 81 (39.6ms), duckdb.top_k 77 (56.1ms), pandas.top_k 71 (49.9ms), duckdb.group_by 49 (116.0ms), duckdb.sort 49 (34.8ms), pandas.group_by 21 (85.3ms), pandas.sort 1 (1.3ms)
Diagnostics: the replacement contract
import kwker.contract
The replacement contract: what a Kwker integration promises for each call it takes over from a host library.
exact the host's result, element by element (NaN equals NaN, -0.0 equals +0.0)
equivalent a result the host could return as well, where the host leaves a choice open - the order of equal keys,
the arrangement on either side of a partition point; the route's check says what must hold
approximate within a tolerance the user chose by asking for a lower precision
Each integration lists its calls as routes (kwker.numpy_ops.routes(), ...). run(routes) feeds every route generated inputs - sizes from 0 to thousands, sorted, reversed and few-valued data, NaN and -0.0 for floats - and compares Kwker's call with the host's under the route's class.
contract.APPROXIMATE Page
contract.APPROXIMATE = 'approximate'
contract.arrays(rng, size, dtypes=('int8', 'uint8', 'int16', 'uint16', 'int32', 'uint32', 'int64', 'uint64', 'float16', 'float32', 'float64', 'bool')) Page
A 1-D test array of the given size: a random dtype and one of several shapes of data (random, few distinct values, sorted, reversed, all equal), floats with some NaN, -0.0 and infinities.
Arguments
rngsize: a random dtype and one of several shapes of data (random, few distinct values, sorted, reversed, all equal), floats with some NaN, -0.0 and infinitiesdtypes
contract.EQUIVALENT Page
contract.EQUIVALENT = 'equivalent'
contract.EXACT Page
contract.EXACT = 'exact'
contract.identical(a, b) -> bool Page
True when two results are equal bit for bit: same() and, for floats, the same bit patterns (-0.0 differs from +0.0, NaN payloads count).
Arguments
ab
class contract.Route(name: str, contract: str, host: object, kwker: object, inputs: object, check: object = None, note: str = '') -> None Page
One call an integration takes over: the host's function, Kwker's, and the promise between them.
Arguments
name: The call, as the host names it (for example "numpy.sort").contract: EXACT, EQUIVALENT or APPROXIMATE.host: The host library's own function.kwker: The function that replaces it (it may hand some inputs back to the host).inputs: inputs(rng, size) -> a tuple of arguments for one call (each side gets its own copy).check: check(args, host_result, kwker_result) -> None, or a message when Kwker's result breaks the promise. The default for EXACT compares the results; EQUIVALENT and APPROXIMATE routes give their own.note: What the class means for this call, in a few words.
contract.run(routes, seconds=5.0, seed=0, sizes=(0, 1, 2, 7, 33, 300, 1500, 5000), retime=True) -> dict Page
Run each route on generated inputs for a share of the time and compare Kwker's result with the host's.
Arguments
routes: Route objects.seconds: The whole run's time budget, shared by the routes.seed: The random seed (a failure names the seed and the case number that reproduce it).sizes: Array sizes to cycle through.retime: False skips re-timing the slowest cases (a correctness-only run).
Returns
A dict: "ok", and per route its contract, the cases run, up to five failures, the seconds spent making inputs, in the host, in Kwker and in the checks, "slowest": the case where Kwker took the most time relative to the host (of the three worst first readings, each re-timed as the best of three runs a side: its arguments, both times, the ratio and the first reading's ratio), and "by_size": per input size the cases and the median host and Kwker seconds.
Examples: Evaluate Kwker on your machine: 3. Run your own program both ways
contract.same(a, b) -> bool Page
True when two results are equal: the same type, shape and dtype, equal elements (NaN equals NaN, -0.0 equals +0.0). Takes NumPy arrays, PyTorch tensors, scalars and tuples / lists of them.
Arguments
ab
Benchmark
import kwker.bench
Benchmark Kwker on your machine against the libraries you have installed, and get a shareable report.
python -m kwker.bench # 1K / 100K / 1M values: sort, argsort, top-k, smallest-k; 4 types, 3 inputs
python -m kwker.bench --quick # 100K values only (about 10 s)
python -m kwker.bench --md report.md --json report.json
python -m kwker.bench --native # also x86-simd-sort and VQSort, compiled here (needs a C++ compiler)
python -m kwker.bench --suite # the input-family suite: 40 key patterns + 3 of sort-research-rs, 6 key types,
# 1K / 100K / 1M keys (add --native for x86-simd-sort and VQSort)
python -m kwker.bench --frames # data frames: group-by, table sorts, ORDER BY .. LIMIT vs pandas / Polars /
# pyarrow / DuckDB on 1M rows (--rows to change)
python -m kwker.bench --models # image classifiers: KwkCNN vs PyTorch eager, Inductor and OpenVINO (float32)
python -m kwker.bench --models --int8 --images photos/
# + KwkCNN int8 vs OpenVINO int8, both calibrated on the same photos
python -m kwker.bench --encoders --int8 # text encoders (BERT, MiniLM): KwkEncoder vs OpenVINO bf16 (float32 without AMX) / int8,
# torch.compile(backend="kwker") and eager PyTorch (KwkEncoder needs AVX2 or AVX-512)
python -m kwker.bench --llm HuggingFaceTB/SmolLM2-135M --text wiki.test.raw
# a language model: KwkDecoder int4 / int8 / bf16 vs transformers eager,
# prompt + generation tokens/s and perplexity (+ your llama.cpp build: --llama-bench)
Competitors, each when installed: NumPy, PyTorch, pyarrow and Polars; with --native also Intel's x86-simd-sort and Google Highway's VQSort, built from the commits behind kwker.io/benchmarks. Single-threaded unless --threads is given (--llm: every core). Every result is checked against NumPy's before it is timed.
bench.main(argv=None) Page
The command line: python -m kwker bench [--quick] [--sizes N,..] [--types int32,..] [--ops sort,..] [--reps N] [--threads N] [--native] [--suite [--families name,..]] [--frames [--rows N]] [--models [name,..] [--no-compile] [--int8 --images DIR]] [--encoders [name,..] [--shapes 1x128,..] [--no-compile] [--int8]] [--llm MODEL [--llm-modes int4,..] [--text FILE [--chunks N] [--bos]] [--llama-bench PATH --gguf FILE,.. [--llama-perplexity PATH]]] [--md FILE] [--json FILE].
Arguments
argv
bench.markdown(env, rows) Page
Return the machine description and run()'s rows as a Markdown report: one row per cell, a column per library, and the geometric-mean speed-up with the number of cells below 1.0x.
Arguments
envrows
bench.run(sizes, dtypes, ops, reps, k=100, seed=1, log=<built-in function print>, native=None) Page
Time every operation, size, type and input pattern, Kwker against each installed library, best of reps.
Arguments
sizes: The array sizes.dtypes: NumPy dtype names.ops: Any of "sort", "argsort", "top-k", "smallest-k".reps: Timed runs per cell; the best counts.k: k for the top-k operations.seed: The data seed.log: Called with progress lines.native: kwker._bench_native.load()'s result, to include x86-simd-sort and VQSort.
Returns
One dict per cell: milliseconds per library and Kwker's speed-up over the fastest other one.