Kwker

SciPy drop-in

kwker.scipy.install() makes SciPy's sparse-matrix conversions, rank statistics and window filters run on Kwker. Your code keeps calling tocsr(), rankdata() or median_filter(), and scikit-learn keeps calling them for you. The results are SciPy's, bit for bit: the same arrays, data types and format flags.

PythonNeeds scipy: runs on your machine.
import numpy as np
import scipy.sparse as sp
import kwker.scipy

kwker.scipy.install()              # from here on, the conversions below run on Kwker
rng = np.random.default_rng(0)
rows, cols = rng.integers(0, 1000, 200_000), rng.integers(0, 1000, 200_000)
A = sp.coo_array((rng.integers(1, 10, 200_000), (rows, cols)), shape=(1000, 1000)).tocsr()   # duplicates summed
print(A.nnz <= 200_000, A.has_canonical_format)
kwker.scipy.uninstall()            # SciPy's own methods back
Output
True True

Which calls

Call With Kwker
coo.tocsr(), coo.tocsc() the same matrix: duplicates summed, indices sorted, SciPy's index type
coo.sum_duplicates() the same entries in the same order, duplicates added in input order
csr.tocsc(), csc.tocsr() the same matrix, entries in the same order
sort_indices(), sum_duplicates() on CSR and CSC the same arrays, changed in place as SciPy changes them
scipy.ndimage.median_filter with 3 x 3, 5 x 5 or 7 x 7 windows (2-D images, or stacks with size=(1, k, k)) the same image, every border mode
scipy.ndimage.rank_filter and scipy.ndimage.percentile_filter with the same windows the same image, every rank or percentile and border mode
scipy.signal.medfilt2d with 3 x 3, 5 x 5 or 7 x 7 windows (uint8, float32, float64) the same zero-padded image
scipy.stats.wasserstein_distance, scipy.stats.energy_distance (unweighted) the same distance, bit for bit
scipy.stats.rankdata and the rank tests built on it (spearmanr, kruskal, mannwhitneyu, wilcoxon, ansari and more) the same ranks, data types and statistics, every method, axis and nan_policy

Other SciPy functions that sort through NumPy's own functions, such as ks_2samp and kendalltau, get Kwker's sorts from the NumPy drop-in instead: kwker.numpy_ops.install(), with the same results.

Kwker takes a sparse call for 2-D matrices of float32, float64, int32, int64, uint32 or uint64 values, rankdata for NumPy arrays (or lists) of any integer or float type, and median_filter, rank_filter and percentile_filter for 8- to 64-bit integer and float32 / float64 images. Other value types (complex, bool, objects) stay with SciPy.

When a call stays with SciPy

Small matrices stay with SciPy, where its own call is faster: coo.tocsr() and coo.tocsc() below 4,096 entries, coo.sum_duplicates() below 1,024, CSR / CSC sum_duplicates() below 20,000, sort_indices() below 65,536, and csr.tocsc() / csc.tocsr() below 131,072. On one Xeon test machine, at a million entries Kwker was about 1.6-1.8x faster for coo.tocsr(), 5-7x for coo.sum_duplicates(), 1.5x for CSR sum_duplicates() and csr.tocsc(), and 1.85x for sort_indices(). rankdata takes every size (SciPy's own call costs about 150 microseconds before it ranks anything), except many short slices along an axis: more than 32 slices of fewer than 256 values stay with SciPy. median_filter takes images from 2,048 pixels (25x SciPy's speed at a million), medfilt2d from 512 (8x), rank_filter and percentile_filter from 32,768, and the two distances from 256 values (5x at 100,000).

coo.tocsr(), coo.tocsc() and CSR / CSC sum_duplicates() of a float matrix with duplicate entries in a row of more than 16 entries stay with SciPy, and so does sort_indices() of any matrix with a repeated index in such a row. SciPy sorts each row with an unstable sort before adding duplicates, so the order of the additions there, and the last bits of a sum, are not defined. Float matrices of whole numbers (counts, implicit-feedback ones) are the exception: while their total stays below 2^24 (float32) or 2^53 (float64), every order gives the same sums, and Kwker takes them. Shorter rows keep their input order in SciPy's sort on Linux, and Kwker adds them in that order. Above the size limits, integer matrices and float matrices without duplicates are always routed. The window filters (median_filter, rank_filter, percentile_filter, medfilt2d) leave float images holding NaN or -0.0 to SciPy: it picks either zero of a tie, and its NaN results are not defined. They also leave 64-bit integer images with values past 2^53 to SciPy, which selects in doubles.

The first install() on a machine checks every routed call against SciPy's own on small inputs. The result is cached for that SciPy, NumPy and Kwker build and CPU, and a call whose result differs stays with SciPy. kwker.scipy.self_check() runs the check, and KWKER_SELF_CHECK=force repeats it. kwker.scipy.report() lists the calls Kwker ran and the calls it left to SciPy, with the reasons. install() routes nothing on a SciPy version that was not verified (kwker.scipy.VERSIONS), and KWKER_DISABLE=1 or KWKER_DISABLE=scipy turns it off.

Measure a program both ways

python -m kwker audit --mode scipy runs your program with and without kwker.scipy.install() and compares the time and the output. The report lists each routed call with the time spent on Kwker and on SciPy. See Evaluate on your machine.

ShellOn your machine.
python -m kwker audit --mode scipy -- build_graph.py --edges edges.npy