SciPy drop-in
kwker.scipy.install() makes SciPy's sparse-matrix conversions, rank statistics and window filters run on Kwker.
Your code keeps calling tocsr(), rankdata() or median_filter(), and scikit-learn keeps calling them for you. The
results are SciPy's, bit for bit: the same arrays, data types and format flags.
import numpy as np
import scipy.sparse as sp
import kwker.scipy
kwker.scipy.install() # from here on, the conversions below run on Kwker
rng = np.random.default_rng(0)
rows, cols = rng.integers(0, 1000, 200_000), rng.integers(0, 1000, 200_000)
A = sp.coo_array((rng.integers(1, 10, 200_000), (rows, cols)), shape=(1000, 1000)).tocsr() # duplicates summed
print(A.nnz <= 200_000, A.has_canonical_format)
kwker.scipy.uninstall() # SciPy's own methods back
True True
Which calls
| Call | With Kwker |
|---|---|
coo.tocsr(), coo.tocsc() |
the same matrix: duplicates summed, indices sorted, SciPy's index type |
coo.sum_ |
the same entries in the same order, duplicates added in input order |
csr.tocsc(), csc.tocsr() |
the same matrix, entries in the same order |
sort_, sum_ on CSR and CSC |
the same arrays, changed in place as SciPy changes them |
scipy.ndimage.median_ with 3 x 3, 5 x 5 or 7 x 7 windows (2-D images, or stacks with size=(1, k, k)) |
the same image, every border mode |
scipy.ndimage.rank_ and scipy.ndimage.percentile_ with the same windows |
the same image, every rank or percentile and border mode |
scipy.signal.medfilt2d with 3 x 3, 5 x 5 or 7 x 7 windows (uint8, float32, float64) |
the same zero-padded image |
scipy.stats.wasserstein_, scipy.stats.energy_ (unweighted) |
the same distance, bit for bit |
scipy.stats.rankdata and the rank tests built on it (spearmanr, kruskal, mannwhitneyu, wilcoxon, ansari and more) |
the same ranks, data types and statistics, every method, axis and nan_ |
Other SciPy functions that sort through NumPy's own functions, such as ks_2samp and kendalltau, get Kwker's sorts
from the NumPy drop-in instead: kwker.numpy_ops.install(), with the same results.
Kwker takes a sparse call for 2-D matrices of float32, float64, int32, int64, uint32 or uint64 values,
rankdata for NumPy arrays (or lists) of any integer or float type, and median_filter, rank_filter and
percentile_filter for 8- to 64-bit integer and float32 / float64 images. Other value types (complex, bool, objects) stay with SciPy.
When a call stays with SciPy
Small matrices stay with SciPy, where its own call is faster: coo.tocsr() and coo.tocsc() below 4,096 entries,
coo.sum_duplicates() below 1,024, CSR / CSC sum_duplicates() below 20,000, sort_indices() below 65,536, and
csr.tocsc() / csc.tocsr() below 131,072. On one Xeon test machine, at a million entries Kwker was about 1.6-1.8x
faster for coo.tocsr(), 5-7x for coo.sum_duplicates(), 1.5x for CSR sum_duplicates() and csr.tocsc(), and
1.85x for sort_indices(). rankdata takes every size (SciPy's own call costs about 150 microseconds before it ranks
anything), except many short slices along an axis: more than 32 slices of fewer than 256 values stay with SciPy.
median_filter takes images from 2,048 pixels (25x SciPy's speed at a million), medfilt2d from 512 (8x),
rank_filter and percentile_filter from 32,768, and the two distances from 256 values (5x at 100,000).
coo.tocsr(), coo.tocsc() and CSR / CSC sum_duplicates() of a float matrix with duplicate entries in a row of
more than 16 entries stay with SciPy, and so does sort_indices() of any matrix with a repeated index in such a row.
SciPy sorts each row with an unstable sort before adding duplicates, so the order of the additions there, and the last
bits of a sum, are not defined. Float matrices of whole numbers (counts, implicit-feedback ones) are the exception:
while their total stays below 2^24 (float32) or 2^53 (float64), every order gives the same sums, and Kwker takes them. Shorter rows keep their input order in SciPy's sort on Linux, and Kwker adds them in
that order. Above the size limits, integer matrices and float matrices without duplicates are always routed.
The window filters (median_filter, rank_filter, percentile_filter, medfilt2d) leave float images holding NaN or
-0.0 to SciPy: it picks either zero of a tie, and its NaN results are not defined. They also leave 64-bit integer images
with values past 2^53 to SciPy, which selects in doubles.
The first install() on a machine checks every routed call against SciPy's own on small inputs. The result is cached
for that SciPy, NumPy and Kwker build and CPU, and a call whose result differs stays with SciPy.
kwker.scipy.self_check() runs the check, and KWKER_SELF_CHECK=force repeats it. kwker.scipy.report() lists the
calls Kwker ran and the calls it left to SciPy, with the reasons. install() routes nothing on a SciPy version that
was not verified (kwker.scipy.VERSIONS), and KWKER_DISABLE=1 or KWKER_DISABLE=scipy turns it off.
Measure a program both ways
python -m kwker audit --mode scipy runs your program with and without kwker.scipy.install() and compares the time
and the output. The report lists each routed call with the time spent on Kwker and on SciPy. See
Evaluate on your machine.
python -m kwker audit --mode scipy -- build_graph.py --edges edges.npy
Related
- NumPy drop-in: NumPy's sort, argsort, unique and set operations.
- Behavior: the promise every drop-in call makes.
- Sparse matrices: the same conversions as explicit calls.