Kwker

Searching sorted data

Once data is sorted, you can find where any value belongs in it quickly. Kwker's searchsorted, bucketize and bucket_counts answer three versions of that question.

Where does a value go? searchsorted

searchsorted(sorted_array, values) returns, for each value, the position where it would be inserted to keep the array sorted. With side="left" (the default) the position is before any equal elements. With side="right" it is after them. It works like numpy.searchsorted.

import numpy as np
import kwker

grades = np.array([60, 70, 80, 90])          # must already be sorted
scores = np.array([55, 70, 85, 99])
print(kwker.searchsorted(grades, scores))
print(kwker.searchsorted(grades, scores, side="right"))
[0 1 3 4]
[0 2 3 4]

Which bucket? bucketize

bucketize(values, boundaries) returns the bucket number of every value, as torch.bucketize does. Bucket 0 holds values up to the first boundary, bucket 1 values up to the second, and so on. The last bucket holds values above every boundary.

import numpy as np
import kwker

age_limits = np.array([12, 19, 64])           # child, teen, adult, senior
ages = np.array([8, 15, 30, 70, 19])
print(kwker.bucketize(ages, age_limits))
[0 1 2 3 1]

How many in each bucket? bucket_counts

bucket_counts(values, boundaries) counts the values in each bucket. It is a histogram with the edges you choose, and it never builds the per-value bucket array.

import numpy as np
import kwker

age_limits = np.array([12, 19, 64])
ages = np.array([8, 15, 30, 70, 19, 45, 3])
print(kwker.bucket_counts(ages, age_limits))
[2 2 2 1]

Things to know

Important

The sorted array must be sorted in the same order you pass (descending, nans_first). Kwker does not check it, because checking would take as long as the search. An unsorted array gives positions that mean nothing.