Kwker

Groups, merges and sets

These calls build on sorting: totals per key, distinct values, merging sorted lists, and set operations such as intersection.

Totals per key: reduce_by_key

reduce_by_key(keys, values, op) groups equal keys and combines their values. It returns the distinct keys in order and one result per key, like SQL SELECT key, SUM(value) ... GROUP BY key. count_by_key(keys) counts the rows of each key.

import numpy as np
import kwker

customer = np.array([3, 1, 3, 2, 1, 3])
spent = np.array([10.0, 5.0, 2.5, 8.0, 1.0, 4.0])
keys, total = kwker.reduce_by_key(customer, spent, op="sum")
print(keys, total)
keys, visits = kwker.reduce_by_key(customer, op="count")
print(keys, visits)
[1 2 3] [ 6.   8.  16.5]
[1 2 3] [2 1 3]

The operations are sum, min, max, first, last, mean (Python and JavaScript; mean_by_key in Rust, C and C++) and count.

Group several key columns: group_codes

group_codes(columns) gives every row a group number. Rows with the same values in all the columns share a number, and the numbers follow the sorted order of the groups. It also returns each group's first row and its size.

import numpy as np
import kwker

country = np.array([2, 1, 2, 1, 2])
year = np.array([2024, 2023, 2024, 2024, 2023])
codes, first_row, size = kwker.group_codes([country, year])
print(codes)
print(size)
[3 0 3 1 2]
[1 1 1 2]

For pandas, Polars and pyarrow tables, kwker.frame.group_by does the whole group-by with each library's own rules. See DataFrames, Arrow and DuckDB.

Merge sorted lists: kway_merge

kway_merge(lists) merges any number of sorted arrays into one sorted array. When values are equal, the one from the earlier list comes first.

import numpy as np
import kwker

monday = np.array([1, 4, 9])
tuesday = np.array([2, 4, 7, 10])
wednesday = np.array([3])
print(kwker.kway_merge([monday, tuesday, wednesday]))
[ 1  2  3  4  4  7  9 10]

Distinct values: unique

unique(a) returns the sorted distinct values of an array, like np.unique. Add return_counts=True to count each value, return_inverse=True for the index of every element's value, and return_index=True for where each value first occurs.

PythonRuns on your machine.
import numpy as np
import kwker

user_ids = np.array([7, 3, 7, 9, 3, 7])
ids, counts = kwker.unique(user_ids, return_counts=True)
print(ids, counts)
Output
[3 7 9] [2 3 1]

Set operations

In Python, intersect1d, union1d, setdiff1d and setxor1d work like NumPy's functions of the same names and take unsorted input.

import numpy as np
import kwker

subscribed = np.array([5, 1, 9, 3, 7])
active = np.array([3, 9, 4, 5])
print(kwker.intersect1d(subscribed, active))   # in both
print(kwker.setdiff1d(subscribed, active))     # subscribed but not active
print(kwker.union1d(subscribed, active))       # in either
[3 5 9]
[1 7]
[1 3 4 5 7 9]

On arrays that are already sorted, set_op (set_op_sorted in Python) skips the sorting step. The operation is intersection, union, difference or symmetric difference. With the multiset option, repeated values count as many times as they occur.

import numpy as np
import kwker

a = np.array([1, 2, 2, 3, 5])
b = np.array([2, 2, 2, 5, 8])
print(kwker.set_op_sorted(a, b, "intersection"))
print(kwker.set_op_sorted(a, b, "intersection", multiset=True))
[2 5]
[2 2 5]

Membership tests

isin(element, test_elements) tells, for each value of element, whether it appears in test_elements. It is NumPy's np.isin and returns a boolean array of element's shape. Use invert=True for the values that do not appear.

PythonRuns on your machine.
import numpy as np
import kwker

tokens = np.array([12, 7, 99, 7, 3])
stopwords = np.array([7, 3])
print(kwker.isin(tokens, stopwords))
print(tokens[~kwker.isin(tokens, stopwords)])
Output
[False  True False  True  True]
[12 99]

Notes