Kwker

4-bit packed values

Quantized models store 4-bit numbers two per byte: value i sits in byte i // 2, the low 4 bits first. This is the layout of GGML, ONNX and PyTorch's quint4x2. Kwker sorts, orders and ranks such data where it is, without unpacking it to one byte per value first.

Sort packed values

sort_int4_packed sorts the values in place. Here the bytes 0x31, 0x02 hold the values 1, 3, 2, 0; sorted they are 0, 1, 2, 3, packed as 0x10, 0x32.

PythonRuns on your machine.
import numpy as np
import kwker

data = np.array([0x31, 0x02], dtype=np.uint8)
kwker.sort_int4_packed(data)
print([hex(b) for b in data])
Output
['0x10', '0x32']

Pass signed=True for values from -8 to 7 instead of 0 to 15, and n= when the last byte holds only one value.

The largest values and their positions

top_k_int4_packed returns the positions of the k largest values; argsort_int4_packed returns the positions of all of them in sorted order. Neither changes the data.

PythonRuns on your machine.
import numpy as np
import kwker

data = np.array([0x31, 0x92], dtype=np.uint8)   # values 1, 3, 2, 9
print(kwker.top_k_int4_packed(data, 2))
print(kwker.argsort_int4_packed(data))
Output
[3 1]
[0 2 1 3]

Notes