4-bit packed values
Quantized models store 4-bit numbers two per byte: value i sits in byte i // 2, the low 4 bits first. This is the
layout of GGML, ONNX and PyTorch's quint4x2. Kwker sorts, orders and ranks such data where it is, without unpacking
it to one byte per value first.
Sort packed values
sort_int4_packed sorts the values in place. Here the bytes 0x31, 0x02 hold the values 1, 3, 2, 0; sorted they are
0, 1, 2, 3, packed as 0x10, 0x32.
import numpy as np
import kwker
data = np.array([0x31, 0x02], dtype=np.uint8)
kwker.sort_int4_packed(data)
print([hex(b) for b in data])
Output
['0x10', '0x32']
Pass signed=True for values from -8 to 7 instead of 0 to 15, and n= when the last byte holds only one value.
The largest values and their positions
top_k_int4_packed returns the positions of the k largest values; argsort_int4_packed returns the positions of all
of them in sorted order. Neither changes the data.
import numpy as np
import kwker
data = np.array([0x31, 0x92], dtype=np.uint8) # values 1, 3, 2, 9
print(kwker.top_k_int4_packed(data, 2))
print(kwker.argsort_int4_packed(data))
Output
[3 1] [0 2 1 3]
Notes
- Equal values keep their position order in
argsort_int4_packedandtop_k_int4_packed. - To read the values as one byte each, unpack them:
np.stack([data & 15, data >> 4], axis=1).ravel().
Related
- Sorting: every other key type.
- Python API reference: every argument of these calls.