Kwker

Strings

Kwker sorts lists of strings and NumPy string arrays in one of four orders, called collations. The names each language uses are at the end of this page.

Collation Order Example
bytes (default) by the bytes of the text (code point order) "B" < "a" < "b"
caseless ignores the case of A-Z "a" = "A" < "b"
natural runs of digits compare as numbers "file2" < "file10"
natural caseless both of the above "File2" < "file10"

Sort a list of strings

sort_strings(strings, collation) sorts a list of strings (Python returns a new list). Strings that the collation treats as equal keep their original order.

import kwker

files = ["file10.txt", "File2.txt", "file1.txt", "file2.txt"]
for collation in ["bytes", "caseless", "natural", "natural_caseless"]:
    print(f"{collation:17}", kwker.sort_strings(files, collation))
bytes             ['File2.txt', 'file1.txt', 'file10.txt', 'file2.txt']
caseless          ['file1.txt', 'file10.txt', 'File2.txt', 'file2.txt']
natural           ['File2.txt', 'file1.txt', 'file2.txt', 'file10.txt']
natural_caseless  ['file1.txt', 'File2.txt', 'file2.txt', 'file10.txt']

The order of strings

argsort_strings(strings, collation) returns the positions that sort the strings. Use it to reorder other data by a text column.

import numpy as np
import kwker

product = ["pear", "Apple", "fig", "banana"]
price = np.array([1.20, 0.80, 2.50, 0.30])
order = kwker.argsort_strings(product, "caseless")
print([product[i] for i in order])
print(price[order])
['Apple', 'banana', 'fig', 'pear']
[0.8 0.3 2.5 1.2]

NumPy string arrays

In Python, arrays of type S (bytes) and U (text) work directly, and the result is an array.

import numpy as np
import kwker

codes = np.array(["B12", "A7", "B2", "A10"])
print(kwker.sort_strings(codes, "natural"))
['A7' 'A10' 'B2' 'B12']

Other alphabets and mainframe data

In the native packages, a collation can also be any table of 256 byte weights: kwker.collation_table("ebcdic037") in Python, TABLE_EBCDIC_037 with sort_strings_table in Rust and kwker_argsort_strings_table in C order text the way IBM mainframes do. Language-specific orders (German, Swedish, ...) are not built in: sort by keys made with a library such as ICU.

Collation names in each language

Language bytes caseless natural natural caseless
Python, JavaScript "bytes" "caseless" "natural" "natural_caseless"
Rust Collation::Bytes AsciiCaseless Natural NaturalCaseless
C KWKER_COLLATE_BYTES KWKER_COLLATE_ASCII_CASELESS KWKER_COLLATE_NATURAL KWKER_COLLATE_NATURAL_CASELESS
C++ kwker::Collation::bytes ascii_caseless natural natural_caseless
Go kwker.CollateBytes CollateCaseless CollateNatural CollateNaturalCaseless
Java Kwker.COLLATE_BYTES COLLATE_CASELESS COLLATE_NATURAL COLLATE_NATURAL_CASELESS
C# Sorter.Collation.Bytes Caseless Natural NaturalCaseless