Glossary
Every page links its first use of a term below to the term's entry here. Point at such a link to read the definition without leaving the page.
Sorting
argsort
The positions that put an array in order: element i of the result is the index of the i-th smallest value. Kwker
breaks ties by index, so the result is the same on every machine.
Stable sort
A sort that keeps equal keys in their input order. It matters when keys carry other data, such as values or row
numbers. Index results are always stable; key-value sorts are stable when you ask (stable=True).
Key-value sort
Sorting pairs by the key alone: each value moves with its key. Kwker sorts a keys array and a values array together, without building pairs first.
Selection
Finding the value a full sort would put at index k and moving it there, without sorting the rest. Values before
index k are smaller or equal, values after it larger or equal, and neither side is in any particular order. It
costs about n steps instead of the n log n of a full sort: enough for a median, a percentile or a cut-off.
Partial sort
Sorting only the k smallest values into the first k positions. The rest of the array stays unordered.
Top-k
The k smallest or largest values of an array, usually with their positions. Kwker returns both, ties broken by
index.
Radix sort
A sort that groups keys by their digits (groups of bits) instead of comparing them. Kwker uses radix passes where they are faster than comparisons.
Total order
An order in which any two keys compare one way: smaller, equal or larger. Kwker gives floats a total order: -0.0
before +0.0, and NaNs as one block at the chosen end.
Numbers
NaN
Not a Number: the floating-point value of undefined results such as 0 / 0. A NaN compares with nothing, so every
library decides where NaNs go. Kwker puts them in one block, last by default (nans_first puts them first).
Signed zero
Floats have two zeros, +0.0 and -0.0, which are equal in arithmetic. Kwker's sorts put -0.0 first; ranks, groups,
set operations and unique treat the two as one value, as NumPy does.
float32
IEEE 754 binary floating point with 32 or 64 bits (float64): about 7 or 16 significant decimal digits.
bfloat16
A 16-bit float with float32's range but only 8 significant bits (about 3 decimal digits). Most language-model checkpoints store their weights in it.
float16
The 16-bit IEEE 754 float: 11 significant bits and a largest value of 65,504.
int8
Signed integers of 8 or 4 bits. As keys they sort like any integer. In the CPU backend they hold quantized model weights.
Quantization
Storing numbers in fewer bits than they were trained in, for example model weights as int8 or int4 with one scale per small group. Less memory traffic makes decoding faster, at a small cost in accuracy.
dtype
The element type of a NumPy array or a PyTorch tensor, such as int32 or float64. Kwker picks its code for the
dtype it gets.
Hardware
SIMD
Single instruction, multiple data: CPU instructions that work on a vector of values at once, for example 16 float32 numbers in one AVX-512 register. Kwker's engines are built on them.
ISA
Instruction set architecture: the instructions a CPU understands. Kwker uses the word for its engines:
kwker.isa() names the one in use.
Engine
One compiled version of Kwker's code for one instruction set: avx512, avx2, sse42 (x86 without AVX2), neon
(with sve), simd128 (WebAssembly) or portable. Kwker picks the fastest one the CPU runs; every engine gives the
same results.
AVX-512
The 512-bit SIMD instructions of Intel Xeon CPUs (Skylake and newer), some Intel Core CPUs, and AMD Zen 4 and newer. Kwker's fastest x86 engine uses them.
AVX2
The 256-bit SIMD instructions of nearly every x86-64 CPU since 2013 (Intel Haswell and newer, AMD Excavator and Zen). Kwker's second x86 engine uses them.
SSE4.2
The 128-bit SIMD instructions that nearly every x86-64 CPU sold since about 2011 has. On x86-64 CPUs without AVX2, Kwker runs its SSE4.2 engine.
NEON
The 128-bit SIMD instructions of every 64-bit ARM CPU, Apple silicon and AWS Graviton included.
SVE
Scalable Vector Extension: ARM SIMD instructions whose vector length depends on the CPU (128 to 2048 bits), on Neoverse cores such as AWS Graviton 3 and 4.
VNNI
Vector Neural Network Instructions: x86 instructions that multiply 8-bit or 16-bit integers and add the products in one step. The CPU backend's int8 and int4 kernels use them.
AMX
Advanced Matrix Extensions: matrix-multiply instructions on tile registers for bfloat16 and int8, in Intel Xeon CPUs from Sapphire Rapids on. The CPU backend uses them for large matrix products.
Machine learning
CPU backend
Kwker's faster CPU kernels for PyTorch: matrix products, attention and other operators.
kwker.torch_ops.install(), torch.compile(backend="kwker") and the whole-model runners use it.
Eager mode
PyTorch running operators one by one, as Python calls them, without compiling the model first.
Inductor
PyTorch's default compiler behind torch.compile: it fuses operators and generates code for them.
backend="kwker" builds on it and puts Kwker's kernels in.
GEMM
General matrix multiply: the matrix product behind linear layers. A GEMV multiplies a matrix by one vector, as each layer does when a model decodes one token.
Prefill
Reading the prompt: the model processes all prompt tokens in one pass and fills the KV cache. Its speed is reported in prompt tokens per second.
Decode
Generating text one token at a time. Each step reads every weight once, so its speed depends mostly on memory bandwidth.
KV cache
The attention layers' keys and values of the tokens so far, kept so that each new token computes only its own.
Speculative decoding
Guessing several next tokens cheaply, with a small draft model or from the prompt, and checking them all in one step of the large model. Greedy output stays the same; it is faster when the guesses are right.
Perplexity
How well a language model predicts a text; lower is better. It shows what quantization costs in quality.
Mixture of experts
A model whose feed-forward layers are split into experts. A router sends each token to a few of them, so each token reads only part of the weights.
Packaging and platforms
ABI
Application binary interface: how compiled code calls a library (symbol names, argument layout). Kwker's C functions keep their signatures and return codes within a minor version.
FFI
Foreign function interface: how one language calls a library written in another. Kwker's language bindings call its C library this way.
Wheel
A prebuilt Python package file (.whl) that pip installs without compiling anything.
pkg-config
A tool that gives C and C++ builds the compiler and linker flags of an installed library:
pkg-config --cflags --libs kwker.
WebAssembly
A portable binary format that browsers and Node.js run at close to native speed. The playground and Kwker's JavaScript package run its WebAssembly build.
GIL
The global interpreter lock: it lets one thread at a time run Python code. Free-threaded CPython builds (3.13t and later) remove it.
Apache Arrow
A columnar memory format shared by pandas, Polars, DuckDB and others. Kwker reads Arrow arrays' buffers directly.
Related
- Core concepts: how Kwker orders keys, uses memory and threads.
- Behavior specification: the exact rules, each with its test.