Kwker

int8 and int4 weights

By default KwkDecoder runs a language model on its own weights (bfloat16 or float32). int8 and int4 weights mean less memory and faster tokens, at a small cost in quality. On SmolLM2-1.7B, Kwker's int4 raised wikitext perplexity by 6.0% over the float32 model at about the size of llama.cpp's Q4_K_M, which raised it by 8.4%.

PythonNeeds torch, transformers: runs on your machine.
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.cpu_backend as scb
from kwker.decode import KwkDecoder, runner_unsupported

model = LlamaForCausalLM(LlamaConfig(vocab_size=1000, hidden_size=128, intermediate_size=256, num_hidden_layers=2,
                                     num_attention_heads=2, num_key_value_heads=1)).eval()
if scb.decoder_available() and runner_unsupported(model, 64) is None:
    dec = KwkDecoder(model, max_cache_len=64, precision="int8")        # or "int4"; "balanced" = int8
    out = dec.generate(torch.tensor([[1, 5, 7, 9]]), max_new_tokens=8)
    print(out.shape)
Output
torch.Size([1, 12])

Pick a format

precision= Bits per weight Quality
"preserve" (the default) 16 or 32, as stored nothing rounded: a bf16 checkpoint runs on bf16 weights, a float32 one on float32
"balanced" or "int8" 8 about +0.1% perplexity on the models Kwker measured
"int4" about 5 the smallest and fastest: +2.4 to +8.4% perplexity; measure it on your text

"bf16" and "float32" pick those weights by name. Mixture-of-experts models need int4 or int8.

Improve int4

Keep the packed weights

cache_dir=True keeps the packed weights on disk (~/.cache/kwker/packs) and loads them next time instead of packing again. SmolLM2-1.7B in int4 then starts in 0.3 s instead of 4.7 s. The cache keeps one file per model, packs changed weights again, and stays under 16 GB (KWKER_PACK_CACHE_GB), least recently used files first.

Measure the quality on your text

ShellOn your machine.
python -m kwker.bench --llm HuggingFaceTB/SmolLM2-135M --text wiki.test.raw

The report gives the perplexity of each weight format next to the original model's, measured with llama.cpp's protocol, and the prompt and generation speed of each. Add --gguf and --llama-bench to put your own llama.cpp in the same table.