int8 and int4 weights
By default KwkDecoder runs a language model on its own weights (bfloat16 or float32). int8 and int4 weights mean
less memory and faster tokens, at a small cost in quality. On SmolLM2-1.7B, Kwker's int4 raised wikitext perplexity by 6.0% over the
float32 model at about the size of llama.cpp's Q4_K_M, which raised it by 8.4%.
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.cpu_backend as scb
from kwker.decode import KwkDecoder, runner_unsupported
model = LlamaForCausalLM(LlamaConfig(vocab_size=1000, hidden_size=128, intermediate_size=256, num_hidden_layers=2,
num_attention_heads=2, num_key_value_heads=1)).eval()
if scb.decoder_available() and runner_unsupported(model, 64) is None:
dec = KwkDecoder(model, max_cache_len=64, precision="int8") # or "int4"; "balanced" = int8
out = dec.generate(torch.tensor([[1, 5, 7, 9]]), max_new_tokens=8)
print(out.shape)
torch.Size([1, 12])
Pick a format
precision= |
Bits per weight | Quality |
|---|---|---|
"preserve" (the default) |
16 or 32, as stored | nothing rounded: a bf16 checkpoint runs on bf16 weights, a float32 one on float32 |
"balanced" or "int8" |
8 | about +0.1% perplexity on the models Kwker measured |
"int4" |
about 5 | the smallest and fastest: +2.4 to +8.4% perplexity; measure it on your text |
"bf16" and "float32" pick those weights by name. Mixture-of-experts models need int4 or int8.
Improve int4
calib=token_idsfits int4 to the model's real activations on your text;gptq=Trueadds GPTQ on top.mix_bits: 6 or 8, the bits of the few tensors int4 keeps at higher precision. 6 is faster and 8 more accurate; the default is 6 from a billion parameters and 8 below. On SmolLM2-1.7B and Qwen2.5-0.5B, 6 was 9-13% faster than 8, with perplexity +6.0% instead of +5.1%.
Keep the packed weights
cache_dir=True keeps the packed weights on disk (~/.cache/kwker/packs) and loads them next time instead of
packing again. SmolLM2-1.7B in int4 then starts in 0.3 s instead of 4.7 s. The cache keeps one file per model, packs
changed weights again, and stays under 16 GB (KWKER_PACK_CACHE_GB), least recently used files first.
Measure the quality on your text
python -m kwker.bench --llm HuggingFaceTB/SmolLM2-135M --text wiki.test.raw
The report gives the perplexity of each weight format next to the original model's, measured with llama.cpp's
protocol, and the prompt and generation speed of each. Add --gguf and --llama-bench to put your own llama.cpp in
the same table.
Related
- Generate text with a Hugging Face model
- Evaluate on your machine: fair comparisons with llama.cpp.