Kwker

Accelerate Hugging Face

Kwker runs Hugging Face language models and encoders as one native call per step, on x86-64 CPUs with AVX2 or AVX-512. You start from the same model object from_pretrained gives you. Against llama.cpp on the same CPU, Kwker's 4-bit models generated 1.75x faster and read prompts 2.65x faster; BERT-family encoders ran ahead of OpenVINO in bfloat16 and int8.

Examples

Generate text

Chat and serve

Embed

Any other model

Next steps