Accelerate Hugging Face
Kwker runs Hugging Face language models and encoders as one native call per step, on x86-64 CPUs with AVX2 or
AVX-512. You start from the same model object from_pretrained gives you. Against llama.cpp on the same CPU, Kwker's
4-bit models generated 1.75x faster and read prompts 2.65x faster; BERT-family encoders ran ahead of OpenVINO in
bfloat16 and int8.
Examples
Generate text
- Generate text with a Hugging Face model: check support, generate, sample, keep
model.generate(). - int8 and int4 weights: smaller and faster models, with a perplexity check.
- Faster generation: prompt lookup, draft models and batches, the same tokens.
Chat and serve
- Chat and serving: reuse earlier chat turns, batch requests, an OpenAI-compatible server.
Embed
- Sentence embeddings: KwkEncoder in float or int8, checked against the original model.
Any other model
- Speed up inference with one line:
install()speeds up a plainmodel.generate(...)too. - Compile with torch.compile:
backend="kwker"for any model.
Next steps
- Quickstart: language models and embeddings.
- Evaluate on your machine: Kwker against llama.cpp, OpenVINO and PyTorch in one report.