Kwker

Chat and serving

A chat sends the whole conversation with every turn, and a server handles many requests at once. KwkDecoder reuses the work it did for earlier turns, and kwker.serve puts many requests into one batched decode step. With 8 requests, SmolLM2-1.7B produced 135 tokens per second in total, 4.7x one sequence alone; on SmolLM2-135M and 360M that was 2.1-2.4x llama.cpp's batching.

Step 1: chat turns reuse the earlier prompt

When a prompt starts with the previous prompt's tokens, such as the next turn of a chat, generate() reuses what it computed for them. The tokens are exactly those of a fresh decoder:

PythonNeeds torch, transformers: runs on your machine.
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.cpu_backend as scb
from kwker.decode import KwkDecoder, runner_unsupported
from kwker.serve import Engine

model = LlamaForCausalLM(LlamaConfig(vocab_size=1000, hidden_size=128, intermediate_size=256, num_hidden_layers=2,
                                     num_attention_heads=2, num_key_value_heads=1)).eval()
print(scb.decoder_available() and runner_unsupported(model, 256) is None)
dec = KwkDecoder(model, max_cache_len=256, precision="int4")
turn1 = torch.randint(0, 1000, (1, 80))                                # system prompt + first message
out = dec.generate(turn1, max_new_tokens=8)
turn2 = torch.cat([out, torch.randint(0, 1000, (1, 20))], dim=1)       # the conversation so far + a new message
dec.generate(turn2, max_new_tokens=8)
print(dec.stats["prompt_reused"])                                      # prompt tokens not read again
Output
True
64

Step 2: batch many requests

Engine puts up to max_batch requests into one KwkDecoder and runs one batched step for all of them, admitting new requests as slots free up. Each request gets the same tokens it would get alone:

PythonRuns on your machine.
batched = KwkDecoder(model, max_cache_len=128, max_batch=4, precision="int4")   # 4 requests share every step
engine = Engine(batched)                       # engine.start() serves on a background thread instead
reqs = [engine.submit(torch.tensor([1, 5, 7, k]), max_new_tokens=8) for k in range(6)]
engine.run_until_done()
print([len(r.tokens) for r in reqs], reqs[0].finish_reason)
Output
[8, 8, 8, 8, 8, 8] length

Iterate a request to stream its tokens, or call result(). In asyncio code, async for t in req and await req.aresult() do the same without blocking the event loop.

Step 3: an OpenAI-compatible server

ShellOn your machine.
python -m kwker.serve --model HuggingFaceTB/SmolLM2-360M-Instruct --max-batch 8

It serves /v1/completions, /v1/chat/completions (with "stream": true) and /v1/models on port 8000. One asyncio event loop serves every connection, so thousands of open streams cost no thread each. Any OpenAI client works against it:

PythonRuns on your machine.
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused")   # any key, unless the server sets one
reply = client.chat.completions.create(model="HuggingFaceTB/SmolLM2-360M-Instruct",
                                       messages=[{"role": "user", "content": "Name three sorting algorithms."}])
print(reply.choices[0].message.content)

Requests take max_tokens, temperature, top_p, seed, stop (up to four strings: the reply ends before the first one) and, when streaming, stream_options: {"include_usage": true} (a last chunk with the token counts).

--precision picks the weights: preserve (the default: the checkpoint's own), balanced (int8) or int4; measure the quality first (int8 and int4 weights).

"response_format": {"type": "json_object"} makes the reply one valid JSON value: each token is checked as it is generated, and the reply ends when the value closes. {"type": "json_schema", "json_schema": {"schema": {...}}} makes it a value that matches your JSON Schema: the right keys, types and enum values, within your bounds. "tools": [...] passes the tool definitions to the model's chat template; when the reply calls a tool, the response carries tool_calls and finish_reason "tool_calls" (Hermes and Qwen <tool_call> blocks, Mistral [TOOL_CALLS] lists and plain JSON calls are recognized). "tool_choice": "required" or {"type": "function", "function": {"name": ...}} makes the reply a call: its arguments are generated against the function's parameters schema, the same way as a json_schema reply ("parallel_tool_calls": false allows one call; an object that lists properties takes only those keys unless it sets additionalProperties).

PythonRuns on your machine.
import json

city = {"type": "object", "properties": {"city": {"type": "string"}, "country": {"type": "string"}},
        "required": ["city", "country"], "additionalProperties": False}
reply = client.chat.completions.create(model="HuggingFaceTB/SmolLM2-360M-Instruct",
                                       response_format={"type": "json_schema", "json_schema": {"name": "city", "schema": city}},
                                       messages=[{"role": "user", "content": "Give a city and its country as JSON."}])
print(json.loads(reply.choices[0].message.content))   # {'city': ..., 'country': ...}

In Python, without a server, pass the schema to chat or generate:

PythonRuns on your machine.
import kwker

llm = kwker.Decoder.from_pretrained("HuggingFaceTB/SmolLM2-360M-Instruct")
print(json.loads(llm.chat("Give a city and its country as JSON.", schema=city)))

To require a key, start the server with --api-key (or set KWKER_API_KEY). Clients then pass the same key; requests without it get HTTP 401:

ShellOn your machine.
python -m kwker.serve --model HuggingFaceTB/SmolLM2-360M-Instruct --api-key "$MY_KEY"
PythonRuns on your machine.
import os

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key=os.environ["MY_KEY"])

Add --embeddings <model> to serve /v1/embeddings from the same server, or give it alone for an embeddings-only server. Each text is pooled and normalized the way the model's sentence-transformers configuration says:

ShellOn your machine.
python -m kwker.serve --embeddings BAAI/bge-small-en-v1.5
PythonRuns on your machine.
vectors = client.embeddings.create(model="BAAI/bge-small-en-v1.5", input=["first text", "second text"])
print(len(vectors.data[0].embedding))   # 384

Notes

Next steps