Chat and serving
A chat sends the whole conversation with every turn, and a server handles many requests at once. KwkDecoder reuses
the work it did for earlier turns, and kwker.serve puts many requests into one batched decode step. With 8 requests,
SmolLM2-1.7B produced 135 tokens per second in total, 4.7x one sequence alone; on SmolLM2-135M and 360M that was
2.1-2.4x llama.cpp's batching.
Step 1: chat turns reuse the earlier prompt
When a prompt starts with the previous prompt's tokens, such as the next turn of a chat, generate() reuses what it
computed for them. The tokens are exactly those of a fresh decoder:
import torch
from transformers import LlamaConfig, LlamaForCausalLM
import kwker.cpu_backend as scb
from kwker.decode import KwkDecoder, runner_unsupported
from kwker.serve import Engine
model = LlamaForCausalLM(LlamaConfig(vocab_size=1000, hidden_size=128, intermediate_size=256, num_hidden_layers=2,
num_attention_heads=2, num_key_value_heads=1)).eval()
print(scb.decoder_available() and runner_unsupported(model, 256) is None)
dec = KwkDecoder(model, max_cache_len=256, precision="int4")
turn1 = torch.randint(0, 1000, (1, 80)) # system prompt + first message
out = dec.generate(turn1, max_new_tokens=8)
turn2 = torch.cat([out, torch.randint(0, 1000, (1, 20))], dim=1) # the conversation so far + a new message
dec.generate(turn2, max_new_tokens=8)
print(dec.stats["prompt_reused"]) # prompt tokens not read again
True 64
Step 2: batch many requests
Engine puts up to max_batch requests into one KwkDecoder and runs one batched step for all of them, admitting new
requests as slots free up. Each request gets the same tokens it would get alone:
batched = KwkDecoder(model, max_cache_len=128, max_batch=4, precision="int4") # 4 requests share every step
engine = Engine(batched) # engine.start() serves on a background thread instead
reqs = [engine.submit(torch.tensor([1, 5, 7, k]), max_new_tokens=8) for k in range(6)]
engine.run_until_done()
print([len(r.tokens) for r in reqs], reqs[0].finish_reason)
[8, 8, 8, 8, 8, 8] length
Iterate a request to stream its tokens, or call result(). In asyncio code, async for t in req and
await req.aresult() do the same without blocking the event loop.
Step 3: an OpenAI-compatible server
python -m kwker.serve --model HuggingFaceTB/SmolLM2-360M-Instruct --max-batch 8
It serves /v1/completions, /v1/chat/completions (with "stream": true) and /v1/models on port 8000. One asyncio
event loop serves every connection, so thousands of open streams cost no thread each. Any OpenAI client works
against it:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="unused") # any key, unless the server sets one
reply = client.chat.completions.create(model="HuggingFaceTB/SmolLM2-360M-Instruct",
messages=[{"role": "user", "content": "Name three sorting algorithms."}])
print(reply.choices[0].message.content)
Requests take max_tokens, temperature, top_p, seed, stop (up to four strings: the reply ends before the first
one) and, when streaming, stream_options: {"include_usage": true} (a last chunk with the token counts).
--precision picks the weights: preserve (the default: the checkpoint's own), balanced (int8) or int4; measure the
quality first (int8 and int4 weights).
"response_format": {"type": "json_object"} makes the reply one valid JSON value: each token is checked as it is
generated, and the reply ends when the value closes. {"type": "json_schema", "json_schema": {"schema": {...}}} makes it
a value that matches your JSON Schema: the right keys, types and enum values, within your bounds. "tools": [...] passes
the tool definitions to the model's chat template; when the reply calls a tool, the response carries tool_calls and
finish_reason "tool_calls" (Hermes and Qwen <tool_call> blocks, Mistral [TOOL_CALLS] lists and plain JSON calls are
recognized). "tool_choice": "required" or {"type": "function", "function": {"name": ...}} makes the reply a call: its
arguments are generated against the function's parameters schema, the same way as a json_schema reply
("parallel_tool_calls": false allows one call; an object that lists properties takes only those keys unless it sets
additionalProperties).
import json
city = {"type": "object", "properties": {"city": {"type": "string"}, "country": {"type": "string"}},
"required": ["city", "country"], "additionalProperties": False}
reply = client.chat.completions.create(model="HuggingFaceTB/SmolLM2-360M-Instruct",
response_format={"type": "json_schema", "json_schema": {"name": "city", "schema": city}},
messages=[{"role": "user", "content": "Give a city and its country as JSON."}])
print(json.loads(reply.choices[0].message.content)) # {'city': ..., 'country': ...}
In Python, without a server, pass the schema to chat or generate:
import kwker
llm = kwker.Decoder.from_pretrained("HuggingFaceTB/SmolLM2-360M-Instruct")
print(json.loads(llm.chat("Give a city and its country as JSON.", schema=city)))
To require a key, start the server with --api-key (or set KWKER_API_KEY). Clients then pass the same key; requests
without it get HTTP 401:
python -m kwker.serve --model HuggingFaceTB/SmolLM2-360M-Instruct --api-key "$MY_KEY"
import os
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key=os.environ["MY_KEY"])
Add --embeddings <model> to serve /v1/embeddings from the same server, or give it alone for an embeddings-only
server. Each text is pooled and normalized the way the model's sentence-transformers configuration says:
python -m kwker.serve --embeddings BAAI/bge-small-en-v1.5
vectors = client.embeddings.create(model="BAAI/bge-small-en-v1.5", input=["first text", "second text"])
print(len(vectors.data[0].embedding)) # 384
Notes
-
A schema can use
type,properties,required,additionalProperties,enum,const,anyOf/oneOf,allOf,$refto$defs,items,prefixItems,minItems/maxItems,minLength/maxLength,pattern,format(date,time,date-time,uuid,email,ipv4),minimum/maximumand their exclusive forms, andmultipleOf. Other keywords are accepted and not enforced. A$refto another document is an error. -
nabove 1 andlogprobsare refused with HTTP 400 rather than ignored;presence_penaltyandfrequency_penaltyare ignored. -
KwkDecoder(..., prompt_cache=False)turns the reuse off. In the server, a request reuses an earlier request's prompt when that prompt's cache slot is free (request.stats["reused"]). -
Engine(dec, lookup=8)(the default) checks copied tokens as extra rows of the batched step when that pays: about 1.2x more tokens per second at low load.
Next steps
- Faster generation: draft models and prompt lookup.
- Command line: every
kwker.serveoption.