# API One server process answers three kinds of request. The GPU server (`serve/server.py`, PyTorch) and the Mac server (`serve/server_mlx.py`, MLX) share the same routes and formats. | Route | What it does | | --- | --- | | `POST /v1/systemone` | Typed questions about a document, answered in one pass each, with probabilities and a "needs generation" flag | | `POST /v1/chat/completions` | Plain chat with the base model (the adapter is off) | | `POST /v1/auto` | One question: decide, and if the router flags it, answer by generating from the same document cache | Start a server: ```bash # From the downloaded Dev-4B folder (see README.md) python -m serve.server --artifacts . # NVIDIA GPU python -m serve.server_mlx --model mlx-8bit --artifacts . # Mac (Dev-4B-MLX-8bit in mlx-8bit) ``` The examples below are real responses from the Mac server with the 4-bit build; GPU numbers differ slightly. ## Decisions: `POST /v1/systemone` The *state* is the document. Each question is one of three types: | Type | Options | Answer fields | | --- | --- | --- | | `noul` | yes / no | `noul`: probability of yes | | `choice` | `criteria`: 1 to 16 named options, each with an optional description | `choice`: the most likely option name | | `score` | `criteria`: 2 to 6 ordered level descriptions, lowest first | `score`: expected level, 0 = first | Every answer also has `probabilities`, a `router_score` (0 to 1, higher means a single pass is more likely to be wrong) and `needs_generation` (the router's flag at its tuned threshold, 0.516). All questions in one request share one encoding of the state. Request: ```json { "state": "Order #1182 was charged twice on 3 March. The customer asks for one charge to be refunded and says the parcel has not arrived yet.", "questions": { "team": {"type": "choice", "instructions": "Which team should handle this?", "criteria": {"billing": "charges and refunds", "shipping": "deliveries", "returns": "exchanges"}}, "urgent": {"type": "noul", "instructions": "Does this need a reply today?"}, "tone": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["calm", "annoyed", "angry"]} } } ``` Response (trimmed): ```json { "answers": { "team": {"type": "choice", "choice": "billing", "probabilities": {"billing": 0.879, "shipping": 0.112, "returns": 0.009}, "router_score": 0.246, "needs_generation": false}, "urgent": {"type": "noul", "noul": 0.615, "probabilities": {"false": 0.385, "true": 0.615}, "router_score": 0.799, "needs_generation": true}, "tone": {"type": "score", "score": 1.46, "probabilities": {"0": 0.038, "1": 0.462, "2": 0.500}, "router_score": 0.363, "needs_generation": false} }, "latency_ms": 861.5 } ``` Reading it: the routing question is clear (88% billing). "Urgent" is close to a coin flip, and the router flags it. The score sits between "annoyed" and "angry". Errors: a question with an unknown type, more than 16 choice options, or other than 2 to 6 score levels returns `422` with a message. ## Generation: `POST /v1/chat/completions` The base model, unchanged: with the adapter off its output is token-identical to the base model's. The request takes `messages` (OpenAI-style roles) and `max_tokens`. Decoding is greedy. ```json {"messages": [{"role": "user", "content": "In one sentence, what is a refund?"}], "max_tokens": 40} ``` ```json { "object": "chat.completion", "choices": [{"index": 0, "finish_reason": "stop", "message": {"role": "assistant", "content": "A refund is the return of money to a customer after a purchase, typically due to a return, cancellation, or dissatisfaction with the product or service."}}], "tokens": [32, 20965, 374, "..."], "usage": {"completion_tokens": 30} } ``` `tokens` carries the raw token ids, which is how the identity checks compare outputs. ## Decide, then generate if needed: `POST /v1/auto` One question. The server encodes the state once and decides. If the router flags the question, it keeps the state's cache, swaps the question for a "think step by step, then give 'Answer: X'" prompt, and generates from that cache. The state is never encoded twice (`state_prefills` is always 1). ```json { "state": "Tara has 17 stamps. She gives 5 to Bo, buys 12, then gives half of what she has to Cy.", "name": "count", "question": {"type": "choice", "instructions": "How many stamps does Tara have now?", "criteria": {"12": null, "13": null, "24": null, "14": null}}, "max_tokens": 300 } ``` ```json { "decision": {"type": "choice", "choice": "14", "probabilities": {"12": 0.047, "13": 0.344, "24": 0.023, "14": 0.587}, "router_score": 0.855, "needs_generation": true}, "escalated": true, "state_prefills": 1, "generation": "Let's go step by step:\n\n1. Tara starts with 17 stamps.\n2. She gives 5 stamps to Bo: 17 - 5 = 12 ...\n4. ... So she now has 24 - 12 = 12 stamps.\n\nAnswer: A", "generation_tokens": ["126 ids"] } ``` Here the one-pass decision was wrong (14), the router flagged it, and the generated answer is right. Options are lettered A, B, C... in the order given, so "Answer: A" is "12". When `escalated` is false the response has only `decision`. ## Speed | | First token after escalation, cache reuse | Encoding the document again | | --- | --- | --- | | Mac, M1 Pro, 8-bit, 1,013-token document | 0.30 s | 3.4 s | | Mac, M1 Pro, 8-bit, 4,091-token document | 0.39 s | 13.9 s | | H100, bf16, 4,091-token document | 33 ms | 98 ms | | H100, bf16, 257-token document | 32 ms | 29 ms | A plain decision costs about 1.1 s at 257 document tokens on the Mac (document reading runs at about 250 tokens a second) and about 0.16 s on an H100. ## Limits - One request on the model at a time; requests queue. - English, and documents up to about 4,000 characters were trained on. - Score questions are the weakest kind; arithmetic, counting and multi-hop questions are what `auto` is for.