# CLAUDE.md Guidance for Claude Code (claude.ai/code) when working in this repository. ## What this is ControlAI is an offline control-systems engineering assistant. A locally held model (`mlx-community/Qwen3-14B-4bit` by default, via MLX) answers control questions, and every number it states comes from a deterministic solver — SciPy/LAPACK/CVXPY behind a validated tool registry — never from the model's own arithmetic. A BM25 + dense hybrid retriever over a local control-theory corpus grounds conceptual answers. A FastAPI + vanilla-JS console and a terminal CLI are the two front ends. Nothing leaves the machine. Apple Silicon only for local use. There is no GGUF or Ollama path — an earlier version carried both plus CUDA and a Spaces deployment, and maintaining four backends for one machine was most of the complexity in the codebase. One CUDA path came back, narrowly, for the public demo Space: see **Deployment** below. It is a second `Engine` implementation behind the same contract, selected by one env var. It does not touch the agent loop, the registry, or any tool. ## Commands ```bash ./run.sh # web console at http://127.0.0.1:8000 ./run.sh --cli # interactive terminal chat ./run.sh --cli "design an LQR for A=[[0,1],[-2,-3]], B=[[0],[1]], Q=eye(2), R=1" ./run.sh --fetch-index # download the prebuilt index from the private Hub dataset repo ./run.sh --build-index # rebuild the dense retrieval index (~45 min, checkpointed/resumable) ./run.sh --ingest-corpus # merge data/processed/ chunks into the index ./run.sh --calibrate # re-measure MIN_COSINE against the current corpus python -m unittest discover tests -v # unittest, not pytest -- pytest is not installed python -m unittest tests.test_agent_tools # deterministic numerics python -m unittest tests.test_agent_core # parsing, stream gating, truncation pip install -r requirements.txt # runtime pip install -r requirements-training.txt # dataset generation / ground truth pip install -r requirements-corpus.txt # corpus crawling and extraction ``` Environment: `CONTROLAI_MODEL`, `CONTROLAI_ADAPTER`, `CONTROLAI_THINKING` (`off|auto|on`), `CONTROLAI_THINK_BUDGET`, `CONTROLAI_EMBED_MODEL`. There is no lint or format command configured — don't invent one. ## Architecture ### Request flow `web/` (SSE console) or `cli.py` → `app.py` (`/api/chat`, `/api/chat/stream`) → `ControlAgent` (`controlai_agent/agent.py`) → `LocalEngine` (`controlai_agent/engine.py`, MLX) and `registry.execute` (`controlai_agent/registry.py`) → tools in `controlai_agent/tools/*.py`. ### `engine.py` — inference One persistent KV cache lives for the process. `_align_cache` finds the longest prefix of the incoming prompt that the cache already holds, trims to that point, and feeds only the remainder, so the fixed system-prompt-plus-tool-schema prefix (~8k tokens) is prefilled once at startup by `prewarm()` rather than once per tool step. Generated tokens are tracked in the cache too, which is what makes continuing a tool-calling turn nearly free. Measured effect: first-token latency fell from ~6.4 s per step to ~0.7 s. `stream()` yields tokens as they are produced. `think_budget` caps tokens spent inside a `` block and closes it by hand on overrun, bounding worst-case latency on a reasoning model. Sampling uses `presence_penalty`, not a flat `repetition_penalty`. A blanket repetition penalty punishes the repeated structural tokens that matrices and JSON are made of (`[`, `0`, `,`) exactly when the model is emitting a tool call. ### `agent.py` — the loop A short tool-calling loop (`MAX_TOOL_STEPS = 2`, `MAX_CALLS_PER_TOOL = 2`) that streams throughout. `_StreamGate` releases text as soon as it cannot be the start of ``, so prose appears immediately while a tool call never leaks into the chat. Retrieved passages are attached to the *user* turn, never spliced into the system prompt, because that keeps the cached prefix byte-identical across questions. **There is deliberately no parameter-provenance check.** An earlier version refused any matrix it could not trace back to the user's message. It blocked the single most useful thing the assistant does — working an example that the user asked for — so it was removed. Schema validation and sandboxed execution in `registry.execute` are the real guarantees; `toolcall.degenerate_reason` catches genuine decoding loops with thresholds set well above any hand-written matrix. ### `prompts.py` 506 tokens, down from 2,587. States what to do rather than what not to do. The previous prompt's "never invent a parameter" clause taught the model to refuse worked examples; the current one explicitly authorises choosing an illustrative system when the user asks for a demonstration. ### Tool registry (`registry.py` + `controlai_agent/tools/`) Tools are plain functions decorated with `@registry.register(name, description, parameters_schema)` across `linear.py`, `frequency.py`, `synthesis.py`, `estimation.py`, `nonlinear.py`, `robust.py`, `allocation.py`, `simulation.py`, `matrix_ops.py`, `plotting.py`, `python_executor.py`, `rag.py`. `registry.execute()` coerces stringified arrays, validates against JSON Schema, runs the function, and rounds every float to 6 significant figures. `controlai_agent/verifier.py` cross-checks results with independent invariants — Riccati residual, pole-placement error, CBF forward invariance. **Prefer putting a fact in a tool result over instructing the model to derive it.** Two observed failures were fixed that way rather than by prompt wording, and it is the pattern to reach for first: - `continuous_lqr`, `discrete_lqr` and `place_state_feedback` return `closed_loop_A`. Given a correct gain, the model wrote $A-BK$ for a double integrator as `[[-6,1],[-5,-6]]` instead of `[[0,1],[-6,-5]]` — right gain, wrong write-up. - `bode_analysis` returns the gain/phase margins as well. Asked for a phase margin the model reached for it instead of `stability_margins`, got back sampled curves, and correctly but uselessly concluded the margin "cannot be determined". ### Retrieval (`controlai_rag/`) Three data bugs were found here and are worth knowing about, because each one silently degraded retrieval rather than failing: - **Chunk ids were not unique.** `chunk_document` is called once per page and restarted its counter each time, so every page's first chunk was `_c0000` — 154 distinct ids across 9,976 chunks. Anything keyed on chunk_id resolved to the wrong row. Fixed in `chunker.py` (ids are now page-qualified); `scripts/repair_chunk_ids.py` migrated the existing index in place. - **Embeddings were pooled without an EOS token.** Qwen3-Embedding pools the last position and was trained with `<|endoftext|>` there. Omitting it dropped the margin between relevant and irrelevant passages from +0.32 to +0.15. Note the checkpoint's own `eos_token_id` is `<|im_end|>`, which is the wrong token here — `controlai_rag/embeddings.py` pins the right one. - **71.5% of the Nise textbook was mojibake.** Its PDF has a broken symbol-font ToUnicode map, so every extractor returns `L½ f ðtÞ/C138 ¼FðsÞ` for `L[f(t)] = F(s)`. `controlai_rag/textfix.py` reverses the substitution (it is deterministic); `scripts/repair_corpus_text.py` migrated the index. `document_loader` now applies it on ingest. `document_loader.py` → `chunker.py` → `index.py` (BM25, persisted to `data/rag_index/`) and `embeddings.py` (Qwen3-Embedding-0.6B under MLX, last-token pooling, instruction-prefixed queries). **The index holds 80,370 chunks.** 9,976 come from `data/user_docs/` (course notes, Nise, Ogata); the other 70,394 were bridged in from `data/processed/*_chunks/` by `scripts/ingest_processed_corpus.py`. That bridge did not exist before: the `scripts/` pipeline (`raw → extracted → processed`) fed *training dataset generation* only, while `ControlRAGIndex` read `user_docs` alone. Retrieval was therefore seeing 12% of the corpus, and none of the canonical texts — Doyle/Francis/Tannenbaum, Åström & Murray, Rawlings/Mayne/Diehl, Sontag, Liberzon, Boyd, Söderström & Stoica — which were downloaded, extracted and chunked on disk, unread. Re-run the bridge after adding anything to `data/processed/`. **The index is not in git.** `chunks.json` (187MB), `embeddings.npz` (144MB) and `bm25.pkl` (124MB) are each past GitHub's 100MB per-file limit, and `chunks.json` holds the full extracted text of 666 documents including commercial textbooks (Nise, Ogata) — fine to keep locally, wrong to redistribute. They live in the **private** Hub dataset repo `atakankahya/controlai-rag-index`; `controlai_rag/fetch_index.py` (`./run.sh --fetch-index`) pulls them with your HF token. It copies out of the Hub cache rather than symlinking, because the index is mutated in place by uploads and by `scripts/repair_*.py`. A public clone has no index and builds its own from `data/user_docs/`. `build()` is checkpointed every 4,000 chunks to `embeddings.partial.npz` and resumes from it. An 80k-chunk build takes ~45 minutes, and writing only at the end meant one interruption discarded all of it (observed at 72,008 of 80,370). `retriever.py` fuses the two rankings with reciprocal rank fusion and gates on cosine similarity (`MIN_COSINE = 0.62`). That number is a property of the model *and* the corpus, so re-measure it with `./run.sh --calibrate` whenever the corpus changes size materially — it moved when the index grew 8x. Current margins: in-domain worst best-match 0.678, off-domain best 0.571. Lexical candidates are scored densely too — without that, a chunk BM25 ranked first but that fell outside the dense top-K was dropped for having no score rather than for being irrelevant. `_is_low_value` discards back-of-book index pages, which match almost any control query because they contain every term in the field, while explaining nothing. The gate matters. BM25 scores are unbounded and corpus-relative, so the old threshold of 2.5 passed essentially everything: a question about the Bode sensitivity integral retrieved Routh-Hurwitz tables at score 19.7 and injected them as authoritative context. Returning nothing is a valid and frequent outcome — the model then answers from its own knowledge. `display_source_name()` maps raw indexed filenames (which carry owner initials, course codes, scan artifacts) to clean citable labels. Never let a raw filename reach an answer. Uploads through `/api/upload` are embedded immediately via `HybridRetriever.add_chunks`, or they would have no vector and the cosine gate would make them permanently unreachable. ### `app.py` — serving Local-only FastAPI. MLX keeps its compute stream in thread-local state, so every inference call goes through one dedicated worker thread for the process lifetime; a lock serialises turns because they all mutate the same KV cache. `_to_wire_events` translates the agent's event vocabulary (`text`/`thinking`/…) into what `web/app.js` consumes (`token`/`thought`/…). ### Model artifacts `adapters/` and `models/` hold earlier fine-tuning output. **They are gitignored — not backed up by git.** The LoRA adapters are not loaded by default: `behavior_v1` emits a spurious empty `` as its first output on essentially every prompt, so no tool ever runs, and it refuses fully-specified problems; `sft_v2` produces empty output when no tools are exposed and string-typed numbers when they are. Both were measured against the base model, which routes and formats correctly. `CONTROLAI_ADAPTER=` loads one anyway for A/B work. `configs/*.yaml` are MLX-LoRA training configs and `scripts/` holds the offline pipeline (corpus discovery → extraction → dataset generation → training → evaluation). That pipeline is independent of the serving path and uses `requirements-training.txt`/`requirements-corpus.txt`. ### Deployment (`app_space.py`, `engine_torch.py`, `requirements-space.txt`) The demo Space is a **Gradio app on ZeroGPU**, not the FastAPI console. That is not a preference: **ZeroGPU assumes the Gradio app is the Space.** It schedules GPU workers by forking the server process, and its startup validation looks for a `@spaces.GPU` function wired to a Gradio event handler. Keeping FastAPI on the public port with Gradio as a hidden side-car was tried at length and does not work — see the list below. `app_space.py` therefore gives Gradio the port and drives `ControlAgent` from inside `@spaces.GPU`. The `web/` console is lost on the Space only; the agent, the 29 solvers, the verifier and the 80,370-chunk retriever are all the same. Seven ZeroGPU constraints, each found the hard way. Read these before changing anything here: 1. **No upper bounds in `requirements-space.txt`.** The platform appends its own `gradio[oauth,mcp]`, `spaces`, `uvicorn` and a `torch` ceiling; one extra constraint can make the resolve impossible. `transformers<4.56` did, because gradio 6 needs `huggingface-hub>=1.16` and every `transformers<4.56` needs `<1.0`. **And do not pin torch down to CUDA 12** — ZeroGPU is backed by Blackwell (sm_120), which needs CUDA 12.8+; an older wheel has no kernels for it. 2. **Build the model at import scope.** ZeroGPU patches torch during the entry module's import and only intercepts CUDA inside that window. Building it in a lifespan or a request reaches real CUDA init and raises `Low-level CUDA init (torch._C._cuda_init) reached`. 3. **No `device_map`, no bitsandbytes.** `device_map` routes transformers through `caching_allocator_warmup`'s direct `torch.empty(..., device="cuda")`, which trips the same guard; bitsandbytes requires `device_map`, so 4-bit is unavailable and the model must fit in bf16. 4. **Inference must be inside `@spaces.GPU`.** Otherwise the packed tensors are never materialised and forward passes return fluent multilingual noise at ~15 min/turn — up, responsive, and confidently wrong. Wrap the whole *turn*: a turn is several generations sharing one KV cache. 5. **Never pass the agent as an argument** to a `@spaces.GPU` function; reach it through a module global. ZeroGPU marshals arguments across a process boundary and tries to share CUDA tensors, hanging with no output. 6. **The embedder is a second model** with the same import-window problem. It loads lazily on the first query, so `_build()` embeds one throwaway string to force it in. Without that the Space starts fine and every answer carries `[agent] retrieval failed`. 7. **Gradio must own the public port.** With all of the above fixed but FastAPI on the port, the GPU was scheduled and acquired and ZeroGPU's own forked worker still died in `torch.init()`, while the platform probed the public port for `/api/predict` and got 404. A no-op `@spaces.GPU` function failed identically, so it was not the model or the payload. `HF_TOKEN` must be a Space secret with read access to the private index dataset repo, or retrieval is silently disabled. `engine_api.py` is an unused-but-working alternative: `CONTROLAI_BACKEND=api` runs generation over Inference Providers with everything else local. It needs a token carrying *Make calls to Inference Providers*. Kept for anyone who wants a demo without a GPU at all. ### Benchmark (`benchmarks/`) `controlbench_v1.jsonl` is the eval set; `SCOPE.md` defines the taxonomy and `README.md` the data-hygiene rules (never let benchmark prompts leak into training data; split by `family`, not by individual question). `scripts/eval_answer_quality.py` is the fast qualitative check across 18 domain cases.