Text Generation
Transformers
Safetensors
GGUF
English
causal-lm
qwen2.5
reasoning
code-generation
Mixture of Experts
qlora
multimodal
tool-use
Eval Results (legacy)
conversational
Instructions to use ram1234598766/Cesium2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ram1234598766/Cesium2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ram1234598766/Cesium2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ram1234598766/Cesium2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ram1234598766/Cesium2 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ram1234598766/Cesium2:Q8_0 # Run inference directly in the terminal: llama cli -hf ram1234598766/Cesium2:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ram1234598766/Cesium2:Q8_0 # Run inference directly in the terminal: llama cli -hf ram1234598766/Cesium2:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ram1234598766/Cesium2:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf ram1234598766/Cesium2:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ram1234598766/Cesium2:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ram1234598766/Cesium2:Q8_0
Use Docker
docker model run hf.co/ram1234598766/Cesium2:Q8_0
- LM Studio
- Jan
- vLLM
How to use ram1234598766/Cesium2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ram1234598766/Cesium2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ram1234598766/Cesium2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ram1234598766/Cesium2:Q8_0
- SGLang
How to use ram1234598766/Cesium2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ram1234598766/Cesium2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ram1234598766/Cesium2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ram1234598766/Cesium2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ram1234598766/Cesium2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use ram1234598766/Cesium2 with Ollama:
ollama run hf.co/ram1234598766/Cesium2:Q8_0
- Unsloth Studio
How to use ram1234598766/Cesium2 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for ram1234598766/Cesium2 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for ram1234598766/Cesium2 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for ram1234598766/Cesium2 to start chatting
- Pi
How to use ram1234598766/Cesium2 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ram1234598766/Cesium2:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ram1234598766/Cesium2:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ram1234598766/Cesium2 with Docker Model Runner:
docker model run hf.co/ram1234598766/Cesium2:Q8_0
- Lemonade
How to use ram1234598766/Cesium2 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ram1234598766/Cesium2:Q8_0
Run and chat with the model
lemonade run user.Cesium2-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use ram1234598766/Cesium2 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ram1234598766/Cesium2:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ram1234598766/Cesium2:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ram1234598766/Cesium2 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ram1234598766/Cesium2:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ram1234598766/Cesium2:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 29,477 Bytes
82f262a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 | # Cesium2 (MORPH-AI) β Unified System-1/System-2 Architecture v5.5
**Lead Architect / Systems Engineering Design**
Integrates machine-learning layers (CLM, MoE, SFT/LoRA, VLM/ViT) with
deterministic logic layers (production rules, FSM, regex gating), live
knowledge layers (Google Search / web crawl via RAG, NER + knowledge graph),
search/optimization layers (best-of-N + verifier, adaptive compute skip),
and structured storage (JSON routing matrices, persistent KV cache, JSON
knowledge graph) into one verifiable pipeline.
---
## 1. Design Goals & Scope
| Goal | Constraint it answers |
|------|----------------------|
| G1. Deterministic safety | No prompt ever reaches the model unvalidated; no output ever leaves unverified. |
| G2. Trained capability | Token prediction, MoE routing, skill conditioning, and visual understanding are learned, not hard-coded. |
| G3. Controllable compute | Easy inputs pay O(1) cost; hard inputs pay O(steps) cost. Adaptive. |
| G4. Verifiable reasoning | Every produced candidate is scored; only the best passes the compliance gate. |
| G5. Persistent context | Cross-turn state survives via KV cache + scratchpad + disk-backed routing + knowledge graph. |
| G6. Hot-swappable skills | New domains = new JSON + LoRA adapter. No retrain of base weights. |
| G7. Live grounded knowledge | Answers can cite fetched web text, parsed into facts, and cross-examined by the verifier. |
| G8. Multimodal input | Text + image both enter the same pipeline; visual analysis is adaptive and skippable. |
**Out of scope:** distributed serving, multi-node training parallelism,
RLHF/RLVR, external vector DBs, cloud auth. This is a single-process,
on-device-first design.
---
## 2. The Five-Layer Pipeline
```
INPUT (text [+ image]) βββΊ L1 PRE-PROCESS GUARDRAILS βββΊ L1.5 VISION & SEARCH
(FSM state, production rules, (VLM/ViT analyze, search
regex/token gating) gate, RAG fetch, NER+graph)
L1.5 βββΊ L2 ROUTING & SKILL LOAD βββΊ L3 ML CORE βββΊ L4 SEARCH & REFINEMENT βββΊ L5 VERIFICATION & OUTPUT
(JSON skill matrix) (CLM + MoE + (adaptive depth, (verifier scoring,
LoRA skill best-of-N decoding, fact cross-examine,
conditioning) reasoning loop) compliance rules,
KV write-back, response)
```
Each layer is a closed function: `layer_n(input, state) -> (output, state')`.
State flows left-to-right through a **runtime FSM** (L0, the governor).
---
## 3. Layer Map to Existing Code (v4 baseline)
The v4 repo already contains ~70% of this design. This section maps every
requested algorithm to its current home, then flags what is **new** in v5.
| Requested algorithm | Existing in v4 | New in v5 / v5.5 |
|---------------------|----------------|-----------|
| CLM token prediction | `MorphModel.forward` β `base.lm_head` (architecture.py:648) | β |
| MoE | `MoEBlock` top-2/4 + aux loss (architecture.py:94) | expert count from routing matrix |
| SFT/LoRA hot-swap skills | `apply_lora`, `SkillTokenModule`, `.skill` files | per-skill LoRA adapter cache |
| Production rules (IF-THEN) | none | `RuleEngine` (L1 in, L5 out) |
| FSM runtime governor | implicit in runtime loop | `RuntimeFSM` state machine |
| Regex / token feature gating | `build_code_features` 4-dim (architecture.py:363), `CodeAwareBias` | regex feature channels: keywords, quotes, braces |
| Best-of-N + Verifier | `generate_best_of_n` (architecture.py:769), `VerifierHead` | normalized score, early-exit when Ξscore high, fact cross-examination |
| Adaptive depth / compute skip | `Coordinator` gates + `steps` (architecture.py:237, 630) | explicit skip algebra + hard floor + VISION/SEARCH skip gates |
| JSON skill routing matrix | `.skill` files, `auto_route` (runtime.py:233) | formal routing matrix with precedence + regex keys |
| Persistent KV cache | `MemoryModule` (in-memory), `ScratchpadMemory`, `_persist_turn` | disk-backed KV store + TTL/eviction |
| **VLM / ViT multimodal** (v5.5) | none | `VisionAnalyzer`: ViT encode + object detection + pixel-fact fallback |
| **Google Search / web crawl + RAG** (v5.5) | none | `SearchClient` (Google CSE + DDG crawl), `RAGPipeline` |
| **NER + Graph Query** (v5.5) | none | `FactExtractor`, `KnowledgeGraph`, `GraphQuery` |
---
## 4. Component Specification
### 4.0 L0 β RuntimeFSM (NEW)
Governing state machine. Every pipeline phase is a state; only legal
transitions exist. The ML model cannot be reached outside its state.
```
States: IDLE β INTAKE β GUARD_IN β VISION β SEARCH_GATE β SEARCH β FACT_EXTRACT β
ROUTED β GEN β REFINE* β VERIFY β GUARD_OUT β RESPOND β IDLE
(FAULT is a global trap state)
Transitions (whitelist only):
IDLE --input present--> INTAKE
INTAKE --rule engine passes--> GUARD_IN
INTAKE --rule blocks--> RESPOND (safe refusal)
GUARD_IN --image present--> VISION (adaptive: skip if trivial)
GUARD_IN --text only--> SEARCH_GATE
VISION --visual facts extracted--> SEARCH_GATE
SEARCH_GATE --live data needed--> SEARCH
SEARCH_GATE --cache hit / trivial--> ROUTED | GEN (compute skip)
SEARCH --web text fetched--> FACT_EXTRACT
FACT_EXTRACT --NER + graph updated--> ROUTED
GUARD_IN --skill matched--> ROUTED
GUARD_IN --no skill--> GEN
ROUTED --skill adapter loaded--> GEN
GEN --need reasoning--> REFINE (steps >= 1)
GEN --depth==0, easy--> VERIFY (skip compute)
REFINE --loop budget consumed--> VERIFY
VERIFY --score passes threshold--> GUARD_OUT
VERIFY --score low, n not spent--> GEN (next candidate)
VERIFY --score low, n spent--> RESPOND (best-effort fallback)
GUARD_OUT --compliance ok--> RESPOND
GUARD_OUT --compliance fail--> RESPOND (masked/refused)
RESPOND --state persisted--> IDLE
*any --exception--> FAULT --> (recover | IDLE)
```
Enforced in code by `RuntimeFSM` (transitions table = dict of allowed next
states); any transition not in the table raises `IllegalTransitionError`
before executing. This makes execution phases provably sequential and
interrupt-safe.
---
### 4.1 L1 β Pre-Process Guardrails (NEW: RuleEngine + regex gating)
Two deterministic layers run *before* any ML forward pass.
**A. `RuleEngine` β production rules (IF-THEN).**
Rules are data (JSON), not code. Each rule:
```json
{
"id": "safety.refuse_harmful",
"phase": "in",
"if": { "regex": "(\\b(kill|bomb|weapon)\\b)", "or": [
{"contains": ["harmful", "exploit"]}
]},
"then": { "action": "block", "reply": "I can't help with that." }
}
```
Schema: `if` = predicate tree (`and`/`or`/`not` over `regex`, `contains`,
`length_gt`, `token_count_gt`); `then` = `{action: block|warn|transform|allow,
reply?, mask_pattern?}`. Engine evaluates rules in order; first matching rule
wins (priority field for ties). Evaluation is string-level β zero model
invocation, guaranteed latency.
**B. `RegexFeatureExtractor` β token feature gating (extends v4).**
v4 builds 4 dims/token (`build_code_features`). v5 extends to an explicit
feature vector per token via regex:
```
dim 0: is_code_like r'[{}()\[\];=<>!&|+\-*/%\'"`#@.,:]'
dim 1: indent_depth r'^( +|\t+)' -> normalized depth
dim 2: bracket_balance r'[{\[(]' = +1, r'[}\])]' = -1, else 0
dim 3: has_newline r'\n'
dim 4: keyword_hit r'\b(def|class|import|return|if|else|for|while|try)\b' (NEW)
dim 5: quote_state r'"|\'' tracked as running state (NEW)
dim 6: numeric_literal r'\b\d+(\.\d+)?\b' (NEW)
```
The extractor emits a **running state** (bracket stack + quote open/closed)
rather than per-token only β so `CodeAwareBias` and the FSM can check
"bracket balance is non-negative" or "string literal still open" as cheap
pre-model sanity gates (e.g. reject inputs with unbalanced braces before the
model sees them, matching the v4 `CodeAwareBias` contract).
Interface:
```python
class RegexFeatureExtractor:
def extract(self, text: str, tokenizer) -> (torch.Tensor, SyntaxState)
def gate(self, syntax_state: SyntaxState) -> GateDecision # pass|warn|block
```
---
### 4.1.5 L1.5 β Vision, Search, and Fact Extraction (NEW in v5.5)
Three subsystems live between the text guardrails (L1) and routing (L2).
They are **adaptive**: each is skipped unless the input needs it (G3, G8).
**A. `VisionAnalyzer` β multimodal VLM / ViT (vision.py).**
Lazy-loads a ViT (`google/vit-base-patch16-224-in21k`) + object-detection
head (`hustvl/yolos-small`) when transformers provides them; otherwise falls
back to pure pixel statistics so the pipeline still produces `ImageFacts`
(zero model download, offline-testable).
```
image βββΊ VISION
ββ ViT encode ββββββββββββββββββββββΊ visual embeddings (patches+1, H)
ββ object detection (YOLOS) ββββββββΊ [{label, conf, box}]
ββ pixel facts (always) ββββββββββββΊ {size, colors, brightness,
edge_density, saliency_regions}
ImageFacts βββΊ image_text (serialized facts for RAG context + graph + verifier)
```
- `VisionAnalyzer.analyze(source)` β `ImageFacts` (with optional `embedding`).
- `ImageFacts.to_text()` β compact string packed into the prompt context.
- Object detections become facts: `{label, conf, box}` β knowledge graph nodes.
- Adaptive skip: `VISION` is only entered when an image path is supplied.
**B. `SearchClient` + `RAGPipeline` β built-in keyless web search (search.py).**
No API key, no quota ceiling. Backends are tried in order until one returns
results; a 90s cooldown suppresses a backend after 3 consecutive failures so
one blocked endpoint can't stall the pipeline:
| Backend | Trigger | Auth |
|---------|---------|------|
| DuckDuckGo HTML | default (built-in) | none |
| Bing HTML | fallback | none |
| Mojeek HTML | fallback (scrape-friendly) | none |
| Google Custom Search JSON API | **optional**, only when `GOOGLE_CSE_API_KEY` + `GOOGLE_CSE_ID` env set | API key |
```
SEARCH_GATE (does this need live data?)
ββ force_search flag / live hints ("today", "latest", "news", "price"...)
ββ cache hit in KV (`rag:<query>`) β skip network entirely (G3)
ββ depth >= 2 β skip search (cost floor)
RAGPipeline.retrieve(query):
query ββΊ SearchClient.search ββΊ [result titles/urls/snippets]
ββΊ fetch pages (HTMLβtext) ββΊ chunk (β€800 chars) ββΊ rank (token overlap)
ββΊ pack context (β€ max_context_chars) ββΊ KV cache (ttl=3600) ββΊ context str
```
Search results are **rule-verified** downstream by the same `RuleEngine`
(out-phase): snippets containing blocked content are dropped before packing.
**C. `FactExtractor` + `KnowledgeGraph` + `GraphQuery` β NER + graph (facts.py).**
```
web text / image_text ββΊ FactExtractor
ββ regex NER: PERSON, ORG, LOCATION, DATE, NUMBER, EMAIL, URL
ββ keyword dims: code / math / data / time / place (routing + search hints)
ββ triples: (subject, relation, object) via trigger verbs
ββΊ KnowledgeGraph (entity β relation β entities; dedup; save/load JSON)
ββΊ GraphQuery.facts_for_question(question) β grounded context string
```
The graph is **persistent** (`cache/graph.json`), so facts gathered this turn
are queryable next turn (G5). Triples are tagged with `source: web|image` so
the verifier can cite grounding. `GraphQuery.facts_for_question` matches the
question's entities against the graph as **subjects AND objects** (reverse
index), and falls back to content-word overlap when the question has no
proper-noun entities β so plain questions like "capital of france" still
surface relevant edges.
**Decision flow (search gate).** `_needs_live_data(prompt, force_search,
depth)` returns True when force_search, or a live hint matches AND
`depth < 2`. This bounds worst-case web latency while keeping live queries
available on demand.
---
### 4.2 L2 β Routing & Skill Load (extend `.skill` β routing matrix)
**JSON routing matrix.** Replace per-skill trigger matching with a single
matrix file with precedence and typed keys:
```json
{
"version": 1,
"skills": [
{
"name": "code_expert",
"token": "<SKILL:code>",
"priority": 10,
"patterns": [{"type": "regex", "value": "\\b(python|function|debug)\\b"},
{"type": "keyword", "value": "algorithm"}],
"adapter": "adapters/code_expert_lora",
"requires_syntax": ["balanced_brackets"]
}
],
"default": {"token": null, "adapter": null}
}
```
Routing algorithm:
1. Score every skill: `match = sum(w[type] * hits)`; normalize by pattern count.
2. Highest priority skill with `match > threshold` wins (ties β priority).
3. Winner's `token` β `SkillTokenModule` embedding index (runtime.py:131 hashes
skill name β index; v5 uses the matrix `index` field instead of `hash()`,
making routing deterministic across runs).
4. Winner's `adapter` β load matching LoRA adapter from KV/disk cache. If
already cached in memory, skip load (hot-swap, no retrain).
5. No winner β default (no skill token, base behavior).
---
### 4.3 L3 β ML Core (existing v4, wiring unchanged)
`MorphModel.forward` (architecture.py:579):
1. `inputs_embeds = base.embed_tokens(input_ids)` (+ `0.1 * skill_emb` when a
skill token is active).
2. `base_out = self.base_model(inputs_embeds, ...)` β frozen Qwen + LoRA
adapters β `base_hidden (B, T, H)`.
3. `Coordinator(base_hidden)` β `(gates, steps_dist, steps)` β the
System-1/System-2 controller.
4. `MultiStepReasoner(base_hidden, steps)` β weight-tied System-2 loop.
5. Subsystem gating by `adaptive_threshold` (architecture.py:619):
`code_bias` if `g_code`, `moe_block` if `g_think`, `memory.read` if
`g_mem`, `scratchpad.read` if `g_scratch`.
6. `depth_module` conditioning + `g * (moe_out + mem_out)` scaling.
7. `logits = base.lm_head(refined)`; `verifier_score = VerifierHead(refined)`.
8. Loss = `CE + moe_aux + step_entropy + verifier_mse` (training only).
**MoE** (`MoEBlock`, architecture.py:94): per-token top-2-of-4 experts,
softmax-normalized routing weights, load-balance aux loss. v5 makes
`num_experts` and `top_k` configurable per-skill via routing matrix
(`moe.top_k` field) so code/math skills can afford a deeper router than a
chat skill.
---
### 4.4 L4 β Search & Refinement
**A. Adaptive Depth / Compute Skip (formalize v4 skip algebra).**
The Coordinator emits `steps β [0, max_steps]` and gates `g β [0,1]^4`.
v5 defines the skip decision explicitly:
```
skip_reasoner = (steps == 0)
skip_moe = (g_think < threshold) or easy_query
skip_mem = (g_mem < threshold) and cache_miss
skip_scratch = (g_scratch < threshold) and no_prior_state
cost = base_forward # always
+ (0 if skip_reasoner else steps * reasoner_layer_cost)
+ (0 if skip_moe else moe_cost)
+ (0 if skip_mem else mem_read_cost)
```
"Easy query" floor: if `Coordinator` says `steps == 0` AND
`g_think.mean() < threshold`, L4 short-circuits to VERIFY immediately β
zero reasoning-loop cost. The hard floor guarantees correctness: an easy
query still gets a full base-model forward (G3 holds, nothing is skipped
below the base forward).
**B. Best-of-N + Verifier (formalize `generate_best_of_n`).**
v4 already generates `n` candidates and keeps max verifier score
(architecture.py:769). v5 adds:
1. **Normalized scoring:** score = `verifier(refined).mean(1)` normalized to
`[0,1]` over the batch so thresholds are stable across candidates.
2. **Early exit:** after each candidate, if `score > accept_threshold` and
`score - best < 0.01`, stop β the marginal value of more candidates is
~0 (heuristic search pruning: explore the reasoning-tree only while
promising).
3. **Diversity sampling:** candidate `i` uses `temperature + i * Ξtemp` so the
n samples cover the search space instead of re-rolling the same mode.
4. **Scoring cache:** per-input KV cache stores `(candidate_hash, score)` so
re-asking the same question reuses the stored score (no re-generate).
Heuristic-search framing: the reasoning tree is searched greedily
best-first by verifier score; `n` bounds the beam width; early-exit bounds
depth. This is beam search over *generated continuations*, scored by a
trained critic rather than a hand-written heuristic.
---
### 4.5 L5 β Verification & Output
1. **Verifier gate:** best candidate's normalized score must be `β₯
verify_min`. Below it: (a) if `n` not exhausted β back to GEN for another
candidate (FSM REFINEβGEN), (b) exhausted β graceful fallback reply.
2. **Outbound `RuleEngine`:** same JSON rules evaluated on the *output*
text (`phase: "out"`). Actions:
- `block` β replace with safe refusal.
- `mask` β regex-substitute sensitive spans (`mask_pattern`) then allow.
- `allow` β pass through.
3. **KV write-back:** response + refined hidden state + verifier score are
written to the persistent KV cache (cross-turn context, G5).
4. **Respond:** FSM β RESPOND, emits final string, returns to IDLE.
---
## 5. End-to-End Data Flow (concrete, step-by-step)
Given a raw user input `s` (text, optionally `+ image` path) and runtime state `st`:
```
Step FSM Action Data produced
ββββ βββββββββββ ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ βββββββββββββββββββββ
1 IDLE receive s; validate encoding/UTF-8 s
2 INTAKE RuntimeFSM.transition(INTAKE); normalize whitespace s_norm
3 GUARD_IN RuleEngine.eval(s_norm, phase="in")
4 GUARD_IN RegexFeatureExtractor.extract(s_norm) β X_feat (B,T,7) X_feat, SyntaxState
5 GUARD_IN syntax gate: unbalanced brackets? β warn|block GateDecision
6 VISION image present? VisionAnalyzer.analyze(path) β ImageFacts image_text, vis_emb
(adaptive skip if text-only or trivial) objects[]
7 SEARCH_GATE _needs_live_data? cache hit? depth floor? search_dec
8 SEARCH RAGPipeline.retrieve(query) β fetched chunks β packed ctx search_context
9 FACT_EXTRACT FactExtractor.triples(web+image_text) β graph.add_many graph facts
GraphQuery.facts_for_question β grounded context graph_context
10 ROUTED load routing matrix; score skills β winner or default skill_id, token_idx
11 ROUTED adapter cache hit? else load LoRA from disk/KV base_model (frozen+LoRA)
12 GEN pack [context]+[image_text]+prompt; tokenize (+skill token) ids (B,T)
13 GEN MorphModel.forward: embed β base β Coordinator β gating base_hidden, gates, steps
14 GEN steps==0 & g_think<thresh? β jump to 19 (compute skip) β
15 REFINE MultiStepReasoner loop Γsteps (weight-tied) scratch, reasoned
16 REFINE code_bias(reasoned, X_feat) | moe_block | memory.read refined
17 VERIFY VerifierHead(refined) β score cand_i, score_i
18 VERIFY score < accept & n not spent? β back to 12 (next candidate) β
19 VERIFY pick argmax score across candidates; cross-examine vs facts best_cand, fact_score
(entity overlap + rule compliance) β verifier note if weak
20 GUARD_OUT RuleEngine.eval(best_cand, phase="out") pass|mask|block
21 GUARD_OUT (mask) apply regex mask; (block) β refusal final_text
22 RESPOND KV write-back: {input, best_cand, score, refined-state}; state'
scratchpad.write; memory.write; graph persisted
23 RESPOND emit final_text; FSM β IDLE response
```
Fault path: any step raising β FSM β FAULT β log, return safe refusal, β IDLE.
No partial state is emitted to the user.
---
## 6. Data Contracts & Interfaces
```python
# L0
class RuntimeFSM:
def transition(self, next_state: str) -> None # raises IllegalTransitionError
def state(self) -> str
# L1
class RuleEngine:
def load(self, path: str) -> None # rules.json
def eval(self, text: str, phase: str) -> RuleDecision # {action, reply?, masked?}
class RegexFeatureExtractor:
def extract(self, text: str, tokenizer) -> (Tensor, SyntaxState)
def gate(self, syn: SyntaxState) -> GateDecision
# L1.5 (v5.5)
class VisionAnalyzer:
def analyze(self, source) -> ImageFacts # {size, colors, objects, embedding?...}
class ImageFacts:
def to_text(self) -> str # compact fact string for context/verifier
def to_dict(self) -> dict
class SearchClient:
def search(self, query: str, num: int = 5) -> List[SearchResult] # Google CSE | DDG
class RAGPipeline:
def retrieve(self, query, num=5, top_k=3, use_cache=True) -> str # packed context
class FactExtractor:
def extract(self, text) -> Dict[str, Set[str]] # NER types
def triples(self, text) -> List[Fact] # (s, r, o)
class KnowledgeGraph:
def add_many(self, facts); def to_text(self, entities, depth) -> str
def save(self, path); def load(self, path)
class GraphQuery:
def facts_for_question(self, graph, question, depth=1) -> str
# L2
class RoutingMatrix:
def load(self, path: str) -> None
def route(self, text: str) -> RouteDecision # {skill, token_idx, adapter}
def load_adapter(self, route: RouteDecision, cache) -> PeftModel | None
# L4
def best_of_n(model, ids, n, accept_threshold, ...) -> (best_ids, score)
def maybe_skip(coordinator_out, threshold) -> SkipDecision # used by FSM at step 10
# L5 / storage
class KVStore: # disk-backed
def get(self, key: str) -> Value | None # TTL-aware
def set(self, key: str, value: Value, ttl: float) -> None
def evict(self) -> None # LRU
```
The runtime composes these: `RuntimeFSM` drives the sequence; `RuleEngine`
and `RegexFeatureExtractor` gate both directions; `VisionAnalyzer` +
`SearchClient`/`RAGPipeline` + `FactExtractor`/`KnowledgeGraph`/`GraphQuery`
produce and ground live context; `RoutingMatrix` feeds `SkillTokenModule` +
adapter load; `best_of_n` + `VerifierHead` + `_cross_examine` + skip algebra
implement search; `KVStore` + `ScratchpadMemory` + `MemoryModule` + graph
persistence carry state across turns.
---
## 7. Trade-offs (solves / worsens / when-to-change)
| Choice | Solves | Worsens | When to change |
|--------|--------|---------|----------------|
| Deterministic rules before model | Safety, compliance, guaranteed cost | Misses model-internal nuance; manual rule maintenance | When rules explode > ~50; move to classifier + rules |
| FSM strict transitions | Provable phase order, no re-entrancy | Boilerplate; refactor cost for new phases | When phases become data-driven β codegen FSM from config |
| Regex token features | Cheap structure signal, no training | Limited semantic depth | When code-like languages beyond Python/JS added β add grammar tokenizer |
| Best-of-N + trained verifier | Strong reasoning/code quality | nΓ inference cost | When latency-bound on device β reduce n or use early-exit (already in design) |
| Adaptive skip | Easy queries are O(1)-ish | Risk of skipping needed compute on mis-gated inputs | Monitor gate calibration on dev set; retrain Coordinator if skip rate drifts |
| JSON routing matrix | Skills hot-swappable, no retrain | Matrix tuning; pattern overlap ambiguity | When >20 skills β add learned router |
| Disk KV + scratchpad | Cross-turn context, resumability | Stale-state risk if eviction wrong | When TTL wrong β tune TTL by session length |
| Vision pixel-fact fallback | VLM works with zero model download | Loses true object semantics | When real ViT weights are bundled β detect() uses YOLOS |
| Web crawl vs API search | Keyless HTML backends need no keys and no quota; multi-backend cooldown survives blocks | Scrape fragility, rate limits, no SLA | When Google CSE keys exist β optional 4th backend (env) |
| Regex NER + triple extraction | No model, instant, deterministic | Over/under-extraction on complex prose | When precision matters β add spaCy/HF NER behind same interface |
| Fact cross-examination | Answers stay grounded; no silent fabrication | Weak-overlap note may annoy on fuzzy queries | Tune overlap threshold; keep note only for fact-dependent turns |
---
## 8. Failure Modes & Mitigation
| Failure | Detection | Mitigation |
|---------|-----------|------------|
| Rule engine blocks a benign prompt (false positive) | user-side complaint / eval set | priority ordering, `warn` tier instead of `block`, allowlist overrides |
| Coordinator skips needed compute (under-gate) | low-quality answers on known-hard eval | adaptive_threshold floor; monitor gate distribution |
| Verifier over-scores garbage (reward hacking) | score deltas don't match human pref | normalize score; keep verifier trained on seq-likelihood target (v4 loss) |
| Skill matrix tie / pattern bleed | two skills claim one prompt | priority field + require_syntax constraints |
| KV cache staleness (cross-turn poisoning) | degraded multi-turn | TTL, LRU eviction, scratchpad rewrite on every turn |
| Adapter load race (two skills same turn) | adapter mismatch | single-threaded FSM (REQUIRED by design) or per-skill locking |
| FSM stuck state on crash | process restart | journal last state; on boot, INTAKE validates and resets to IDLE |
| ViT / detector download fails at runtime | vision degraded to pixel facts | catch per-backend, never crash the pipeline (already implemented) |
| Search network timeout | empty context, no grounding | RAG returns `_fallback_context` (snippets) or ""; FSM continues; never blocks hard |
| NER over-extraction (garbage triples) | graph noise pollutes context | `max_facts` cap in to_text; sources tagged; dedup by subject |
| Stale web facts served from cache | outdated answers | KV ttl=3600 on `rag:<query>`; force_search bypasses cache |
---
## 9. Proposed Implementation Order
1. `RuntimeFSM` + transitions table (`src/fsm.py`) β pure logic, no ML dep.
2. `RuleEngine` + `rules.json` (phase-in/out defaults) (`src/rules.py`).
3. `RegexFeatureExtractor` extending `build_code_features` (7-dim + SyntaxState)
(`src/regex_features.py`).
4. `RoutingMatrix` replacing `auto_route` + hash-based skill index
(`src/routing.py`); keep `.skill` file schema for back-compat.
5. `KVStore` (disk, TTL+LRU) wrapping in-memory `MemoryModule` (`src/kvstore.py`).
6. Wire into `runtime.py`: FSM guards each call; best-of-n early-exit +
normalized scoring added to `MorphModel.generate_best_of_n`.
7. Optional: per-skill `top_k`/`num_experts` in routing matrix β `MorphConfig`.
8. `VisionAnalyzer` (vision.py) β ViT + detector lazy-load, pixel-fact fallback.
9. `SearchClient` + `RAGPipeline` (search.py) β Google CSE + DDG crawl,
chunk/rank/pack, KV cache.
10. `FactExtractor` + `KnowledgeGraph` + `GraphQuery` (facts.py) β regex NER,
triples, persistent graph.
11. Extend `RuntimeFSM` with `VISION` / `SEARCH_GATE` / `SEARCH` /
`FACT_EXTRACT`; wire `_ingest` into `chat` / `chat_best_of_n`;
`_cross_examine` re-scores the winner against image+web facts.
Steps 8β11 are offline-testable (pixel-fact vision, mocked search,
regex-only NER) and require no model or API keys.
Each step is independently testable (rules and FSM have zero model
dependencies and can be unit-tested offline).
---
## 10. Key Takeaway
The ML layers (CLM/MoE/SFT-LoRA/VLM-ViT) provide **trained intelligence**;
the deterministic layers (rules/FSM/regex) provide **guaranteed structure and
safety**; the live-knowledge layers (search/RAG/NER/graph) provide **grounded,
fresh context** that the verifier cross-examines; the search layer (best-of-N +
verifier + skip algebra) provides **controlled cost and verifiable quality**;
the storage layer (JSON matrix + KV cache + knowledge graph) provides
**persistence and hot-swappability**. The FSM is the spine that makes the whole
composition provable: the ML model only ever runs inside legal states, and no
output leaves without passing a verifier score, a fact cross-examination, and
an outbound rule gate. |