# Cesium2 (MORPH-AI) — Unified System-1/System-2 Architecture v5.5 **Lead Architect / Systems Engineering Design** Integrates machine-learning layers (CLM, MoE, SFT/LoRA, VLM/ViT) with deterministic logic layers (production rules, FSM, regex gating), live knowledge layers (Google Search / web crawl via RAG, NER + knowledge graph), search/optimization layers (best-of-N + verifier, adaptive compute skip), and structured storage (JSON routing matrices, persistent KV cache, JSON knowledge graph) into one verifiable pipeline. --- ## 1. Design Goals & Scope | Goal | Constraint it answers | |------|----------------------| | G1. Deterministic safety | No prompt ever reaches the model unvalidated; no output ever leaves unverified. | | G2. Trained capability | Token prediction, MoE routing, skill conditioning, and visual understanding are learned, not hard-coded. | | G3. Controllable compute | Easy inputs pay O(1) cost; hard inputs pay O(steps) cost. Adaptive. | | G4. Verifiable reasoning | Every produced candidate is scored; only the best passes the compliance gate. | | G5. Persistent context | Cross-turn state survives via KV cache + scratchpad + disk-backed routing + knowledge graph. | | G6. Hot-swappable skills | New domains = new JSON + LoRA adapter. No retrain of base weights. | | G7. Live grounded knowledge | Answers can cite fetched web text, parsed into facts, and cross-examined by the verifier. | | G8. Multimodal input | Text + image both enter the same pipeline; visual analysis is adaptive and skippable. | **Out of scope:** distributed serving, multi-node training parallelism, RLHF/RLVR, external vector DBs, cloud auth. This is a single-process, on-device-first design. --- ## 2. The Five-Layer Pipeline ``` INPUT (text [+ image]) ──► L1 PRE-PROCESS GUARDRAILS ──► L1.5 VISION & SEARCH (FSM state, production rules, (VLM/ViT analyze, search regex/token gating) gate, RAG fetch, NER+graph) L1.5 ──► L2 ROUTING & SKILL LOAD ──► L3 ML CORE ──► L4 SEARCH & REFINEMENT ──► L5 VERIFICATION & OUTPUT (JSON skill matrix) (CLM + MoE + (adaptive depth, (verifier scoring, LoRA skill best-of-N decoding, fact cross-examine, conditioning) reasoning loop) compliance rules, KV write-back, response) ``` Each layer is a closed function: `layer_n(input, state) -> (output, state')`. State flows left-to-right through a **runtime FSM** (L0, the governor). --- ## 3. Layer Map to Existing Code (v4 baseline) The v4 repo already contains ~70% of this design. This section maps every requested algorithm to its current home, then flags what is **new** in v5. | Requested algorithm | Existing in v4 | New in v5 / v5.5 | |---------------------|----------------|-----------| | CLM token prediction | `MorphModel.forward` → `base.lm_head` (architecture.py:648) | — | | MoE | `MoEBlock` top-2/4 + aux loss (architecture.py:94) | expert count from routing matrix | | SFT/LoRA hot-swap skills | `apply_lora`, `SkillTokenModule`, `.skill` files | per-skill LoRA adapter cache | | Production rules (IF-THEN) | none | `RuleEngine` (L1 in, L5 out) | | FSM runtime governor | implicit in runtime loop | `RuntimeFSM` state machine | | Regex / token feature gating | `build_code_features` 4-dim (architecture.py:363), `CodeAwareBias` | regex feature channels: keywords, quotes, braces | | Best-of-N + Verifier | `generate_best_of_n` (architecture.py:769), `VerifierHead` | normalized score, early-exit when Δscore high, fact cross-examination | | Adaptive depth / compute skip | `Coordinator` gates + `steps` (architecture.py:237, 630) | explicit skip algebra + hard floor + VISION/SEARCH skip gates | | JSON skill routing matrix | `.skill` files, `auto_route` (runtime.py:233) | formal routing matrix with precedence + regex keys | | Persistent KV cache | `MemoryModule` (in-memory), `ScratchpadMemory`, `_persist_turn` | disk-backed KV store + TTL/eviction | | **VLM / ViT multimodal** (v5.5) | none | `VisionAnalyzer`: ViT encode + object detection + pixel-fact fallback | | **Google Search / web crawl + RAG** (v5.5) | none | `SearchClient` (Google CSE + DDG crawl), `RAGPipeline` | | **NER + Graph Query** (v5.5) | none | `FactExtractor`, `KnowledgeGraph`, `GraphQuery` | --- ## 4. Component Specification ### 4.0 L0 — RuntimeFSM (NEW) Governing state machine. Every pipeline phase is a state; only legal transitions exist. The ML model cannot be reached outside its state. ``` States: IDLE → INTAKE → GUARD_IN → VISION → SEARCH_GATE → SEARCH → FACT_EXTRACT → ROUTED → GEN → REFINE* → VERIFY → GUARD_OUT → RESPOND → IDLE (FAULT is a global trap state) Transitions (whitelist only): IDLE --input present--> INTAKE INTAKE --rule engine passes--> GUARD_IN INTAKE --rule blocks--> RESPOND (safe refusal) GUARD_IN --image present--> VISION (adaptive: skip if trivial) GUARD_IN --text only--> SEARCH_GATE VISION --visual facts extracted--> SEARCH_GATE SEARCH_GATE --live data needed--> SEARCH SEARCH_GATE --cache hit / trivial--> ROUTED | GEN (compute skip) SEARCH --web text fetched--> FACT_EXTRACT FACT_EXTRACT --NER + graph updated--> ROUTED GUARD_IN --skill matched--> ROUTED GUARD_IN --no skill--> GEN ROUTED --skill adapter loaded--> GEN GEN --need reasoning--> REFINE (steps >= 1) GEN --depth==0, easy--> VERIFY (skip compute) REFINE --loop budget consumed--> VERIFY VERIFY --score passes threshold--> GUARD_OUT VERIFY --score low, n not spent--> GEN (next candidate) VERIFY --score low, n spent--> RESPOND (best-effort fallback) GUARD_OUT --compliance ok--> RESPOND GUARD_OUT --compliance fail--> RESPOND (masked/refused) RESPOND --state persisted--> IDLE *any --exception--> FAULT --> (recover | IDLE) ``` Enforced in code by `RuntimeFSM` (transitions table = dict of allowed next states); any transition not in the table raises `IllegalTransitionError` before executing. This makes execution phases provably sequential and interrupt-safe. --- ### 4.1 L1 — Pre-Process Guardrails (NEW: RuleEngine + regex gating) Two deterministic layers run *before* any ML forward pass. **A. `RuleEngine` — production rules (IF-THEN).** Rules are data (JSON), not code. Each rule: ```json { "id": "safety.refuse_harmful", "phase": "in", "if": { "regex": "(\\b(kill|bomb|weapon)\\b)", "or": [ {"contains": ["harmful", "exploit"]} ]}, "then": { "action": "block", "reply": "I can't help with that." } } ``` Schema: `if` = predicate tree (`and`/`or`/`not` over `regex`, `contains`, `length_gt`, `token_count_gt`); `then` = `{action: block|warn|transform|allow, reply?, mask_pattern?}`. Engine evaluates rules in order; first matching rule wins (priority field for ties). Evaluation is string-level — zero model invocation, guaranteed latency. **B. `RegexFeatureExtractor` — token feature gating (extends v4).** v4 builds 4 dims/token (`build_code_features`). v5 extends to an explicit feature vector per token via regex: ``` dim 0: is_code_like r'[{}()\[\];=<>!&|+\-*/%\'"`#@.,:]' dim 1: indent_depth r'^( +|\t+)' -> normalized depth dim 2: bracket_balance r'[{\[(]' = +1, r'[}\])]' = -1, else 0 dim 3: has_newline r'\n' dim 4: keyword_hit r'\b(def|class|import|return|if|else|for|while|try)\b' (NEW) dim 5: quote_state r'"|\'' tracked as running state (NEW) dim 6: numeric_literal r'\b\d+(\.\d+)?\b' (NEW) ``` The extractor emits a **running state** (bracket stack + quote open/closed) rather than per-token only — so `CodeAwareBias` and the FSM can check "bracket balance is non-negative" or "string literal still open" as cheap pre-model sanity gates (e.g. reject inputs with unbalanced braces before the model sees them, matching the v4 `CodeAwareBias` contract). Interface: ```python class RegexFeatureExtractor: def extract(self, text: str, tokenizer) -> (torch.Tensor, SyntaxState) def gate(self, syntax_state: SyntaxState) -> GateDecision # pass|warn|block ``` --- ### 4.1.5 L1.5 — Vision, Search, and Fact Extraction (NEW in v5.5) Three subsystems live between the text guardrails (L1) and routing (L2). They are **adaptive**: each is skipped unless the input needs it (G3, G8). **A. `VisionAnalyzer` — multimodal VLM / ViT (vision.py).** Lazy-loads a ViT (`google/vit-base-patch16-224-in21k`) + object-detection head (`hustvl/yolos-small`) when transformers provides them; otherwise falls back to pure pixel statistics so the pipeline still produces `ImageFacts` (zero model download, offline-testable). ``` image ──► VISION ├─ ViT encode ─────────────────────► visual embeddings (patches+1, H) ├─ object detection (YOLOS) ───────► [{label, conf, box}] └─ pixel facts (always) ───────────► {size, colors, brightness, edge_density, saliency_regions} ImageFacts ──► image_text (serialized facts for RAG context + graph + verifier) ``` - `VisionAnalyzer.analyze(source)` → `ImageFacts` (with optional `embedding`). - `ImageFacts.to_text()` → compact string packed into the prompt context. - Object detections become facts: `{label, conf, box}` → knowledge graph nodes. - Adaptive skip: `VISION` is only entered when an image path is supplied. **B. `SearchClient` + `RAGPipeline` — built-in keyless web search (search.py).** No API key, no quota ceiling. Backends are tried in order until one returns results; a 90s cooldown suppresses a backend after 3 consecutive failures so one blocked endpoint can't stall the pipeline: | Backend | Trigger | Auth | |---------|---------|------| | DuckDuckGo HTML | default (built-in) | none | | Bing HTML | fallback | none | | Mojeek HTML | fallback (scrape-friendly) | none | | Google Custom Search JSON API | **optional**, only when `GOOGLE_CSE_API_KEY` + `GOOGLE_CSE_ID` env set | API key | ``` SEARCH_GATE (does this need live data?) ├─ force_search flag / live hints ("today", "latest", "news", "price"...) ├─ cache hit in KV (`rag:`) → skip network entirely (G3) └─ depth >= 2 → skip search (cost floor) RAGPipeline.retrieve(query): query ─► SearchClient.search ─► [result titles/urls/snippets] ─► fetch pages (HTML→text) ─► chunk (≤800 chars) ─► rank (token overlap) ─► pack context (≤ max_context_chars) ─► KV cache (ttl=3600) ─► context str ``` Search results are **rule-verified** downstream by the same `RuleEngine` (out-phase): snippets containing blocked content are dropped before packing. **C. `FactExtractor` + `KnowledgeGraph` + `GraphQuery` — NER + graph (facts.py).** ``` web text / image_text ─► FactExtractor ├─ regex NER: PERSON, ORG, LOCATION, DATE, NUMBER, EMAIL, URL ├─ keyword dims: code / math / data / time / place (routing + search hints) └─ triples: (subject, relation, object) via trigger verbs ─► KnowledgeGraph (entity → relation → entities; dedup; save/load JSON) ─► GraphQuery.facts_for_question(question) → grounded context string ``` The graph is **persistent** (`cache/graph.json`), so facts gathered this turn are queryable next turn (G5). Triples are tagged with `source: web|image` so the verifier can cite grounding. `GraphQuery.facts_for_question` matches the question's entities against the graph as **subjects AND objects** (reverse index), and falls back to content-word overlap when the question has no proper-noun entities — so plain questions like "capital of france" still surface relevant edges. **Decision flow (search gate).** `_needs_live_data(prompt, force_search, depth)` returns True when force_search, or a live hint matches AND `depth < 2`. This bounds worst-case web latency while keeping live queries available on demand. --- ### 4.2 L2 — Routing & Skill Load (extend `.skill` → routing matrix) **JSON routing matrix.** Replace per-skill trigger matching with a single matrix file with precedence and typed keys: ```json { "version": 1, "skills": [ { "name": "code_expert", "token": "", "priority": 10, "patterns": [{"type": "regex", "value": "\\b(python|function|debug)\\b"}, {"type": "keyword", "value": "algorithm"}], "adapter": "adapters/code_expert_lora", "requires_syntax": ["balanced_brackets"] } ], "default": {"token": null, "adapter": null} } ``` Routing algorithm: 1. Score every skill: `match = sum(w[type] * hits)`; normalize by pattern count. 2. Highest priority skill with `match > threshold` wins (ties → priority). 3. Winner's `token` → `SkillTokenModule` embedding index (runtime.py:131 hashes skill name → index; v5 uses the matrix `index` field instead of `hash()`, making routing deterministic across runs). 4. Winner's `adapter` → load matching LoRA adapter from KV/disk cache. If already cached in memory, skip load (hot-swap, no retrain). 5. No winner → default (no skill token, base behavior). --- ### 4.3 L3 — ML Core (existing v4, wiring unchanged) `MorphModel.forward` (architecture.py:579): 1. `inputs_embeds = base.embed_tokens(input_ids)` (+ `0.1 * skill_emb` when a skill token is active). 2. `base_out = self.base_model(inputs_embeds, ...)` — frozen Qwen + LoRA adapters → `base_hidden (B, T, H)`. 3. `Coordinator(base_hidden)` → `(gates, steps_dist, steps)` — the System-1/System-2 controller. 4. `MultiStepReasoner(base_hidden, steps)` — weight-tied System-2 loop. 5. Subsystem gating by `adaptive_threshold` (architecture.py:619): `code_bias` if `g_code`, `moe_block` if `g_think`, `memory.read` if `g_mem`, `scratchpad.read` if `g_scratch`. 6. `depth_module` conditioning + `g * (moe_out + mem_out)` scaling. 7. `logits = base.lm_head(refined)`; `verifier_score = VerifierHead(refined)`. 8. Loss = `CE + moe_aux + step_entropy + verifier_mse` (training only). **MoE** (`MoEBlock`, architecture.py:94): per-token top-2-of-4 experts, softmax-normalized routing weights, load-balance aux loss. v5 makes `num_experts` and `top_k` configurable per-skill via routing matrix (`moe.top_k` field) so code/math skills can afford a deeper router than a chat skill. --- ### 4.4 L4 — Search & Refinement **A. Adaptive Depth / Compute Skip (formalize v4 skip algebra).** The Coordinator emits `steps ∈ [0, max_steps]` and gates `g ∈ [0,1]^4`. v5 defines the skip decision explicitly: ``` skip_reasoner = (steps == 0) skip_moe = (g_think < threshold) or easy_query skip_mem = (g_mem < threshold) and cache_miss skip_scratch = (g_scratch < threshold) and no_prior_state cost = base_forward # always + (0 if skip_reasoner else steps * reasoner_layer_cost) + (0 if skip_moe else moe_cost) + (0 if skip_mem else mem_read_cost) ``` "Easy query" floor: if `Coordinator` says `steps == 0` AND `g_think.mean() < threshold`, L4 short-circuits to VERIFY immediately — zero reasoning-loop cost. The hard floor guarantees correctness: an easy query still gets a full base-model forward (G3 holds, nothing is skipped below the base forward). **B. Best-of-N + Verifier (formalize `generate_best_of_n`).** v4 already generates `n` candidates and keeps max verifier score (architecture.py:769). v5 adds: 1. **Normalized scoring:** score = `verifier(refined).mean(1)` normalized to `[0,1]` over the batch so thresholds are stable across candidates. 2. **Early exit:** after each candidate, if `score > accept_threshold` and `score - best < 0.01`, stop — the marginal value of more candidates is ~0 (heuristic search pruning: explore the reasoning-tree only while promising). 3. **Diversity sampling:** candidate `i` uses `temperature + i * Δtemp` so the n samples cover the search space instead of re-rolling the same mode. 4. **Scoring cache:** per-input KV cache stores `(candidate_hash, score)` so re-asking the same question reuses the stored score (no re-generate). Heuristic-search framing: the reasoning tree is searched greedily best-first by verifier score; `n` bounds the beam width; early-exit bounds depth. This is beam search over *generated continuations*, scored by a trained critic rather than a hand-written heuristic. --- ### 4.5 L5 — Verification & Output 1. **Verifier gate:** best candidate's normalized score must be `≥ verify_min`. Below it: (a) if `n` not exhausted → back to GEN for another candidate (FSM REFINE→GEN), (b) exhausted → graceful fallback reply. 2. **Outbound `RuleEngine`:** same JSON rules evaluated on the *output* text (`phase: "out"`). Actions: - `block` → replace with safe refusal. - `mask` → regex-substitute sensitive spans (`mask_pattern`) then allow. - `allow` → pass through. 3. **KV write-back:** response + refined hidden state + verifier score are written to the persistent KV cache (cross-turn context, G5). 4. **Respond:** FSM → RESPOND, emits final string, returns to IDLE. --- ## 5. End-to-End Data Flow (concrete, step-by-step) Given a raw user input `s` (text, optionally `+ image` path) and runtime state `st`: ``` Step FSM Action Data produced ──── ─────────── ──────────────────────────────────────────────────────────── ───────────────────── 1 IDLE receive s; validate encoding/UTF-8 s 2 INTAKE RuntimeFSM.transition(INTAKE); normalize whitespace s_norm 3 GUARD_IN RuleEngine.eval(s_norm, phase="in") 4 GUARD_IN RegexFeatureExtractor.extract(s_norm) → X_feat (B,T,7) X_feat, SyntaxState 5 GUARD_IN syntax gate: unbalanced brackets? → warn|block GateDecision 6 VISION image present? VisionAnalyzer.analyze(path) → ImageFacts image_text, vis_emb (adaptive skip if text-only or trivial) objects[] 7 SEARCH_GATE _needs_live_data? cache hit? depth floor? search_dec 8 SEARCH RAGPipeline.retrieve(query) → fetched chunks → packed ctx search_context 9 FACT_EXTRACT FactExtractor.triples(web+image_text) → graph.add_many graph facts GraphQuery.facts_for_question → grounded context graph_context 10 ROUTED load routing matrix; score skills → winner or default skill_id, token_idx 11 ROUTED adapter cache hit? else load LoRA from disk/KV base_model (frozen+LoRA) 12 GEN pack [context]+[image_text]+prompt; tokenize (+skill token) ids (B,T) 13 GEN MorphModel.forward: embed → base → Coordinator → gating base_hidden, gates, steps 14 GEN steps==0 & g_think None # raises IllegalTransitionError def state(self) -> str # L1 class RuleEngine: def load(self, path: str) -> None # rules.json def eval(self, text: str, phase: str) -> RuleDecision # {action, reply?, masked?} class RegexFeatureExtractor: def extract(self, text: str, tokenizer) -> (Tensor, SyntaxState) def gate(self, syn: SyntaxState) -> GateDecision # L1.5 (v5.5) class VisionAnalyzer: def analyze(self, source) -> ImageFacts # {size, colors, objects, embedding?...} class ImageFacts: def to_text(self) -> str # compact fact string for context/verifier def to_dict(self) -> dict class SearchClient: def search(self, query: str, num: int = 5) -> List[SearchResult] # Google CSE | DDG class RAGPipeline: def retrieve(self, query, num=5, top_k=3, use_cache=True) -> str # packed context class FactExtractor: def extract(self, text) -> Dict[str, Set[str]] # NER types def triples(self, text) -> List[Fact] # (s, r, o) class KnowledgeGraph: def add_many(self, facts); def to_text(self, entities, depth) -> str def save(self, path); def load(self, path) class GraphQuery: def facts_for_question(self, graph, question, depth=1) -> str # L2 class RoutingMatrix: def load(self, path: str) -> None def route(self, text: str) -> RouteDecision # {skill, token_idx, adapter} def load_adapter(self, route: RouteDecision, cache) -> PeftModel | None # L4 def best_of_n(model, ids, n, accept_threshold, ...) -> (best_ids, score) def maybe_skip(coordinator_out, threshold) -> SkipDecision # used by FSM at step 10 # L5 / storage class KVStore: # disk-backed def get(self, key: str) -> Value | None # TTL-aware def set(self, key: str, value: Value, ttl: float) -> None def evict(self) -> None # LRU ``` The runtime composes these: `RuntimeFSM` drives the sequence; `RuleEngine` and `RegexFeatureExtractor` gate both directions; `VisionAnalyzer` + `SearchClient`/`RAGPipeline` + `FactExtractor`/`KnowledgeGraph`/`GraphQuery` produce and ground live context; `RoutingMatrix` feeds `SkillTokenModule` + adapter load; `best_of_n` + `VerifierHead` + `_cross_examine` + skip algebra implement search; `KVStore` + `ScratchpadMemory` + `MemoryModule` + graph persistence carry state across turns. --- ## 7. Trade-offs (solves / worsens / when-to-change) | Choice | Solves | Worsens | When to change | |--------|--------|---------|----------------| | Deterministic rules before model | Safety, compliance, guaranteed cost | Misses model-internal nuance; manual rule maintenance | When rules explode > ~50; move to classifier + rules | | FSM strict transitions | Provable phase order, no re-entrancy | Boilerplate; refactor cost for new phases | When phases become data-driven → codegen FSM from config | | Regex token features | Cheap structure signal, no training | Limited semantic depth | When code-like languages beyond Python/JS added → add grammar tokenizer | | Best-of-N + trained verifier | Strong reasoning/code quality | n× inference cost | When latency-bound on device → reduce n or use early-exit (already in design) | | Adaptive skip | Easy queries are O(1)-ish | Risk of skipping needed compute on mis-gated inputs | Monitor gate calibration on dev set; retrain Coordinator if skip rate drifts | | JSON routing matrix | Skills hot-swappable, no retrain | Matrix tuning; pattern overlap ambiguity | When >20 skills → add learned router | | Disk KV + scratchpad | Cross-turn context, resumability | Stale-state risk if eviction wrong | When TTL wrong → tune TTL by session length | | Vision pixel-fact fallback | VLM works with zero model download | Loses true object semantics | When real ViT weights are bundled → detect() uses YOLOS | | Web crawl vs API search | Keyless HTML backends need no keys and no quota; multi-backend cooldown survives blocks | Scrape fragility, rate limits, no SLA | When Google CSE keys exist → optional 4th backend (env) | | Regex NER + triple extraction | No model, instant, deterministic | Over/under-extraction on complex prose | When precision matters → add spaCy/HF NER behind same interface | | Fact cross-examination | Answers stay grounded; no silent fabrication | Weak-overlap note may annoy on fuzzy queries | Tune overlap threshold; keep note only for fact-dependent turns | --- ## 8. Failure Modes & Mitigation | Failure | Detection | Mitigation | |---------|-----------|------------| | Rule engine blocks a benign prompt (false positive) | user-side complaint / eval set | priority ordering, `warn` tier instead of `block`, allowlist overrides | | Coordinator skips needed compute (under-gate) | low-quality answers on known-hard eval | adaptive_threshold floor; monitor gate distribution | | Verifier over-scores garbage (reward hacking) | score deltas don't match human pref | normalize score; keep verifier trained on seq-likelihood target (v4 loss) | | Skill matrix tie / pattern bleed | two skills claim one prompt | priority field + require_syntax constraints | | KV cache staleness (cross-turn poisoning) | degraded multi-turn | TTL, LRU eviction, scratchpad rewrite on every turn | | Adapter load race (two skills same turn) | adapter mismatch | single-threaded FSM (REQUIRED by design) or per-skill locking | | FSM stuck state on crash | process restart | journal last state; on boot, INTAKE validates and resets to IDLE | | ViT / detector download fails at runtime | vision degraded to pixel facts | catch per-backend, never crash the pipeline (already implemented) | | Search network timeout | empty context, no grounding | RAG returns `_fallback_context` (snippets) or ""; FSM continues; never blocks hard | | NER over-extraction (garbage triples) | graph noise pollutes context | `max_facts` cap in to_text; sources tagged; dedup by subject | | Stale web facts served from cache | outdated answers | KV ttl=3600 on `rag:`; force_search bypasses cache | --- ## 9. Proposed Implementation Order 1. `RuntimeFSM` + transitions table (`src/fsm.py`) — pure logic, no ML dep. 2. `RuleEngine` + `rules.json` (phase-in/out defaults) (`src/rules.py`). 3. `RegexFeatureExtractor` extending `build_code_features` (7-dim + SyntaxState) (`src/regex_features.py`). 4. `RoutingMatrix` replacing `auto_route` + hash-based skill index (`src/routing.py`); keep `.skill` file schema for back-compat. 5. `KVStore` (disk, TTL+LRU) wrapping in-memory `MemoryModule` (`src/kvstore.py`). 6. Wire into `runtime.py`: FSM guards each call; best-of-n early-exit + normalized scoring added to `MorphModel.generate_best_of_n`. 7. Optional: per-skill `top_k`/`num_experts` in routing matrix → `MorphConfig`. 8. `VisionAnalyzer` (vision.py) — ViT + detector lazy-load, pixel-fact fallback. 9. `SearchClient` + `RAGPipeline` (search.py) — Google CSE + DDG crawl, chunk/rank/pack, KV cache. 10. `FactExtractor` + `KnowledgeGraph` + `GraphQuery` (facts.py) — regex NER, triples, persistent graph. 11. Extend `RuntimeFSM` with `VISION` / `SEARCH_GATE` / `SEARCH` / `FACT_EXTRACT`; wire `_ingest` into `chat` / `chat_best_of_n`; `_cross_examine` re-scores the winner against image+web facts. Steps 8–11 are offline-testable (pixel-fact vision, mocked search, regex-only NER) and require no model or API keys. Each step is independently testable (rules and FSM have zero model dependencies and can be unit-tested offline). --- ## 10. Key Takeaway The ML layers (CLM/MoE/SFT-LoRA/VLM-ViT) provide **trained intelligence**; the deterministic layers (rules/FSM/regex) provide **guaranteed structure and safety**; the live-knowledge layers (search/RAG/NER/graph) provide **grounded, fresh context** that the verifier cross-examines; the search layer (best-of-N + verifier + skip algebra) provides **controlled cost and verifiable quality**; the storage layer (JSON matrix + KV cache + knowledge graph) provides **persistence and hot-swappability**. The FSM is the spine that makes the whole composition provable: the ML model only ever runs inside legal states, and no output leaves without passing a verifier score, a fact cross-examination, and an outbound rule gate.