fsi-anomaly / docs /harness_research.md
FerrellSyntheticIntelligence's picture
backup all: 100 files (batch)
d83b47a verified
|
Raw
History Blame Contribute Delete
8.31 kB

Harness research β€” what big tech recommends for a small-model "suit" (2026-08-09)

Multi-source digest (OpenAI, Anthropic, Google, Meta, Microsoft, DeepSeek, MIT). Purpose: encode the battle-tested system-layer doctrine behind skills/ tiny-model-harness. Full claim: every rule below cites a source.

1. Agent = model + tools + instructions, with guardrails (OpenAI)

  • "A practical guide to building agents" (OpenAI, 2025): an agent is three components β€” MODEL, TOOLS, INSTRUCTIONS β€” plus GUARDRAILS. Tools should have standardized definitions and fall into data/action/compute types. Start with ONE agent and add tools; go multi-agent only when it pays. Guardrails are first-class: relevance classifier, safety classifier, PII filter, rules-based protections (blocklists, regex), output validation, and human-in-the-loop.
  • Map to FSI: model = TinyLiquid; tools = BM25, web/tor, workspace, helix memory; instructions = SOP library (the Codex-style "task bar" agent.py already implements: plan + ledger + guardrails). MISSING: input guardrails (relevance/safety/injection) β€” built 2026-08-09 in research/guardrails.py.

2. Simple composable patterns; workflows before agents (Anthropic)

  • "Building effective agents" (Anthropic, 2024-12): the most successful implementations use simple, composable patterns, not frameworks. Use deterministic WORKFLOWS (prompt chaining, routing, parallelization, evaluator-optimizer, orchestrator-workers) where the path is predictable; use model-driven AGENTS only where decisions are needed at runtime. Add guardrails (off-topic detection, output validation) and plan for human intervention. "Find the simplest solution and only increase complexity when needed."
  • Map: our rule spine (verify.py) + constrained decode + calibration table are deterministic workflows; the head only runs on fall-through. Correct shape.

3. Context engineering: context is a finite resource (Anthropic)

  • "Effective context engineering for AI agents" (Anthropic, 2025-09): curate what goes into the window β€” structure it, retrieve instead of stuffing, compact/summarize old material, dedupe, write-before-you-read, keep the goal visible. Long sessions degrade; the system must manage its own context.
  • Map: helix memory + workspace artifacts + SOP blocks are our context curation. GAP: no budget/compaction guard in the TUI for long rabbit-hole sessions β€” planned next.

4. Calibration: format matters; verbalized confidence is anti-calibrated

  • "Language Models (Mostly) Know What They Know" (Anthropic, 2022): models are well-calibrated only in the RIGHT format; smaller models are worse.
  • arXiv 2408.11774 (2024): verbalized confidence is less honest than probabilistic confidence; ORCE (2026), Direct Confidence Alignment (2025): verbalized labels need alignment to measured accuracy.
  • Rule (applied 2026-08-09): confidence labels are inputs to a CALIBRATION TABLE (research/calibration.py), never trusted directly; the decision spine (research/decision.py) votes with measured per-bucket accuracy and abstains below threshold (selective prediction β€” conformal selective prediction arXiv 2026; governance-ready SLM recipe 2025).

5. Verification must be EXTERNAL, not self-critique

  • "Large Language Models Cannot Self-Correct Reasoning Yet" (arXiv 2310.01798, 2023): intrinsic self-correction fails; feedback must be external.
  • "Small Language Models Need Strong Verifiers to Self-Correct" (arXiv 2404.09931, 2024): small models especially need STRONG (external/learned) verifiers, not weak self-critique.
  • CRITIC (Microsoft, arXiv 2305.11738): tool-interactive critiquing β€” search and code execution act as the critic. Chain-of-Verification (Meta, arXiv 2309.09308): draft -> verification questions -> answer independently -> revise.
  • Map: deterministic rule spine (research/verify.py) + retrieval = the external critic. Built 2026-08-09: research/verify_loop.py (draft -> plan checks -> retrieve -> deterministic compare -> revise + trace).

6. Self-consistency: sample and vote β€” but weight it

  • Self-Consistency (Google, arXiv 2203.11171): sample N reasoning paths, majority vote beats greedy. Universal SC (2311.08110): no manual CoT needed.
  • Rule (applied): keep N small (3-5) on the fall-through path; weight by CALIBRATED reliability (research/decision.py), not naive majority.

7. Multi-agent debate improves factuality (MIT/Google, arXiv 2305.14325)

  • Independent agents debate/aggregate; improves factuality and reasoning.
  • Map: dual-mind fusion (research/fusion.py) + 4-angle orchestrator (research/orchestrator.py). GAP: fusion raises confidence naively on agreement β€” replace with calibrated merge (decision.py) when the champion calibration lands.

8. RL with verifiable rewards (DeepSeek-R1, arXiv 2501.12948; Reasoning Gym

arXiv 2025-05)

  • Outcome supervision on verifiable problems incentivizes correct reasoning. Reasoning Gym provides a verifier library.
  • Map: our constrained verdict space + decision spine are a verifier; RLVR on verdicts is recorded as the training-side unlock (not this week).

9. Distillation: small CAN beat big with the right curriculum

  • "Distilling Step-by-Step" (Google, arXiv 2305.02301): CoT distillation lets smaller models OUTPERFORM larger with less data. Supports the KD pipeline (tiny-model-kd) and the handcrafted-gold rule.

10. Training pipeline research β€” what big tech actually does (2026-08-10)

Multi-source digest for the tiny-model-pipeline skill. Every claim sourced.

  1. SmolLM (HF, 2024) β€” huggingface.co/blog/smollm: curated corpus (Cosmopedia v2 synthetic textbooks/stories by Mixtral + FineWeb-Edu + code), trapezoidal LR with cooldown = 20% of training, 49,152 vocab tokenizer trained on the corpus, embedding tying, context 2048. Instruct: SFT lr 3e-4 (Zephyr-Gemma alignment handbook recipe) then DPO 1 epoch (HelpSteer 135M/1.7B, dpo-mix-7k 360M).
  2. SmolLM2 (HF, 2025) β€” arXiv 2502.02737: 135M on 2T tokens, multi-stage training mixing web/math/code/instruction data; data-mix rates refined at each stage from previous-stage eval (small-scale ablations + manual refinement). Instruct: SFT (SmolTalk) then DPO (UltraFeedback).
  3. Zephyr (2023) β€” arXiv 2310.16944: dSFT then dDPO, few hours, no sampling during fine-tune.
  4. Phi-3 (MS, 2024) β€” arXiv 2404.14219: 3.8B on 3.3T tokens of heavily filtered web + synthetic; data quality is the lever at small scale.
  5. DeepSeek-R1 (2025) β€” arXiv 2501.12948: RL with verifiable rewards (RLVR) incentivizes reasoning without human-annotated traces.
  6. LIMA (2023) β€” arXiv 2305.11206: 1,000 hand-curated examples shape style; supports the handcrafted-gold directive.
  7. Forgetting β€” arXiv 2401.05605 (LoRA still forgets; inverse-linear perf/forgetting tradeoff, not fixable by early stop/rank) and 2502.06042 (injecting ~1%+ pretraining data into the finetune mixture prevents drift/overfit).
  8. Liquid LFM2 (2025) β€” liquid.ai/research/lfm2-technical-report (the architecture family we use): training pipeline = tempered decoupled Top-K distillation objective (avoids support mismatch); curriculum learning with difficulty-ordered data; three-stage post-training = supervised fine-tuning -> length-normalized preference optimization -> model merging. 350M-8.3B, 32K context.

Audit of OUR pipeline vs the recipe (2026-08-10)

MATCH: AdamW (0.9,0.95) wd 0.1 grad-clip 1.0; cosine LR+warmup (GPT-3/LLaMA standard; SmolLM uses trapezoidal+cooldown); BF16; embedding tying; SFT then DPO 1 epoch; handcrafted gold (LIMA/phi-1, better than synthetic here); replay 0.5; curriculum stages; eval gates every stage. GAP: 8k tokenizer (merge artifacts; SmolLM 49k, 16k retrain queued). GAP: model merging after preference (LFM2 3rd stage) β€” naive averaging failed at 25M; TIES/task-vector untested at 50M. GAP: length-normalized preference optimization (LFM2) β€” our DPO full-epoch collapsed to "abstain"; length norm / IPO may fix. GAP: RLVR (DeepSeek-R1) β€” verifier exists (constrained verdicts + decision spine); recorded unlock. GAP: explicit cooldown tail (SmolLM 20%) β€” cosine covers it; optional.