fugu-lite / docs /V1_PROGRESS.md
tahsinsoyak's picture
Upload private Fugu-Lite V1 snapshot
88e15cd verified
|
Raw
History Blame Contribute Delete
4.79 kB

V1 Experiment and Progress Record

Snapshot date: 2026-08-05

Completed work

  • Built a Python package and fugu-lite CLI with mock, OpenRouter, and generic OpenAI-compatible providers.
  • Implemented deterministic graders, optional LLM judging, quality/cost/latency utility, resumable reward generation, routing, worker calls, and an OpenAI-compatible API.
  • Implemented SFT, contextual-bandit RL, Sep-CMA-ES, checkpoint evaluation, and eight tests.
  • Proved the no-cost mock route across math, code, and general specialists.
  • Generated a real 300-task benchmark set and repeated every worker call three times.
  • Saved the full reward matrix, SFT and ES routing heads, tokenizers, reports, evaluations, and Python distribution packages.

Saved V1 data

Item Value
Tasks 300
Train / validation / test 240 / 30 / 30
MMLU / ARC-Challenge / MMLU-Pro 150 / 100 / 50
Reasoning / math / code 278 / 16 / 6
Workers 3
Repetitions per task and worker 3
Total recorded worker calls 2,700
Recorded call errors 0
Empty responses 88, all from the DeepSeek worker
Recorded total API cost $0.453761

The recorded worker pool was:

  • qwen/qwen3-8b
  • deepseek/deepseek-v4-flash-0731
  • google/gemini-3.1-flash-lite

Model availability and pricing can change. These identifiers describe the saved experiment, not a promise that the same OpenRouter routes are currently available.

Worker observations

Worker Calls Mean quality Mean cost/call Mean latency
Qwen 900 0.8200 $0.00040353 14,797 ms
DeepSeek 900 0.8544 $0.00005951 5,821 ms
Gemini 900 0.8600 $0.00004114 1,371 ms

These are descriptive values from the saved response records. They are not a controlled provider benchmark and should not be generalized beyond this dataset and run configuration.

Training and evaluation

The SFT run completed five epochs and 150 optimizer steps. Its held-out 30-task evaluation scored 0.8111 router utility, below the best fixed worker at 0.8667.

The Sep-CMA-ES run completed 30 generations with population size 16. Its best training utility reached 0.9014. On the held-out set it scored:

Metric ES result
Router utility 0.8667
Oracle utility 0.9000
Best fixed utility 0.8667
Random utility 0.8370
Regret 0.0333
Routes: Qwen / DeepSeek / Gemini 3 / 3 / 24

The ES router matched the best fixed worker but did not beat it. This is a working pipeline, not yet evidence of a useful learned production router.

Why V1 stopped here

The reward matrix has weak and heavily imbalanced routing signal:

Signal check Count Share
Any tie for best reward 269 89.7%
All workers equal 226 75.3%
Informative reward spread 74 24.7%
Unique winner 31 10.3%

Reasoning accounts for 92.7% of the tasks, while code has only six examples. More training on the same matrix would mostly reinforce the imbalance. The correct next step is better data, not additional V1 optimization.

Environment lesson

The server initially had a project .venv using Python 3.14 nested inside an activated Conda Python 3.12 environment. The inner .venv remained first on PATH, so installation correctly failed the package's <3.14 requirement. The recovery was to deactivate both layers, move the old .venv aside, activate only the Python 3.12 environment, verify sys.executable, and then reinstall the project.

V2 continuation plan

  1. Replace the current domain mapping with a deliberately balanced, harder task builder.
  2. Create a 90-task pilot: 30 math, 30 code, and 30 reasoning tasks.
  3. Run three workers with three repetitions: 810 worker calls.
  4. Measure domain balance, ties, unique winners, empty responses, and held-out baselines.
  5. Scale to at least 600 tasks only if the pilot has healthy disagreement.
  6. Retrain SFT first, add RL or ES only when held-out utility improves, and require the router to beat the best fixed worker before adding multi-agent delegation.

Where progress is stored

Path Contents
data/tasks.real-300.jsonl Task prompts, sources, labels, domains, and splits
data/rewards.real-300.jsonl Repeated responses, rewards, costs, latency, tokens, and IDs
artifacts/router-sft-real-v1/ SFT routing head, tokenizer, configuration, and report
artifacts/router-es-real-v1/ ES routing head, tokenizer, configuration, and report
artifacts/eval-*-real-v1.json Held-out evaluation summaries
dist/ Version 0.1.0 wheel and source distribution

The real OPENROUTER_API_KEY is not present in these files. Keep it only in an ignored local .env file.