fugu-lite / docs /V1_PROGRESS.md
tahsinsoyak's picture
Upload private Fugu-Lite V1 snapshot
88e15cd verified
|
Raw
History Blame Contribute Delete
4.79 kB
# V1 Experiment and Progress Record
Snapshot date: 2026-08-05
## Completed work
- Built a Python package and `fugu-lite` CLI with mock, OpenRouter, and generic
OpenAI-compatible providers.
- Implemented deterministic graders, optional LLM judging, quality/cost/latency utility,
resumable reward generation, routing, worker calls, and an OpenAI-compatible API.
- Implemented SFT, contextual-bandit RL, Sep-CMA-ES, checkpoint evaluation, and eight tests.
- Proved the no-cost mock route across math, code, and general specialists.
- Generated a real 300-task benchmark set and repeated every worker call three times.
- Saved the full reward matrix, SFT and ES routing heads, tokenizers, reports, evaluations,
and Python distribution packages.
## Saved V1 data
| Item | Value |
|---|---:|
| Tasks | 300 |
| Train / validation / test | 240 / 30 / 30 |
| MMLU / ARC-Challenge / MMLU-Pro | 150 / 100 / 50 |
| Reasoning / math / code | 278 / 16 / 6 |
| Workers | 3 |
| Repetitions per task and worker | 3 |
| Total recorded worker calls | 2,700 |
| Recorded call errors | 0 |
| Empty responses | 88, all from the DeepSeek worker |
| Recorded total API cost | $0.453761 |
The recorded worker pool was:
- `qwen/qwen3-8b`
- `deepseek/deepseek-v4-flash-0731`
- `google/gemini-3.1-flash-lite`
Model availability and pricing can change. These identifiers describe the saved experiment,
not a promise that the same OpenRouter routes are currently available.
## Worker observations
| Worker | Calls | Mean quality | Mean cost/call | Mean latency |
|---|---:|---:|---:|---:|
| Qwen | 900 | 0.8200 | $0.00040353 | 14,797 ms |
| DeepSeek | 900 | 0.8544 | $0.00005951 | 5,821 ms |
| Gemini | 900 | 0.8600 | $0.00004114 | 1,371 ms |
These are descriptive values from the saved response records. They are not a controlled
provider benchmark and should not be generalized beyond this dataset and run configuration.
## Training and evaluation
The SFT run completed five epochs and 150 optimizer steps. Its held-out 30-task evaluation
scored 0.8111 router utility, below the best fixed worker at 0.8667.
The Sep-CMA-ES run completed 30 generations with population size 16. Its best training
utility reached 0.9014. On the held-out set it scored:
| Metric | ES result |
|---|---:|
| Router utility | 0.8667 |
| Oracle utility | 0.9000 |
| Best fixed utility | 0.8667 |
| Random utility | 0.8370 |
| Regret | 0.0333 |
| Routes: Qwen / DeepSeek / Gemini | 3 / 3 / 24 |
The ES router matched the best fixed worker but did not beat it. This is a working pipeline,
not yet evidence of a useful learned production router.
## Why V1 stopped here
The reward matrix has weak and heavily imbalanced routing signal:
| Signal check | Count | Share |
|---|---:|---:|
| Any tie for best reward | 269 | 89.7% |
| All workers equal | 226 | 75.3% |
| Informative reward spread | 74 | 24.7% |
| Unique winner | 31 | 10.3% |
Reasoning accounts for 92.7% of the tasks, while code has only six examples. More training on
the same matrix would mostly reinforce the imbalance. The correct next step is better data,
not additional V1 optimization.
## Environment lesson
The server initially had a project `.venv` using Python 3.14 nested inside an activated
Conda Python 3.12 environment. The inner `.venv` remained first on `PATH`, so installation
correctly failed the package's `<3.14` requirement. The recovery was to deactivate both
layers, move the old `.venv` aside, activate only the Python 3.12 environment, verify
`sys.executable`, and then reinstall the project.
## V2 continuation plan
1. Replace the current domain mapping with a deliberately balanced, harder task builder.
2. Create a 90-task pilot: 30 math, 30 code, and 30 reasoning tasks.
3. Run three workers with three repetitions: 810 worker calls.
4. Measure domain balance, ties, unique winners, empty responses, and held-out baselines.
5. Scale to at least 600 tasks only if the pilot has healthy disagreement.
6. Retrain SFT first, add RL or ES only when held-out utility improves, and require the router
to beat the best fixed worker before adding multi-agent delegation.
## Where progress is stored
| Path | Contents |
|---|---|
| `data/tasks.real-300.jsonl` | Task prompts, sources, labels, domains, and splits |
| `data/rewards.real-300.jsonl` | Repeated responses, rewards, costs, latency, tokens, and IDs |
| `artifacts/router-sft-real-v1/` | SFT routing head, tokenizer, configuration, and report |
| `artifacts/router-es-real-v1/` | ES routing head, tokenizer, configuration, and report |
| `artifacts/eval-*-real-v1.json` | Held-out evaluation summaries |
| `dist/` | Version 0.1.0 wheel and source distribution |
The real `OPENROUTER_API_KEY` is not present in these files. Keep it only in an ignored local
`.env` file.