| # V1 Experiment and Progress Record |
|
|
| Snapshot date: 2026-08-05 |
|
|
| ## Completed work |
|
|
| - Built a Python package and `fugu-lite` CLI with mock, OpenRouter, and generic |
| OpenAI-compatible providers. |
| - Implemented deterministic graders, optional LLM judging, quality/cost/latency utility, |
| resumable reward generation, routing, worker calls, and an OpenAI-compatible API. |
| - Implemented SFT, contextual-bandit RL, Sep-CMA-ES, checkpoint evaluation, and eight tests. |
| - Proved the no-cost mock route across math, code, and general specialists. |
| - Generated a real 300-task benchmark set and repeated every worker call three times. |
| - Saved the full reward matrix, SFT and ES routing heads, tokenizers, reports, evaluations, |
| and Python distribution packages. |
|
|
| ## Saved V1 data |
|
|
| | Item | Value | |
| |---|---:| |
| | Tasks | 300 | |
| | Train / validation / test | 240 / 30 / 30 | |
| | MMLU / ARC-Challenge / MMLU-Pro | 150 / 100 / 50 | |
| | Reasoning / math / code | 278 / 16 / 6 | |
| | Workers | 3 | |
| | Repetitions per task and worker | 3 | |
| | Total recorded worker calls | 2,700 | |
| | Recorded call errors | 0 | |
| | Empty responses | 88, all from the DeepSeek worker | |
| | Recorded total API cost | $0.453761 | |
|
|
| The recorded worker pool was: |
|
|
| - `qwen/qwen3-8b` |
| - `deepseek/deepseek-v4-flash-0731` |
| - `google/gemini-3.1-flash-lite` |
|
|
| Model availability and pricing can change. These identifiers describe the saved experiment, |
| not a promise that the same OpenRouter routes are currently available. |
|
|
| ## Worker observations |
|
|
| | Worker | Calls | Mean quality | Mean cost/call | Mean latency | |
| |---|---:|---:|---:|---:| |
| | Qwen | 900 | 0.8200 | $0.00040353 | 14,797 ms | |
| | DeepSeek | 900 | 0.8544 | $0.00005951 | 5,821 ms | |
| | Gemini | 900 | 0.8600 | $0.00004114 | 1,371 ms | |
|
|
| These are descriptive values from the saved response records. They are not a controlled |
| provider benchmark and should not be generalized beyond this dataset and run configuration. |
|
|
| ## Training and evaluation |
|
|
| The SFT run completed five epochs and 150 optimizer steps. Its held-out 30-task evaluation |
| scored 0.8111 router utility, below the best fixed worker at 0.8667. |
|
|
| The Sep-CMA-ES run completed 30 generations with population size 16. Its best training |
| utility reached 0.9014. On the held-out set it scored: |
|
|
| | Metric | ES result | |
| |---|---:| |
| | Router utility | 0.8667 | |
| | Oracle utility | 0.9000 | |
| | Best fixed utility | 0.8667 | |
| | Random utility | 0.8370 | |
| | Regret | 0.0333 | |
| | Routes: Qwen / DeepSeek / Gemini | 3 / 3 / 24 | |
|
|
| The ES router matched the best fixed worker but did not beat it. This is a working pipeline, |
| not yet evidence of a useful learned production router. |
|
|
| ## Why V1 stopped here |
|
|
| The reward matrix has weak and heavily imbalanced routing signal: |
|
|
| | Signal check | Count | Share | |
| |---|---:|---:| |
| | Any tie for best reward | 269 | 89.7% | |
| | All workers equal | 226 | 75.3% | |
| | Informative reward spread | 74 | 24.7% | |
| | Unique winner | 31 | 10.3% | |
|
|
| Reasoning accounts for 92.7% of the tasks, while code has only six examples. More training on |
| the same matrix would mostly reinforce the imbalance. The correct next step is better data, |
| not additional V1 optimization. |
|
|
| ## Environment lesson |
|
|
| The server initially had a project `.venv` using Python 3.14 nested inside an activated |
| Conda Python 3.12 environment. The inner `.venv` remained first on `PATH`, so installation |
| correctly failed the package's `<3.14` requirement. The recovery was to deactivate both |
| layers, move the old `.venv` aside, activate only the Python 3.12 environment, verify |
| `sys.executable`, and then reinstall the project. |
|
|
| ## V2 continuation plan |
|
|
| 1. Replace the current domain mapping with a deliberately balanced, harder task builder. |
| 2. Create a 90-task pilot: 30 math, 30 code, and 30 reasoning tasks. |
| 3. Run three workers with three repetitions: 810 worker calls. |
| 4. Measure domain balance, ties, unique winners, empty responses, and held-out baselines. |
| 5. Scale to at least 600 tasks only if the pilot has healthy disagreement. |
| 6. Retrain SFT first, add RL or ES only when held-out utility improves, and require the router |
| to beat the best fixed worker before adding multi-agent delegation. |
|
|
| ## Where progress is stored |
|
|
| | Path | Contents | |
| |---|---| |
| | `data/tasks.real-300.jsonl` | Task prompts, sources, labels, domains, and splits | |
| | `data/rewards.real-300.jsonl` | Repeated responses, rewards, costs, latency, tokens, and IDs | |
| | `artifacts/router-sft-real-v1/` | SFT routing head, tokenizer, configuration, and report | |
| | `artifacts/router-es-real-v1/` | ES routing head, tokenizer, configuration, and report | |
| | `artifacts/eval-*-real-v1.json` | Held-out evaluation summaries | |
| | `dist/` | Version 0.1.0 wheel and source distribution | |
|
|
| The real `OPENROUTER_API_KEY` is not present in these files. Keep it only in an ignored local |
| `.env` file. |
|
|