V1 Experiment and Progress Record
Snapshot date: 2026-08-05
Completed work
- Built a Python package and
fugu-liteCLI with mock, OpenRouter, and generic OpenAI-compatible providers. - Implemented deterministic graders, optional LLM judging, quality/cost/latency utility, resumable reward generation, routing, worker calls, and an OpenAI-compatible API.
- Implemented SFT, contextual-bandit RL, Sep-CMA-ES, checkpoint evaluation, and eight tests.
- Proved the no-cost mock route across math, code, and general specialists.
- Generated a real 300-task benchmark set and repeated every worker call three times.
- Saved the full reward matrix, SFT and ES routing heads, tokenizers, reports, evaluations, and Python distribution packages.
Saved V1 data
| Item | Value |
|---|---|
| Tasks | 300 |
| Train / validation / test | 240 / 30 / 30 |
| MMLU / ARC-Challenge / MMLU-Pro | 150 / 100 / 50 |
| Reasoning / math / code | 278 / 16 / 6 |
| Workers | 3 |
| Repetitions per task and worker | 3 |
| Total recorded worker calls | 2,700 |
| Recorded call errors | 0 |
| Empty responses | 88, all from the DeepSeek worker |
| Recorded total API cost | $0.453761 |
The recorded worker pool was:
qwen/qwen3-8bdeepseek/deepseek-v4-flash-0731google/gemini-3.1-flash-lite
Model availability and pricing can change. These identifiers describe the saved experiment, not a promise that the same OpenRouter routes are currently available.
Worker observations
| Worker | Calls | Mean quality | Mean cost/call | Mean latency |
|---|---|---|---|---|
| Qwen | 900 | 0.8200 | $0.00040353 | 14,797 ms |
| DeepSeek | 900 | 0.8544 | $0.00005951 | 5,821 ms |
| Gemini | 900 | 0.8600 | $0.00004114 | 1,371 ms |
These are descriptive values from the saved response records. They are not a controlled provider benchmark and should not be generalized beyond this dataset and run configuration.
Training and evaluation
The SFT run completed five epochs and 150 optimizer steps. Its held-out 30-task evaluation scored 0.8111 router utility, below the best fixed worker at 0.8667.
The Sep-CMA-ES run completed 30 generations with population size 16. Its best training utility reached 0.9014. On the held-out set it scored:
| Metric | ES result |
|---|---|
| Router utility | 0.8667 |
| Oracle utility | 0.9000 |
| Best fixed utility | 0.8667 |
| Random utility | 0.8370 |
| Regret | 0.0333 |
| Routes: Qwen / DeepSeek / Gemini | 3 / 3 / 24 |
The ES router matched the best fixed worker but did not beat it. This is a working pipeline, not yet evidence of a useful learned production router.
Why V1 stopped here
The reward matrix has weak and heavily imbalanced routing signal:
| Signal check | Count | Share |
|---|---|---|
| Any tie for best reward | 269 | 89.7% |
| All workers equal | 226 | 75.3% |
| Informative reward spread | 74 | 24.7% |
| Unique winner | 31 | 10.3% |
Reasoning accounts for 92.7% of the tasks, while code has only six examples. More training on the same matrix would mostly reinforce the imbalance. The correct next step is better data, not additional V1 optimization.
Environment lesson
The server initially had a project .venv using Python 3.14 nested inside an activated
Conda Python 3.12 environment. The inner .venv remained first on PATH, so installation
correctly failed the package's <3.14 requirement. The recovery was to deactivate both
layers, move the old .venv aside, activate only the Python 3.12 environment, verify
sys.executable, and then reinstall the project.
V2 continuation plan
- Replace the current domain mapping with a deliberately balanced, harder task builder.
- Create a 90-task pilot: 30 math, 30 code, and 30 reasoning tasks.
- Run three workers with three repetitions: 810 worker calls.
- Measure domain balance, ties, unique winners, empty responses, and held-out baselines.
- Scale to at least 600 tasks only if the pilot has healthy disagreement.
- Retrain SFT first, add RL or ES only when held-out utility improves, and require the router to beat the best fixed worker before adding multi-agent delegation.
Where progress is stored
| Path | Contents |
|---|---|
data/tasks.real-300.jsonl |
Task prompts, sources, labels, domains, and splits |
data/rewards.real-300.jsonl |
Repeated responses, rewards, costs, latency, tokens, and IDs |
artifacts/router-sft-real-v1/ |
SFT routing head, tokenizer, configuration, and report |
artifacts/router-es-real-v1/ |
ES routing head, tokenizer, configuration, and report |
artifacts/eval-*-real-v1.json |
Held-out evaluation summaries |
dist/ |
Version 0.1.0 wheel and source distribution |
The real OPENROUTER_API_KEY is not present in these files. Keep it only in an ignored local
.env file.