| # RUN NOTES β agentic SWE/terminal post-training (100h) |
|
|
| Deadline epoch: 1786991334 (`date -d @1786991334`) β boot 2026-08-13T14:28Z, ends 2026-08-17T18:28Z. |
| Check remaining: `echo $(( $(cat DEADLINE | cut -d. -f1) - $(date +%s) ))` seconds. |
|
|
| ## Environment map (verified) |
|
|
| | thing | path | |
| |---|---| |
| | workspace | `/mnt/pvc/users/simon/agentptb/runs/a-opus-max/workspace` (= `$AGENTPTB_WORKSPACE`) | |
| | prime-rl | `/root/work/a/prime-rl` β real `/mnt/pvc/users/simon/agentptb/work/a/prime-rl` | |
| | venv | `/root/work/a/prime-rl/.venv` (bins: `sft`, `rl`, `eval`, `orchestrator`, `trainer`, `inference`) | |
| | verifiers src (editable) | `/root/work/a/prime-rl/deps/verifiers` | |
| | research-environments (editable tasksets) | `/root/work/a/prime-rl/deps/research-environments/environments/{swe,terminal,tool_use,code,...}` | |
| | base model | `$HF_HOME/hub/models--Qwen--Qwen3.5-9B-Base/snapshots/68c46c4b...` | |
| | eval tasksets on disk | `/root/work/shared/tasksets/{terminal-bench-2,swe-bench-verified}` (symlinked into `~/.cache/harbor/`) | |
| | my GPUs | `CUDA_VISIBLE_DEVICES=0,1,2,3` (B200 183GB each) | |
| | sandbox broker | `$SANDBOX_BASE_URL` + `X-API-Key: $SANDBOX_API_KEY` | |
| | TMPDIR / caches | `/var/lib/agentptb-cache/a/tmp` (local disk β keep them there) | |
|
|
| ## Model facts |
|
|
| - `Qwen3.5-9B-Base`: `Qwen3_5ForConditionalGeneration`, hybrid linear+full attention (full every 4th layer), |
| 32 layers, hidden 4096, head_dim 256, vocab 248320, **has a vision tower** (β `limit_mm_per_prompt={image=0,video=0}`). |
| - **It already ships the full Qwen3.5 ChatML chat template** (`<|im_start|>`/`<|im_end|>`, `<tool_call>`, `<think>`). |
| - `eos_token = <|endoftext|>` (248044) but turns end with `<|im_end|>` (248046). No `generation_config.json`. |
| β vLLM stops only on 248044 β base model runs to the token cap every turn. **Fix: add |
| `generation_config.json` with `eos_token_id:[248044,248046]`** (and/or stop strings in sampling). |
| |
| ## Non-obvious operational rules (from eval-kit README) |
| |
| - serve with `--enable-auto-tool-choice --tool-call-parser qwen3_coder` else 400s |
| - runtime `block_network = false` (pi installs itself in the task container at rollout time) |
| - `export PRIME_API_KEY="$(cat "$AGENTPTB_PRIME_KEY_FILE")"` β already in env too |
| - trainer `attn = "flash_attention_2"` (auto β fa4 β grad norm inf β silently no-op) |
| - `batch_size` must be several multiples of `group_size` |
| - kill orphan trainers/vllm (they hold GPU mem + ports) |
| |
| ## Rules I must respect |
| |
| - No training on terminal-bench-2 / swe-bench-verified (reading/running them = fine). No test-item-clustered synthetic data. |
| - Submitted weights must derive from `Qwen3.5-9B-Base` by my own training. **Do not touch `Qwen/Qwen3.5-9B` |
| (post-trained sibling) as a weights source** β that voided a prior run. |
| - Distillation from public models I run myself IS allowed (rule 7). Env API keys are NOT a data source (rule 5). |
| - Do not read: operator notes, other cells' workspaces, prior runs' logs, benchmark repo. |
| (`/root/work/operator-debug`, `$AGENTPTB_WORKSPACE_{A,B}` β leave alone.) |
| NOTE: `$HF_HOME` is a *shared* cache and contains datasets other cells pulled. I use it only as a |
| model/dataset cache for things I decide on independently; I do not treat its contents as a strategy signal. |
|
|
| ## Key findings |
|
|
| 1. **`generation_config.json` is missing on the base and it matters a lot.** Chat template ends turns |
| with `<|im_end|>`; without it as an eos id vLLM never stops there, so the base model *continues past |
| its own turn and hallucinates the user/tool turns* (16384-token completions, `stop=context_length`). |
| The official post-trained sibling ships `eos_token_id:[248046,248044]` β so this is correct packaging, |
| not a trick. Added to `/var/lib/agentptb-cache/a/models/base/generation_config.json`; mean completion |
| tokens/turn dropped from ~16k to ~250. **Must be in every checkpoint I serve/submit.** |
| 2. **Template is thinking-ON**: generation prompt ends `<|im_start|>assistant\n<think>\n`. |
| Tool-call syntax is Qwen3-Coder XML (`<tool_call><function=name><parameter=x>`), matching |
| `--tool-call-parser qwen3_coder`. Don't fight it β train short `<think>` blocks. |
| 3. **Renderer must be pinned**: `Qwen/Qwen3.5-9B-Base` is NOT in `MODEL_RENDERER_MAP`, so `auto` falls |
| back to the default renderer (no tool support). Always set `renderer.name = "qwen3.5"`. |
| 4. `Qwen/Qwen3.5-35B-A3B` (post-trained MoE, 3B active, 72GB) has a **byte-identical chat template** to |
| the base β ideal distillation teacher. Downloaded to `/var/lib/agentptb-cache/a/models/teacher35`. |
| Used ONLY as a data generator (rule 7), never as weights (rule 3). |
| 5. pi defaults `maxTokens=16384` per turn (`models.json` knob) β harness-side lever for my own harness. |
| 6. Trace tool_calls are FLAT (`{id,name,arguments}`); prime-rl SFT needs OAI nested |
| (`{id,type,function:{name,arguments}}`) or `deserialize_tool_calls` breaks the renderer. |
| 7. Broker pulls arbitrary public Docker Hub images fine (tested ubuntu, swebench/*, alexgshaw/*). |
| It does NOT pull Prime-platform refs (`prime/primeintellect/...`) β `tmax-v1`, `terminal-lego-v1`, |
| `openthoughts-tblite-v1` are unusable. `swesmith-v1` images also fail (TaskError/SandboxError). |
| 8. **`BrokerRuntime` has no `workdir` field.** `utils/compile.py` only applies `task.data.workdir` |
| when the runtime config *has* that field, so with the broker every exec runs in `/workspace` |
| (a scratch dir), never the task's `/testbed`. Two consequences: |
| - the agent lands outside the repo and must find it (only 43/103 teacher rollouts ever |
| mentioned `/testbed`, median 7 turns in) β a real, large capability tax on swe-bench; |
| - **`r2e-gym-v1` scores 0 for everyone**: its `solved()` runs `sh -c "/bin/bash run_tests.sh"` |
| with a *relative* path, which does not exist in `/workspace`. Same class of bug makes its |
| `capture_patch` fail with `fatal: not a git repository`. This is why the 35B teacher scored |
| 0/52 on r2e-gym β not a capability result. |
| `swebench-verified-v1` (harbor verifier, absolute paths) is unaffected and does score. |
| 9. Round-trip verified: my converted rows render **byte-identical** to |
| `tokenizer.apply_chat_template` under the `qwen3.5` renderer, loss mask = assistant spans |
| (`<think>β¦</think>β¦<tool_call>β¦<|im_end|>`). 8/8 exact on a sample. |
| |
| 10. **No local container escape hatch.** `docker`/`podman` exist on this node but rootless podman |
| cannot start a container: pulls work once `~/.config/containers/storage.conf` sets |
| `ignore_chown_errors=true` and `USER` is corrected from `root` to `agentptb` (that env var |
| being wrong is what produces the misleading "no subuid ranges for user root"), but every run |
| then dies at `crun: mount 'proc' to 'proc': Operation not permitted` β the pod has |
| `CapEff=0` and mount(2) is blocked. The broker really is the only way to run tasks. Don't |
| retry this. |
| 11. The broker exposes **`GET /resources`** (undocumented in the runbook): node capacity plus |
| every sandbox pod and its phase. This is the way to tell "my job is slow" from "the cluster |
| is full" β `{'Pending': 222, 'Running': 7}` means the queue is wedged, not that I am doing |
| something wrong. Pod size does not help: 1 CPU / 4 GiB requests queue behind the same wall. |
| |
| 12. **Three traps between a finished SFT and a running RL job**, all of which look like |
| unrelated failures: |
| - `/app` holds a **different prime-rl build** and the default `PATH` puts `/app/.venv` |
| first, so the launcher spawns its orchestrator/trainer from there. Its schema disagrees |
| (`max_inflight_rollouts` vs `max_inflight_episodes`, `train.env` vs `train.source`) and |
| **its verifiers has no broker runtime at all**. Symptom: `Extra inputs are not permitted |
| --train.source`. Fix: `scripts/run_rl.sh` prepends `/root/work/a/prime-rl/.venv/bin`. |
| - The trainer's saved checkpoint has **no `preprocessor_config.json`**, so prime-rl's |
| inference server dies with `Can't load image processor for <ckpt>`. (Plain `vllm serve` |
| does not care.) Fix: copy `preprocessor_config.json`, `video_preprocessor_config.json`, |
| `merges.txt`, `vocab.json` from the base dir into every checkpoint β **do this before |
| submitting too**. |
| - prime-rl's inference does not pass `limit_mm_per_prompt`, so vLLM profiles the vision |
| tower and crashes in the CUTE/FA4 kernel with `TypeError: fmax() missing 1 required |
| positional argument`. Fix: `[inference.vllm_extra] limit_mm_per_prompt = {image=0,video=0}`. |
| |
| 13. **GRPO through the pi harness does not work in this stack.** After clearing the three traps |
| above, every rollout dies with `ACP agent produced no visible reply`, and the env log shows |
| the real cause: `model call failed: TrainClient does not support streaming`. pi always |
| streams; the interception server routes streaming requests to the client, and prime-rl's |
| `TrainClient` (the one that returns token ids + logprobs for the RL loss) raises on stream. |
| The eval client proxies streams fine, which is why evaluation works and training does not. |
| Fixing it means changing shared verifiers code, so RL was replaced with **expert iteration** |
| through the same pi harness: sample β keep reward-1 trajectories β SFT. Same tool surface, |
| same prompt, every component already validated. |
| |
| 14. **Both RL blockers are now cleared, in local plugin code, without touching the shared |
| install.** Finding 13 was blocker one (streaming); blocker two was render parity. |
|
|
| - *Streaming*: `pkgs/pi_rl` writes a ~90-line Node shim into the task container, points pi |
| at `127.0.0.1:8899`, and the shim forwards each call upstream with `stream` stripped, then |
| re-emits the JSON reply as the SSE chunks pi expects. Transport-only β same agent, same |
| tools, same prompt. Verified transparent: 6/6 episodes, same turn count and token usage. |
| - *Render parity*: the renderer serialises whatever tool objects it is handed, and verifiers |
| hands it its own **flat** `ToolSpec`, so the rollout's tool block came out as |
| `{"name":β¦,"description":β¦,"parameters":β¦}` while the serving chat template emits |
| `{"type":"function","function":{β¦}}` β 68 chars different in the system prompt of *every* |
| turn. `pi_rl.enable_render_parity()` normalises tools to the OAI-nested form before the |
| renderer sees them, in the training process only. |
| |
| **Settled against the live server, not against a reconstruction.** POST the same |
| messages+tools to vLLM's `/tokenize` with `add_generation_prompt=true` and compare token ids |
| with `renderer.render_ids(...)`: |
| |
| | | tokens | identical to served prompt | |
| |---|---|---| |
| | unpatched renderer | 916 | no | |
| | patched renderer | 952 | **yes** | |
| | vLLM `/tokenize` | 952 | β | |
| |
| That also settles the one open question in the patch: pi sends `strict` in each tool schema |
| and the served template keeps it, so `_nest_tools` must preserve it. Dropping it costs 36 |
| tokens and puts the policy off-distribution again. `enable_render_parity()` is called at |
| `pi_rl` import; `pi_rl` is training-only and nothing in the evaluation path imports it. |
| |
| **Where the patch has to live, and why the obvious placements both fail silently.** |
| |
| - *Wrong process.* The renderer is built in prime-rl's **orchestrator** |
| (`orchestrator/utils.py::setup_policy_inference_pool`); the harness β and therefore |
| `pi_rl` β is imported in the **env-server**. Patching at `pi_rl` import time lands in a |
| process that never renders. The patch now lives in `pkgs/render_parity.py` and is |
| installed by `pkgs/sitecustomize.py` (which `PYTHONPATH` already reaches) under |
| `VF_RENDER_PARITY=1`, set by `scripts/run_rl.sh` and nothing else. Evaluation shares that |
| `PYTHONPATH`, never sets the variable, and is untouched. |
| - *Wrong moment.* The cheap version β wrap `builtins.__import__`, check `sys.modules` after |
| each call β does not work. Python inserts a module into `sys.modules` **before** running |
| its body, so the first sighting of `renderers.qwen35` (an inner import from inside that |
| very module) finds no `Qwen35Renderer` yet; fire-once logic then retires the hook having |
| patched nothing. `install()` uses a `sys.meta_path` finder that wraps the real loader and |
| patches inside `exec_module`, after the body has run. |
| - `pkgs/sitecustomize.py` **shadows** `/usr/lib/python3.12/sitecustomize.py`, so it redoes |
| that file's only job (installing Ubuntu's apport hook) before its own. |
| |
| **Verified live, in the running job**, not just on a bench: |
| `[render-parity] Qwen35Renderer patched` appears in `orchestrator.log` at the |
| `Initializing policy inference pool` line, and a real rollout's first call logs |
| `prompt_tokens=2562` β which is exactly what the *patched* renderer produces for that |
| conversation (unpatched: 2526). |
| |
|
|
| 15. **A stale `vllm::router` silently blocks the next RL launch.** Symptom is |
| `Error: Inference failed with exit code -15` about 13 s after startup, which reads like an |
| OOM or an external kill. The real cause is in `ckpt/<run>/logs/inference.log`: |
| `PanicException: failed to install Prometheus metrics exporter: |
| FailedToCreateHTTPListener("Address already in use")` β a router from a previous, killed run |
| still holds ports 8000 and 29000. **`pkill -f "vllm serve"` does not match it**: the router |
| renames its own process to `vllm::router`, so only `pkill -x "vllm::router"` finds it. |
| `scripts/launch_rl.sh` clears it before every launch. |
| |
| Related, and the same trap as the old `pgrep -f` self-match: **use `pkill -x` (process name), |
| never `pkill -f` (full command line), in any script written from a heredoc.** A `-f` pattern |
| matches every shell whose command line quotes that string β including the shell writing the |
| script β so the launcher kills itself and produces no output at all to explain why. |
| |
| 16. **An RL run cannot be found, or killed, by the names you launched it with.** Every process |
| renames itself: the launcher becomes `PRIME-RL::Launcher`, then `PRIME-RL::Orchestrator`, |
| `PRIME-RL::Trainer`, `PRIME-RL::EnvServer`, `vllm::router`. Consequences, all of which cost |
| time here: |
| - `ps | grep run_rl` / `grep orchestrator` shows **nothing**, so a live run looks dead. |
| - The pidfile is useless: `setsid nohup bash β¦ &` records the transient `setsid` PID, which |
| exits at once. `kill $(cat rl.pid)` reports success and kills nothing. |
| - `pkill -x "PRIME-RL::Launcher"` also fails β `-x` matches the kernel's `comm`, capped at |
| **15 characters**, so the real name is `PRIME-RL::Launc`. |
| |
| A previous run therefore survives invisibly, and the next launch dies on its ports ~80 s in. |
| `scripts/launch_rl.sh` sweeps by *comm prefix* through `ps -eo pid=,comm=` (matching |
| `^(PRIME-RL|vllm|VLLM)::`), kills by PID, then re-checks and refuses to launch if anything |
| survived. Kill the launcher first or it restarts its children. Watch for an orphaned |
| `PRIME-RL::Train` holding ~40 GB β if it outlives its launcher it silently starves the next |
| trainer of GPU memory. |
| |
| 18. **RL ran, and it is closed β for a fourth reason, which is structural rather than a bug.** |
| With blockers 1 and 2 cleared and verified live, GRPO produced **0 reward across 99 scored |
| episodes** (r2e-gym-ws 0/50, swelego-v1 0/33). Ruling things out, in order: |
| |
| - *Not the renderer.* Verified in the running job: a live rollout's first call logged |
| `prompt_tokens=2562`, exactly the patched render (unpatched: 2526). |
| - *Not tool parsing.* The train path parses completions with the renderer instead of vLLM's |
| `qwen3_coder` parser, so this was a real risk β but 2926 of 2949 assistant messages carry |
| `tool_calls`. Rollouts read like competent work ("The fix works", "issue resolved"). |
| - *Not the grader.* Every scored episode carries |
| `info.patch_error: fatal: not a git repository`, which looks damning and is a red herring: |
| `capture_patch`'s own docstring says a failed capture still lets the rollout score. |
| **Gold-patch validation of `swelego-v1` on a free pool: 6/6 valid.** A correct patch does |
| score. `runs/probe_shim` is *not* a baseline for these tasksets β it is swe-bench. |
| - *Not my harness.* `pi_ws.run_acp_in` uses a per-command `cd`, so it cannot alter the cwd |
| of the grader's later `runtime.run` calls. |
| |
| What it actually is β two config mismatches of mine, and one wall: |
| |
| | | eval (scores 0.208) | RL as configured | |
| |---|---|---| |
| | `max_turns` | **None** (unlimited; pi stops when done) | 40 β and **61% of episodes hit it** | |
| | `system_prompt` | `cfg/agent_prompt.txt` | none set on either train source | |
| |
| The wall: pi's episodes on these tasks are long. Peak prompt tokens per episode are |
| **p50 20k, p90 42k, max 101k**, so **24 of 113 episodes already exceed the trainer's |
| `seq_len = 32768` at 40 turns**. Raising `max_turns` to match eval makes *more* of the batch |
| untrainable, not less. Fixing it properly means `seq_len` 65536 on one B200 plus ~6 min of |
| sandbox time per episode β at batch 96 and 24 inflight that is β³40 min/step, so the clock |
| buys ~50 steps with a third of each batch discarded. |
| |
| That is a real project, not a fix, so I stopped rather than start it with ~34 h left and a |
| fully-measured checkpoint to protect. **`pkgs/pi_rl` + `pkgs/render_parity.py` are left |
| working and documented; they are not part of the submission.** For anyone resuming: set |
| `max_turns` unlimited, add the system prompt to both sources, raise `seq_len` to 65536, and |
| expect the sandbox pool β not the GPUs β to be the bottleneck. |
| |
|
|
|
|
|
|
| 19. **Failure-mode audit of my own traces (4,496 episodes), and why it did not lead to another |
| training run.** Three signatures, measured rather than guessed: |
|
|
| | signature | prevalence (SFT model) | prevalence (base) | solve rate with / without | |
| |---|---|---|---| |
| | leaked tool-call XML in `content` (`</parameter></function></tool_call>`) | 12.9% | 0.8β2.0% | **5.7% vs 16.5%** | |
| | "I apologize for the difficulty/confusion" | 41.7% | ~0% | 13.5% vs 16.3% | |
| | SWE-agent harness artifacts ("Exit due to cost limit", "submit button") | 0.5% | 0% | β | |
| |
| The apology habit and the harness artifacts are **my SFT corpus talking** β the base model |
| essentially never produces them. But the apology is style, not damage (13.5 vs 16.3 is weak |
| and confounded: hard tasks cause both), and the harness artifacts are too rare to matter. I |
| had expected the artifacts to be the story; they are not. Worth stating plainly because the |
| first example I looked at was an "Exit due to cost limit", which is exactly how a 0.5% effect |
| gets mistaken for the main one. |
| |
| The leaked-XML signature is real and 3Γ predictive, but it caps out small: eliminating it |
| entirely moves 12.9% of episodes from 5.7% to 16.5%, i.e. **+1.4 points absolute** β a quarter |
| of the Β±5-point interval at n=250. Not worth an 8 h retrain plus an 8 h re-measurement. |
| |
| A separate cut, on the verification behaviour the earlier failure analysis flagged: |
| |
| | swe-bench, episodes that made an edit | ran a test after the **last** edit | |
| |---|---| |
| | solved (n=51) | **63%** | |
| | failed (n=174) | **37%** | |
| |
| Same direction as `cfg/agent_prompt_v2.txt`, which was built for exactly this and measured at |
| **0.233 vs 0.208, McNemar p=0.30** β directionally right, not separable. Note only 12% of |
| failures never edited at all, so "nudge it to act" is the wrong lever; "make it verify" is the |
| right one, and it is already in the prompt. |
| |
| **Conclusion, and the reason this section ends here:** every remaining lever I can identify is |
| worth ~1β3 points, and my measurement floor is Β±5 at n=250. Chasing them would produce changes |
| I could not distinguish from noise, which is how a benchmark run talks itself into a |
| regression. The honest use of the remaining time was to make the *reported* number better, so |
| the last measurement runs the **complete 500-task suite** on both harnesses instead of a |
| 250-task sample, which is the one improvement that does not depend on my guessing right. |
| |
| 20. **Fairness audit of the two arms β one asymmetry found, and it did not bite.** The `pi-ws` |
| configs set `ready_timeout_seconds = 1200`; the stock configs did not, and the broker default |
| is **600** (`v1/runtimes/broker.py:59`). Half the readiness budget means roughly twice the |
| sandbox-timeout rate under a busy pool, and the logs do show more error lines on the stock |
| side (`big_swe_stock` 71 vs `big_swe_ws` 29) β an asymmetry pointing in *my* favour, which is |
| the direction that must never go unchecked. |
| |
| Checked against the traces rather than the logs, because the two count different things: an |
| episode that has no `rewards` record at all is one that never got graded, i.e. genuinely lost |
| to infrastructure. |
| |
| | run | n | solved | ungraded | score | score excluding ungraded | |
| |---|---|---|---|---|---| |
| | `big_swe_stock` | 250 | 33 | **0** | 0.132 | 0.132 | |
| | `big_swe_ws` | 250 | 52 | **0** | 0.208 | 0.208 | |
| |
| **Zero ungraded episodes in either arm** β the log error lines were transient and retried, no |
| episode was lost, and the headline comparison is unaffected. `cfg/final500_stock.toml` now |
| sets 1200 explicitly anyway, and the 500-task stock arm was restarted after the fix so that |
| both arms are identical in everything except the harness and the appended prompt. |
| |
| 21. **The base-weights numbers were the weakest link in the attribution, so they are being |
| re-measured too.** The submission claims two separate gains β weights and scaffold β but they |
| rested on different-quality evidence: |
| |
| | | n | quality | |
| |---|---|---| |
| | `big_swe_base` (base + stock) | 153 | **176 errored episodes** in the same run | |
| | `base_swe_ws` (base + pi-ws) | 64 | taken when pi-ws still had the **no-op `cd`** and the `maxTokens`/`contextWindow` caps that are now **off** β a different harness | |
| | `big_swe_stock` / `big_swe_ws` (submitted) | 250 / 250 | clean, 0 ungraded | |
| |
| Taken at face value the old pair says base+pi-ws = 0.188 [0.11, 0.30] against submitted+pi-ws |
| = 0.208 [0.16, 0.26] β heavily overlapping, i.e. *on the submitted harness I could not show |
| the weights help at all*. That reading is not sound (different harness, n=64), but it is the |
| right worry, and the honest fix is power, not argument. |
| |
| So the final measurement is the full **2Γ2 at n=500**: {base, submitted} Γ {stock, pi-ws}, |
| same 500 tasks, same protocol, base served on port 8001 so it can never be confused with the |
| submitted checkpoint on 8000. `cfg/final500_{stock,ws,base_stock,base_ws}.toml` differ only |
| in harness id, the appended prompt, and the port. |
| |
| 22. **Sandbox pod sizing, corrected β and what a saturated cluster looks like.** Each sandbox pod |
| requests **4 CPU**, not 1 as finding 17 assumed. On a ~250β330 CPU shared cluster that means |
| the *whole cluster* holds only ~60β80 concurrent sandboxes across all tenants, so 40 |
| concurrent from me was over half of it. |
| |
| On 2026-08-16 from ~04:40Z the pool went to **371 Pending / 34 Running with `available.cpu` |
| at 0** β a ~1,500 CPU backlog against a cluster that has none β while only ~20 of those |
| requests were mine. Throughput went to roughly zero regardless of what I did. Symptoms to |
| recognise next time: |
| |
| - `avail_cpu` falling to 0 while `Pending` climbs into the hundreds β the queue is global, so |
| dropping *my* concurrency does not clear it and does not restore *my* throughput. |
| - A single hand-made sandbox going ready in 3 s is **not** evidence of headroom (finding 17); |
| when the cluster is truly full even that stalls. |
| - `/resources` has no owner field, so pods cannot be attributed. Judge by arithmetic: my |
| outstanding requests versus total Pending. |
| |
| Practical rule: size runs to *finish*, because a partial run is not usable β with |
| `shuffle=false` a prefix is the dataset's own ordering, and with `shuffle=true` it is |
| ascending difficulty (Results section). Both bias the score. When the pool is saturated, |
| prefer the measurement that fills a genuine gap over the one that merely adds precision. |
| |
| 23. **The 2Γ2 lands at n=250, and it does not just correct the attribution β it reverses part of |
| it.** All four cells complete, **zero ungraded**, identical 250 tasks (`SEED=0` is pinned, so |
| `shuffle=true` draws the same sample every run). |
| |
| | weights | stock pi | pi-ws | |
| |---|---|---| |
| | base (+eos fix) | 0.044 [0.025, 0.077] | **0.276** [0.224, 0.334] | |
| | submitted (SFT) | 0.132 [0.096, 0.180] | 0.208 [0.162, 0.263] | |
| |
| | paired comparison (n=250) | only-A | only-B | delta | McNemar | |
| |---|---|---|---|---| |
| | base: stock β pi-ws | 4 | 62 | **+0.232** | **<0.0001** | |
| | submitted: stock β pi-ws | 12 | 31 | **+0.076** | **0.0054** | |
| | stock: submitted β base | 29 | 7 | **β0.088** | **0.0003** | |
| | **pi-ws: submitted β base** | 21 | 38 | **+0.068** | **0.036** | |
| |
| - The **scaffold** is the robust result: +23 points on base weights, +7.6 on the submitted |
| checkpoint. The larger intervention on both, by a wide margin. |
| - The **SFT weights** are worth +8.8 points **under the stock harness** (p=0.0003) β real, and |
| the reason they are still what I submit. |
| - **Under pi-ws the SFT weights are significantly *worse* than base** (0.208 vs 0.276, |
| p=0.036). Not "no difference" β a measured 6.8-point regression. The two interventions |
| partly **conflict**. |
| |
| **How the number moved with n β the brief's warning, demonstrated.** Same comparison: |
| n=100 β +0.000, p=1.00; n=225 β +0.062, p=0.076; n=250 β +0.068, p=0.036. The point estimate |
| barely moved; the interval closed. At no n was there evidence the SFT weights *help* under |
| pi-ws, which is exactly why the base arms were extended rather than stopped at the first clean |
| number. Had I stopped at n=100 I would have reported "exactly tied" as a finding. |
| |
| Only visible because the 2Γ2 was finished rather than argued about. It retro-justifies finding |
| 21's worry β `base_swe_ws` 0.188 at n=64 was pointing straight at this and I explained it away |
| as a stale-harness artifact. |
| |
| **It does not change what I submit.** The brief requires weights derived from the base by my |
| own training, so the base itself is not eligible, and `sft_v5/step_900` is worth +8.8 under the |
| stock harness. It changes what I claim, and it motivates finding 25. |
| |
| 24. **There is a hard cap of 32 concurrent host tunnels per API token, and `kill -9` leaks them.** |
| This masqueraded as a sandbox-pool problem for over an hour and is invisible in `/resources`. |
| |
| Every eval episode opens a host tunnel so the sandbox can reach the local vLLM server. |
| `PrimeTunnel.expose` (`v1/interception/tunnel/prime.py`) tears its tunnel down in a `finally` |
| that is shielded against cancellation β so a normal exit or **SIGTERM** cleans up. **SIGKILL |
| does not.** My own `topup_loop.sh` was doing `kill -9` on runs with 24 episodes in flight |
| every 15 minutes; within a few passes all 32 slots were leaked and every subsequent episode |
| died instantly with |
| |
| TunnelError: TunnelLimitReachedError: Maximum number of tunnels (32) reached |
| |
| at `turns=0`. Confirmed by listing: 32/32 held with **no eval process running at all**. |
| |
| Two rules follow, and both are now enforced in `scripts/topup_loop.sh`: |
| - Stop eval runs with **SIGTERM**, never `kill -9`, unless you reclaim tunnels afterwards. |
| - **Total `max_concurrent` across every simultaneously running eval process must stay under |
| 32** β it is per *token*, not per run. Two arms at 24 each is 48 and cannot work, however |
| healthy the pool looks. Two arms at 14 each (28) is fine. |
| |
| `scripts/tunnels.py` lists the token's tunnels and `--delete` reclaims them; only run the |
| delete when no eval is active, since it cannot tell a live tunnel from a leaked one. |
| |
| Worth noting how this presented: `stop=TunnelError` with `turns=0`, arriving in bursts, while |
| the pool showed hundreds of free CPU. Any diagnosis that stops at "the pool is busy" misses |
| it entirely. |
| |
| 25. **The scaffold's two halves, separated β and a pre-registered hypothesis that failed.** |
| `pi-ws` is the working-directory fix plus the appended `cfg/agent_prompt.txt`, and the two had |
| never been separated on the SFT weights. `cfg/ws_noprompt.toml` β `runs/ws_noprompt`, same |
| seeded 250, zero ungraded: |
| |
| | submitted weights, harness | score | vs stock | vs full pi-ws | |
| |---|---|---|---| |
| | stock pi | 0.132 | β | | |
| | pi-ws, workdir only | 0.164 | +0.032, p=0.26 | β0.044, p=0.099 | |
| | pi-ws, workdir + prompt | **0.208** | **+0.076, p=0.005** | β | |
| |
| Both halves contribute, neither clears significance alone, together they do. The scaffold is |
| the *pair*. |
| |
| **The prediction was wrong.** Registered in advance: since `public_to_sft.py` put pi's own |
| system prompt on every training row, an appended operating-procedure block should be |
| off-distribution for the SFT model but pure gain for the base β which would explain why the |
| scaffold is worth +23 points on base and only +7.6 on the SFT weights. If so, dropping the |
| prompt should have *recovered* most of the gap. Instead dropping it **costs** 4.4 points. |
| Prompt off-distribution is not the explanation for the SFT weights' deficit against base under |
| pi-ws, and I do not have a replacement explanation β only the measurement. |
| |
| **And a second demonstration of the partial-read trap**, worse than finding 23's: this same |
| comparison read β0.075, p=0.096 at n=107 (looking like a strong refutation *in the opposite |
| direction*), β0.039, p=0.31 at n=152, and β0.044, p=0.099 at n=250. Three different stories |
| from one run. Only the completed number means anything. |
| |
| 26. **Pre-registered: is the training/scaffold conflict monotone in training length?** |
| (Registered 2026-08-16 18:30Z, before the run.) |
| |
| Finding 23 leaves the submitted checkpoint in an awkward place: 900 steps of SFT buys +8.8 |
| points under the stock harness and **loses** 6.8 under pi-ws versus the untrained base. |
| Hypothesis with a mechanism: the stock-harness gain is mostly **format** competence β emit a |
| well-formed tool call, use the four tools, stop when done β which saturates within a couple of |
| hundred steps; the pi-ws regression is **stylistic over-specialisation** that accumulates over |
| the full run. If so, an early *fully annealed* checkpoint should keep most of the +8.8 while |
| giving back most of the β6.8, and would dominate `step_900` on both harnesses. |
| |
| `cfg/sft_short.toml`: same corpus, same recipe, cosine annealed over **200 steps** (not step |
| 200 of the 900-step run β a complete run, the thing I would actually submit). 2 GPUs, |
| batch 8 β ~52M tokens. |
| |
| **Decision rule, fixed in advance:** |
| - Measure `sft_short` on **both** harnesses, n=250, same seeded sample |
| (`cfg/short_{ws,stock}.toml`, served on port 8002). |
| - The decisive cell is **pi-ws** vs `step_900`'s 0.208 on a paired McNemar. |
| - *Scheduling note (changed after registering, and only the order):* I originally planned to |
| run pi-ws first and gate the stock arm on it, to save sandbox budget. With ~24 h left and the |
| pool delivering 30β60 episodes/hour, sequential gating risks finishing with only half the |
| pair measured β the failure mode I have been guarding against all run. Both arms therefore |
| run concurrently (14 each = 28 tunnels, under the cap). The rule below is unchanged. |
| - Submit `sft_short` **only** if it wins under pi-ws *and* its stock number is not |
| significantly worse than 0.132. Any other outcome: `step_900` ships, unchanged. |
| - `step_900` stays fully measured and is never at risk. Nothing here can leave an unmeasured |
| checkpoint as the submission. |
| |
| Expected value is honest about itself: the most likely single outcome is that the trade-off is |
| monotone, no dominating checkpoint exists, and this returns a *curve* (0 / 200 / 900 steps) |
| rather than a better submission. That curve is worth having either way. |
| |
| **Trained 18:22β21:01Z** (200 steps, 2 GPUs, ~47 s/step, loss 0.245 β ~0.14; |
| `ckpt/sft_short/weights/step_200`, 760 tensors, eos ids correct, `step_100` kept as a spare). |
| |
| **RESULT β the hypothesis is refuted, and the curve is not the shape I guessed.** Both arms |
| complete at n=250, zero ungraded, same seeded sample as every other cell: |
| |
| | weights | stock pi | pi-ws | |
| |---|---|---| |
| | base (0 steps) | 0.044 [0.025, 0.077] | **0.276** [0.224, 0.334] | |
| | sft_short (200 steps) | 0.104 [0.072, 0.148] | 0.168 [0.127, 0.219] | |
| | sft_v5 (900 steps, submitted) | 0.132 [0.096, 0.180] | **0.208** [0.162, 0.263] | |
| |
| | paired comparison (n=250) | only-A | only-B | delta | McNemar | |
| |---|---|---|---|---| |
| | stock: base β short | 7 | 22 | +0.060 | **0.0081** | |
| | stock: base β sub900 | 7 | 29 | +0.088 | **0.0003** | |
| | stock: short β sub900 | 12 | 19 | +0.028 | 0.28 | |
| | pi-ws: base β short | 41 | 14 | **β0.108** | **0.0004** | |
| | pi-ws: base β sub900 | 38 | 21 | **β0.068** | **0.036** | |
| | pi-ws: short β sub900 | 18 | 28 | +0.040 | 0.18 | |
| |
| Two clean statements: |
| - **Under the stock harness, SFT helps monotonically** with training length: |
| 0.044 β 0.104 β 0.132, each significant against base. |
| - **Under pi-ws, *every* amount of SFT is significantly worse than the untrained base** |
| (200 steps: p=0.0004; 900 steps: p=0.036). And it is **not monotone** β 200 steps (0.168) is |
| *worse* than 900 (0.208), not better. My prediction was that a shorter anneal would sit |
| closer to base under the scaffold. It sits further away. |
| |
| So the conflict is not "too much training". Whatever SFT does that the scaffold does not like, |
| it does early, and more training partially undoes it. I do not have a mechanism for that, and |
| with ~16 h left I am not going to get one honestly β three points on a curve is what this buys. |
| |
| **Decision rule applied: `sft_short` did not win under pi-ws (0.168 vs 0.208, p=0.18 favouring |
| step_900), so `ckpt/sft_v5/weights/step_900` remains the submission, unchanged.** `sft_short` |
| is kept on disk and reported; it is not submitted. |
| |
| 27. **Never run `eval @ config` and `eval --resume` against the same run directory at once.** |
| They both own `traces.jsonl`, and the second one does not merge β the graded count went |
| *backwards* (57 β 44) while both were live, because each process wrote the file from its own |
| view of what was done. No corruption, but ~30 completed episodes were silently lost and the |
| surviving file had mixed provenance. |
| |
| My own sequencing error: I launched both arms with `scripts/run_eval.sh` (which runs |
| `eval @ cfg -o dir`) and *then* started `scripts/topup_loop.sh` on the same dirs, whose |
| `stop_runs` only matches `eval --resume` β so it never stopped the original writers. |
| |
| Rules: one writer per run directory, ever. If you want the top-up loop, either start the run |
| with it from the beginning, or SIGTERM the original process first and confirm it is gone. |
| A `graded` count that decreases is the signature β check for two writers before anything else. |
| When it happens, restart the affected runs from empty rather than resuming: a trace file of |
| mixed provenance is not something to base a submission decision on. |
| |
| 28. **terminal-bench-2, the same 2Γ2, completed β and it measures nothing.** My own reporting had |
| been inconsistent: I demanded completed runs for swe-bench while quoting tb2 at n=83β88 of 89. |
| All four cells are complete 89-task runs with **zero ungraded**: |
| |
| | weights | stock pi | pi-ws | |
| |---|---|---| |
| | base (+eos fix) | 5/89 = 0.056 | 5/89 = 0.056 | |
| | submitted (SFT) | 3/89 = 0.034 | 1/89 = 0.011 | |
| |
| One task nearly did not make it, and the cause is worth recording: |
| `terminal-bench/qemu-alpine-ssh` boots a QEMU VM and takes **longer than 15 minutes**, which is |
| exactly the interval at which `scripts/topup_loop.sh` stops and re-resumes its runs. Every pass |
| started that task and killed it before it could finish, forever. The loop is right for |
| pool-stalled runs and wrong for genuinely long tasks β when a single task is all that is left, |
| stop the loop and run one plain `eval --resume` with nothing killing it. It then completed on |
| the first attempt. |
| |
| Every paired McNemar is non-significant: p from 0.125 (base+pi-ws vs submitted+pi-ws, 4 vs 0) |
| to 1.00. **No effect claimed in any direction.** With 1β5 solves out of 89 the interval swamps |
| everything. |
| |
| Two observations, explicitly not results: |
| - **The scaffold that is worth +23 points on swe-bench does nothing here** β base scores 5/88 |
| under *both* harnesses (5/89 each). That fits its content: a working-directory fix and a repo-oriented |
| operating procedure have nothing to grip on in tasks that are not repo fixes and already |
| start in `/app`. The pi-ws gain is specific to the SWE-repo setting, not general agentic |
| competence. |
| - The directional ordering matches swe-bench (base β₯ submitted everywhere, submitted+pi-ws |
| lowest), but at these counts that is a coincidence I would not defend. |
| |
| Also fixed here: `cfg/base_tb2.toml` had the same missing `ready_timeout_seconds` asymmetry |
| caught on swe-bench in finding 20 (broker default 600 against pi-ws's 1200, in my favour). |
| Both tb2 configs now pin 3600. It had no effect on the numbers β zero ungraded either side β |
| but it should not have been there. |
| |
| 29. **Replicate of the headline cell β the intervals hold, and the noise floor is now measured.** |
| Every comparison in this file is a paired McNemar, which treats *which tasks were sampled* as |
| the only source of variation and says nothing about run-to-run variance from sampling |
| temperature and the environment. The brief warns that repeat reads of identical weights have |
| differed by more than ten points, so I re-ran the headline cell: same checkpoint, same harness, |
| same seeded 250 tasks, same config, nothing different but the run. |
| |
| | run | n | solved | score | ci95 | |
| |---|---|---|---|---| |
| | original (`big_swe_ws`) | 250 | 52 | 0.208 | [0.162, 0.263] | |
| | replicate (`replicate_ws`) | 250 | 49 | **0.196** | [0.152, 0.250] | |
| | pooled | 500 | 101 | **0.202** | [0.169, 0.239] | |
| |
| - **The replicate lands well inside the original interval.** The intervals quoted throughout |
| are trustworthy as stated; this system does not swing ten points between reads. |
| - **Noise floor, measured: two runs of the identical system disagree on 39 of 250 tasks |
| (15.6%) β 21 one way, 18 the other.** That is the yardstick for every McNemar here. The null |
| run is *symmetric*; the real effects are lopsided (12 vs 31 scaffold, 21 vs 38 |
| base-over-submitted). McNemar tests exactly that asymmetry, which is why it is the right |
| test β but anything near 22 vs 30 is within reach of this floor. |
| - **Both headline conclusions reproduce on the fresh run**: scaffold 14 vs 30, p=0.023 (was |
| 12 vs 31, p=0.005); base-beats-submitted under pi-ws 16 vs 36, p=0.008 (was 21 vs 38, |
| p=0.036) β the second one is *stronger* on the replicate. |
| |
| This is the measurement I would have wanted first if I had known how much the report would end |
| up resting on paired tests. It costs one run and it tells the reader what a difference has to |
| look like before it means anything. |
| |
| 30. **Pre-registered, final experiment: weight-space interpolation between base and my SFT.** |
| (Registered 2026-08-17 07:15Z, before building or measuring anything.) |
| |
| The standing problem from finding 23: my SFT weights are **6.8 points worse than the untrained |
| base under my own scaffold** (0.208 vs 0.276, p=0.036; p=0.008 on the replicate), while being |
| 8.8 points *better* under the stock harness. That is the classic shape of fine-tuning trading |
| away out-of-distribution robustness. |
| |
| Finding 26 already tested the obvious axis β *less training* (a 200-step anneal) β and it was |
| **worse** under pi-ws, not better (0.168). So "closer to base along the training trajectory" |
| does not help. **Weight-space interpolation is a different axis**: a convex combination of the |
| base and my fine-tune, which is the standard remedy for exactly this fine-tuning/robustness |
| trade (WiSE-FT). It is cheap β `scripts/soup.py` averages two checkpoints in ~20 minutes β and |
| it has not been tried. |
| |
| `ckpt/soup_base_sft/` = **0.5 x base + 0.5 x sft_v5/step_900**, produced by `scripts/soup.py`. |
| |
| **Rules check, stated explicitly because it deserves scrutiny:** every parameter still traces |
| to `Qwen/Qwen3.5-9B-Base` through training I ran β it is a blend of the designated base with my |
| own SFT checkpoint and nothing else. The prohibition is on initializing from, continuing |
| training on, merging in, or submitting *another party's* post-trained checkpoint; the base |
| itself is the designated starting point, not another party's post-training. No third-party |
| weights are involved at any point. |
| |
| **Decision rule, fixed in advance:** |
| - Measure the soup on **both** harnesses, n=250, same seeded sample, both to completion. |
| - Submit it **only** if it beats `step_900` under pi-ws on a paired McNemar at **p<0.05** |
| *and* is not significantly worse than `step_900` under the stock harness. |
| - **Hard stop 15:30Z.** If both cells are not complete by then, `step_900` ships. A partial run |
| is not evidence and will not be used for this decision β three times this run a partial read |
| told a different story than the finished one. |
| - `step_900` remains fully measured throughout and is never at risk. |
| |
| Honest prior: maybe one chance in three. The 200-step result shows the pi-ws deficit is not a |
| simple function of distance-from-base, so interpolation may not touch it either. Reported |
| either way. |
| |
| **RESULT (both arms complete, n=250 each, zero ungraded).** Interpolation *works* on the axis |
| that training length did not β and it buys the pi-ws points by giving back stock points: |
| |
| | weights | stock pi | pi-ws | |
| |---|---|---| |
| | base | 0.044 [0.025, 0.077] | 0.276 [0.224, 0.334] | |
| | **soup (0.5/0.5)** | **0.084** [0.056, 0.125] | **0.280** [0.228, 0.339] | |
| | submitted step_900 | 0.132 [0.096, 0.180] | 0.208 / 0.196 (two runs) | |
| |
| | paired vs step_900 (n=250) | only-A | only-B | delta | McNemar | |
| |---|---|---|---|---| |
| | pi-ws: step_900 β soup | 16 | 34 | **+0.072** | **0.0153** | |
| | pi-ws: replicate β soup | 17 | 38 | **+0.084** | **0.0065** | |
| | stock: step_900 β soup | 23 | 11 | **β0.048** | 0.0576 | |
| |
| - **Condition (a) is met clearly**: the soup beats `step_900` under pi-ws, and beats the |
| *replicate* more strongly still, so this is not one lucky run. |
| - **Condition (b) passes only on a technicality**: p=0.0576 is "not significant at 0.05", but a |
| 4.8-point drop (0.132 β 0.084, a 36% relative fall) on 23-vs-11 discordant pairs is almost |
| certainly a real cost. Reading that as "no effect" would be exactly the rule-lawyering the |
| noise-floor work in finding 29 was meant to prevent. |
| |
| So it is a genuine **trade**, not a dominating win: roughly +7 to +8 points on the harness I |
| submit, roughly β5 on the stock control. And the mechanism is not subtle β the soup's pi-ws |
| score (0.280) is statistically indistinguishable from the untrained base's (0.276), so its |
| advantage there comes from being *closer to base*, i.e. from partially undoing my own training. |
| |
| **Cannot decide on swe-bench alone**: the submission publishes terminal-bench-2 as well, and |
| the soup has no tb2 numbers. Both tb2 arms launched 11:30Z; decision at 15:30Z at the latest, |
| on completed runs only. |
| |
| 31. **Replicate of the SUBMITTED system's headline cell (running from 2026-08-17 12:55Z).** |
| The submission was switched to the soup on the strength of **one** measurement of |
| soup+pi-ws (0.280), while the checkpoint it displaced has **two** (0.208, 0.196). Finding 29 |
| established that two runs of an identical system disagree on 15.6% of tasks. The submitted |
| system's headline number should therefore carry the same replicated status as the number it |
| replaced β otherwise I have held the alternative to a higher evidentiary standard than my own |
| choice, which is precisely backwards. |
| |
| `cfg/soup_ws_rep.toml` β `runs/soup_ws_rep`: same checkpoint, same harness, same seeded 250, |
| nothing different but the run. |
| |
| **Pre-committed reading.** The switch rested on soup 0.280 vs step_900 0.202 pooled, paired |
| p=0.015 (p=0.0065 vs the replicate). |
| - If this lands near 0.28 β or anywhere that keeps the paired comparison against `step_900` |
| significant β the switch stands and the headline gets a pooled figure. |
| - If it lands near 0.21, the single 0.280 was a lucky run, the switch is not supported, and |
| **I revert the submission to `ckpt/sft_v5/weights/step_900`**, which is intact and fully |
| measured. |
| - **Hard stop 17:00Z.** If the run is not complete by then it is not used at all β a partial |
| read has told a different story than the finished one three times in this run β and the |
| submission stays as it is, with `SUBMISSION.md` stating the headline rests on a single |
| measurement. |
| |
| **RESULT (complete, n=250, zero ungraded): 0.228, not 0.280. The first run flattered the soup.** |
| |
| | system | run 1 | run 2 | pooled (n=500) | |
| |---|---|---|---| |
| | soup + pi-ws | 0.280 | **0.228** | **0.254** [0.218, 0.294] | |
| | step_900 + pi-ws | 0.208 | 0.196 | 0.202 [0.169, 0.239] | |
| |
| - **Run 2 alone does not beat `step_900`** (27 vs 22, p=0.57; against the step_900 replicate, |
| 32 vs 24, p=0.35). The p=0.015 that triggered the switch was the better of two draws. |
| - **What survives is the pooled comparison across all four runs (1,000 episodes):** per-task |
| sign test, soup better on 54 tasks, step_900 better on 30, 166 ties, **p=0.0116**. So the |
| gain is real but is **+5 points, not +8**. |
| - **The soup is the noisier system**: its two runs differ by 5.2 points (32 vs 19 discordant, |
| p=0.09) against step_900's 1.2 (21 vs 18, p=0.75). |
| - Pooled, the soup **no longer exceeds the base** under pi-ws (0.254 vs 0.276, one run) β it |
| matches it, which is the same story as every other checkpoint on this harness. |
| |
| **Decision: the switch stands, but only just, and the headline is restated as 0.254.** The |
| pre-registered revert trigger was "lands near 0.21, i.e. the 0.280 was a lucky run and the soup |
| is not actually better". The pooled data says the soup *is* better than `step_900` under pi-ws |
| (p=0.0116), so the substance of that trigger is not met β but 0.280 was clearly a high draw and |
| quoting it would misrepresent the system. `SUBMISSION.md` now leads with the pooled 0.254, says |
| the replicate alone fails significance, and keeps `step_900` named as the alternative for |
| anyone weighting the stock-harness number higher. The trade is now roughly +5 pi-ws / β5 stock: |
| close to even, with more evidence behind the gain (four runs) than the cost (two). |
| |
| 32. **Final measurement: the stock cost is confirmed, and the two checkpoints end in a dead heat.** |
| The submission decision had become a trade β about +5 points under pi-ws against about β5 under |
| the stock harness β with four runs behind the pi-ws side and only one per system behind the |
| stock side. That asymmetry was the weakest link, so the last hours went to a replicate of the |
| soup's stock cell (`cfg/soup_stock_rep.toml`, n=100, a strict prefix of the 250). |
| |
| **Complete, 100/100, zero ungraded: 0.070**, against the first run's 0.120 on those same 100 |
| tasks (10 vs 5 discordant, p=0.30 β consistent, not a contradiction). Pooled over all 350 stock |
| episodes the soup scores **0.080 [0.056, 0.113]**, essentially the 0.084 first measured. The |
| cost as weights alone is real, not a low draw. |
| |
| | | stock | pi-ws | sum | |
| |---|---|---|---| |
| | soup (submitted) | 0.080 | **0.254** | **0.334** | |
| | step_900 | **0.132** | 0.202 | **0.334** | |
| |
| **Interpolating along the baseβSFT line trades the two published numbers off almost exactly |
| one-for-one.** There is no free lunch on that axis. Which checkpoint is "better" is therefore a |
| question about which number is primary, not something the data settles β and `SUBMISSION.md` |
| says that plainly rather than implying the soup dominates. |
| |
| **Decision: no revert.** The pre-committed trigger was "pooled stock cost *clearly larger* than |
| the pi-ws gain"; it is equal, not larger. The soup stays because the brief asks to make *the |
| system* score higher and the pi-ws number is the system, and because that side rests on four |
| runs and 1,000 episodes (p=0.0116) against three runs and 600 on the stock side (p=0.058). |
| `step_900` is named in `SUBMISSION.md` for anyone weighting stock higher. |
| |
| 33. **Closing the last evidence asymmetry (running from 2026-08-17 16:45Z).** After finding 32 the |
| submitted soup's stock cell had **350** episodes (250 + a 100-task replicate) while the |
| alternative it displaced, `step_900`, had **250**. The extra scrutiny went to my own preferred |
| checkpoint on the axis where it looks worst β the right direction β but a published comparison |
| should not rest on unequal evidence in *either* direction. |
| |
| `cfg/s900_stock_rep.toml` β `runs/s900_stock_rep`: `step_900`'s stock cell, same seeded 100-task |
| prefix, served on 8000 (GPUs 0/1; the soup stays up on 8002). |
| |
| **Decision-relevant, and pre-committed before looking:** the current trade is +5.2 points pi-ws |
| against β5.2 stock, a dead heat on the sum (0.334 each). |
| - If `step_900`'s stock number comes in **higher** than the 0.132 measured once, the stock cost |
| becomes *clearly larger* than the pi-ws gain, which is exactly the revert trigger from |
| finding 30 β **submit `step_900`**. |
| - If it comes in at or below 0.132, the dead heat stands and the soup stays. |
| - **Hard stop 18:00Z**; incomplete β not used, nothing changes. |
| |
| **RESULT β the trigger fires; submission reverts to `step_900`.** Complete, 100/100, zero |
| ungraded: `step_900` scored **0.200** on the replicate prefix against **0.150** on those same |
| 100 tasks the first time (5 vs 10 discordant, p=0.30 β consistent). Pooled over 350 stock |
| episodes `step_900` scores **0.151 [0.118, 0.193]**, not the single-run 0.132. |
| |
| | | stock (pooled n=350) | pi-ws (pooled n=500) | sum | |
| |---|---|---|---| |
| | **step_900 (submitted)** | **0.151** | 0.202 | **0.353** | |
| | soup | 0.080 | **0.254** | 0.334 | |
| |
| The soup buys **+5.2** under pi-ws and costs **β7.1** under stock. The pre-committed trigger |
| ("pooled stock cost clearly larger than the pi-ws gain") is met, so **the submission reverts to |
| `ckpt/sft_v5/weights/step_900`**. `ckpt/soup_base_sft` stays in the workspace, fully measured, |
| and is still ahead if the my-harness number is the only one that counts. |
| |
| **The lesson, and it is the same one twice:** both replicates moved the answer *against* |
| whichever checkpoint I preferred at the time β 0.280 β 0.254 for the soup when I had just |
| switched to it, 0.132 β 0.151 for `step_900` when I had just switched away. A single run |
| flattered whichever system I had most recently chosen. Nothing about that is mysterious given |
| finding 29's 15.6% task-level noise floor; what saved the conclusion was replicating *both* |
| sides rather than only the one whose result I liked. |
| |
| 34. **Final audit: every published number recomputed from disk, and one magnitude corrected.** |
| With the deadline close and both documents heavily edited, I recomputed every figure in |
| `SUBMISSION.md` directly from `runs/*/traces.jsonl`. All matched exactly (step_900 53/350 and |
| 101/500; soup 28/350 and 127/500; base 11/250 and 69/250; sft_short 26/250 and 42/250; all six |
| tb2 cells). |
| |
| One thing the audit did change. The scaffold gain on the submitted weights had been quoted as |
| "+7.6 points", which was correct for the paired 250-task runs but is not the only estimate now |
| that both arms have replicates: |
| |
| | comparison on `step_900` | stock | pi-ws | gain | test | |
| |---|---|---|---|---| |
| | paired 250-task runs (one each) | 0.132 | 0.208 | +7.6 | McNemar 12 vs 31, p=0.005 | |
| | the same, replicate pair | 0.132 | 0.196 | +6.4 | 14 vs 30, p=0.023 | |
| | 100 tasks where both arms have two runs | 0.175 | 0.270 | +9.5 | sign test 23 vs 11, p=0.058 | |
| | pooled over all episodes (350 vs 500) | 0.151 | 0.202 | +5.1 | β | |
| |
| Every one is positive, so the *direction* is not in doubt; the honest **range is +5 to +9 |
| points**, and `SUBMISSION.md` now prints the whole table instead of the single best figure. |
| (On base weights the effect is +23 and on the soup +19.6, both p<0.0001 β no such ambiguity.) |
| |
| Worth recording as a general point: once you replicate arms unevenly, "the" effect size stops |
| being one number. Quoting the most favourable pooling would have been indefensible after |
| spending the whole run insisting on completed runs and measured noise floors. |
| |
| 35. **The submission's own reproduction commands did not reproduce its numbers.** Caught in the |
| last hour by actually reading the command block as a stranger would. `SUBMISSION.md` named |
| `cfg/ws_swe.toml` and `cfg/base_swe.toml` as the swe-bench arms β those are `num_tasks = 100`, |
| left over from the early 100-task reads β while every published swe-bench figure comes from the |
| 250-task runs (`cfg/big_swe_*.toml`) and their replicates. Anyone following the instructions |
| would have got a different sample and quietly different numbers, and concluded the submission |
| was overstated. |
| |
| Fixed: the block now names the 250-task configs, lists the replicate configs that the pooled |
| figures add in, annotates each line with the count it should produce (52/250, 33/250, 1/89, |
| 3/89, 49/250, 20/100), and says explicitly that `ws_swe`/`base_swe` are the 100-task variants |
| that do *not* reproduce the headline. Every config and script named in the file was then |
| checked to exist. |
| |
| Worth generalising: a reproduction section is code, and it was the only part of this submission |
| never executed as written. Read it against the numbers it claims to produce before shipping. |
| |
| **Then actually ran it.** `cfg/smoke_repro.toml` is `cfg/big_swe_ws.toml` with `num_tasks = 3` |
| and nothing else changed, executed against the submitted checkpoint served on `:8000` exactly |
| as the block documents. Result: **3/3 graded, zero ungraded, `errors: []`**, all three |
| `agent_completed`, mean 21.7 turns / 39 s. So the whole chain works end to end β |
| `PYTHONPATH` β `pkgs/pi_ws` resolves β harness launches β broker runtime β grading returns. |
| |
| Two details worth having on record, because they are the harness's central claims and this is |
| the first time they were confirmed on the *submitted* artifact rather than during development: |
| - **The agent really is in the task workdir.** All three episodes work in `/testbed` from their |
| first command (`find /testbed -type f -name "*.py" ...`, and observations returning |
| `/testbed/django/core/handlers/asgi.py`). |
| - **The orientation training shows up in behaviour.** One episode's very first tool call is |
| `pwd; ls -a; echo ---; ls -d /testbed /app /workspace /repo /code /srv 2>/dev/null` β the |
| exact pattern `scripts/add_orientation.py` put on 40% of training rows. |
| |
| 0/3 solved is the expected outcome at a ~20% rate (0.8Β³ β 51% chance of zero) and is not a |
| signal either way; the point of the run was that it *executes*, not what it scores. |
| |
| ## Measurement protocol (fixed β do not change mid-run) |
|
|
| The exact McNemar in `scripts/final_2x2.py` was checked against `scipy.stats.binomtest` on every |
| discordant pair reported in this file (12/31, 5/18, 11/2, 13/13, 2/24, 0/17, 3/17): identical to |
| 1e-9. Worth doing once β every significance claim here rests on that one function. |
|
|
|
|
|
|
| `cfg/base_tb2.toml` (all 89 tasks) and `cfg/base_swe.toml` (100 tasks, `shuffle=true` fixed seed), |
| 1 rollout/task, stock `pi` harness, broker runtime `block_network=false`, **no sampling overrides**. |
| Score with `scripts/score.py <outdir>` (solve rate + Wilson 95% CI). |
|
|
| ## Results |
|
|
| **Submitted checkpoint: `ckpt/sft_v5/weights/step_900`.** |
|
|
| ### Headline β 250-task swe-bench-verified sample (the number to quote) |
|
|
| | ckpt | harness | n | solved | score | ci95 | |
| |---|---|---|---|---|---| |
| | base (+eos fix) | stock pi | 153 | 7 | 0.046 | [0.022, 0.091] | |
| | SFT v5 step_900 | stock pi | 250 | 33 | **0.132** | [0.096, 0.180] | |
| | SFT v5 step_900 | **pi-ws** | 250 | 52 | **0.208** | [0.162, 0.263] | |
|
|
| Paired tests on the shared tasks (the powered comparison β the subset is seeded, so runs line |
| up task-for-task): |
|
|
| | comparison | only-A | only-B | McNemar | |
| |---|---|---|---| |
| | base β SFT, stock harness | 3 | 17 | **p=0.003** | |
| | base β SFT, pi-ws | 3 | 23 | **p<0.001** | |
| | stock β pi-ws, same weights | 12 | 31 | **p=0.005** | |
|
|
| So on swe-bench-verified: the weights are worth **+8.6 points** (0.046 β 0.132) and the scaffold |
| a further **+7.6** (0.132 β 0.208), each independently significant. |
|
|
| ### Fixed 100-task protocol (used for every checkpoint comparison in this run) |
|
|
| | ckpt | suite | harness | n | solved | score | ci95 | |
| |---|---|---|---|---|---|---| |
| | base (+eos fix) | swe-bench-verified | stock pi | 97 | 4 | 0.041 | [0.016, 0.101] | |
| | SFT v5 step_900 | swe-bench-verified | stock pi | 88 | 11 | 0.125 | [0.071, 0.210] | |
| | SFT v5 step_900 | swe-bench-verified | pi-ws | 98 | 31 | 0.316 | [0.233, 0.414] | |
| | base (+eos fix) | terminal-bench-2 | stock pi | 83 | 2 | 0.024 | [0.007, 0.084] | |
| | SFT v5 step_900 | terminal-bench-2 | stock pi | 88 | 3 | 0.034 | [0.012, 0.096] | |
| | SFT v5 step_900 | terminal-bench-2 | pi-ws | 83 | 2 | 0.024 | [0.007, 0.084] | |
|
|
| * **Weights**: swe-bench 0.041 β 0.125 on the stock harness. Paired, the SFT model solves 9 the |
| base does not and loses 1 (McNemar **p=0.021**). The stock number reproduces almost exactly on |
| the independent 250-task sample (0.125 β 0.123), so it is solid. |
| * **Harness**: +7.6 points on the complete 250-task sample, on identical weights. |
| * **terminal-bench-2 does not move** under any combination β 2β3 solves out of ~85 throughout. |
| It is out of reach for a 9B here, and no harness variant separates from another (Fisher p=0.62 |
| for the widest gap). |
|
|
| β **The 100-task pi-ws figure (0.316) is the optimistic tail of the run-to-run spread.** The same |
| harness and weights measured 0.208 on 250 tasks, and paired on their 95 shared tasks the two runs |
| agree (McNemar p=0.48) β the gap is which tasks landed in the first 100, not a real difference. |
| This is exactly the failure the brief warns about; **0.208 is the number I stand behind.** |
|
|
| ### Candidates that did not pan out (all measured, all kept in `runs/`) |
|
|
| | tried | result | |
| |---|---| |
| | 900 more SFT steps on the unseen 65% of the corpus (`sft_cont2`) | 7/64 swe stock, 14/58 pi-ws β no gain; supervised signal saturated | |
| | checkpoint averaging of the three (`soup1`) | 7/63, 13/58 β no gain | |
| | `temperature 0.2` on pi-ws | 21/81 = .259 vs .253, but ~2Γ episode length β rejected | |
| | pi-ws per-turn caps (`maxTokens 4096`, `contextWindow 60000`) | **actively harmful now**: 23/91 vs 31/98 uncapped (paired 5 vs 11) β removed | |
| | Dockerfile-derived workdir for terminal-bench | 0/79 and 1/86 vs 3/88 stock β left off by default | |
| | a sharper prompt that insists on verifying after the last edit (`cfg/agent_prompt_v2.txt`) | 53/227 = .233 vs .208 for the shipped prompt, paired 27 vs 19, **McNemar p=0.30** β not separable, so the shipped prompt (the one every headline number was measured with) stays | |
| | GRPO through pi | blocked by **two independent** things, both now diagnosed exactly. (1) pi hard-codes `stream: true` and prime-rl's `TrainClient` β the only client returning the token ids/logprobs the RL loss needs β raises on streaming. **Solved**: `pkgs/pi_rl` starts a ~90-line Node shim in the task container, points pi's `baseUrl` at it, and it forwards non-streaming then re-emits SSE. Verified transparent under the eval client (6/6 episodes, same turn count and token usage as un-shimmed). (2) **The RL rollout path renders a different system prompt than the serving path.** The renderer serialises whatever tool objects it is handed; verifiers hands it its own flat `ToolSpec`, so the tool block comes out as `{"name":β¦,"description":β¦,"parameters":β¦}`, whereas the served chat template emits `{"type":"function","function":{β¦}}`. Verified byte-for-byte on a real pi conversation: `renderer(flat ToolSpec) == served template` β **False**; `renderer(OAI-nested) == served template` β **True** (68 chars of difference in the system prompt, on every single turn). So the policy is off-distribution for the entire rollout: **0/85 solved** on tasks the identical harness solves ~16% of the time under the eval client, and after temperature 1.0 β 0.7, 57 of 76 episodes ran to the 40-turn cap instead of finishing. The fix is a one-line change in how verifiers hands tools to the renderer β shared-install code, so out of scope here. | |
| | GRPO through the verifiers-native `bash` harness (which does *not* stream, so the stack runs) | the harness runs and batches fill, but **every reward is 0**: the policy scores 0/24 on r2e-gym-ws under `bash` even with pi's own system prompt injected, against ~16% under pi. An SFT'd agent is tightly coupled to its harness's tool schema and loop, so there is no signal to learn from. | |
| | expert iteration on r2e-gym-ws | 203 successes but only **32 distinct tasks** β too narrow to train on | |
|
|
| ### What the failures actually look like (250-task stock run) |
|
|
| Among **solved** episodes the model runs a check after its last edit 36% of the time; among |
| **failed** ones, 17% (any check at all: 58% vs 45%). Verifying is the behaviour most associated |
| with succeeding and the policy does it unreliably β but **72% of the training corpus already |
| verifies**, so this is a generalisation gap, not a coverage gap, which is consistent with more |
| SFT buying nothing. It is the kind of gap RL closes, and RL is the thing this stack would not |
| run. 35 of 217 failed episodes never made an edit at all; ~10% of tool calls in failed episodes |
| name a tool that does not exist (`submit`, `submission`, `output`). |
|
|
| | ckpt | suite | harness | n | solved | score | ci95 | |
| |---|---|---|---|---|---|---| |
| | base (+eos fix) | swe-bench-verified | stock pi | 97 | 4 | 0.041 | [0.016, 0.101] | |
| | base (+eos fix) | swe-bench-verified | pi-ws | 64 | 12 | 0.188 | [0.111, 0.300] | |
| | base (+eos fix) | terminal-bench-2 | stock pi | 83 | 2 | 0.024 | [0.007, 0.084] | |
| | **SFT v5 step_900** | swe-bench-verified | stock pi | 88 | 11 | **0.125** | [0.071, 0.210] | |
| | **SFT v5 step_900** | swe-bench-verified | pi-ws | 91 | 23 | **0.253** | [0.175, 0.351] | |
| | **SFT v5 step_900** | terminal-bench-2 | stock pi | 88 | 3 | **0.034** | [0.012, 0.096] | |
| | **SFT v5 step_900** | terminal-bench-2 | pi-ws | 85 | 1 | 0.012 | [0.002, 0.064] | |
|
|
| Paired (McNemar, same seeded task subset): |
| * swe stock, base β SFT: solved 9 tasks base did not, lost 1. **p=0.021** β the weights alone |
| roughly triple the stock number (0.041 β 0.125). |
| * swe pi-ws, base β SFT: +7 / β3, p=0.34 β not separable at this n; the scaffold had already |
| captured much of what SFT teaches (find the repo, use absolute paths). |
| * terminal-bench-2 moves ~nothing either way. It is simply out of reach for a 9B here: the |
| base solves 2/83, the SFT 3/88. pi-ws is 1/85 β no evidence it helps on this suite, and the |
| starting-in-`/app` change removes the orientation step the model was trained to perform. |
|
|
| β Sandbox losses matter for these numbers: episodes that die as `SandboxError` drop out of `n`. |
| Use `eval --resume <run_dir>` (`scripts/resume_eval.sh`, or `scripts/topup.sh` which retries in |
| free windows) to re-run them β that is how the swe runs got from nβ74 to nβ90. Resume needs the |
| vLLM server **still up**; `measure.sh` kills it at the end, so restart it first or every resumed |
| rollout dies with `ProviderError`. |
|
|
| β **Never read a partial run as a score.** Episodes finish in roughly ascending order of |
| difficulty, so a run at 30% completion reads far too high β `big_swe_stock` showed 8/33 = 0.24 |
| early against its true β0.17, and `sft5_swe_ws_t02` showed 3/8 = 0.38 against β0.26. Only |
| compare completed runs. |
|
|
| **The scaffold alone is worth ~15 points on swe-bench-verified** (Fisher p=0.005; paired on the |
| 63 shared tasks, pi-ws solved 11 that stock did not and stock solved 0 that pi-ws did not, |
| McNemar p=0.001), on frozen base weights. |
|
|
| β **Attribution correction.** When those numbers were taken, pi-ws's `cd` was a *no-op*, so the |
| gain came from the appended `cfg/agent_prompt.txt` (which names `/testbed`) plus the |
| `maxTokens`/`contextWindow` caps β not from the working directory. The reason: pi resolves every |
| tool call against the ACP **session** cwd, which `verifiers/v1/acp/_runner.py` sets from |
| `os.getcwd()` of the *runner* process; `cd`-ing the agent command underneath it changes nothing. |
| `pi_ws.run_acp_in` now starts the runner itself inside the workdir (re-anchoring the config path |
| and exporting `VF_LAUNCH_BASE` so `PI_ACP_PI_COMMAND` / `PI_CODING_AGENT_DIR` / skills still |
| resolve against the launch dir). Verified live: `pwd` returns `/testbed` on swe-bench and `/app` |
| on terminal-bench, with the task's files right there. The measured 0.188 therefore *understates* |
| the current harness. |
|
|
| **Gold-patch validation of the RL taskset** (`validate --only-gold`, 6 tasks, broker runtime): |
|
|
| | taskset | valid | |
| |---|---| |
| | `r2e-gym-v1` (stock) | **0/6** β `gold apply failed: error: Orange/data/util.py: No such file or directory` | |
| | `r2e-gym-ws` (my wrapper) | **6/6** valid, ~80 s each | |
| | `swelego-v1` (stock) | **6/6** valid, ~25 s each β works as shipped, and grades ~3Γ faster | |
|
|
| The stock r2e taskset cannot score a *correct* patch under this runtime; the wrapper is what |
| makes RL on it possible at all. RL trains on both sources. |
|
|
| ## β HARD-LEARNED RULE |
|
|
| **`/var/lib/agentptb-cache/a/` DOES NOT SURVIVE A NODE RESTART.** At ~04:46 on 2026-08-14 the |
| node was recreated and everything there except `prime-rl/` and `tmp/` was erased: the local base |
| model copy, the 72 GB teacher, the raw trajectory corpora, and **all SFT v1 checkpoints** (the run |
| had reached step 600/940). ~9 h of GPU work lost. The workspace on the PVC was untouched. |
|
|
| β Model checkpoints go to `$AGENTPTB_WORKSPACE/ckpt/...`. Local disk is for the read-only model |
| copy (fast weight loading) and TMPDIR only β anything reproducible in <15 min. |
|
|
| ## Where things are (update as you go) |
|
|
| | artifact | path | |
| |---|---| |
| | base + generation_config fix | `/var/lib/agentptb-cache/a/models/base` | |
| | teacher (data-gen only, never weights) | `/var/lib/agentptb-cache/a/models/teacher35` | |
| | SFT corpus v1 (15,112 traj, 244M tok) | `data/sft_v1` | |
| | SFT corpus v2 (30,583 traj, real trajectory endings) | `data/sft_v2` | |
| | SFT corpus v3 (46,286 traj, ~750M tok) | `data/sft_v3` | |
| | **SFT corpus v4 (v3 + 30% path-relocated) β the one in use** | `data/sft_v4` | |
| | SFT v2 run + checkpoints | `ckpt/sft_v2/weights/step_N` | |
| | local plugins (PYTHONPATH) | `pkgs/{r2e_gym_ws,pi_ws}` | |
| | measure driver | `scripts/measure.sh` β args: `<model_dir> <tag>`, then suite (both/swe/tb2), arm (stock/ws/both), concurrency | |
| | **status at a glance (run this first on resume)** | `bash scripts/status.sh` | |
| | scoring | `scripts/score.py <run_dir>` | |
| | raw-trajectory corpora | `/var/lib/agentptb-cache/a/data/{klear66k,swesmith_traj,r2e_sft,swegym_oh,nebius,deepswe_k2}` | |
|
|
| Checkpoints are HF-servable as saved (config + chat_template.jinja + tokenizer + the correct |
| `generation_config.json` with `eos_token_id:[248044,248046]`) β serve the `weights/step_N` dir directly. |
| |
| - [h20] Orientation training is taking. At an equal 180 steps, under the **stock** harness: |
| |
| | corpus | mentions `/testbed` | emits an early `pwd`/`ls -d` probe | solved | |
| |---|---|---|---| |
| | v4 (no orientation turns) | 3/8 | 0/8 | 0/8 | |
| | v5 (40% orientation turns) | 8/16 | 5/16 | 1/15 | |
| |
| Still only 20% through training. `scripts/compare.py A B` does Wilson + Fisher for run pairs; |
| the shuffled 100-task swe subset is seeded, so runs are paired on identical tasks. |
| |
| ## Status: COMPLETE (2026-08-17 17:40Z). Submission = `ckpt/sft_v5/weights/step_900`. |
| |
| **Submission: `ckpt/sft_v5/weights/step_900` + `pkgs/pi_ws`**, evaluated with |
| `cfg/big_swe_ws.toml` / `cfg/ws_tb2.toml` (mine) against `cfg/big_swe_stock.toml` / |
| `cfg/base_tb2.toml` (stock). |
| Verified: 15 files, 4 shards, 760 tensors, 18.8 GB, `eos_token_id=[248044,248046]`, cold-serve |
| tested and answering correctly. Every published figure was recomputed from `runs/*/traces.jsonl` |
| (finding 34) and every config/script named in the reproduction block was checked to exist and to |
| produce the number claimed for it (finding 35). |
|
|
| **Machine state:** no eval, training or RL jobs are running and no tunnels are held. One vLLM |
| server is deliberately left up β the **submitted** checkpoint on `localhost:8000` (GPUs 0,1), so a |
| verifier can hit it immediately; GPUs 2,3 are free. 19 measurement runs are on disk and **every one |
| of them is complete with zero ungraded episodes**. |
|
|
| **Final numbers β every swe-bench cell replicated:** |
|
|
| | | stock pi | pi-ws (submitted system) | |
| |---|---|---| |
| | swe-bench-verified | **0.151** [0.118, 0.193] (n=350) | **0.202** [0.169, 0.239] (n=500) | |
| | terminal-bench-2 | 3/89 = 0.034 | 1/89 = 0.011 | |
|
|
| **The interpolated checkpoint `ckpt/soup_base_sft` was submitted for ~6 h and withdrawn on its own |
| pre-registered rule** (findings 30β33): pooled, it buys +5.2 points under pi-ws and costs β7.1 under |
| the stock harness, so `step_900` leads on the sum, 0.353 to 0.334. It stays in the workspace fully |
| measured, and is still ahead on the my-harness number alone. |
|
|
| **What the run established (all replicated or at n=250+):** |
| 1. **The scaffold is the result** β +23 points on base weights (p<0.0001), +7.6 on the submitted |
| weights (p=0.005, reproduced p=0.023), +19.6 on the soup. Both halves of it contribute; neither |
| is significant alone (finding 25). |
| 2. **SFT helps monotonically under the stock harness** (0.044 β 0.104 β 0.132/0.151) and **hurts |
| under pi-ws at every training length**, non-monotonically (findings 23, 26). |
| 3. **Weight-space interpolation trades the two published numbers roughly one-for-one** β no free |
| lunch on that axis (findings 30β33). |
| 4. **terminal-bench-2 measures nothing** (all pβ₯0.125); the scaffold is worth zero there, so its |
| gain is specific to the SWE-repo setting (finding 28). |
| 5. **Noise floor: 15.6% of tasks flip between two runs of an identical system**, symmetric |
| (finding 29). Every headline cell is replicated because of it β and both replicates moved the |
| answer *against* whichever checkpoint I preferred at the time. |
|
|
| Six pre-registered hypotheses: four refuted (prompt off-distribution 25, shorter anneal 26, RL |
| through the real harness 18, tb2 corroboration 28), one confirmed then withdrawn on its own rule |
| (interpolation 30β33), plus the replicate protocol (29) which changed how everything is quoted. |
|
|
| **The uncomfortable summary, kept at the top of `SUBMISSION.md`:** on the harness I submit, my |
| training is a net liability; the scaffold, not the training, is what carries this result. I could |
| only see that by finishing the 2Γ2 instead of arguing about it. |
|
|
| ## Next actions (if resumed) |
|
|
| The submission is `ckpt/sft_v5/weights/step_900` + `pkgs/pi_ws`, complete, measured on both suites, |
| with every swe-bench cell replicated and every published number recomputed from disk (finding 34). |
| Nothing is unfinished. If someone picks this up, the things actually worth doing, in order: |
|
|
| 1. **Explain the SFT/scaffold conflict.** This is the central unexplained result and everything |
| else in the run circles it: SFT helps monotonically under the *stock* harness |
| (0.044 β 0.104 β 0.151) and **hurts under pi-ws at every training length**, non-monotonically β |
| 200 steps (0.168) is worse than 900 (0.202), and both are below the untrained base (0.276) |
| (findings 23, 26). Three mechanisms were tested and refuted: prompt off-distribution (25), |
| shorter anneal (26), and weight-space interpolation, which trades the two published numbers |
| roughly one-for-one rather than fixing anything (30β33). |
| The place to start is a *behavioural* diff, not another sweep: the traces are all in `runs/`, |
| and `base250_ws` solves 38 tasks that `big_swe_ws` does not. Compare turn counts, tool mix, and |
| where the two diverge on exactly those tasks. |
|
|
| 2. **Sweep the interpolation coefficient, if the my-harness number is what matters.** Only Ξ±=0.5 |
| was measured, chosen a priori to avoid selection noise. It lands at pi-ws 0.254 / stock 0.080 |
| against step_900's 0.202 / 0.151 β i.e. it moves along a roughly one-for-one trade line |
| (finding 32). Ξ± β {0.25, 0.75} would show whether that line is really straight, but budget for |
| the noise floor: Β±5 points needs 250 tasks *and* a replicate, or you will measure nothing |
| (findings 29, 31, 33). |
| |
| 3. **RL**, if there are days rather than hours. Both known blockers are cleared and verified in |
| `pkgs/pi_rl` + `pkgs/render_parity.py` (findings 14, 18), but episode prompts reach p90 42k / |
| max 101k tokens against `seq_len=32768`, so a third of every batch is untrainable. Needs 65536, |
| `max_turns` unlimited to match eval, the system prompt on the train sources, and a sandbox pool |
| that can sustain ~100 episodes/hour. |
|
|
| **Do not bother with:** more harness knobs (finding 19 β everything left is 1β3 points against a |
| 15.6% task-level noise floor); more of the same SFT (saturated, finding 26); or reading any partial |
| run as a result. That last one misled me four separate times (findings 23, 25, 31, 33), twice in |
| the direction of the checkpoint I happened to prefer at the moment. |
|
|
| **If you change the submission, replicate both sides before you do.** Every time I measured one arm |
| and switched, the next replicate moved the answer back: 0.280 β 0.254 for the interpolation right |
| after I adopted it, 0.132 β 0.151 for step_900 right after I abandoned it. |
| |