RUN NOTES β agentic SWE/terminal post-training (100h)
Deadline epoch: 1786991334 (date -d @1786991334) β boot 2026-08-13T14:28Z, ends 2026-08-17T18:28Z.
Check remaining: echo $(( $(cat DEADLINE | cut -d. -f1) - $(date +%s) )) seconds.
Environment map (verified)
| thing | path |
|---|---|
| workspace | /mnt/pvc/users/simon/agentptb/runs/a-opus-max/workspace (= $AGENTPTB_WORKSPACE) |
| prime-rl | /root/work/a/prime-rl β real /mnt/pvc/users/simon/agentptb/work/a/prime-rl |
| venv | /root/work/a/prime-rl/.venv (bins: sft, rl, eval, orchestrator, trainer, inference) |
| verifiers src (editable) | /root/work/a/prime-rl/deps/verifiers |
| research-environments (editable tasksets) | /root/work/a/prime-rl/deps/research-environments/environments/{swe,terminal,tool_use,code,...} |
| base model | $HF_HOME/hub/models--Qwen--Qwen3.5-9B-Base/snapshots/68c46c4b... |
| eval tasksets on disk | /root/work/shared/tasksets/{terminal-bench-2,swe-bench-verified} (symlinked into ~/.cache/harbor/) |
| my GPUs | CUDA_VISIBLE_DEVICES=0,1,2,3 (B200 183GB each) |
| sandbox broker | $SANDBOX_BASE_URL + X-API-Key: $SANDBOX_API_KEY |
| TMPDIR / caches | /var/lib/agentptb-cache/a/tmp (local disk β keep them there) |
Model facts
Qwen3.5-9B-Base:Qwen3_5ForConditionalGeneration, hybrid linear+full attention (full every 4th layer), 32 layers, hidden 4096, head_dim 256, vocab 248320, has a vision tower (βlimit_mm_per_prompt={image=0,video=0}).- It already ships the full Qwen3.5 ChatML chat template (
<|im_start|>/<|im_end|>,<tool_call>,<think>). eos_token = <|endoftext|>(248044) but turns end with<|im_end|>(248046). Nogeneration_config.json. β vLLM stops only on 248044 β base model runs to the token cap every turn. Fix: addgeneration_config.jsonwitheos_token_id:[248044,248046](and/or stop strings in sampling).
Non-obvious operational rules (from eval-kit README)
- serve with
--enable-auto-tool-choice --tool-call-parser qwen3_coderelse 400s - runtime
block_network = false(pi installs itself in the task container at rollout time) export PRIME_API_KEY="$(cat "$AGENTPTB_PRIME_KEY_FILE")"β already in env too- trainer
attn = "flash_attention_2"(auto β fa4 β grad norm inf β silently no-op) batch_sizemust be several multiples ofgroup_size- kill orphan trainers/vllm (they hold GPU mem + ports)
Rules I must respect
- No training on terminal-bench-2 / swe-bench-verified (reading/running them = fine). No test-item-clustered synthetic data.
- Submitted weights must derive from
Qwen3.5-9B-Baseby my own training. Do not touchQwen/Qwen3.5-9B(post-trained sibling) as a weights source β that voided a prior run. - Distillation from public models I run myself IS allowed (rule 7). Env API keys are NOT a data source (rule 5).
- Do not read: operator notes, other cells' workspaces, prior runs' logs, benchmark repo.
(
/root/work/operator-debug,$AGENTPTB_WORKSPACE_{A,B}β leave alone.) NOTE:$HF_HOMEis a shared cache and contains datasets other cells pulled. I use it only as a model/dataset cache for things I decide on independently; I do not treat its contents as a strategy signal.
Key findings
generation_config.jsonis missing on the base and it matters a lot. Chat template ends turns with<|im_end|>; without it as an eos id vLLM never stops there, so the base model continues past its own turn and hallucinates the user/tool turns (16384-token completions,stop=context_length). The official post-trained sibling shipseos_token_id:[248046,248044]β so this is correct packaging, not a trick. Added to/var/lib/agentptb-cache/a/models/base/generation_config.json; mean completion tokens/turn dropped from ~16k to ~250. Must be in every checkpoint I serve/submit.Template is thinking-ON: generation prompt ends
<|im_start|>assistant\n<think>\n. Tool-call syntax is Qwen3-Coder XML (<tool_call><function=name><parameter=x>), matching--tool-call-parser qwen3_coder. Don't fight it β train short<think>blocks.Renderer must be pinned:
Qwen/Qwen3.5-9B-Baseis NOT inMODEL_RENDERER_MAP, soautofalls back to the default renderer (no tool support). Always setrenderer.name = "qwen3.5".Qwen/Qwen3.5-35B-A3B(post-trained MoE, 3B active, 72GB) has a byte-identical chat template to the base β ideal distillation teacher. Downloaded to/var/lib/agentptb-cache/a/models/teacher35. Used ONLY as a data generator (rule 7), never as weights (rule 3).pi defaults
maxTokens=16384per turn (models.jsonknob) β harness-side lever for my own harness.Trace tool_calls are FLAT (
{id,name,arguments}); prime-rl SFT needs OAI nested ({id,type,function:{name,arguments}}) ordeserialize_tool_callsbreaks the renderer.Broker pulls arbitrary public Docker Hub images fine (tested ubuntu, swebench/, alexgshaw/). It does NOT pull Prime-platform refs (
prime/primeintellect/...) βtmax-v1,terminal-lego-v1,openthoughts-tblite-v1are unusable.swesmith-v1images also fail (TaskError/SandboxError).BrokerRuntimehas noworkdirfield.utils/compile.pyonly appliestask.data.workdirwhen the runtime config has that field, so with the broker every exec runs in/workspace(a scratch dir), never the task's/testbed. Two consequences:- the agent lands outside the repo and must find it (only 43/103 teacher rollouts ever
mentioned
/testbed, median 7 turns in) β a real, large capability tax on swe-bench; r2e-gym-v1scores 0 for everyone: itssolved()runssh -c "/bin/bash run_tests.sh"with a relative path, which does not exist in/workspace. Same class of bug makes itscapture_patchfail withfatal: not a git repository. This is why the 35B teacher scored 0/52 on r2e-gym β not a capability result.swebench-verified-v1(harbor verifier, absolute paths) is unaffected and does score.
- the agent lands outside the repo and must find it (only 43/103 teacher rollouts ever
mentioned
Round-trip verified: my converted rows render byte-identical to
tokenizer.apply_chat_templateunder theqwen3.5renderer, loss mask = assistant spans (<think>β¦</think>β¦<tool_call>β¦<|im_end|>). 8/8 exact on a sample.No local container escape hatch.
docker/podmanexist on this node but rootless podman cannot start a container: pulls work once~/.config/containers/storage.confsetsignore_chown_errors=trueandUSERis corrected fromroottoagentptb(that env var being wrong is what produces the misleading "no subuid ranges for user root"), but every run then dies atcrun: mount 'proc' to 'proc': Operation not permittedβ the pod hasCapEff=0and mount(2) is blocked. The broker really is the only way to run tasks. Don't retry this.The broker exposes
GET /resources(undocumented in the runbook): node capacity plus every sandbox pod and its phase. This is the way to tell "my job is slow" from "the cluster is full" β{'Pending': 222, 'Running': 7}means the queue is wedged, not that I am doing something wrong. Pod size does not help: 1 CPU / 4 GiB requests queue behind the same wall.Three traps between a finished SFT and a running RL job, all of which look like unrelated failures:
/appholds a different prime-rl build and the defaultPATHputs/app/.venvfirst, so the launcher spawns its orchestrator/trainer from there. Its schema disagrees (max_inflight_rolloutsvsmax_inflight_episodes,train.envvstrain.source) and its verifiers has no broker runtime at all. Symptom:Extra inputs are not permitted --train.source. Fix:scripts/run_rl.shprepends/root/work/a/prime-rl/.venv/bin.- The trainer's saved checkpoint has no
preprocessor_config.json, so prime-rl's inference server dies withCan't load image processor for <ckpt>. (Plainvllm servedoes not care.) Fix: copypreprocessor_config.json,video_preprocessor_config.json,merges.txt,vocab.jsonfrom the base dir into every checkpoint β do this before submitting too. - prime-rl's inference does not pass
limit_mm_per_prompt, so vLLM profiles the vision tower and crashes in the CUTE/FA4 kernel withTypeError: fmax() missing 1 required positional argument. Fix:[inference.vllm_extra] limit_mm_per_prompt = {image=0,video=0}.
GRPO through the pi harness does not work in this stack. After clearing the three traps above, every rollout dies with
ACP agent produced no visible reply, and the env log shows the real cause:model call failed: TrainClient does not support streaming. pi always streams; the interception server routes streaming requests to the client, and prime-rl'sTrainClient(the one that returns token ids + logprobs for the RL loss) raises on stream. The eval client proxies streams fine, which is why evaluation works and training does not. Fixing it means changing shared verifiers code, so RL was replaced with expert iteration through the same pi harness: sample β keep reward-1 trajectories β SFT. Same tool surface, same prompt, every component already validated.Both RL blockers are now cleared, in local plugin code, without touching the shared install. Finding 13 was blocker one (streaming); blocker two was render parity.
- Streaming:
pkgs/pi_rlwrites a ~90-line Node shim into the task container, points pi at127.0.0.1:8899, and the shim forwards each call upstream withstreamstripped, then re-emits the JSON reply as the SSE chunks pi expects. Transport-only β same agent, same tools, same prompt. Verified transparent: 6/6 episodes, same turn count and token usage. - Render parity: the renderer serialises whatever tool objects it is handed, and verifiers
hands it its own flat
ToolSpec, so the rollout's tool block came out as{"name":β¦,"description":β¦,"parameters":β¦}while the serving chat template emits{"type":"function","function":{β¦}}β 68 chars different in the system prompt of every turn.pi_rl.enable_render_parity()normalises tools to the OAI-nested form before the renderer sees them, in the training process only.
Settled against the live server, not against a reconstruction. POST the same messages+tools to vLLM's
/tokenizewithadd_generation_prompt=trueand compare token ids withrenderer.render_ids(...):tokens identical to served prompt unpatched renderer 916 no patched renderer 952 yes vLLM /tokenize952 β That also settles the one open question in the patch: pi sends
strictin each tool schema and the served template keeps it, so_nest_toolsmust preserve it. Dropping it costs 36 tokens and puts the policy off-distribution again.enable_render_parity()is called atpi_rlimport;pi_rlis training-only and nothing in the evaluation path imports it.Where the patch has to live, and why the obvious placements both fail silently.
- Wrong process. The renderer is built in prime-rl's orchestrator
(
orchestrator/utils.py::setup_policy_inference_pool); the harness β and thereforepi_rlβ is imported in the env-server. Patching atpi_rlimport time lands in a process that never renders. The patch now lives inpkgs/render_parity.pyand is installed bypkgs/sitecustomize.py(whichPYTHONPATHalready reaches) underVF_RENDER_PARITY=1, set byscripts/run_rl.shand nothing else. Evaluation shares thatPYTHONPATH, never sets the variable, and is untouched. - Wrong moment. The cheap version β wrap
builtins.__import__, checksys.modulesafter each call β does not work. Python inserts a module intosys.modulesbefore running its body, so the first sighting ofrenderers.qwen35(an inner import from inside that very module) finds noQwen35Rendereryet; fire-once logic then retires the hook having patched nothing.install()uses asys.meta_pathfinder that wraps the real loader and patches insideexec_module, after the body has run. pkgs/sitecustomize.pyshadows/usr/lib/python3.12/sitecustomize.py, so it redoes that file's only job (installing Ubuntu's apport hook) before its own.
Verified live, in the running job, not just on a bench:
[render-parity] Qwen35Renderer patchedappears inorchestrator.logat theInitializing policy inference poolline, and a real rollout's first call logsprompt_tokens=2562β which is exactly what the patched renderer produces for that conversation (unpatched: 2526).- Streaming:
A stale
vllm::routersilently blocks the next RL launch. Symptom isError: Inference failed with exit code -15about 13 s after startup, which reads like an OOM or an external kill. The real cause is inckpt/<run>/logs/inference.log:PanicException: failed to install Prometheus metrics exporter: FailedToCreateHTTPListener("Address already in use")β a router from a previous, killed run still holds ports 8000 and 29000.pkill -f "vllm serve"does not match it: the router renames its own process tovllm::router, so onlypkill -x "vllm::router"finds it.scripts/launch_rl.shclears it before every launch.Related, and the same trap as the old
pgrep -fself-match: usepkill -x(process name), neverpkill -f(full command line), in any script written from a heredoc. A-fpattern matches every shell whose command line quotes that string β including the shell writing the script β so the launcher kills itself and produces no output at all to explain why.An RL run cannot be found, or killed, by the names you launched it with. Every process renames itself: the launcher becomes
PRIME-RL::Launcher, thenPRIME-RL::Orchestrator,PRIME-RL::Trainer,PRIME-RL::EnvServer,vllm::router. Consequences, all of which cost time here:ps | grep run_rl/grep orchestratorshows nothing, so a live run looks dead.- The pidfile is useless:
setsid nohup bash β¦ &records the transientsetsidPID, which exits at once.kill $(cat rl.pid)reports success and kills nothing. pkill -x "PRIME-RL::Launcher"also fails β-xmatches the kernel'scomm, capped at 15 characters, so the real name isPRIME-RL::Launc.
A previous run therefore survives invisibly, and the next launch dies on its ports ~80 s in.
scripts/launch_rl.shsweeps by comm prefix throughps -eo pid=,comm=(matching^(PRIME-RL|vllm|VLLM)::), kills by PID, then re-checks and refuses to launch if anything survived. Kill the launcher first or it restarts its children. Watch for an orphanedPRIME-RL::Trainholding ~40 GB β if it outlives its launcher it silently starves the next trainer of GPU memory.RL ran, and it is closed β for a fourth reason, which is structural rather than a bug. With blockers 1 and 2 cleared and verified live, GRPO produced 0 reward across 99 scored episodes (r2e-gym-ws 0/50, swelego-v1 0/33). Ruling things out, in order:
- Not the renderer. Verified in the running job: a live rollout's first call logged
prompt_tokens=2562, exactly the patched render (unpatched: 2526). - Not tool parsing. The train path parses completions with the renderer instead of vLLM's
qwen3_coderparser, so this was a real risk β but 2926 of 2949 assistant messages carrytool_calls. Rollouts read like competent work ("The fix works", "issue resolved"). - Not the grader. Every scored episode carries
info.patch_error: fatal: not a git repository, which looks damning and is a red herring:capture_patch's own docstring says a failed capture still lets the rollout score. Gold-patch validation ofswelego-v1on a free pool: 6/6 valid. A correct patch does score.runs/probe_shimis not a baseline for these tasksets β it is swe-bench. - Not my harness.
pi_ws.run_acp_inuses a per-commandcd, so it cannot alter the cwd of the grader's laterruntime.runcalls.
What it actually is β two config mismatches of mine, and one wall:
eval (scores 0.208) RL as configured max_turnsNone (unlimited; pi stops when done) 40 β and 61% of episodes hit it system_promptcfg/agent_prompt.txtnone set on either train source The wall: pi's episodes on these tasks are long. Peak prompt tokens per episode are p50 20k, p90 42k, max 101k, so 24 of 113 episodes already exceed the trainer's
seq_len = 32768at 40 turns. Raisingmax_turnsto match eval makes more of the batch untrainable, not less. Fixing it properly meansseq_len65536 on one B200 plus ~6 min of sandbox time per episode β at batch 96 and 24 inflight that is β³40 min/step, so the clock buys ~50 steps with a third of each batch discarded.That is a real project, not a fix, so I stopped rather than start it with ~34 h left and a fully-measured checkpoint to protect.
pkgs/pi_rl+pkgs/render_parity.pyare left working and documented; they are not part of the submission. For anyone resuming: setmax_turnsunlimited, add the system prompt to both sources, raiseseq_lento 65536, and expect the sandbox pool β not the GPUs β to be the bottleneck.- Not the renderer. Verified in the running job: a live rollout's first call logged
Failure-mode audit of my own traces (4,496 episodes), and why it did not lead to another training run. Three signatures, measured rather than guessed:
signature prevalence (SFT model) prevalence (base) solve rate with / without leaked tool-call XML in content(</parameter></function></tool_call>)12.9% 0.8β2.0% 5.7% vs 16.5% "I apologize for the difficulty/confusion" 41.7% ~0% 13.5% vs 16.3% SWE-agent harness artifacts ("Exit due to cost limit", "submit button") 0.5% 0% β The apology habit and the harness artifacts are my SFT corpus talking β the base model essentially never produces them. But the apology is style, not damage (13.5 vs 16.3 is weak and confounded: hard tasks cause both), and the harness artifacts are too rare to matter. I had expected the artifacts to be the story; they are not. Worth stating plainly because the first example I looked at was an "Exit due to cost limit", which is exactly how a 0.5% effect gets mistaken for the main one.
The leaked-XML signature is real and 3Γ predictive, but it caps out small: eliminating it entirely moves 12.9% of episodes from 5.7% to 16.5%, i.e. +1.4 points absolute β a quarter of the Β±5-point interval at n=250. Not worth an 8 h retrain plus an 8 h re-measurement.
A separate cut, on the verification behaviour the earlier failure analysis flagged:
swe-bench, episodes that made an edit ran a test after the last edit solved (n=51) 63% failed (n=174) 37% Same direction as
cfg/agent_prompt_v2.txt, which was built for exactly this and measured at 0.233 vs 0.208, McNemar p=0.30 β directionally right, not separable. Note only 12% of failures never edited at all, so "nudge it to act" is the wrong lever; "make it verify" is the right one, and it is already in the prompt.Conclusion, and the reason this section ends here: every remaining lever I can identify is worth ~1β3 points, and my measurement floor is Β±5 at n=250. Chasing them would produce changes I could not distinguish from noise, which is how a benchmark run talks itself into a regression. The honest use of the remaining time was to make the reported number better, so the last measurement runs the complete 500-task suite on both harnesses instead of a 250-task sample, which is the one improvement that does not depend on my guessing right.
Fairness audit of the two arms β one asymmetry found, and it did not bite. The
pi-wsconfigs setready_timeout_seconds = 1200; the stock configs did not, and the broker default is 600 (v1/runtimes/broker.py:59). Half the readiness budget means roughly twice the sandbox-timeout rate under a busy pool, and the logs do show more error lines on the stock side (big_swe_stock71 vsbig_swe_ws29) β an asymmetry pointing in my favour, which is the direction that must never go unchecked.Checked against the traces rather than the logs, because the two count different things: an episode that has no
rewardsrecord at all is one that never got graded, i.e. genuinely lost to infrastructure.run n solved ungraded score score excluding ungraded big_swe_stock250 33 0 0.132 0.132 big_swe_ws250 52 0 0.208 0.208 Zero ungraded episodes in either arm β the log error lines were transient and retried, no episode was lost, and the headline comparison is unaffected.
cfg/final500_stock.tomlnow sets 1200 explicitly anyway, and the 500-task stock arm was restarted after the fix so that both arms are identical in everything except the harness and the appended prompt.The base-weights numbers were the weakest link in the attribution, so they are being re-measured too. The submission claims two separate gains β weights and scaffold β but they rested on different-quality evidence:
n quality big_swe_base(base + stock)153 176 errored episodes in the same run base_swe_ws(base + pi-ws)64 taken when pi-ws still had the no-op cdand themaxTokens/contextWindowcaps that are now off β a different harnessbig_swe_stock/big_swe_ws(submitted)250 / 250 clean, 0 ungraded Taken at face value the old pair says base+pi-ws = 0.188 [0.11, 0.30] against submitted+pi-ws = 0.208 [0.16, 0.26] β heavily overlapping, i.e. on the submitted harness I could not show the weights help at all. That reading is not sound (different harness, n=64), but it is the right worry, and the honest fix is power, not argument.
So the final measurement is the full 2Γ2 at n=500: {base, submitted} Γ {stock, pi-ws}, same 500 tasks, same protocol, base served on port 8001 so it can never be confused with the submitted checkpoint on 8000.
cfg/final500_{stock,ws,base_stock,base_ws}.tomldiffer only in harness id, the appended prompt, and the port.Sandbox pod sizing, corrected β and what a saturated cluster looks like. Each sandbox pod requests 4 CPU, not 1 as finding 17 assumed. On a ~250β330 CPU shared cluster that means the whole cluster holds only ~60β80 concurrent sandboxes across all tenants, so 40 concurrent from me was over half of it.
On 2026-08-16 from ~04:40Z the pool went to 371 Pending / 34 Running with
available.cpuat 0 β a ~1,500 CPU backlog against a cluster that has none β while only ~20 of those requests were mine. Throughput went to roughly zero regardless of what I did. Symptoms to recognise next time:avail_cpufalling to 0 whilePendingclimbs into the hundreds β the queue is global, so dropping my concurrency does not clear it and does not restore my throughput.- A single hand-made sandbox going ready in 3 s is not evidence of headroom (finding 17); when the cluster is truly full even that stalls.
/resourceshas no owner field, so pods cannot be attributed. Judge by arithmetic: my outstanding requests versus total Pending.
Practical rule: size runs to finish, because a partial run is not usable β with
shuffle=falsea prefix is the dataset's own ordering, and withshuffle=trueit is ascending difficulty (Results section). Both bias the score. When the pool is saturated, prefer the measurement that fills a genuine gap over the one that merely adds precision.The 2Γ2 lands at n=250, and it does not just correct the attribution β it reverses part of it. All four cells complete, zero ungraded, identical 250 tasks (
SEED=0is pinned, soshuffle=truedraws the same sample every run).weights stock pi pi-ws base (+eos fix) 0.044 [0.025, 0.077] 0.276 [0.224, 0.334] submitted (SFT) 0.132 [0.096, 0.180] 0.208 [0.162, 0.263] paired comparison (n=250) only-A only-B delta McNemar base: stock β pi-ws 4 62 +0.232 <0.0001 submitted: stock β pi-ws 12 31 +0.076 0.0054 stock: submitted β base 29 7 β0.088 0.0003 pi-ws: submitted β base 21 38 +0.068 0.036 - The scaffold is the robust result: +23 points on base weights, +7.6 on the submitted checkpoint. The larger intervention on both, by a wide margin.
- The SFT weights are worth +8.8 points under the stock harness (p=0.0003) β real, and the reason they are still what I submit.
- Under pi-ws the SFT weights are significantly worse than base (0.208 vs 0.276, p=0.036). Not "no difference" β a measured 6.8-point regression. The two interventions partly conflict.
How the number moved with n β the brief's warning, demonstrated. Same comparison: n=100 β +0.000, p=1.00; n=225 β +0.062, p=0.076; n=250 β +0.068, p=0.036. The point estimate barely moved; the interval closed. At no n was there evidence the SFT weights help under pi-ws, which is exactly why the base arms were extended rather than stopped at the first clean number. Had I stopped at n=100 I would have reported "exactly tied" as a finding.
Only visible because the 2Γ2 was finished rather than argued about. It retro-justifies finding 21's worry β
base_swe_ws0.188 at n=64 was pointing straight at this and I explained it away as a stale-harness artifact.It does not change what I submit. The brief requires weights derived from the base by my own training, so the base itself is not eligible, and
sft_v5/step_900is worth +8.8 under the stock harness. It changes what I claim, and it motivates finding 25.There is a hard cap of 32 concurrent host tunnels per API token, and
kill -9leaks them. This masqueraded as a sandbox-pool problem for over an hour and is invisible in/resources.Every eval episode opens a host tunnel so the sandbox can reach the local vLLM server.
PrimeTunnel.expose(v1/interception/tunnel/prime.py) tears its tunnel down in afinallythat is shielded against cancellation β so a normal exit or SIGTERM cleans up. SIGKILL does not. My owntopup_loop.shwas doingkill -9on runs with 24 episodes in flight every 15 minutes; within a few passes all 32 slots were leaked and every subsequent episode died instantly withTunnelError: TunnelLimitReachedError: Maximum number of tunnels (32) reachedat
turns=0. Confirmed by listing: 32/32 held with no eval process running at all.Two rules follow, and both are now enforced in
scripts/topup_loop.sh:- Stop eval runs with SIGTERM, never
kill -9, unless you reclaim tunnels afterwards. - Total
max_concurrentacross every simultaneously running eval process must stay under 32 β it is per token, not per run. Two arms at 24 each is 48 and cannot work, however healthy the pool looks. Two arms at 14 each (28) is fine.
scripts/tunnels.pylists the token's tunnels and--deletereclaims them; only run the delete when no eval is active, since it cannot tell a live tunnel from a leaked one.Worth noting how this presented:
stop=TunnelErrorwithturns=0, arriving in bursts, while the pool showed hundreds of free CPU. Any diagnosis that stops at "the pool is busy" misses it entirely.- Stop eval runs with SIGTERM, never
The scaffold's two halves, separated β and a pre-registered hypothesis that failed.
pi-wsis the working-directory fix plus the appendedcfg/agent_prompt.txt, and the two had never been separated on the SFT weights.cfg/ws_noprompt.tomlβruns/ws_noprompt, same seeded 250, zero ungraded:submitted weights, harness score vs stock vs full pi-ws stock pi 0.132 β pi-ws, workdir only 0.164 +0.032, p=0.26 β0.044, p=0.099 pi-ws, workdir + prompt 0.208 +0.076, p=0.005 β Both halves contribute, neither clears significance alone, together they do. The scaffold is the pair.
The prediction was wrong. Registered in advance: since
public_to_sft.pyput pi's own system prompt on every training row, an appended operating-procedure block should be off-distribution for the SFT model but pure gain for the base β which would explain why the scaffold is worth +23 points on base and only +7.6 on the SFT weights. If so, dropping the prompt should have recovered most of the gap. Instead dropping it costs 4.4 points. Prompt off-distribution is not the explanation for the SFT weights' deficit against base under pi-ws, and I do not have a replacement explanation β only the measurement.And a second demonstration of the partial-read trap, worse than finding 23's: this same comparison read β0.075, p=0.096 at n=107 (looking like a strong refutation in the opposite direction), β0.039, p=0.31 at n=152, and β0.044, p=0.099 at n=250. Three different stories from one run. Only the completed number means anything.
Pre-registered: is the training/scaffold conflict monotone in training length? (Registered 2026-08-16 18:30Z, before the run.)
Finding 23 leaves the submitted checkpoint in an awkward place: 900 steps of SFT buys +8.8 points under the stock harness and loses 6.8 under pi-ws versus the untrained base. Hypothesis with a mechanism: the stock-harness gain is mostly format competence β emit a well-formed tool call, use the four tools, stop when done β which saturates within a couple of hundred steps; the pi-ws regression is stylistic over-specialisation that accumulates over the full run. If so, an early fully annealed checkpoint should keep most of the +8.8 while giving back most of the β6.8, and would dominate
step_900on both harnesses.cfg/sft_short.toml: same corpus, same recipe, cosine annealed over 200 steps (not step 200 of the 900-step run β a complete run, the thing I would actually submit). 2 GPUs, batch 8 β ~52M tokens.Decision rule, fixed in advance:
- Measure
sft_shorton both harnesses, n=250, same seeded sample (cfg/short_{ws,stock}.toml, served on port 8002). - The decisive cell is pi-ws vs
step_900's 0.208 on a paired McNemar. - Scheduling note (changed after registering, and only the order): I originally planned to run pi-ws first and gate the stock arm on it, to save sandbox budget. With ~24 h left and the pool delivering 30β60 episodes/hour, sequential gating risks finishing with only half the pair measured β the failure mode I have been guarding against all run. Both arms therefore run concurrently (14 each = 28 tunnels, under the cap). The rule below is unchanged.
- Submit
sft_shortonly if it wins under pi-ws and its stock number is not significantly worse than 0.132. Any other outcome:step_900ships, unchanged. step_900stays fully measured and is never at risk. Nothing here can leave an unmeasured checkpoint as the submission.
Expected value is honest about itself: the most likely single outcome is that the trade-off is monotone, no dominating checkpoint exists, and this returns a curve (0 / 200 / 900 steps) rather than a better submission. That curve is worth having either way.
Trained 18:22β21:01Z (200 steps, 2 GPUs, ~47 s/step, loss 0.245 β ~0.14;
ckpt/sft_short/weights/step_200, 760 tensors, eos ids correct,step_100kept as a spare).RESULT β the hypothesis is refuted, and the curve is not the shape I guessed. Both arms complete at n=250, zero ungraded, same seeded sample as every other cell:
weights stock pi pi-ws base (0 steps) 0.044 [0.025, 0.077] 0.276 [0.224, 0.334] sft_short (200 steps) 0.104 [0.072, 0.148] 0.168 [0.127, 0.219] sft_v5 (900 steps, submitted) 0.132 [0.096, 0.180] 0.208 [0.162, 0.263] paired comparison (n=250) only-A only-B delta McNemar stock: base β short 7 22 +0.060 0.0081 stock: base β sub900 7 29 +0.088 0.0003 stock: short β sub900 12 19 +0.028 0.28 pi-ws: base β short 41 14 β0.108 0.0004 pi-ws: base β sub900 38 21 β0.068 0.036 pi-ws: short β sub900 18 28 +0.040 0.18 Two clean statements:
- Under the stock harness, SFT helps monotonically with training length: 0.044 β 0.104 β 0.132, each significant against base.
- Under pi-ws, every amount of SFT is significantly worse than the untrained base (200 steps: p=0.0004; 900 steps: p=0.036). And it is not monotone β 200 steps (0.168) is worse than 900 (0.208), not better. My prediction was that a shorter anneal would sit closer to base under the scaffold. It sits further away.
So the conflict is not "too much training". Whatever SFT does that the scaffold does not like, it does early, and more training partially undoes it. I do not have a mechanism for that, and with ~16 h left I am not going to get one honestly β three points on a curve is what this buys.
Decision rule applied:
sft_shortdid not win under pi-ws (0.168 vs 0.208, p=0.18 favouring step_900), sockpt/sft_v5/weights/step_900remains the submission, unchanged.sft_shortis kept on disk and reported; it is not submitted.- Measure
Never run
eval @ configandeval --resumeagainst the same run directory at once. They both owntraces.jsonl, and the second one does not merge β the graded count went backwards (57 β 44) while both were live, because each process wrote the file from its own view of what was done. No corruption, but ~30 completed episodes were silently lost and the surviving file had mixed provenance.My own sequencing error: I launched both arms with
scripts/run_eval.sh(which runseval @ cfg -o dir) and then startedscripts/topup_loop.shon the same dirs, whosestop_runsonly matcheseval --resumeβ so it never stopped the original writers.Rules: one writer per run directory, ever. If you want the top-up loop, either start the run with it from the beginning, or SIGTERM the original process first and confirm it is gone. A
gradedcount that decreases is the signature β check for two writers before anything else. When it happens, restart the affected runs from empty rather than resuming: a trace file of mixed provenance is not something to base a submission decision on.terminal-bench-2, the same 2Γ2, completed β and it measures nothing. My own reporting had been inconsistent: I demanded completed runs for swe-bench while quoting tb2 at n=83β88 of 89. All four cells are complete 89-task runs with zero ungraded:
weights stock pi pi-ws base (+eos fix) 5/89 = 0.056 5/89 = 0.056 submitted (SFT) 3/89 = 0.034 1/89 = 0.011 One task nearly did not make it, and the cause is worth recording:
terminal-bench/qemu-alpine-sshboots a QEMU VM and takes longer than 15 minutes, which is exactly the interval at whichscripts/topup_loop.shstops and re-resumes its runs. Every pass started that task and killed it before it could finish, forever. The loop is right for pool-stalled runs and wrong for genuinely long tasks β when a single task is all that is left, stop the loop and run one plaineval --resumewith nothing killing it. It then completed on the first attempt.Every paired McNemar is non-significant: p from 0.125 (base+pi-ws vs submitted+pi-ws, 4 vs 0) to 1.00. No effect claimed in any direction. With 1β5 solves out of 89 the interval swamps everything.
Two observations, explicitly not results:
- The scaffold that is worth +23 points on swe-bench does nothing here β base scores 5/88
under both harnesses (5/89 each). That fits its content: a working-directory fix and a repo-oriented
operating procedure have nothing to grip on in tasks that are not repo fixes and already
start in
/app. The pi-ws gain is specific to the SWE-repo setting, not general agentic competence. - The directional ordering matches swe-bench (base β₯ submitted everywhere, submitted+pi-ws lowest), but at these counts that is a coincidence I would not defend.
Also fixed here:
cfg/base_tb2.tomlhad the same missingready_timeout_secondsasymmetry caught on swe-bench in finding 20 (broker default 600 against pi-ws's 1200, in my favour). Both tb2 configs now pin 3600. It had no effect on the numbers β zero ungraded either side β but it should not have been there.- The scaffold that is worth +23 points on swe-bench does nothing here β base scores 5/88
under both harnesses (5/89 each). That fits its content: a working-directory fix and a repo-oriented
operating procedure have nothing to grip on in tasks that are not repo fixes and already
start in
Replicate of the headline cell β the intervals hold, and the noise floor is now measured. Every comparison in this file is a paired McNemar, which treats which tasks were sampled as the only source of variation and says nothing about run-to-run variance from sampling temperature and the environment. The brief warns that repeat reads of identical weights have differed by more than ten points, so I re-ran the headline cell: same checkpoint, same harness, same seeded 250 tasks, same config, nothing different but the run.
run n solved score ci95 original ( big_swe_ws)250 52 0.208 [0.162, 0.263] replicate ( replicate_ws)250 49 0.196 [0.152, 0.250] pooled 500 101 0.202 [0.169, 0.239] - The replicate lands well inside the original interval. The intervals quoted throughout are trustworthy as stated; this system does not swing ten points between reads.
- Noise floor, measured: two runs of the identical system disagree on 39 of 250 tasks (15.6%) β 21 one way, 18 the other. That is the yardstick for every McNemar here. The null run is symmetric; the real effects are lopsided (12 vs 31 scaffold, 21 vs 38 base-over-submitted). McNemar tests exactly that asymmetry, which is why it is the right test β but anything near 22 vs 30 is within reach of this floor.
- Both headline conclusions reproduce on the fresh run: scaffold 14 vs 30, p=0.023 (was 12 vs 31, p=0.005); base-beats-submitted under pi-ws 16 vs 36, p=0.008 (was 21 vs 38, p=0.036) β the second one is stronger on the replicate.
This is the measurement I would have wanted first if I had known how much the report would end up resting on paired tests. It costs one run and it tells the reader what a difference has to look like before it means anything.
Pre-registered, final experiment: weight-space interpolation between base and my SFT. (Registered 2026-08-17 07:15Z, before building or measuring anything.)
The standing problem from finding 23: my SFT weights are 6.8 points worse than the untrained base under my own scaffold (0.208 vs 0.276, p=0.036; p=0.008 on the replicate), while being 8.8 points better under the stock harness. That is the classic shape of fine-tuning trading away out-of-distribution robustness.
Finding 26 already tested the obvious axis β less training (a 200-step anneal) β and it was worse under pi-ws, not better (0.168). So "closer to base along the training trajectory" does not help. Weight-space interpolation is a different axis: a convex combination of the base and my fine-tune, which is the standard remedy for exactly this fine-tuning/robustness trade (WiSE-FT). It is cheap β
scripts/soup.pyaverages two checkpoints in ~20 minutes β and it has not been tried.ckpt/soup_base_sft/= 0.5 x base + 0.5 x sft_v5/step_900, produced byscripts/soup.py.Rules check, stated explicitly because it deserves scrutiny: every parameter still traces to
Qwen/Qwen3.5-9B-Basethrough training I ran β it is a blend of the designated base with my own SFT checkpoint and nothing else. The prohibition is on initializing from, continuing training on, merging in, or submitting another party's post-trained checkpoint; the base itself is the designated starting point, not another party's post-training. No third-party weights are involved at any point.Decision rule, fixed in advance:
- Measure the soup on both harnesses, n=250, same seeded sample, both to completion.
- Submit it only if it beats
step_900under pi-ws on a paired McNemar at p<0.05 and is not significantly worse thanstep_900under the stock harness. - Hard stop 15:30Z. If both cells are not complete by then,
step_900ships. A partial run is not evidence and will not be used for this decision β three times this run a partial read told a different story than the finished one. step_900remains fully measured throughout and is never at risk.
Honest prior: maybe one chance in three. The 200-step result shows the pi-ws deficit is not a simple function of distance-from-base, so interpolation may not touch it either. Reported either way.
RESULT (both arms complete, n=250 each, zero ungraded). Interpolation works on the axis that training length did not β and it buys the pi-ws points by giving back stock points:
weights stock pi pi-ws base 0.044 [0.025, 0.077] 0.276 [0.224, 0.334] soup (0.5/0.5) 0.084 [0.056, 0.125] 0.280 [0.228, 0.339] submitted step_900 0.132 [0.096, 0.180] 0.208 / 0.196 (two runs) paired vs step_900 (n=250) only-A only-B delta McNemar pi-ws: step_900 β soup 16 34 +0.072 0.0153 pi-ws: replicate β soup 17 38 +0.084 0.0065 stock: step_900 β soup 23 11 β0.048 0.0576 - Condition (a) is met clearly: the soup beats
step_900under pi-ws, and beats the replicate more strongly still, so this is not one lucky run. - Condition (b) passes only on a technicality: p=0.0576 is "not significant at 0.05", but a 4.8-point drop (0.132 β 0.084, a 36% relative fall) on 23-vs-11 discordant pairs is almost certainly a real cost. Reading that as "no effect" would be exactly the rule-lawyering the noise-floor work in finding 29 was meant to prevent.
So it is a genuine trade, not a dominating win: roughly +7 to +8 points on the harness I submit, roughly β5 on the stock control. And the mechanism is not subtle β the soup's pi-ws score (0.280) is statistically indistinguishable from the untrained base's (0.276), so its advantage there comes from being closer to base, i.e. from partially undoing my own training.
Cannot decide on swe-bench alone: the submission publishes terminal-bench-2 as well, and the soup has no tb2 numbers. Both tb2 arms launched 11:30Z; decision at 15:30Z at the latest, on completed runs only.
Replicate of the SUBMITTED system's headline cell (running from 2026-08-17 12:55Z). The submission was switched to the soup on the strength of one measurement of soup+pi-ws (0.280), while the checkpoint it displaced has two (0.208, 0.196). Finding 29 established that two runs of an identical system disagree on 15.6% of tasks. The submitted system's headline number should therefore carry the same replicated status as the number it replaced β otherwise I have held the alternative to a higher evidentiary standard than my own choice, which is precisely backwards.
cfg/soup_ws_rep.tomlβruns/soup_ws_rep: same checkpoint, same harness, same seeded 250, nothing different but the run.Pre-committed reading. The switch rested on soup 0.280 vs step_900 0.202 pooled, paired p=0.015 (p=0.0065 vs the replicate).
- If this lands near 0.28 β or anywhere that keeps the paired comparison against
step_900significant β the switch stands and the headline gets a pooled figure. - If it lands near 0.21, the single 0.280 was a lucky run, the switch is not supported, and
I revert the submission to
ckpt/sft_v5/weights/step_900, which is intact and fully measured. - Hard stop 17:00Z. If the run is not complete by then it is not used at all β a partial
read has told a different story than the finished one three times in this run β and the
submission stays as it is, with
SUBMISSION.mdstating the headline rests on a single measurement.
RESULT (complete, n=250, zero ungraded): 0.228, not 0.280. The first run flattered the soup.
system run 1 run 2 pooled (n=500) soup + pi-ws 0.280 0.228 0.254 [0.218, 0.294] step_900 + pi-ws 0.208 0.196 0.202 [0.169, 0.239] - Run 2 alone does not beat
step_900(27 vs 22, p=0.57; against the step_900 replicate, 32 vs 24, p=0.35). The p=0.015 that triggered the switch was the better of two draws. - What survives is the pooled comparison across all four runs (1,000 episodes): per-task sign test, soup better on 54 tasks, step_900 better on 30, 166 ties, p=0.0116. So the gain is real but is +5 points, not +8.
- The soup is the noisier system: its two runs differ by 5.2 points (32 vs 19 discordant, p=0.09) against step_900's 1.2 (21 vs 18, p=0.75).
- Pooled, the soup no longer exceeds the base under pi-ws (0.254 vs 0.276, one run) β it matches it, which is the same story as every other checkpoint on this harness.
Decision: the switch stands, but only just, and the headline is restated as 0.254. The pre-registered revert trigger was "lands near 0.21, i.e. the 0.280 was a lucky run and the soup is not actually better". The pooled data says the soup is better than
step_900under pi-ws (p=0.0116), so the substance of that trigger is not met β but 0.280 was clearly a high draw and quoting it would misrepresent the system.SUBMISSION.mdnow leads with the pooled 0.254, says the replicate alone fails significance, and keepsstep_900named as the alternative for anyone weighting the stock-harness number higher. The trade is now roughly +5 pi-ws / β5 stock: close to even, with more evidence behind the gain (four runs) than the cost (two).- If this lands near 0.28 β or anywhere that keeps the paired comparison against
Final measurement: the stock cost is confirmed, and the two checkpoints end in a dead heat. The submission decision had become a trade β about +5 points under pi-ws against about β5 under the stock harness β with four runs behind the pi-ws side and only one per system behind the stock side. That asymmetry was the weakest link, so the last hours went to a replicate of the soup's stock cell (
cfg/soup_stock_rep.toml, n=100, a strict prefix of the 250).Complete, 100/100, zero ungraded: 0.070, against the first run's 0.120 on those same 100 tasks (10 vs 5 discordant, p=0.30 β consistent, not a contradiction). Pooled over all 350 stock episodes the soup scores 0.080 [0.056, 0.113], essentially the 0.084 first measured. The cost as weights alone is real, not a low draw.
stock pi-ws sum soup (submitted) 0.080 0.254 0.334 step_900 0.132 0.202 0.334 Interpolating along the baseβSFT line trades the two published numbers off almost exactly one-for-one. There is no free lunch on that axis. Which checkpoint is "better" is therefore a question about which number is primary, not something the data settles β and
SUBMISSION.mdsays that plainly rather than implying the soup dominates.Decision: no revert. The pre-committed trigger was "pooled stock cost clearly larger than the pi-ws gain"; it is equal, not larger. The soup stays because the brief asks to make the system score higher and the pi-ws number is the system, and because that side rests on four runs and 1,000 episodes (p=0.0116) against three runs and 600 on the stock side (p=0.058).
step_900is named inSUBMISSION.mdfor anyone weighting stock higher.Closing the last evidence asymmetry (running from 2026-08-17 16:45Z). After finding 32 the submitted soup's stock cell had 350 episodes (250 + a 100-task replicate) while the alternative it displaced,
step_900, had 250. The extra scrutiny went to my own preferred checkpoint on the axis where it looks worst β the right direction β but a published comparison should not rest on unequal evidence in either direction.cfg/s900_stock_rep.tomlβruns/s900_stock_rep:step_900's stock cell, same seeded 100-task prefix, served on 8000 (GPUs 0/1; the soup stays up on 8002).Decision-relevant, and pre-committed before looking: the current trade is +5.2 points pi-ws against β5.2 stock, a dead heat on the sum (0.334 each).
- If
step_900's stock number comes in higher than the 0.132 measured once, the stock cost becomes clearly larger than the pi-ws gain, which is exactly the revert trigger from finding 30 β submitstep_900. - If it comes in at or below 0.132, the dead heat stands and the soup stays.
- Hard stop 18:00Z; incomplete β not used, nothing changes.
RESULT β the trigger fires; submission reverts to
step_900. Complete, 100/100, zero ungraded:step_900scored 0.200 on the replicate prefix against 0.150 on those same 100 tasks the first time (5 vs 10 discordant, p=0.30 β consistent). Pooled over 350 stock episodesstep_900scores 0.151 [0.118, 0.193], not the single-run 0.132.stock (pooled n=350) pi-ws (pooled n=500) sum step_900 (submitted) 0.151 0.202 0.353 soup 0.080 0.254 0.334 The soup buys +5.2 under pi-ws and costs β7.1 under stock. The pre-committed trigger ("pooled stock cost clearly larger than the pi-ws gain") is met, so the submission reverts to
ckpt/sft_v5/weights/step_900.ckpt/soup_base_sftstays in the workspace, fully measured, and is still ahead if the my-harness number is the only one that counts.The lesson, and it is the same one twice: both replicates moved the answer against whichever checkpoint I preferred at the time β 0.280 β 0.254 for the soup when I had just switched to it, 0.132 β 0.151 for
step_900when I had just switched away. A single run flattered whichever system I had most recently chosen. Nothing about that is mysterious given finding 29's 15.6% task-level noise floor; what saved the conclusion was replicating both sides rather than only the one whose result I liked.- If
Final audit: every published number recomputed from disk, and one magnitude corrected. With the deadline close and both documents heavily edited, I recomputed every figure in
SUBMISSION.mddirectly fromruns/*/traces.jsonl. All matched exactly (step_900 53/350 and 101/500; soup 28/350 and 127/500; base 11/250 and 69/250; sft_short 26/250 and 42/250; all six tb2 cells).One thing the audit did change. The scaffold gain on the submitted weights had been quoted as "+7.6 points", which was correct for the paired 250-task runs but is not the only estimate now that both arms have replicates:
comparison on step_900stock pi-ws gain test paired 250-task runs (one each) 0.132 0.208 +7.6 McNemar 12 vs 31, p=0.005 the same, replicate pair 0.132 0.196 +6.4 14 vs 30, p=0.023 100 tasks where both arms have two runs 0.175 0.270 +9.5 sign test 23 vs 11, p=0.058 pooled over all episodes (350 vs 500) 0.151 0.202 +5.1 β Every one is positive, so the direction is not in doubt; the honest range is +5 to +9 points, and
SUBMISSION.mdnow prints the whole table instead of the single best figure. (On base weights the effect is +23 and on the soup +19.6, both p<0.0001 β no such ambiguity.)Worth recording as a general point: once you replicate arms unevenly, "the" effect size stops being one number. Quoting the most favourable pooling would have been indefensible after spending the whole run insisting on completed runs and measured noise floors.
The submission's own reproduction commands did not reproduce its numbers. Caught in the last hour by actually reading the command block as a stranger would.
SUBMISSION.mdnamedcfg/ws_swe.tomlandcfg/base_swe.tomlas the swe-bench arms β those arenum_tasks = 100, left over from the early 100-task reads β while every published swe-bench figure comes from the 250-task runs (cfg/big_swe_*.toml) and their replicates. Anyone following the instructions would have got a different sample and quietly different numbers, and concluded the submission was overstated.Fixed: the block now names the 250-task configs, lists the replicate configs that the pooled figures add in, annotates each line with the count it should produce (52/250, 33/250, 1/89, 3/89, 49/250, 20/100), and says explicitly that
ws_swe/base_sweare the 100-task variants that do not reproduce the headline. Every config and script named in the file was then checked to exist.Worth generalising: a reproduction section is code, and it was the only part of this submission never executed as written. Read it against the numbers it claims to produce before shipping.
Then actually ran it.
cfg/smoke_repro.tomliscfg/big_swe_ws.tomlwithnum_tasks = 3and nothing else changed, executed against the submitted checkpoint served on:8000exactly as the block documents. Result: 3/3 graded, zero ungraded,errors: [], all threeagent_completed, mean 21.7 turns / 39 s. So the whole chain works end to end βPYTHONPATHβpkgs/pi_wsresolves β harness launches β broker runtime β grading returns.Two details worth having on record, because they are the harness's central claims and this is the first time they were confirmed on the submitted artifact rather than during development:
- The agent really is in the task workdir. All three episodes work in
/testbedfrom their first command (find /testbed -type f -name "*.py" ..., and observations returning/testbed/django/core/handlers/asgi.py). - The orientation training shows up in behaviour. One episode's very first tool call is
pwd; ls -a; echo ---; ls -d /testbed /app /workspace /repo /code /srv 2>/dev/nullβ the exact patternscripts/add_orientation.pyput on 40% of training rows.
0/3 solved is the expected outcome at a ~20% rate (0.8Β³ β 51% chance of zero) and is not a signal either way; the point of the run was that it executes, not what it scores.
- The agent really is in the task workdir. All three episodes work in
Measurement protocol (fixed β do not change mid-run)
The exact McNemar in scripts/final_2x2.py was checked against scipy.stats.binomtest on every
discordant pair reported in this file (12/31, 5/18, 11/2, 13/13, 2/24, 0/17, 3/17): identical to
1e-9. Worth doing once β every significance claim here rests on that one function.
cfg/base_tb2.toml (all 89 tasks) and cfg/base_swe.toml (100 tasks, shuffle=true fixed seed),
1 rollout/task, stock pi harness, broker runtime block_network=false, no sampling overrides.
Score with scripts/score.py <outdir> (solve rate + Wilson 95% CI).
Results
Submitted checkpoint: ckpt/sft_v5/weights/step_900.
Headline β 250-task swe-bench-verified sample (the number to quote)
| ckpt | harness | n | solved | score | ci95 |
|---|---|---|---|---|---|
| base (+eos fix) | stock pi | 153 | 7 | 0.046 | [0.022, 0.091] |
| SFT v5 step_900 | stock pi | 250 | 33 | 0.132 | [0.096, 0.180] |
| SFT v5 step_900 | pi-ws | 250 | 52 | 0.208 | [0.162, 0.263] |
Paired tests on the shared tasks (the powered comparison β the subset is seeded, so runs line up task-for-task):
| comparison | only-A | only-B | McNemar |
|---|---|---|---|
| base β SFT, stock harness | 3 | 17 | p=0.003 |
| base β SFT, pi-ws | 3 | 23 | p<0.001 |
| stock β pi-ws, same weights | 12 | 31 | p=0.005 |
So on swe-bench-verified: the weights are worth +8.6 points (0.046 β 0.132) and the scaffold a further +7.6 (0.132 β 0.208), each independently significant.
Fixed 100-task protocol (used for every checkpoint comparison in this run)
| ckpt | suite | harness | n | solved | score | ci95 |
|---|---|---|---|---|---|---|
| base (+eos fix) | swe-bench-verified | stock pi | 97 | 4 | 0.041 | [0.016, 0.101] |
| SFT v5 step_900 | swe-bench-verified | stock pi | 88 | 11 | 0.125 | [0.071, 0.210] |
| SFT v5 step_900 | swe-bench-verified | pi-ws | 98 | 31 | 0.316 | [0.233, 0.414] |
| base (+eos fix) | terminal-bench-2 | stock pi | 83 | 2 | 0.024 | [0.007, 0.084] |
| SFT v5 step_900 | terminal-bench-2 | stock pi | 88 | 3 | 0.034 | [0.012, 0.096] |
| SFT v5 step_900 | terminal-bench-2 | pi-ws | 83 | 2 | 0.024 | [0.007, 0.084] |
- Weights: swe-bench 0.041 β 0.125 on the stock harness. Paired, the SFT model solves 9 the base does not and loses 1 (McNemar p=0.021). The stock number reproduces almost exactly on the independent 250-task sample (0.125 β 0.123), so it is solid.
- Harness: +7.6 points on the complete 250-task sample, on identical weights.
- terminal-bench-2 does not move under any combination β 2β3 solves out of ~85 throughout. It is out of reach for a 9B here, and no harness variant separates from another (Fisher p=0.62 for the widest gap).
β The 100-task pi-ws figure (0.316) is the optimistic tail of the run-to-run spread. The same harness and weights measured 0.208 on 250 tasks, and paired on their 95 shared tasks the two runs agree (McNemar p=0.48) β the gap is which tasks landed in the first 100, not a real difference. This is exactly the failure the brief warns about; 0.208 is the number I stand behind.
Candidates that did not pan out (all measured, all kept in runs/)
| tried | result |
|---|---|
900 more SFT steps on the unseen 65% of the corpus (sft_cont2) |
7/64 swe stock, 14/58 pi-ws β no gain; supervised signal saturated |
checkpoint averaging of the three (soup1) |
7/63, 13/58 β no gain |
temperature 0.2 on pi-ws |
21/81 = .259 vs .253, but ~2Γ episode length β rejected |
pi-ws per-turn caps (maxTokens 4096, contextWindow 60000) |
actively harmful now: 23/91 vs 31/98 uncapped (paired 5 vs 11) β removed |
| Dockerfile-derived workdir for terminal-bench | 0/79 and 1/86 vs 3/88 stock β left off by default |
a sharper prompt that insists on verifying after the last edit (cfg/agent_prompt_v2.txt) |
53/227 = .233 vs .208 for the shipped prompt, paired 27 vs 19, McNemar p=0.30 β not separable, so the shipped prompt (the one every headline number was measured with) stays |
| GRPO through pi | blocked by two independent things, both now diagnosed exactly. (1) pi hard-codes stream: true and prime-rl's TrainClient β the only client returning the token ids/logprobs the RL loss needs β raises on streaming. Solved: pkgs/pi_rl starts a ~90-line Node shim in the task container, points pi's baseUrl at it, and it forwards non-streaming then re-emits SSE. Verified transparent under the eval client (6/6 episodes, same turn count and token usage as un-shimmed). (2) The RL rollout path renders a different system prompt than the serving path. The renderer serialises whatever tool objects it is handed; verifiers hands it its own flat ToolSpec, so the tool block comes out as {"name":β¦,"description":β¦,"parameters":β¦}, whereas the served chat template emits {"type":"function","function":{β¦}}. Verified byte-for-byte on a real pi conversation: renderer(flat ToolSpec) == served template β False; renderer(OAI-nested) == served template β True (68 chars of difference in the system prompt, on every single turn). So the policy is off-distribution for the entire rollout: 0/85 solved on tasks the identical harness solves ~16% of the time under the eval client, and after temperature 1.0 β 0.7, 57 of 76 episodes ran to the 40-turn cap instead of finishing. The fix is a one-line change in how verifiers hands tools to the renderer β shared-install code, so out of scope here. |
GRPO through the verifiers-native bash harness (which does not stream, so the stack runs) |
the harness runs and batches fill, but every reward is 0: the policy scores 0/24 on r2e-gym-ws under bash even with pi's own system prompt injected, against ~16% under pi. An SFT'd agent is tightly coupled to its harness's tool schema and loop, so there is no signal to learn from. |
| expert iteration on r2e-gym-ws | 203 successes but only 32 distinct tasks β too narrow to train on |
What the failures actually look like (250-task stock run)
Among solved episodes the model runs a check after its last edit 36% of the time; among
failed ones, 17% (any check at all: 58% vs 45%). Verifying is the behaviour most associated
with succeeding and the policy does it unreliably β but 72% of the training corpus already
verifies, so this is a generalisation gap, not a coverage gap, which is consistent with more
SFT buying nothing. It is the kind of gap RL closes, and RL is the thing this stack would not
run. 35 of 217 failed episodes never made an edit at all; ~10% of tool calls in failed episodes
name a tool that does not exist (submit, submission, output).
| ckpt | suite | harness | n | solved | score | ci95 |
|---|---|---|---|---|---|---|
| base (+eos fix) | swe-bench-verified | stock pi | 97 | 4 | 0.041 | [0.016, 0.101] |
| base (+eos fix) | swe-bench-verified | pi-ws | 64 | 12 | 0.188 | [0.111, 0.300] |
| base (+eos fix) | terminal-bench-2 | stock pi | 83 | 2 | 0.024 | [0.007, 0.084] |
| SFT v5 step_900 | swe-bench-verified | stock pi | 88 | 11 | 0.125 | [0.071, 0.210] |
| SFT v5 step_900 | swe-bench-verified | pi-ws | 91 | 23 | 0.253 | [0.175, 0.351] |
| SFT v5 step_900 | terminal-bench-2 | stock pi | 88 | 3 | 0.034 | [0.012, 0.096] |
| SFT v5 step_900 | terminal-bench-2 | pi-ws | 85 | 1 | 0.012 | [0.002, 0.064] |
Paired (McNemar, same seeded task subset):
- swe stock, base β SFT: solved 9 tasks base did not, lost 1. p=0.021 β the weights alone roughly triple the stock number (0.041 β 0.125).
- swe pi-ws, base β SFT: +7 / β3, p=0.34 β not separable at this n; the scaffold had already captured much of what SFT teaches (find the repo, use absolute paths).
- terminal-bench-2 moves ~nothing either way. It is simply out of reach for a 9B here: the
base solves 2/83, the SFT 3/88. pi-ws is 1/85 β no evidence it helps on this suite, and the
starting-in-
/appchange removes the orientation step the model was trained to perform.
β Sandbox losses matter for these numbers: episodes that die as SandboxError drop out of n.
Use eval --resume <run_dir> (scripts/resume_eval.sh, or scripts/topup.sh which retries in
free windows) to re-run them β that is how the swe runs got from nβ74 to nβ90. Resume needs the
vLLM server still up; measure.sh kills it at the end, so restart it first or every resumed
rollout dies with ProviderError.
β Never read a partial run as a score. Episodes finish in roughly ascending order of
difficulty, so a run at 30% completion reads far too high β big_swe_stock showed 8/33 = 0.24
early against its true β0.17, and sft5_swe_ws_t02 showed 3/8 = 0.38 against β0.26. Only
compare completed runs.
The scaffold alone is worth ~15 points on swe-bench-verified (Fisher p=0.005; paired on the 63 shared tasks, pi-ws solved 11 that stock did not and stock solved 0 that pi-ws did not, McNemar p=0.001), on frozen base weights.
β Attribution correction. When those numbers were taken, pi-ws's cd was a no-op, so the
gain came from the appended cfg/agent_prompt.txt (which names /testbed) plus the
maxTokens/contextWindow caps β not from the working directory. The reason: pi resolves every
tool call against the ACP session cwd, which verifiers/v1/acp/_runner.py sets from
os.getcwd() of the runner process; cd-ing the agent command underneath it changes nothing.
pi_ws.run_acp_in now starts the runner itself inside the workdir (re-anchoring the config path
and exporting VF_LAUNCH_BASE so PI_ACP_PI_COMMAND / PI_CODING_AGENT_DIR / skills still
resolve against the launch dir). Verified live: pwd returns /testbed on swe-bench and /app
on terminal-bench, with the task's files right there. The measured 0.188 therefore understates
the current harness.
Gold-patch validation of the RL taskset (validate --only-gold, 6 tasks, broker runtime):
| taskset | valid |
|---|---|
r2e-gym-v1 (stock) |
0/6 β gold apply failed: error: Orange/data/util.py: No such file or directory |
r2e-gym-ws (my wrapper) |
6/6 valid, ~80 s each |
swelego-v1 (stock) |
6/6 valid, ~25 s each β works as shipped, and grades ~3Γ faster |
The stock r2e taskset cannot score a correct patch under this runtime; the wrapper is what makes RL on it possible at all. RL trains on both sources.
β HARD-LEARNED RULE
/var/lib/agentptb-cache/a/ DOES NOT SURVIVE A NODE RESTART. At ~04:46 on 2026-08-14 the
node was recreated and everything there except prime-rl/ and tmp/ was erased: the local base
model copy, the 72 GB teacher, the raw trajectory corpora, and all SFT v1 checkpoints (the run
had reached step 600/940). ~9 h of GPU work lost. The workspace on the PVC was untouched.
β Model checkpoints go to $AGENTPTB_WORKSPACE/ckpt/.... Local disk is for the read-only model
copy (fast weight loading) and TMPDIR only β anything reproducible in <15 min.
Where things are (update as you go)
| artifact | path |
|---|---|
| base + generation_config fix | /var/lib/agentptb-cache/a/models/base |
| teacher (data-gen only, never weights) | /var/lib/agentptb-cache/a/models/teacher35 |
| SFT corpus v1 (15,112 traj, 244M tok) | data/sft_v1 |
| SFT corpus v2 (30,583 traj, real trajectory endings) | data/sft_v2 |
| SFT corpus v3 (46,286 traj, ~750M tok) | data/sft_v3 |
| SFT corpus v4 (v3 + 30% path-relocated) β the one in use | data/sft_v4 |
| SFT v2 run + checkpoints | ckpt/sft_v2/weights/step_N |
| local plugins (PYTHONPATH) | pkgs/{r2e_gym_ws,pi_ws} |
| measure driver | scripts/measure.sh β args: <model_dir> <tag>, then suite (both/swe/tb2), arm (stock/ws/both), concurrency |
| status at a glance (run this first on resume) | bash scripts/status.sh |
| scoring | scripts/score.py <run_dir> |
| raw-trajectory corpora | /var/lib/agentptb-cache/a/data/{klear66k,swesmith_traj,r2e_sft,swegym_oh,nebius,deepswe_k2} |
Checkpoints are HF-servable as saved (config + chat_template.jinja + tokenizer + the correct
generation_config.json with eos_token_id:[248044,248046]) β serve the weights/step_N dir directly.
[h20] Orientation training is taking. At an equal 180 steps, under the stock harness:
corpus mentions /testbedemits an early pwd/ls -dprobesolved v4 (no orientation turns) 3/8 0/8 0/8 v5 (40% orientation turns) 8/16 5/16 1/15 Still only 20% through training.
scripts/compare.py A Bdoes Wilson + Fisher for run pairs; the shuffled 100-task swe subset is seeded, so runs are paired on identical tasks.
Status: COMPLETE (2026-08-17 17:40Z). Submission = ckpt/sft_v5/weights/step_900.
Submission: ckpt/sft_v5/weights/step_900 + pkgs/pi_ws, evaluated with
cfg/big_swe_ws.toml / cfg/ws_tb2.toml (mine) against cfg/big_swe_stock.toml /
cfg/base_tb2.toml (stock).
Verified: 15 files, 4 shards, 760 tensors, 18.8 GB, eos_token_id=[248044,248046], cold-serve
tested and answering correctly. Every published figure was recomputed from runs/*/traces.jsonl
(finding 34) and every config/script named in the reproduction block was checked to exist and to
produce the number claimed for it (finding 35).
Machine state: no eval, training or RL jobs are running and no tunnels are held. One vLLM
server is deliberately left up β the submitted checkpoint on localhost:8000 (GPUs 0,1), so a
verifier can hit it immediately; GPUs 2,3 are free. 19 measurement runs are on disk and every one
of them is complete with zero ungraded episodes.
Final numbers β every swe-bench cell replicated:
| stock pi | pi-ws (submitted system) | |
|---|---|---|
| swe-bench-verified | 0.151 [0.118, 0.193] (n=350) | 0.202 [0.169, 0.239] (n=500) |
| terminal-bench-2 | 3/89 = 0.034 | 1/89 = 0.011 |
The interpolated checkpoint ckpt/soup_base_sft was submitted for ~6 h and withdrawn on its own
pre-registered rule (findings 30β33): pooled, it buys +5.2 points under pi-ws and costs β7.1 under
the stock harness, so step_900 leads on the sum, 0.353 to 0.334. It stays in the workspace fully
measured, and is still ahead on the my-harness number alone.
What the run established (all replicated or at n=250+):
- The scaffold is the result β +23 points on base weights (p<0.0001), +7.6 on the submitted weights (p=0.005, reproduced p=0.023), +19.6 on the soup. Both halves of it contribute; neither is significant alone (finding 25).
- SFT helps monotonically under the stock harness (0.044 β 0.104 β 0.132/0.151) and hurts under pi-ws at every training length, non-monotonically (findings 23, 26).
- Weight-space interpolation trades the two published numbers roughly one-for-one β no free lunch on that axis (findings 30β33).
- terminal-bench-2 measures nothing (all pβ₯0.125); the scaffold is worth zero there, so its gain is specific to the SWE-repo setting (finding 28).
- Noise floor: 15.6% of tasks flip between two runs of an identical system, symmetric (finding 29). Every headline cell is replicated because of it β and both replicates moved the answer against whichever checkpoint I preferred at the time.
Six pre-registered hypotheses: four refuted (prompt off-distribution 25, shorter anneal 26, RL through the real harness 18, tb2 corroboration 28), one confirmed then withdrawn on its own rule (interpolation 30β33), plus the replicate protocol (29) which changed how everything is quoted.
The uncomfortable summary, kept at the top of SUBMISSION.md: on the harness I submit, my
training is a net liability; the scaffold, not the training, is what carries this result. I could
only see that by finishing the 2Γ2 instead of arguing about it.
Next actions (if resumed)
The submission is ckpt/sft_v5/weights/step_900 + pkgs/pi_ws, complete, measured on both suites,
with every swe-bench cell replicated and every published number recomputed from disk (finding 34).
Nothing is unfinished. If someone picks this up, the things actually worth doing, in order:
Explain the SFT/scaffold conflict. This is the central unexplained result and everything else in the run circles it: SFT helps monotonically under the stock harness (0.044 β 0.104 β 0.151) and hurts under pi-ws at every training length, non-monotonically β 200 steps (0.168) is worse than 900 (0.202), and both are below the untrained base (0.276) (findings 23, 26). Three mechanisms were tested and refuted: prompt off-distribution (25), shorter anneal (26), and weight-space interpolation, which trades the two published numbers roughly one-for-one rather than fixing anything (30β33). The place to start is a behavioural diff, not another sweep: the traces are all in
runs/, andbase250_wssolves 38 tasks thatbig_swe_wsdoes not. Compare turn counts, tool mix, and where the two diverge on exactly those tasks.Sweep the interpolation coefficient, if the my-harness number is what matters. Only Ξ±=0.5 was measured, chosen a priori to avoid selection noise. It lands at pi-ws 0.254 / stock 0.080 against step_900's 0.202 / 0.151 β i.e. it moves along a roughly one-for-one trade line (finding 32). Ξ± β {0.25, 0.75} would show whether that line is really straight, but budget for the noise floor: Β±5 points needs 250 tasks and a replicate, or you will measure nothing (findings 29, 31, 33).
RL, if there are days rather than hours. Both known blockers are cleared and verified in
pkgs/pi_rl+pkgs/render_parity.py(findings 14, 18), but episode prompts reach p90 42k / max 101k tokens againstseq_len=32768, so a third of every batch is untrainable. Needs 65536,max_turnsunlimited to match eval, the system prompt on the train sources, and a sandbox pool that can sustain ~100 episodes/hour.
Do not bother with: more harness knobs (finding 19 β everything left is 1β3 points against a 15.6% task-level noise floor); more of the same SFT (saturated, finding 26); or reading any partial run as a result. That last one misled me four separate times (findings 23, 25, 31, 33), twice in the direction of the checkpoint I happened to prefer at the moment.
If you change the submission, replicate both sides before you do. Every time I measured one arm and switched, the next replicate moved the answer back: 0.280 β 0.254 for the interpolation right after I adopted it, 0.132 β 0.151 for step_900 right after I abandoned it.