Hi dipankarsarkar,
Thank you for taking the time to give my work a proper look.
I agree, the token cut should be the primary focus. I led with Terminal-Bench scores because it's been the "headline" result since the first 3.6 pi-tune release. My main goal was to target the effort to success imbalance observed in initial evals of the base model.
Both arms did not get the same number of attempts. Evaluations were performed on cloud infrastructure which led to task errors related to sandbox/compute infrastructure as well as runtime/task enviroment failures. For these tasks that errored and I considered to be beyond the agents control I would queue a recovery attempt. 2 was the largest number of recovery attempts so it became the upper bound for number of attempts. Valid failures were not retried and tasks were scored on first completion.
Turn/Token means are over only kept attempts. Yes I agree, I realized TB2.1 could no longer be a clean evaluation set after its use in the SFT evaluations. To compensate I introduced GPQA and SciCode for post-rl evaluation. These sets had the benefit of generally being reliable and faster and allowed for additional points of comparison.
Ideally I would like to rerun evaluations consistently under a controlled environment, but as a full-time student I am constrained in compute budget as well as time and felt compelled to release the model before the semester gets busy. This concern also extends to the GGUF evaluations but results were reported under a strict pass@1 score.
These are my main priorities after release and I'll update the results when they're done.
Thanks again for the feedback!