Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
bytkim 
posted an update 2 days ago
Post
153
Introducing Qwen3.8-27B-pi! Fine-tuned for the Pi coding harness to turn plans into working, checked implementations.

Trained on curated Pi coding sessions, then refined through reinforcement learning to balance task success with more economical reasoning at low and medium effort, while keeping xhigh focused on correctness.

79.78% on Terminal-Bench 2.1, versus Base model's 75.28%, with approximately 12% fewer output tokens.

Available in BF16, FP8, and GGUF, with serving guides and MTP/DFlash2 companions!

https://huggingface.co/collections/bytkim/qwen38-pi
bytkim/Qwen3.8-27B-pi-FP8
bytkim/Qwen3.8-27B-pi
bytkim/Qwen3.8-27B-pi-GGUF

The token cut is the result I'd lean on. The score gap is thinner than the percentages make it look.

Terminal-Bench 2.1 is 89 tasks, so your table in task counts:

low: 64 vs 63
medium: 67 vs 62
xhigh: 71 vs 67

That's +1, +5 and +4 tasks. And the footnote says best-observed across at most 3 attempts, so a task counts as solved if any attempt passed. That's closer to pass@3 than a pass rate.

Two things decide what the +4 means.

Did both arms get the same number of attempts on every task? "At most 3" reads like either an early stop on a pass, or uneven budgets.

And are the turns and token means over all attempts, or only the kept one? Best-of-3 plus kept-attempt tokens would flatter whichever arm needed more retries.

Pi medium tying Base xhigh at 67 of 89 with 41% fewer tokens is the claim I'd most want to see hold. Do you have first-attempt pass rates for both arms?

·

Hi dipankarsarkar,

Thank you for taking the time to give my work a proper look.

I agree, the token cut should be the primary focus. I led with Terminal-Bench scores because it's been the "headline" result since the first 3.6 pi-tune release. My main goal was to target the effort to success imbalance observed in initial evals of the base model.

Both arms did not get the same number of attempts. Evaluations were performed on cloud infrastructure which led to task errors related to sandbox/compute infrastructure as well as runtime/task enviroment failures. For these tasks that errored and I considered to be beyond the agents control I would queue a recovery attempt. 2 was the largest number of recovery attempts so it became the upper bound for number of attempts. Valid failures were not retried and tasks were scored on first completion.

Turn/Token means are over only kept attempts. Yes I agree, I realized TB2.1 could no longer be a clean evaluation set after its use in the SFT evaluations. To compensate I introduced GPQA and SciCode for post-rl evaluation. These sets had the benefit of generally being reliable and faster and allowed for additional points of comparison.

Ideally I would like to rerun evaluations consistently under a controlled environment, but as a full-time student I am constrained in compute budget as well as time and felt compelled to release the model before the semester gets busy. This concern also extends to the GGUF evaluations but results were reported under a strict pass@1 score.

These are my main priorities after release and I'll update the results when they're done.

Thanks again for the feedback!

That answer makes your number stronger than the footnote does.

If valid failures were never retried, the score is first valid attempt, not best-of-3. "Best-observed across at most 3 attempts" undersells it. Something like "first valid attempt; infra-errored runs re-queued, max 2 recoveries" says what you actually did.

It also settles the token question. If the kept attempt is the only valid one, kept-attempt means are the honest means.

What's left is the infra call itself. A sandbox OOM or a timeout can be the agent's doing, say a runaway build or a loop that never exits.

So the one number I'd add is recoveries per arm. If Base and Pi needed about the same count, the +4 is clean. If one arm needed noticeably more, the labelling is doing some of the work.

Releasing before the semester and tightening evals after is a fair order.

Was the infra-or-valid call made before you knew which arm the run came from?