Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
BoldingBuilds 
posted an update about 9 hours ago
Post
61
I released a one-row weight edit for Qwen3.8 27B and Flash-Next to address a frustrating failure: using the entire output budget thinking, then returning no final answer.

The edit changes only the output-layer row that scores </think> and is packaged into ordinary GGUF files.

On 200 MATH-500 problems at a 4,096-token output limit:

• 27B: 31 empty answers → 0
• Flash-Next: 29 → 0

For comparison, llama.cpp's --reasoning-budget 2048 also eliminated blanks, with similar accuracy. The aim here is to put the control in the model file.

The report includes correctness scores, regressions, compatibility checks, evaluation limitations, and links to both models.

Report and downloads:
BoldingBuilds/helping-qwen-finish-thinking

Built on Qwen's models and Unsloth's GGUF conversions. If you've worked with local reasoning models, I'd like to hear what I got wrong.
In this post