AI & ML interests

Independent measurement of what post-training edits actually cost. Quantization, abliteration, distillation, reasoning budgets. Take nobody's word for it.

BoldingBuilds 
posted an update about 5 hours ago
view post
Post
37
I released a one-row weight edit for Qwen3.8 27B and Flash-Next to address a frustrating failure: using the entire output budget thinking, then returning no final answer.

The edit changes only the output-layer row that scores </think> and is packaged into ordinary GGUF files.

On 200 MATH-500 problems at a 4,096-token output limit:

• 27B: 31 empty answers → 0
• Flash-Next: 29 → 0

For comparison, llama.cpp's --reasoning-budget 2048 also eliminated blanks, with similar accuracy. The aim here is to put the control in the model file.

The report includes correctness scores, regressions, compatibility checks, evaluation limitations, and links to both models.

Report and downloads:
BoldingBuilds/helping-qwen-finish-thinking

Built on Qwen's models and Unsloth's GGUF conversions. If you've worked with local reasoning models, I'd like to hear what I got wrong.
BoldingBuilds 
posted an update 5 days ago
view post
Post
238
I tested every uncensored Ternary Bonsai 2 27B on the Hub, all 11 builds from 6 uploaders, including my own. Same prompts, same judge, same GPUs, every file pinned by sha256.

• Best answers: @Hikari07jp and @dealignai (~0.94 answer quality with thinking off)
• Only edit with no measurable MMLU cost: mine (±0.15 pp). Every other edit loses 0.64–2.68 pp, all p < 0.001
• Heretic and Blackfrost still refuse 11–12% of harmful prompts
• Thinking mode at 4,096 tokens: 5–31% of harmful prompts get no answer. PrismML recommends 16,384+, and I'm rerunning at that budget

Full report, charts and model-card checks: BoldingBuilds/bonsai-2-uncensored-shootout