catching up on some bookmarked reads from the summer, reading Antidoom from @liquidai
small reasoning models get stuck more easily when the task involves a long thinking trace and a hard problem. It starts repeating the same word over and over again ("Wait", "Alternatively"โฆ), each repetition makes the next one likelier, and the generation is spent before it reaches an answer
they measured it, 10.2% of completions for an early LFM2.5-2.6B checkpoint and 22.9% for Qwen3.5-4B at greedy. After training those drop to 1.4% and 1.0%
the fix is FTPO (final token preference optimization). What I like is how narrow it is, it only touches the single token where the loop starts
three ways it differs from DPO: > trains one token position, mid-generation, instead of whole sequences > spreads probability across ~20 plausible alternatives instead of swapping one overtrained token for another > keeps the regularizer in logit space, no softmax, so the rest of the vocabulary stays put
the third one is what makes it usable. If you want to edit one position without disturbing the model, you can't have a loss that reshuffles the other 150k logits on the way
and their explanation abt the result: the training teaches the model nothing new about math or code, it clears the failure mode that was blocking answers the model could already produce
I've spent some time reproducing, in the open, Surya N's idea of training a model to paint with code. It's a coding model that learns to paint watercolours by writing JS code, trained with GRPO. I used TRL and OpenEnv for this, with the whole pipeline running on Hugging Face.
The interesting part is that the reward has no correct answer, unlike a math problem. In this case it's based on the artistic preferences of the person who builds the dataset.
Everything is published: the environment, the reference pool, the trained adapters, every painting of every run with the code that made it, and a write-up with all the decisions, including the ones that went wrong.
while preparing the last class of the Training Agents live series during the summer, i spent some time reading the post-training sections of many frontier model reports, to learn how they use RL environments to improve their models, and wrote a blog about it
if you use any kind of coding harness, or you saw the Blender scenes that went viral recently, this might be interesting to you
ICYMI, Async GRPO in TRL now supports LoRA and we wrote a looong blog testing it
> the adapter is a few megabytes, so the weight sync is a file instead of an NCCL transfer > 3 HF Jobs: 1 trainer and 2 vLLM replicas > the adapter travels through an HF Storage Bucket mounted in all 3 at the same path > a proxy in front of the replicas routes each rollout to the one already holding its KV prefix