On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training
Abstract
The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generalization of other paradigms such as supervised fine-tuning (SFT). To address this limitation, we investigate whether there exists a specific on-policy update behavior that can achieve such improvements. First, our analyses reveal that SFT updates parameters along consistent directions, while the on-policy paradigm continuously adjusts the direction during training. This difference inspires us to focus on the cumulative update direction of each parameter as a promising behavior. Then, we evaluate its effectiveness for improving generalization by proposing On-Policy direction-constrained Supervised Fine-Tuning (OPSFT), which constrains SFT updates to the direction identified by on-policy paradigms. The strong performance of OPSFT indicates that the generalization advantage of on-policy paradigms can be transferred to SFT through the parameter update direction. Once such a direction is identified, even SFT can generalize with its updates constrained to this direction. This finding offers two practical benefits by combining the strong generalization of on-policy paradigms with the advantages of SFT, including the high training efficiency and ability to leverage high-quality trajectories. For efficiency, we identify update directions that support strong generalization using a few on-policy training steps, and subsequently apply OPSFT to achieve high training efficiency. For leveraging high-quality trajectories, OPSFT can utilize these trajectories to continue improving a post-trained model along its update direction without disrupting the ability learned from on-policy training.
Community
Why does on-policy training generalize better than SFT? We show that the direction of parameter updates plays a crucial role. By extracting update directions from a few on-policy steps and constraining subsequent SFT updates to these directions, our OPSFT transfers the generalization benefits of on-policy training to SFT while preserving its efficiency and ability to learn from high-quality trajectories.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- MInTRL: Off-policy Intervention can boost On-policy RL (2026)
- SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation (2026)
- Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement (2026)
- GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs (2026)
- On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics (2026)
- Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR (2026)
- Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.36659 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper