Papers
arxiv:2609.33391

Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents

Published on Sep 27
· Submitted by
Mingju Chen
on Sep 29
Authors:
,
,
,
,

Abstract

Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify Decision--Timestamp Mismatch: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce AlignOPSD, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate AlignOPSD with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. AlignOPSD outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 \% and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at https://github.com/mingju-c/Align-OPSD

Community

Paper author Paper submitter

We introduce AlignOPSD, a framework for improving on-policy distillation in long-horizon agents.

The key insight is that temporal alignment does not necessarily imply decision alignment. Existing approaches typically assign supervision based on timestamps, while real agent decisions often unfold across multiple interaction steps. This creates a hidden mismatch between where supervision is provided and where decisions are actually made.

AlignOPSD first aligns supervision with functional decisions and then performs hierarchical credit assignment over decision spans. By moving beyond timestamp-level alignment, our approach provides more reliable training signals for long-horizon agent learning.

We evaluate AlignOPSD on several interactive agent benchmarks and show consistent improvements over existing baselines.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33391
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.33391 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33391 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.33391 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.