Papers
arxiv:2609.36653

Scheduling Recursive Reasoning in Looped Transformers

Published on Sep 29
· Submitted by
Elvis Wang
on Sep 30
Authors:
,
,
,

Abstract

Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with shared parameters. However, these models apply each learned update with a fixed unit scale, which can be conservative when updates make persistent progress and overly aggressive when they fluctuate, limiting the benefit of additional loops. To understand how the scale should vary along the trajectory, we first analyze the sensitivity of terminal loss to recurrent update scale. We show that its temporal average admits an exact decomposition into persistent-progress and centered-fluctuation contributions. Based on this, we introduce the Trajectory Adaptive Progress-Fluctuation Scheduler (TAPS), which tracks their balance across recurrent updates and adapts the step size online. Theoretically, we establish sufficient conditions under which TAPS reduces expected terminal loss and reaches a target quality in fewer recurrent loops. Empirically, we show that TAPS improves terminal accuracy across structured reasoning tasks without retraining. By further incorporating the progress-fluctuation principle into training, TAPS yields additional accuracy gains with up to 1.56 times wall-clock speedup at matched baseline accuracy. The broad applicability of TAPS is supported by its effectiveness across diverse recurrent architectures and inference strategies. Together, these results establish update scale as complementary control axis of recurrent inference alongside architecture and depth.

Community

Paper author Paper submitter

taps_副本2

Scheduling Recursive Reasoning in Looped Transformers

Recurrent reasoning models scale test-time computation by repeatedly refining latent states with shared parameters. However, they typically apply every learned update with a fixed unit scale. This can be too conservative when updates make persistent progress and too aggressive when they fluctuate, limiting what additional loops can achieve.

This suggests that recurrent reasoning should not only decide how long to think, but also how strongly to apply each update.

截屏2026-09-30 23.04.05

We study this question through the sensitivity of terminal loss to recurrent update scale. Its temporal average admits an exact decomposition into two terms: persistent progress and centered fluctuation. Their balance tells us whether the current trajectory supports a larger or smaller step.

This leads to TAPS, the Trajectory Adaptive Progress–Fluctuation Scheduler. TAPS reads this signal directly from the recurrent trajectory and adjusts the step size online, without changing the model architecture or weights.

Across structured reasoning tasks, TAPS improves terminal accuracy without retraining. Incorporating the same principle during training brings further gains and up to 1.56× wall-clock speedup at matched baseline accuracy, with the effect extending across recurrent architectures and inference strategies.

截屏2026-09-30 23.05.27

截屏2026-09-30 23.05.56

截屏2026-09-30 23.07.27

Architecture determines the update, depth determines how often it is applied, and TAPS controls its scale.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.36653
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 2

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.36653 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.36653 in a Space README.md to link it from this page.

Collections including this paper 1