Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
Abstract
A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, tokens with different numbers of loops cannot share a uniform forward pass and therefore cannot be handled by standard batching systems such as vLLM. The practical value of depth-adaptive inference thus hinges on whether batching can be made efficient. We introduce the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batches between loop steps. Our method dynamically schedules looped and non-looped parts of the architecture, manages looped KV-caching, and predicts which tokens will exit the loop in advance so it can prepare batches asynchronously. Experiments on Ouro 1.4B and Huginn 3.5B show that fully looped architectures are best suited to depth-adaptive inference, as large non-looped layers outside the recurrent core (e.g., token embedding, LM head, and unshared transformer blocks) slow down and complicate scheduling. Overall, CDB realizes up to 99% of the estimated maximum speedup available, leaving further gains primarily dependent on model architecture and exit behavior.
Community
Ever wondered why LLMs are using the same forward pass and thus the same compute for every single token no matter how hard the prediction is?
The solution: looped LMs can use less compute for "easy" tokens and more for "hard" ones.
The problem: tokens with different numbers of loops cannot share a uniform forward pass and therefore cannot be handled by standard batching systems such as vLLM.
This paper: introduces the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batches between loop steps.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models (2026)
- CoRun: Padding is Simple and Efficient for Deterministic LLM Inference (2026)
- RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory (2026)
- ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference (2026)
- T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning with Dynamic Routing (2026)
- FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates (2026)
- Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper