Nereus: Adaptive Parallelism for LLM Post-Training
Abstract
Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating GPU transfers across models and stages. Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model calibrated against the running job. To estimate and execute a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time. Nereus improves end-to-end 8B PPO throughput by 2.14--7.27times over OpenRLHF and by 1.10--1.47times over Verl across diverse clusters.
Community
Nereus adapts the parallel execution plan of an RL post-training job while the job runs.
Resource availability, sequence length, memory pressure, and stage bottlenecks change during a run, so a plan that was good at the start can become slow or even infeasible. Nereus's controller selects a memory-feasible global plan and admits a transition only when the current plan is infeasible or the savings repay the transition cost. It represents each model-stage replica as an Elastic Model Unit and uses a global transition graph to order the transformations and GPU transfers across all models and stages. It achieves 27.7% lower average step latency than a fixed TP/PP layout with DP scaling, on a trace built from real data
Deploying this means the scheduler has to decide whether the remaining run is long enough to pay for a reconfiguration, and that's a bet on the horizon, not on throughput. Most auto-parallelism numbers get reported after the new plan is warm, which quietly prices the transition at zero. What I'd want on the plot is time-to-target including resharding, against how far into the run the re-plan fires โ a switch at 80% through a job is almost always a loss, and the scheduler needs to know that before it says yes. The other thing I'd poke at is optimizer state: changing the TP/PP split means resharding Adam's two moments plus the master weights, so if "reuse the job's distributed state" means re-sharding those tensors over the new mesh, the transition cost is a full extra checkpoint round-trip and the decision threshold moves a lot. If it's something cheaper than that, that's the interesting part and I'd want it spelled out. A run where sequence length grows monotonically would be the clean test โ static plans are provably wrong there, so adaptation should show up clearly or not at all.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Arachne: Learning to Plan Parallel Training on Dynamic Heterogeneous Clusters (2026)
- Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training (2026)
- PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning (2026)
- Scheduling Mixed RL Rollouts Beyond Prefix Locality (2026)
- AInfer-PD: Communication-Safe In-Place Prefill-Decode Multiplexing for Distributed MoE Rollouts (2026)
- TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling (2026)
- Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUs (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.34645 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper