Papers
arxiv:2609.34645

Nereus: Adaptive Parallelism for LLM Post-Training

Published on Sep 28
ยท Submitted by
Songlin Jiang
on Sep 29
Authors:
,
,
,

Abstract

Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating GPU transfers across models and stages. Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model calibrated against the running job. To estimate and execute a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time. Nereus improves end-to-end 8B PPO throughput by 2.14--7.27times over OpenRLHF and by 1.10--1.47times over Verl across diverse clusters.

Community

Paper author Paper submitter

Nereus adapts the parallel execution plan of an RL post-training job while the job runs.
Resource availability, sequence length, memory pressure, and stage bottlenecks change during a run, so a plan that was good at the start can become slow or even infeasible. Nereus's controller selects a memory-feasible global plan and admits a transition only when the current plan is infeasible or the savings repay the transition cost. It represents each model-stage replica as an Elastic Model Unit and uses a global transition graph to order the transformations and GPU transfers across all models and stages. It achieves 27.7% lower average step latency than a fixed TP/PP layout with DP scaling, on a trace built from real data

Deploying this means the scheduler has to decide whether the remaining run is long enough to pay for a reconfiguration, and that's a bet on the horizon, not on throughput. Most auto-parallelism numbers get reported after the new plan is warm, which quietly prices the transition at zero. What I'd want on the plot is time-to-target including resharding, against how far into the run the re-plan fires โ€” a switch at 80% through a job is almost always a loss, and the scheduler needs to know that before it says yes. The other thing I'd poke at is optimizer state: changing the TP/PP split means resharding Adam's two moments plus the master weights, so if "reuse the job's distributed state" means re-sharding those tensors over the new mesh, the transition cost is a full extra checkpoint round-trip and the decision threshold moves a lot. If it's something cheaper than that, that's the interesting part and I'd want it spelled out. A run where sequence length grows monotonically would be the clean test โ€” static plans are provably wrong there, so adaptation should show up clearly or not at all.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.34645
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.34645 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.34645 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.34645 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.