NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale
Abstract
Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout clusters before the next batch. Transferring a full 1T checkpoint for such weight synchronization (refit) takes 87.5 min between two AWS regions. Measurements of BF16 training show that about 1% of weights change their stored values per step. Recent systems exploit this sparsity but fall short on placement, exactness, or efficiency: they reimplement placement rules, assemble full tensors, rebuild values arithmetically, or use a cross-cluster collective, and none fully recovers from mid-refit failures. We present NeMo-DCR (Delta-Compressed Refit), which sends only changes yet is bit-exact: receivers obtain the same parameter and buffer bits as a dense refit. For placement, fixed affine mappings project changes from training shards into the checkpoint's canonical coordinates, residual conversion covers the other changes, and the serving runtime's native loader places all changes in receiver storage. For exactness, compressible XOR masks carry affine changes whose projection and loader preserve stored bits, and overwrites carry the others. Receivers apply both in place, retries overwrite partial writes, and a joint commit binds the policy to the baseline for the next delta. For efficiency, object storage or a relay tree streams payloads during delta construction, without a cross-cluster collective. Even at 3% and 5% change rates, NeMo-DCR refits of 30B-1T models are 12-40times faster than a transport-only full-checkpoint reference. A 1T relay-tree refit at 3% takes 150 s instead of 87.5 min, making refits practical for cross-cluster agentic RL at trillion-parameter scale.
Community
I'm excited to share NeMo-DCR, a project I contributed to during my internship at NVIDIA this summer!
π Paper: https://arxiv.org/abs/2610.08430
π The Problem: In mega-scale Agentic RL, each policy update on the training cluster must reach the rollout cluster that serves requests. At a trillion-parameter scale, transferring a full model checkpoint is too slow (taking over an hour).
π‘ Insights: Since only about 1% of weight elements change their stored values per update, sending a delta saves significant bandwidth. However, this introduces a new challenge: training and rollout clusters use different layouts. How can each change be placed efficiently while keeping the final result bitwise identical to a full synchronization?
π οΈ The NeMo-DCR Solution leverages information from both the training shards and the rollout layout:
- π― Coordinate Projection: It projects ~96% of the changes directly into the checkpoint's canonical coordinates, while the remaining 4% are handled via conversion.
- π XOR Masks + Overwrites: For the rollout side, ~96% of the changes that preserve stored bits during loading use XOR masks for better compression. The rest use absolute overwrites.
- β‘ In-place Memory Interception: We intercept the memory copy operations of the rollout side's native loader to apply the updates in place correctly, eliminating the need for model-specific placement logic or assembling the full model in memory first.
- π Architectural Optimizations: Weight version control utilizes a transactional commit mechanism (safe retries, elastic scaling). It supports pipeline overlapping, and under asynchronous training, allows delta transfer to run concurrently with rollout request generation. It also works seamlessly via object storage or relay trees, bypassing cross-cluster collective communication.
π The Impact: In a 3% element-change rate stress test on a 1-trillion-parameter model
- β Full-checkpoint transfer: Took 87.5 minutes
- β NeMo-DCR weight sync: Completed in just 2.5 minutes!
- π This represents a massive ~35Γ speedup!
In addition, check out our documentation and code for more details: π
π Documentation: https://docs.nvidia.com/nemo/rl/nightly/guides/refit.html
π» Code: NVIDIA-NeMo/RL#2444
Get this paper in your agent:
hf papers read 2610.08430 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper