Papers
arxiv:2610.08430

NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale

Published on Oct 6
Β· Submitted by
Songlin Jiang
on Oct 7
Authors:
,
,
,
,
,
,

Abstract

Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout clusters before the next batch. Transferring a full 1T checkpoint for such weight synchronization (refit) takes 87.5 min between two AWS regions. Measurements of BF16 training show that about 1% of weights change their stored values per step. Recent systems exploit this sparsity but fall short on placement, exactness, or efficiency: they reimplement placement rules, assemble full tensors, rebuild values arithmetically, or use a cross-cluster collective, and none fully recovers from mid-refit failures. We present NeMo-DCR (Delta-Compressed Refit), which sends only changes yet is bit-exact: receivers obtain the same parameter and buffer bits as a dense refit. For placement, fixed affine mappings project changes from training shards into the checkpoint's canonical coordinates, residual conversion covers the other changes, and the serving runtime's native loader places all changes in receiver storage. For exactness, compressible XOR masks carry affine changes whose projection and loader preserve stored bits, and overwrites carry the others. Receivers apply both in place, retries overwrite partial writes, and a joint commit binds the policy to the baseline for the next delta. For efficiency, object storage or a relay tree streams payloads during delta construction, without a cross-cluster collective. Even at 3% and 5% change rates, NeMo-DCR refits of 30B-1T models are 12-40times faster than a transport-only full-checkpoint reference. A 1T relay-tree refit at 3% takes 150 s instead of 87.5 min, making refits practical for cross-cluster agentic RL at trillion-parameter scale.

Community

Paper author Paper submitter

I'm excited to share NeMo-DCR, a project I contributed to during my internship at NVIDIA this summer!

πŸ“‘ Paper: https://arxiv.org/abs/2610.08430

πŸ”„ The Problem: In mega-scale Agentic RL, each policy update on the training cluster must reach the rollout cluster that serves requests. At a trillion-parameter scale, transferring a full model checkpoint is too slow (taking over an hour).

πŸ’‘ Insights: Since only about 1% of weight elements change their stored values per update, sending a delta saves significant bandwidth. However, this introduces a new challenge: training and rollout clusters use different layouts. How can each change be placed efficiently while keeping the final result bitwise identical to a full synchronization?

πŸ› οΈ The NeMo-DCR Solution leverages information from both the training shards and the rollout layout:

  1. 🎯 Coordinate Projection: It projects ~96% of the changes directly into the checkpoint's canonical coordinates, while the remaining 4% are handled via conversion.
  2. 🎭 XOR Masks + Overwrites: For the rollout side, ~96% of the changes that preserve stored bits during loading use XOR masks for better compression. The rest use absolute overwrites.
  3. ⚑ In-place Memory Interception: We intercept the memory copy operations of the rollout side's native loader to apply the updates in place correctly, eliminating the need for model-specific placement logic or assembling the full model in memory first.
  4. 🌐 Architectural Optimizations: Weight version control utilizes a transactional commit mechanism (safe retries, elastic scaling). It supports pipeline overlapping, and under asynchronous training, allows delta transfer to run concurrently with rollout request generation. It also works seamlessly via object storage or relay trees, bypassing cross-cluster collective communication.

πŸš€ The Impact: In a 3% element-change rate stress test on a 1-trillion-parameter model

  • ❌ Full-checkpoint transfer: Took 87.5 minutes
  • βœ… NeMo-DCR weight sync: Completed in just 2.5 minutes!
  • πŸ“ˆ This represents a massive ~35Γ— speedup!

In addition, check out our documentation and code for more details: πŸ‘‡
πŸ“– Documentation: https://docs.nvidia.com/nemo/rl/nightly/guides/refit.html
πŸ’» Code: NVIDIA-NeMo/RL#2444

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.08430
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.08430 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.08430 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.08430 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.