RLinf
/

RLinf-Pi0-LIBERO-Spatial-Object-Goal-SFT

Safetensors

Model card Files Files and versions

xet

Community

Add model card for $\pi_\texttt{RL}$

by nielsr HF Staff - opened Nov 4, 2025

base: refs/heads/main

←

from: refs/pr/1

Discussion Files changed

+39

-0

Files changed (1) hide show

README.md +39 -0

README.md ADDED Viewed

	@@ -0,0 +1,39 @@

+---
+license: apache-2.0
+pipeline_tag: robotics
+---
+# $\pi_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
+This repository contains resources and artifacts related to the $\pi_\texttt{RL}$ model, an open-source framework for training flow-based Vision-Language-Action (VLA) models through online reinforcement learning, as described in the paper below.
+- **Paper**: [$\pi_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models](https://huggingface.co/papers/2510.25889)
+- **GitHub Repository**: [https://github.com/RLinf/RLinf](https://github.com/RLinf/RLinf)
+- **Project Page / Documentation**: [https://rlinf.readthedocs.io/en/latest/](https://rlinf.readthedocs.io/en/latest/)
+## Abstract
+Vision-Language-Action (VLA) models enable robots to understand and perform complex tasks from multimodal input. Although recent work explores using reinforcement learning (RL) to automate the laborious data collection process in scaling supervised fine-tuning (SFT), applying large-scale RL to flow-based VLAs (e.g., $\pi_0$, $\pi_{0.5}$) remains challenging due to intractable action log-likelihoods from iterative denoising. We address this challenge with $\pi_{\text{RL}}$, an open-source framework for training flow-based VLAs in parallel simulation. $\pi_{\text{RL}}$ implements two RL algorithms: (1) {Flow-Noise} models the denoising process as a discrete-time MDP with a learnable noise network for exact log-likelihood computation. (2) {Flow-SDE} integrates denoising with agent-environment interaction, formulating a two-layer MDP that employs ODE-to-SDE conversion for efficient RL exploration. We evaluate $\pi_{\text{RL}}$ on LIBERO and ManiSkill benchmarks. On LIBERO, $\pi_{\text{RL}}$ boosts few-shot SFT models $\pi_0$ and $\pi_{0.5}$ from 57.6% to 97.6% and from 77.1% to 98.3%, respectively. In ManiSkill, we train $\pi_{\text{RL}}$ in 320 parallel environments, improving $\pi_0$ from 41.6% to 85.7% and $\pi_{0.5}$ from 40.0% to 84.8% across 4352 pick-and-place tasks, demonstrating scalable multitask RL under heterogeneous simulation. Overall, $\pi_{\text{RL}}$ achieves significant performance gains and stronger generalization over SFT-models, validating the effectiveness of online RL for flow-based VLAs.
+## RLinf Framework Overview
+$\pi_\texttt{RL}$ is implemented as part of **RLinf**, a flexible and scalable open-source infrastructure designed for post-training foundation models via reinforcement learning. The 'inf' in RLinf stands for `Infrastructure`, highlighting its role as a robust backbone for next-generation training. It also stands for `Infinite`, symbolizing the system’s support for open-ended learning, continuous generalization, and limitless possibilities in intelligence development.
+<div align="center">
+  <img src="https://github.com/RLinf/RLinf/raw/main/docs/source-en/_static/svg/overview.svg" alt="RLinf-overview"/>
+</div>
+RLinf supports mainstream VLA models, mainstream CPU & GPU-based simulators, and enables the first RL fine-tuning of the $\pi_{0}$ and $\pi_{0.5}$ model family with a flow-matching action expert. For detailed installation, usage, and examples, please refer to the [RLinf GitHub repository](https://github.com/RLinf/RLinf) and the [official documentation](https://rlinf.readthedocs.io/en/latest/).
+## Citation
+If you find $\pi_\texttt{RL}$ helpful, please cite the paper:
+```bibtex
+@misc{chen2025pitextttrlonlinerlfinetuning,
+      title={$\pi_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models},
+      author={Kang Chen and Zhihao Liu and Tonghe Zhang and Zhen Guo and Si Xu and Hao Lin and Hongzhi Zang and Quanlu Zhang and Zhaofei Yu and Guoliang Fan and Tiejun Huang and Yu Wang and Chao Yu},
+      year={2025},
+      eprint={2510.25889},
+      archivePrefix={arXiv},
+      primaryClass={cs.LG},
+      url={https://arxiv.org/abs/2510.25889},
+}
+```