---
library_name: transformers
datasets:
- PIPer-iclr/envbench-zeroshot-rl
base_model:
- PIPer-iclr/Qwen3-8B-am
pipeline_tag: text-generation
license: mit
---
# π PIPer: On-Device Environment Setup via Online Reinforcement Learning
[](https://huggingface.co/PIPer-iclr)
[](https://huggingface.co/datasets/PIPer-iclr/envbench-zeroshot-rl)
[](LICENSE)
*Democratizing environment setup with on-device sized models that match the performance of much larger proprietary systems*
## π― Overview
Environment setupβthe process of configuring systems to work with specific software projectsβremains a persistent challenge in software engineering. **PIPer** addresses this by training specialized on-device models that can automatically generate correct Bash scripts for environment configuration.
Our approach combines:
- π **Supervised Fine-Tuning (SFT)** with executable scripts from larger models
- π― **Reinforcement Learning with Verifiable Rewards (RLVR)** using lightweight proxy LLM-reward
## π Key Results
| Model | Size | EnvBench avg@5 | Cost per 1M tokens |
|-------|------|----------------|-------------------|
| **PIPer** | 8B | **19.4** | $0.60 |
| GPT-4o | - | 19.4 | $15.00 |
| Qwen3-32B | 32B | 16.2 | $2.00 |
| Qwen3-8B | 8B | 2.6 | $0.60 |
> π **PIPer achieves 9Γ improvement** over its base model while **matching GPT-4o performance** at **25x lower cost**

## π¦ Available Artifacts
### π€ Model Checkpoints
| Model | Description | HuggingFace Link |
|-------|-------------|------------------|
| **π
PIPer (Full)** | Complete SFT+RL trained model | [PIPer-iclr/PIPer-8B](https://huggingface.co/PIPer-iclr/PIPer-8B) |
| π― PIPer (RL-only) | RLVR checkpoint only | [PIPer-iclr/PIPer-8B-RL-only](https://huggingface.co/PIPer-iclr/PIPer-8B-RL-only) |
| π PIPer (SFT-only) | Supervised fine-tuning only | [PIPer-iclr/PIPer-8B-SFT-only](https://huggingface.co/PIPer-iclr/PIPer-8B-SFT-only) |
### π Datasets
| Dataset | Description | HuggingFace Link |
|---------|-------------|------------------|
| **EnvBench Zero-shot RL** | Training prompts and evaluation data | [PIPer-iclr/envbench-zeroshot-rl](https://huggingface.co/datasets/PIPer-iclr/envbench-zeroshot-rl) |
## π Evaluation Benchmarks
| Benchmark | Description | Metric | Our Result |
|-----------|-------------|---------|------------|
| **EnvBench-Python** | 329 Python repositories | pass@5 | π **27/329** |
| **Repo2Run** | 420 Python repositories | pass@5 | π **103/420** |
| **Terminal-Bench** | 80 terminal tasks | pass@10 | **4/80** |
## π Acknowledgments
- Built on top of [EnvBench](https://github.com/princeton-nlp/EnvBench) evaluation framework
- Uses [VeRL](https://github.com/volcengine/verl) for efficient RL training
- Leverages [Qwen3](https://huggingface.co/Qwen) model family as base architecture
## π License
This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.