Instructions to use PIPer-iclr/PIPer-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PIPer-iclr/PIPer-8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="PIPer-iclr/PIPer-8B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("PIPer-iclr/PIPer-8B") model = AutoModelForCausalLM.from_pretrained("PIPer-iclr/PIPer-8B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PIPer-iclr/PIPer-8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PIPer-iclr/PIPer-8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PIPer-iclr/PIPer-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/PIPer-iclr/PIPer-8B
- SGLang
How to use PIPer-iclr/PIPer-8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PIPer-iclr/PIPer-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PIPer-iclr/PIPer-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PIPer-iclr/PIPer-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PIPer-iclr/PIPer-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use PIPer-iclr/PIPer-8B with Docker Model Runner:
docker model run hf.co/PIPer-iclr/PIPer-8B
File size: 3,358 Bytes
b0e9d5b b754058 b0e9d5b bb5aed1 b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 b0e9d5b b754058 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 | ---
library_name: transformers
datasets:
- PIPer-iclr/envbench-zeroshot-rl
- PIPer-iclr/PIPer-SFT-2500-sharegpt
base_model:
- PIPer-iclr/Qwen3-8B-am
pipeline_tag: text-generation
license: mit
---
# π PIPer: On-Device Environment Setup via Online Reinforcement Learning
<div align="center">
[](https://huggingface.co/PIPer-iclr)
[](https://huggingface.co/datasets/PIPer-iclr/envbench-zeroshot-rl)
[](LICENSE)
*Democratizing environment setup with on-device sized models that match the performance of much larger proprietary systems*
</div>
## π― Overview
Environment setupβthe process of configuring systems to work with specific software projectsβremains a persistent challenge in software engineering. **PIPer** addresses this by training specialized on-device models that can automatically generate correct Bash scripts for environment configuration.
Our approach combines:
- π **Supervised Fine-Tuning (SFT)** with executable scripts from larger models
- π― **Reinforcement Learning with Verifiable Rewards (RLVR)** using lightweight proxy LLM-reward
## π Key Results
| Model | Size | EnvBench avg@5 | Cost per 1M tokens |
|-------|------|----------------|-------------------|
| **PIPer** | 8B | **19.4** | $0.60 |
| GPT-4o | - | 19.4 | $15.00 |
| Qwen3-32B | 32B | 16.2 | $2.00 |
| Qwen3-8B | 8B | 2.6 | $0.60 |
> π **PIPer achieves 9Γ improvement** over its base model while **matching GPT-4o performance** at **25x lower cost**

## π¦ Available Artifacts
### π€ Model Checkpoints
| Model | Description | HuggingFace Link |
|-------|-------------|------------------|
| **π
PIPer (Full)** | Complete SFT+RL trained model | [PIPer-iclr/PIPer-8B](https://huggingface.co/PIPer-iclr/PIPer-8B) |
| π― PIPer (RL-only) | RLVR checkpoint only | [PIPer-iclr/PIPer-8B-RL-only](https://huggingface.co/PIPer-iclr/PIPer-8B-RL-only) |
| π PIPer (SFT-only) | Supervised fine-tuning only | [PIPer-iclr/PIPer-8B-SFT-only](https://huggingface.co/PIPer-iclr/PIPer-8B-SFT-only) |
### π Datasets
| Dataset | Description | HuggingFace Link |
|---------|-------------|------------------|
| **EnvBench Zero-shot RL** | Training prompts and evaluation data | [PIPer-iclr/envbench-zeroshot-rl](https://huggingface.co/datasets/PIPer-iclr/envbench-zeroshot-rl) |
## π Evaluation Benchmarks
| Benchmark | Description | Metric | Our Result |
|-----------|-------------|---------|------------|
| **EnvBench-Python** | 329 Python repositories | pass@5 | π **27/329** |
| **Repo2Run** | 420 Python repositories | pass@5 | π **103/420** |
| **Terminal-Bench** | 80 terminal tasks | pass@10 | **4/80** |
## π Acknowledgments
- Built on top of [EnvBench](https://github.com/princeton-nlp/EnvBench) evaluation framework
- Uses [VeRL](https://github.com/volcengine/verl) for efficient RL training
- Leverages [Qwen3](https://huggingface.co/Qwen) model family as base architecture
## π License
This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details. |