Spaces:
Sleeping
title: PM-Ops RL Environment
emoji: 📋
colorFrom: blue
colorTo: purple
sdk: docker
app_port: 8000
pinned: false
tags:
- openenv
- reinforcement-learning
PM-Ops: Procedural Project-Management RL Environment
A multi-app simulated environment (Ticketing, Codebase, Chat) featuring procedural organizational drift, designed for reinforcement learning training.
What makes it RL-shaped
Every episode samples a fresh OrgConfig -- different label names, priority schemes, team-to-service mappings, and oncall channel names. Frontier LLMs fail because they rely on memorized conventions; a trained agent learns to read the runbook first.
Tasks
| Task | Brief | Max Steps |
|---|---|---|
| triage | Bug report -- correct label, priority, team, channel | 25 |
| incident_routing | Alert -- identify owner via commits, page oncall | 25 |
| release_notes | Compile + post release notes in org style | 40 |
| dep_update | Coordinate dependency update across all owning teams | 40 |
Action space
The agent calls actions like:
meta.read_runbook -- learn this org's conventions ticketing.create_ticket -- file a ticket with label + priority ticketing.assign_ticket -- assign to correct team chat.post_message -- notify the right oncall channel meta.finish -- end episode, trigger scoring
Reward
Delayed to episode end. Deterministic Python verifier -- no LLM-as-judge. Partial credit for partial correctness.
Sponsor alignment
- AI Labs: direct PM-in-the-loop positioning
- Patronus: schema drift -- env tests whether agents track convention changes
- Scale AI: enterprise workflow simulation
- Meta RFC 004: delayed trajectory reward is the native reward mode