Pm-ops / README.md
AdityaGuntur's picture
Update README.md
f2699c7 verified
|
Raw
History Blame Contribute Delete
13.9 kB
metadata
title: PM-Ops RL Environment
emoji: πŸ“‹
colorFrom: blue
colorTo: purple
sdk: docker
app_port: 8000
pinned: false
tags:
  - openenv
  - reinforcement-learning

PM-Ops πŸ› οΈ

Can a small model learn to operate inside a company it has never seen before?

PM-Ops is a reinforcement learning benchmark and training environment where an LLM agent operates as a product manager inside a fully simulated software organization β€” navigating ticketing, codebases, and chat β€” using only the organization's own runbook as its guide.


Links

Resource Link
πŸ€— Live Environment TheCrustaceans/Pm-ops β€” HuggingFace Space
πŸ“ Blog Post BlogPost.mdx
πŸŽ₯ inference image Google Drive
πŸ“Š train image Google Drive

Table of Contents


Motivation

Frontier models achieve near-perfect scores on standard benchmarks. Yet when deployed inside a real organization, they routinely fail tasks that a junior employee handles on day one.

The reason: organizational context is not on the internet.

A model knows that database failures are serious. It does not know that your company outsources DB infrastructure to a vendor and that the correct response is an email, not an internal incident ticket. No benchmark measures this gap. PM-Ops does.

Why system prompts are not the answer:

Problem Detail
Attention decay As conversation grows, the model attends to early context instructions with diminishing weight. Org conventions buried in a long system prompt get effectively ignored.
Static snapshots Channels get renamed. Teams get reorganized. A system prompt written last quarter is already wrong.
No gradient The model is told what to do, not trained on the consequences of ignoring it. There is no learning signal.

PM-Ops addresses this by making the agent experience the consequences of skipping the runbook across thousands of varied organizational configurations during training.


Environment

PM-Ops runs as a WebSocket-based environment server (OpenEnv-compatible) hosted on HuggingFace Spaces. Each episode presents the agent with a freshly generated organization β€” different team names, channel names, label taxonomies, and priority levels every time.

The Three Apps

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    PM-Ops Environment                    β”‚
β”‚                                                         β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚  β”‚  Ticketing   β”‚  β”‚   Codebase   β”‚  β”‚     Chat     β”‚  β”‚
β”‚  β”‚  (Jira-like) β”‚  β”‚ (GitHub-like)β”‚  β”‚ (Slack-like) β”‚  β”‚
β”‚  β”‚              β”‚  β”‚              β”‚  β”‚              β”‚  β”‚
β”‚  β”‚ tickets      β”‚  β”‚ repositories β”‚  β”‚ channels     β”‚  β”‚
β”‚  β”‚ projects     β”‚  β”‚ commits      β”‚  β”‚ threads      β”‚  β”‚
β”‚  β”‚ teams        β”‚  β”‚ pull requestsβ”‚  β”‚ DMs          β”‚  β”‚
β”‚  β”‚ labels       β”‚  β”‚ file authors β”‚  β”‚ user profilesβ”‚  β”‚
β”‚  β”‚ priorities   β”‚  β”‚ changed filesβ”‚  β”‚              β”‚  β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚                                                         β”‚
β”‚  Ground truth verified against DB β€” not LLM-judged      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Ticketing App: The source of truth for org state. Agent creates tickets, assigns teams, sets labels and priorities. Correctness is verified by checking the database, not asking another LLM.

Codebase App: Pure detective work. No code is written. When a bug report arrives, the agent traces commit history, identifies the responsible change, and finds the author.

Chat App: Episode-specific channels. The channel #oncall-payments in episode 1 may be #urgent-billing in episode 2. The agent must read the runbook to know which one is live β€” not guess from training data.

Action Space

{ "action_type": "meta.read_runbook",       "args": {} }
{ "action_type": "meta.finish",             "args": {} }
{ "action_type": "meta.noop",               "args": {} }
{ "action_type": "ticketing.create_ticket", "args": { "summary": "...", "label": "...", "priority": "...", "assignee": "..." } }
{ "action_type": "ticketing.assign_ticket", "args": { "ticket_id": "...", "team": "..." } }
{ "action_type": "chat.list_channels",      "args": {} }
{ "action_type": "chat.post_message",       "args": { "channel": "...", "text": "..." } }
{ "action_type": "codebase.get_commits",    "args": { "repo": "..." } }

All actions are emitted as chain-of-thought reasoning followed by a ```json block. Three-pass extraction handles malformed outputs: code block β†’ raw JSON β†’ regex fallback.

Observation Schema

Each step returns:

{
    "task_brief":             str,   # the incident/task description
    "last_action_result":     dict,  # {"ok": bool, "data": ..., "error": ...}
    "step":                   int,
    "steps_remaining":        int,
    "reward":                 float, # always 0.0 until meta.finish
    "done":                   bool,
    "token_budget_remaining": int,
}

Delayed reward: reward is always 0.0 until the agent calls meta.finish. This forces the agent to commit to a plan and execute it β€” step-by-step reward hacking is not possible.

Org Generator

LLM-generated scenarios sound plausible but fail formal verification. PM-Ops uses a deterministic, parameterized, seedable org generator instead.

Each org is generated from a seed and produces:

  • label_taxonomy β€” org-specific bug/feature labels
  • priority_levels β€” org-specific severity scale
  • team_map β€” service β†’ owning team
  • oncall_channels β€” team β†’ notification channel
  • required_ticket_fields β€” what fields must be set for a valid ticket

The same seed always produces the same org. Evaluation is reproducible. Training sees a different org every episode.

Task Types

Task Description Key Challenge
Triage Bug report arrives. Create ticket with correct label, priority, assignee. Notify correct channel. Using org taxonomy, not generic labels.
Incident Routing Production incident. Identify affected services, find owning team via codebase + team map, escalate. Cross-system reasoning: chat β†’ codebase β†’ ticketing.
Release Notes Compile and post release notes for resolved tickets in the org's specific format. Format compliance, not just content.
Dependency Update Breaking library change hits multiple services. Notify each team through their own oncall channel, create per-team tickets. Multi-target coordination without collapsing to single-service logic.

Difficulty Tiers

Tier Services Labels Priorities Notes
Easy 2 3 2 No ambiguity. Runbook is complete.
Medium 3 4 4 Partial runbook. Agent must infer across systems. Multi-owner configs.
Hard β‰₯5 5+ 5+ Noise channels designed to distract. Outdated/missing docs. Simulated panicking users in general chat β€” correct behavior is to ignore them.

The hard tier specifically tests resistance to NLP pressure. A model that responds to "THE SITE IS DOWN PLEASE HELP" in #general instead of reading the runbook and acting on verified info fails the episode.


Reward Design

Rewards are delayed β€” revealed only at meta.finish.

Scoring Breakdown

Ticket created correctly     +0.25
Correct label                +0.20
Correct priority             +0.20
Correct team assignment      +0.20
Correct channel notification +0.15
─────────────────────────────────
Maximum per episode           1.00

Penalties

Wrong channel post           -0.05 each
                             (>2 wrong posts β†’ net negative; defeats channel-spray)
Duplicate ticket             -0.10 each
Zero valid actions (inaction)-1.00

The -0.05 per wrong channel is calibrated: posting to every channel to guarantee hitting the right one becomes net negative after 2 wrong posts. The agent must read the runbook to know the correct channel.

Combined Training Reward

if valid_action_count == 0:
    combined = -1.0                              # never output valid JSON
elif final_score == 0.0:
    combined = valid_json_ratio * 0.15           # tried but failed
            + read_runbook_reward * 0.10
            - 0.30                               # hard penalty for zero completion
else:
    combined = final_score         * 0.45        # task correctness
            + no_wrong_channels    * 0.15        # anti channel-spray
            + valid_json_ratio     * 0.15        # format discipline
            + read_runbook_reward  * 0.15        # process compliance
            + efficiency           * 0.10        # steps saved

Training Pipeline

SFT Warmup

Before GRPO, a short supervised fine-tuning phase runs on baseline agent demonstrations. Without this, the model outputs prose instead of structured JSON actions and the GRPO gradient is zero β€” the model needs to learn the output format before it can learn the task.

GRPO

We use Group Relative Policy Optimization to train a 7B parameter model directly against PM-Ops episode rewards.

Each training step is a full PM-Ops episode:

  1. Model plays through up to 15 turns
  2. Environment scores the final state
  3. Gradient updates weights based on relative performance across the generation group

No intermediate reward signal is given. The model must learn to plan across multiple steps.

Reward Shaping

Component Weight Purpose
final_score 0.45 Primary correctness from env grader
no_wrong_channels 0.15 Anti-hack: penalise channel spray
valid_json_ratio 0.15 Format discipline
read_runbook 0.15 Process compliance
efficiency 0.10 Steps saved (only when task succeeds)

PMOpsGRPOTrainer subclasses TRL's GRPOTrainer and overrides _calculate_rewards() to inject pre-computed rewards from the rollout directly, bypassing TRL's broken kwargs flow for multi-turn rollouts.


Results

Evaluated on 13 matched episodes, comparing the heuristic baseline agent against the GRPO-trained model:

Model Avg Score (13 eps)
Heuristic Baseline 0.269
GRPO Trained (early) 0.362
Improvement +34.6%

Training reward trend across 14 steps shows a positive slope of +0.0024/step with KL divergence remaining stable and controlled throughout.


Project Structure

pm_ops/
β”œβ”€β”€ server/
β”‚   β”œβ”€β”€ pm_ops_environment.py   # OpenEnv Environment subclass
β”‚   β”œβ”€β”€ org_generator.py        # Deterministic, seedable org factory
β”‚   β”œβ”€β”€ grader.py               # Ground-truth reward computation
β”‚   └── app.py                  # FastAPI server entry point
β”œβ”€β”€ inference.py                # Heuristic baseline agent + eval runner
└── pyproject.toml

training/
β”œβ”€β”€ rollout.py                  # Multi-turn rollout: build_messages, extract_json_action
β”œβ”€β”€ rewards.py                  # Reward functions + weight constants
β”œβ”€β”€ pm_ops_trainer.py           # PMOpsGRPOTrainer (_calculate_rewards override)
β”œβ”€β”€ prompts.py                  # SYSTEM_PROMPT, format_observation
β”œβ”€β”€ dataset.py                  # Dataset loading + seed parsing
└── train_v3.ipynb              # SFT warmup + GRPO training notebook

Setup & Usage

Run the environment server

cd pm_ops
pip install -e .
server  # starts FastAPI on :7860

Run the baseline agent

export API_BASE_URL=https://thecrustaceans-pm-ops.hf.space
python pm_ops/inference.py

Connect via OpenEnv client

from openenv.core import GenericEnvClient

env = GenericEnvClient(base_url="https://thecrustaceans-pm-ops.hf.space").sync()

with env:
    result = env.reset()
    obs = result.observation

    result = env.step({
        "action_type": "meta.read_runbook",
        "args": {}
    })

Training

Open training/train_v3.ipynb in a Colab instance with a GPU. The notebook handles SFT warmup, GRPO setup, and checkpointing.


Live Environment: https://huggingface.co/spaces/TheCrustaceans/Pm-ops