| --- |
| title: B2B Support Triage OpenEnv |
| emoji: "🧭" |
| colorFrom: blue |
| colorTo: indigo |
| sdk: docker |
| pinned: false |
| app_port: 8000 |
| base_path: /web |
| tags: |
| - openenv |
| --- |
| |
| # B2B Support Triage OpenEnv Benchmark |
|
|
| A deterministic, real-world OpenEnv environment that simulates B2B SaaS customer support triage. |
|
|
| ## Why this environment |
|
|
| Support teams repeatedly perform structured triage decisions under policy constraints: |
| - classify ticket domain |
| - assign urgency |
| - route to the right queue |
| - set SLA targets |
| - draft policy-compliant replies |
|
|
| This environment is designed for agent evaluation with deterministic graders and shaped rewards, not toy gameplay. |
|
|
| ## OpenEnv API |
|
|
| The server exposes standard OpenEnv simulation endpoints: |
| - `POST /reset` |
| - `POST /step` |
| - `GET /state` |
|
|
| `reset()` supports `task_id` and `seed`. |
|
|
| ## Action Space |
|
|
| Typed action model: `B2BSupportTriageAction` |
|
|
| - `action_type`: `classify | set_priority | route | draft_reply | submit` |
| - `ticket_id`: required for all actions except `submit` |
| - `payload`: |
| - `category` |
| - `priority` |
| - `route_queue` |
| - `sla_minutes` |
| - `escalate` |
| - `reply_text` |
|
|
| Validation is deterministic and action-specific. |
|
|
| ## Observation Space |
|
|
| Typed observation model: `B2BSupportTriageObservation` |
|
|
| - `task_id`, `step_index`, `max_steps` |
| - `visible_ticket` |
| - `current_plan` |
| - `applied_decisions` |
| - `last_action_error` |
| - `progress_score` in `[0, 1]` |
| - `reward_breakdown`: |
| - `correctness_delta` |
| - `policy_bonus` |
| - `repeat_penalty` |
| - `invalid_penalty` |
| - `terminal_bonus` |
|
|
| ## State Space |
|
|
| Typed state model: `B2BSupportTriageState` |
|
|
| - `episode_id`, `step_count` |
| - `task_id`, `seed`, `max_steps` |
| - `cumulative_reward` |
| - `applied_decisions` |
| - `action_history` |
| - `completion_flags` |
|
|
| ## Tasks and Difficulty |
|
|
| Three deterministic fixtures are bundled in `fixtures/tasks.json`. |
|
|
| 1. `easy` |
| - Goal: classify + prioritize one billing ticket. |
| - Grader: category + priority. |
|
|
| 2. `medium` |
| - Goal: classify + prioritize + route + SLA for enterprise billing anomaly. |
| - Grader: category + priority + route + SLA. |
|
|
| 3. `hard` |
| - Goal: security incident triage with escalation and compliant customer response. |
| - Grader: category + priority + route + SLA + escalate + reply policy phrase coverage. |
|
|
| Each grader returns a normalized deterministic score in `[0.0, 1.0]`. |
|
|
| ## Reward Design |
|
|
| Step reward uses shaped components: |
| - positive reward for incremental correctness gains (no double counting) |
| - penalties for invalid actions |
| - penalties for repeated/no-progress loops and contradictory edits |
| - terminal bonus tied to final task score |
|
|
| ## Baseline Inference |
|
|
| `inference.py` is at repository root (hackathon requirement). |
|
|
| It: |
| - uses OpenAI client calls (`OpenAI(base_url=API_BASE_URL, api_key=HF_TOKEN)`) |
| - runs all tasks in fixed order with fixed seeds |
| - emits required structured logs: |
| - `[START] ...` |
| - `[STEP] ...` |
| - `[END] ...` |
|
|
| Environment variables: |
| - `API_BASE_URL` (default: `https://router.huggingface.co/v1`) |
| - `MODEL_NAME` (default: `Qwen/Qwen2.5-72B-Instruct`) |
| - `HF_TOKEN` (required) |
| - `LOCAL_IMAGE_NAME` (default: `b2b_support_triage_env-env:latest`) |
|
|
| Reference deterministic policy target (with valid execution): |
| - easy: `1.000` |
| - medium: `1.000` |
| - hard: `1.000` |
| - aggregate: `1.000` |
|
|
| ## Setup |
|
|
| ```bash |
| python3 -m venv .venv |
| source .venv/bin/activate |
| pip install -e . |
| pip install pytest |
| ``` |
|
|
| ## Local Run |
|
|
| ```bash |
| uvicorn server.app:app --host 0.0.0.0 --port 8000 |
| ``` |
|
|
| ## Validation |
|
|
| ```bash |
| pytest -q |
| openenv validate -v |
| ``` |
|
|
| ## Docker |
|
|
| ```bash |
| docker build -t b2b_support_triage_env-env:latest -f server/Dockerfile . |
| docker run --rm -p 8000:8000 b2b_support_triage_env-env:latest |
| ``` |
|
|
| ## Hugging Face Space Deployment |
|
|
| ```bash |
| openenv push |
| ``` |
|
|
| Then run validator script provided by the hackathon: |
|
|
| ```bash |
| ./validate-submission.sh <your-space-url> . |
| ``` |
|
|
| ## Inlined Documentation |
|
|
| ### local_setup.md |
| |
| # Local Setup and Run Guide |
| |
| This guide explains how to set up, run, test, and execute inference for this project end-to-end on your local machine. |
| |
| ## 1) Prerequisites |
| |
| Install and verify: |
| |
| - Python 3.10+ |
| - Docker Desktop (running) |
| - `openenv` CLI |
| - `curl` |
| |
| Quick checks: |
| |
| ```bash |
| python3 --version |
| docker --version |
| openenv --help >/dev/null && echo "openenv OK" |
| curl --version | head -n 1 |
| ``` |
| |
| ## 2) Project Directory |
| |
| ```bash |
| cd /Users/aksudhak/Documents/Akhil/POC/Scaler/OpenENV |
| ``` |
| |
| ## 3) Optional: Local Python test dependencies |
| |
| If `pytest` is not installed: |
| |
| ```bash |
| python3 -m pip install pytest |
| ``` |
| |
| ## 4) Environment Variables for Inference |
| |
| Set these before running `inference.py`: |
| |
| ```bash |
| export HF_TOKEN="<YOUR_NEW_HF_TOKEN>" |
| export API_BASE_URL="https://router.huggingface.co/v1" |
| export MODEL_NAME="Qwen/Qwen2.5-72B-Instruct" |
| export LOCAL_IMAGE_NAME="b2b_support_triage_env-env:latest" |
| ``` |
| |
| Notes: |
| - `HF_TOKEN` is required. |
| - `API_BASE_URL`, `MODEL_NAME`, and `LOCAL_IMAGE_NAME` have defaults in code, but export them explicitly for submission clarity. |
| |
| ## 5) Build Docker Image |
| |
| ```bash |
| docker build -t b2b_support_triage_env-env:latest -f server/Dockerfile . |
| ``` |
| |
| ## 6) Run Application Locally |
| |
| ```bash |
| docker run --rm -p 8000:8000 b2b_support_triage_env-env:latest |
| ``` |
| |
| Keep this terminal running. Open a second terminal for API checks. |
| |
| ## 7) API Smoke Test |
| |
| ### Health |
| |
| ```bash |
| curl -s http://127.0.0.1:8000/health |
| ``` |
| |
| ### Reset |
| |
| ```bash |
| curl -s -X POST http://127.0.0.1:8000/reset \ |
| -H "Content-Type: application/json" \ |
| -d '{"task_id":"easy","seed":1}' |
| ``` |
| |
| ### Step |
| |
| ```bash |
| curl -s -X POST http://127.0.0.1:8000/step \ |
| -H "Content-Type: application/json" \ |
| -d '{"action":{"action_type":"classify","ticket_id":"T-EASY-1001","payload":{"category":"billing"}}}' |
| ``` |
| |
| ### State |
| |
| ```bash |
| curl -s http://127.0.0.1:8000/state |
| ``` |
| |
| ## 8) Run Unit + Spec Validation |
| |
| ```bash |
| pytest -q |
| openenv validate -v |
| ``` |
| |
| ## 9) Run Inference Directly |
| |
| ```bash |
| python3 inference.py |
| ``` |
| |
| Expected log pattern in stdout: |
| - `[START] ...` |
| - multiple `[STEP] ...` |
| - `[END] ...` |
| - final aggregate score line |
| |
| ## 10) Recommended Helper Scripts |
| |
| ### A) Full local checks (tests + validate + docker + endpoint checks + optional inference) |
| |
| ```bash |
| ./run_all_checks.sh |
| ``` |
| |
| Options: |
| |
| ```bash |
| RUN_INFERENCE=no ./run_all_checks.sh |
| RUN_INFERENCE=yes ./run_all_checks.sh |
| ``` |
| |
| ### B) Inference helper with token prompt |
| |
| If `HF_TOKEN` is missing, this script prompts securely for it. |
|
|
| ```bash |
| ./run_inference.sh |
| ``` |
|
|
| It also sets defaults for: |
| - `API_BASE_URL` |
| - `MODEL_NAME` |
| - `LOCAL_IMAGE_NAME` |
|
|
| ## 11) Common Issues |
|
|
| ### Docker daemon not running |
| Start Docker Desktop and retry. |
|
|
| ### Token error |
| Use a fresh valid Hugging Face token and re-export `HF_TOKEN`. |
|
|
| ### Port 8000 already in use |
| Run container on a different host port: |
|
|
| ```bash |
| docker run --rm -p 8001:8000 b2b_support_triage_env-env:latest |
| ``` |
|
|
| Then use `http://127.0.0.1:8001` in curl commands. |
|
|
| ### project.md |
|
|
| # B2B Support Triage OpenEnv Benchmark - Complete Project Overview |
|
|
| ## 1) What this project is |
|
|
| This project is a complete OpenEnv simulation environment designed for **real-world B2B SaaS support triage**. |
|
|
| Instead of a toy game, it models a workflow support engineers actually perform: |
| - classify an incoming ticket |
| - set urgency/priority |
| - route to the correct team queue |
| - define SLA expectation |
| - (hard task) provide policy-compliant customer communication |
|
|
| It exposes the standard OpenEnv contract: |
| - `reset(...)` |
| - `step(action)` |
| - `state()` |
|
|
| with typed Pydantic models and deterministic graders. |
|
|
| --- |
|
|
| ## 2) Why this environment is useful |
|
|
| ### Practical utility |
| This environment can evaluate agent quality on an enterprise operation where correctness, consistency, and policy compliance matter. |
|
|
| ### Deterministic benchmarking |
| All tasks are fixture-driven and deterministic, so model comparisons are reproducible. |
|
|
| ### Multi-step reasoning |
| Agents must make a sequence of structured decisions, not just one-shot classification. |
|
|
| --- |
|
|
| ## 3) High-level architecture |
|
|
| ### Main components |
|
|
| 1. **Environment runtime** |
| - File: `server/support_triage_environment.py` |
| - Implements environment state machine and reward shaping. |
|
|
| 2. **Typed schemas** |
| - File: `models.py` |
| - Defines action, observation, state, and reward-breakdown models. |
|
|
| 3. **Task fixtures** |
| - File: `fixtures/tasks.json` |
| - Defines easy/medium/hard scenarios, allowed values, answer keys, and policy hints. |
|
|
| 4. **Grader logic** |
| - File: `graders.py` |
| - Computes normalized score in `[0,1]` using weighted criteria per task. |
|
|
| 5. **OpenEnv API server wiring** |
| - File: `server/app.py` |
| - Binds environment + typed models into OpenEnv HTTP/WebSocket endpoints. |
|
|
| 6. **Inference baseline** |
| - File: `inference.py` |
| - Runs all tasks using OpenAI client and emits required structured logs. |
|
|
| 7. **Client adapter** |
| - File: `client.py` |
| - Converts typed Python actions/observations to/from wire payloads. |
|
|
| --- |
|
|
| ## 4) OpenEnv contract and endpoint behavior |
|
|
| ### Manifest |
| - File: `openenv.yaml` |
| - Declares: |
| - `name: b2b_support_triage_env` |
| - `runtime: fastapi` |
| - `app: server.app:app` |
| - `port: 8000` |
|
|
| ### Runtime endpoints |
| - `POST /reset` |
| - `POST /step` |
| - `GET /state` |
| - `GET /health` |
| - `GET /schema` |
|
|
| ### Environment lifecycle |
|
|
| 1. `reset(task_id=..., seed=...)` |
| - loads selected task fixture |
| - clears previous episode state |
| - returns initial observation with visible ticket + dynamic plan hints |
|
|
| 2. `step(action)` |
| - validates action semantics and allowed values |
| - updates decision state |
| - computes score delta and shaped reward |
| - checks done conditions (`submit` or max steps) |
|
|
| 3. `state()` |
| - returns current internal state object (episode info, step count, cumulative reward, history) |
|
|
| --- |
|
|
| ## 5) Data model design (typed) |
|
|
| ## 5.1 Action model |
| File: `models.py` |
|
|
| `B2BSupportTriageAction` fields: |
| - `action_type`: enum |
| - `classify` |
| - `set_priority` |
| - `route` |
| - `draft_reply` |
| - `submit` |
| - `ticket_id`: required for non-`submit` |
| - `payload`: action-specific content |
|
|
| `payload` supports: |
| - `category` |
| - `priority` |
| - `route_queue` |
| - `sla_minutes` |
| - `escalate` |
| - `reply_text` |
|
|
| ### Validation rules |
| - non-submit requires `ticket_id` |
| - required fields per action: |
| - `classify -> category` |
| - `set_priority -> priority` |
| - `route -> route_queue + sla_minutes` |
| - `draft_reply -> reply_text` |
| - `submit -> none` |
|
|
| This enforces structured behavior and prevents ambiguous free-text action control. |
|
|
| ## 5.2 Observation model |
| `B2BSupportTriageObservation` includes: |
| - `task_id`, `step_index`, `max_steps` |
| - `visible_ticket` |
| - `current_plan` (next actions + policy hints) |
| - `applied_decisions` |
| - `last_action_error` |
| - `progress_score` (grader score in `[0,1]`) |
| - `reward_breakdown`: |
| - `correctness_delta` |
| - `policy_bonus` |
| - `repeat_penalty` |
| - `invalid_penalty` |
| - `terminal_bonus` |
| - inherited OpenEnv fields `reward`, `done`, `metadata` |
|
|
| ## 5.3 State model |
| `B2BSupportTriageState` tracks: |
| - episode metadata (`episode_id`, `step_count`, `task_id`, `seed`) |
| - `max_steps` |
| - `cumulative_reward` |
| - `applied_decisions` |
| - `action_history` |
| - `completion_flags` |
|
|
| --- |
|
|
| ## 6) Task system (easy -> medium -> hard) |
|
|
| Defined in `fixtures/tasks.json`. |
|
|
| ## 6.1 Easy |
| Goal: |
| - correct category + priority for one billing ticket |
|
|
| Hidden answer key includes: |
| - `category=billing` |
| - `priority=medium` |
| - plus routing defaults |
|
|
| ## 6.2 Medium |
| Goal: |
| - category + priority + queue routing + SLA |
|
|
| Hidden answer key includes: |
| - `category=billing` |
| - `priority=high` |
| - `route_queue=billing-l2` |
| - `sla_minutes=120` |
|
|
| ## 6.3 Hard |
| Goal: |
| - security incident handling with escalation + compliant customer response |
|
|
| Hidden answer key includes: |
| - `category=security` |
| - `priority=urgent` |
| - `route_queue=security-incident-response` |
| - `sla_minutes=120` |
| - `escalate=true` |
| - required reply phrases |
|
|
| --- |
|
|
| ## 7) Grading mechanics |
|
|
| File: `graders.py` |
|
|
| Each task has deterministic weighted scoring. |
|
|
| ## 7.1 Easy weighting |
| - category: 0.6 |
| - priority: 0.4 |
|
|
| ## 7.2 Medium weighting |
| - category: 0.35 |
| - priority: 0.25 |
| - route_queue: 0.25 |
| - sla_minutes: 0.15 |
|
|
| ## 7.3 Hard weighting |
| - category: 0.15 |
| - priority: 0.15 |
| - route_queue: 0.15 |
| - sla_minutes: 0.10 |
| - escalate: 0.20 |
| - reply_policy phrase coverage: 0.25 |
| |
| The result is clamped and normalized to `[0,1]`. |
| |
| ### Determinism |
| The grader is pure: |
| - no network calls |
| - no random operations |
| - same input decisions always yield same score |
| |
| --- |
| |
| ## 8) Reward shaping logic |
| |
| File: `server/support_triage_environment.py` |
| |
| Reward is step-wise, dense enough for learning signals, and punishes bad behavior. |
| |
| ### Positive components |
| - `correctness_delta`: only when progress score increases |
| - `policy_bonus`: bonus for substantive policy-style draft replies |
| - `terminal_bonus`: `0.2 * final_score` when episode ends |
|
|
| ### Penalty components |
| - `invalid_penalty`: invalid ticket/value/action |
| - `repeat_penalty`: |
| - repeated identical action |
| - contradiction penalties (changing already-set value) |
| - no-progress streak penalties |
|
|
| ### Episode boundaries |
| - done on `submit` |
| - done on max step limit |
| - steps after done receive penalty (`episode_already_done`) |
|
|
| This discourages looping and rewards incremental correctness. |
|
|
| --- |
|
|
| ## 9) Inference baseline design |
|
|
| File: `inference.py` |
|
|
| ## 9.1 Runtime setup |
| Reads: |
| - `HF_TOKEN` (required) |
| - `API_BASE_URL` (default set) |
| - `MODEL_NAME` (default set) |
| - `LOCAL_IMAGE_NAME` (default set) |
|
|
| ## 9.2 Task execution |
| - runs tasks in fixed order: easy, medium, hard |
| - fixed seeds per task for reproducibility |
|
|
| ## 9.3 Action selection strategy |
| - requests JSON action from model (OpenAI client call) |
| - validates/coerces action |
| - falls back to deterministic policy if output invalid or misaligned |
|
|
| This keeps baseline robust and reproducible while still exercising LLM calls. |
|
|
| ## 9.4 Required stdout format |
| Per task episode: |
| - `[START] ...` |
| - multiple `[STEP] ...` |
| - `[END] ...` |
|
|
| Then aggregate score summary line. |
|
|
| --- |
|
|
| ## 10) Deployment and packaging |
|
|
| ## 10.1 Containerization |
| - Dockerfile: `server/Dockerfile` |
| - builds runnable image exposing uvicorn on port 8000 |
| - includes healthcheck on `/health` |
|
|
| ## 10.2 Python packaging |
| - `pyproject.toml` defines project metadata and dependencies |
| - installs environment package editable in image |
|
|
| ## 10.3 OpenEnv compatibility |
| - `openenv validate -v` passes |
| - simulation API surface is compliant |
|
|
| --- |
|
|
| ## 11) Testing strategy |
|
|
| Files under `tests/` verify behavior: |
|
|
| - `test_environment.py` |
| - reset clean state |
| - invalid ticket penalty |
| - hard task can reach full score |
| - max-step termination behavior |
|
|
| - `test_graders.py` |
| - perfect answer gives 1.0 |
| - weighted partial score check |
| - deterministic hard-task reply scoring |
|
|
| - `test_inference_logging.py` |
| - `[START]/[STEP]/[END]` formatting checks |
|
|
| This coverage ensures both task semantics and evaluation fidelity. |
|
|
| --- |
|
|
| ## 12) How to use the project locally |
|
|
| ## 12.1 End-to-end helper |
| ```bash |
| ./run_all_checks.sh |
| ``` |
| Runs tests, validation, docker build/run smoke checks, and optional inference. |
|
|
| ## 12.2 Inference helper (prompts for token if missing) |
| ```bash |
| ./run_inference.sh |
| ``` |
|
|
| ## 12.3 Manual core commands |
| ```bash |
| pytest -q |
| openenv validate -v |
| docker build -t b2b_support_triage_env-env:latest -f server/Dockerfile . |
| docker run --rm -p 8000:8000 b2b_support_triage_env-env:latest |
| ``` |
|
|
| Detailed runbook: `local_setup.md`. |
|
|
| --- |
|
|
| ## 13) Usability summary for your hackathon goal |
|
|
| What this project gives you: |
| - clear real-world domain |
| - structured and typed agent action space |
| - deterministic multi-level tasks |
| - strong reward shaping (partial progress + anti-loop penalties) |
| - reproducible baseline script with required logs |
| - Docker + OpenEnv validation support |
|
|
| This aligns well with your round requirements and helps reduce submission-time surprises. |
|
|
| --- |
|
|
| ## 14) Current limitations and next improvements |
|
|
| ### Current limitations |
| - single-ticket episode scope (per task) rather than multi-ticket queue backlog |
| - deterministic phrase matching for hard reply policy (simple lexical strategy) |
| - baseline includes deterministic fallback policy (useful for reliability, less pure model-only behavior) |
|
|
| ### Next upgrades (optional) |
| - multi-ticket routing episodes with queue-level capacity constraints |
| - richer policy engine with explicit rule graph and violation tags |
| - adversarial or noisy customer phrasing variants per seed |
| - benchmark report exporter with per-criterion analytics |
|
|
| --- |
|
|
| ## 15) Security and publication posture |
|
|
| - fixtures are synthetic |
| - no hardcoded credentials |
| - token usage is environment-variable based only |
| - suitable for external publication, assuming you do not commit local secret files (`.env`, shell history, etc.) |
|
|