synapse-x / openenv.yaml
Nithin1026's picture
Initial submission Synapse-X
cb330aa
Raw
History Blame Contribute Delete
5.79 kB
name: synapse-x
version: "1.0.0"
description: >
SYNAPSE-X is a reproducible benchmark for real-world operational decision-making
under uncertainty, deadlines, and constrained resources. It models choices an
autonomous agent faces during incident response, escalation handling, and
moderation triage: not just what has highest value, but what should be done now,
with limited time, imperfect information, and failure spillover across dependent tasks.
domain: decision-intelligence
tags:
- openenv
- scheduling
- risk-management
- uncertainty
- resource-allocation
- real-world
- content-moderation
- incident-response
author: "SYNAPSE-X Team"
license: MIT
entry_point: env.environment:SynapseXEnvironment
# ---------------------------------------------------------------------------
# Environment class
# ---------------------------------------------------------------------------
environment:
class: SynapseXEnvironment
module: env.environment
observation:
type: Observation
description: >
A snapshot of all tasks and episode-level state. Tasks are sorted by id.
fields:
tasks:
type: List[Task]
description: Sorted list of tasks with risk, deadlines, ECHO predictions and cascade state
time:
type: int
range: [0, 30]
description: Current episode timestep
resources:
type: float
range: [0.0, 1.0]
description: Available resource pool (normalized)
episode_done:
type: bool
description: True when episode has terminated
action:
type: Action
description: JSON object selecting an action type and target task id
fields:
action_type:
type: str
choices: [execute, delay, reallocate]
description: >
execute — attempt task completion via PRISM;
delay — defer work, reduce deadline, accumulate delay penalty;
reallocate — recover up to 0.2 resources before a future execution
task_id:
type: int
description: ID of the target task (matches Task.id in observation)
reward:
type: float
description: >
Dense, continuous, multi-objective reward. Clamped to [-2.0, 2.0] per step.
Terminal episode bonus of +2.0 when all tasks complete successfully.
Grader score is always in [0.0, 1.0].
range: [-2.0, 2.0]
terminal_bonus: 2.0
# ---------------------------------------------------------------------------
# Tasks with agent graders
# ---------------------------------------------------------------------------
tasks:
- name: easy
description: >
Low-risk tasks with generous deadlines. Tests basic execution ordering
and resource management. Baseline score ~ 0.978.
difficulty: easy
num_tasks: 3
grader_range: [0.0, 1.0]
reward_bounds: {min: -10.0, max: 8.0}
seed: 42
- name: medium
description: >
Moderate risk and tighter deadlines. One task is intentionally designed
to likely fail under canonical seed — models realistic partial-completion
scenarios. Baseline score ~ 0.679.
difficulty: medium
num_tasks: 3
grader_range: [0.0, 1.0]
reward_bounds: {min: -6.0, max: 7.0}
seed: 42
- name: hard
description: >
High uncertainty, tight deadlines, and CASCADE-X dependency propagation.
Tasks form a dependency chain (task 1 depends on task 0; task 2 depends
on both). Failed upstream tasks amplify risk across all dependents via
phase transitions. Requires dependency-aware planning under compounding
pressure. Baseline score ~ 0.988.
difficulty: hard
num_tasks: 3
grader_range: [0.0, 1.0]
reward_bounds: {min: -10.0, max: 7.0}
seed: 19
- name: triage
description: >
Wave-based content moderation with no task dependencies. Five moderation
items (abuse escalations, spam waves, graphic content, VIP appeals) arrive
in three release waves. Tests pure prioritization under time pressure.
Baseline score ~ 1.000.
difficulty: triage
num_tasks: 5
grader_range: [0.0, 1.0]
reward_bounds: {min: -8.0, max: 8.0}
seed: 31
# ---------------------------------------------------------------------------
# API surface
# ---------------------------------------------------------------------------
api:
base_url: "http://localhost:7860"
endpoints:
reset:
method: POST
path: /reset
returns: Observation
step:
method: POST
path: /step
returns: StepResult
state:
method: GET
path: /state
returns: dict
grade:
method: POST
path: /grade
returns: GradeResult
health:
method: GET
path: /health
validate:
method: GET
path: /validate
tasks:
method: GET
path: /tasks
# ---------------------------------------------------------------------------
# Inference script configuration
# ---------------------------------------------------------------------------
inference:
script: inference.py
tasks_evaluated: [easy, medium, hard, triage]
env_vars:
required:
- API_BASE_URL
- MODEL_NAME
optional:
- HF_TOKEN
- API_KEY
- LOCAL_IMAGE_NAME
- SEED
runtime_limit_minutes: 20
compute: vcpu=2, memory=8gb
fallback: baseline-fallback # used when HF_TOKEN is absent
# ---------------------------------------------------------------------------
# Observation and action space summary (for openenv validate)
# ---------------------------------------------------------------------------
observation_space:
tasks: list
time: int
resources: float
episode_done: bool
action_space:
action_type: string # execute | delay | reallocate
task_id: int