Spaces:
Sleeping
Sleeping
| title: SYNAPSE-X | |
| emoji: 🧠 | |
| colorFrom: blue | |
| colorTo: purple | |
| sdk: docker | |
| pinned: false | |
| app_port: 7860 | |
| # SYNAPSE-X | |
| **SYNAPSE-X** is a reproducible benchmark for real-world operational decision-making under uncertainty, deadlines, and constrained resources. It models the choices an autonomous assistant or ops agent faces in incident response, escalation handling, and moderation triage — not just "what has highest value," but *what should be done now*, with limited time, imperfect information, and failure spillover across dependent tasks. | |
| --- | |
| ## What Makes SYNAPSE-X Distinct | |
| Three coupled mechanisms produce benchmark-relevant, non-linear behaviour: | |
| | Mechanism | Role | | |
| |-----------|------| | |
| | **ECHO** | Forecasts how task risk evolves over time, exposing `future_risk` and `deadline_pressure` so agents can reason ahead | | |
| | **PRISM** | Injects seeded Gaussian noise into execution outcomes, preserving reproducibility while preventing trivial deterministic policies | | |
| | **CASCADE-X** | Links failures and delays across dependency chains, turning hard-mode into a true order-sensitive planning problem | | |
| --- | |
| ## Tasks | |
| Four structurally different task presets, each with a deterministic grader returning scores in **[0.0, 1.0]**: | |
| | Task | Difficulty | Tasks | Description | | |
| |------|-----------|-------|-------------| | |
| | `easy` | Easy | 3 | Low-risk, generous deadlines — basic execution ordering | | |
| | `medium` | Medium | 3 | Moderate risk; partial completion is expected and intentional | | |
| | `hard` | Hard | 3 | Dependency chain + CASCADE-X dynamics | | |
| | `triage` | Wave-based | 5 | Content moderation; tasks arrive in release waves | | |
| --- | |
| ## Baseline Scores | |
| Measured with `python scripts/evaluate.py` (canonical seed): | |
| | Task | Baseline | Random | Difficulty | | |
| |------|:--------:|:------:|------------| | |
| | Easy | 0.9778 | 0.6779 | Easy | | |
| | Medium | 0.6786 | 0.0712 | Medium | | |
| | Hard | 0.9882 | 0.6905 | Hard | | |
| | Triage | 1.0000 | 0.8400 | Wave-based | | |
| --- | |
| ## Quick Start | |
| ### Install | |
| ```bash | |
| pip install -r requirements.txt | |
| ``` | |
| ### Run inference (all tasks) | |
| ```bash | |
| python inference.py | |
| ``` | |
| Expected output (baseline-fallback mode): | |
| ``` | |
| [START] task=easy env=synapse-x model=baseline-fallback | |
| [STEP] step=1 action={"action_type":"execute","task_id":0} reward=2.00 done=false error=null | |
| [STEP] step=2 action={"action_type":"execute","task_id":1} reward=2.00 done=false error=null | |
| [STEP] step=3 action={"action_type":"execute","task_id":2} reward=2.00 done=true error=null | |
| [END] success=true steps=3 score=0.978 rewards=2.00,2.00,2.00 | |
| [START] task=medium env=synapse-x model=baseline-fallback | |
| ... | |
| [START] task=hard env=synapse-x model=baseline-fallback | |
| ... | |
| [START] task=triage env=synapse-x model=baseline-fallback | |
| ... | |
| ``` | |
| ### Run with LLM | |
| ```bash | |
| export HF_TOKEN=your_token | |
| export API_BASE_URL=https://router.huggingface.co/v1 | |
| export MODEL_NAME=meta-llama/Meta-Llama-3-8B-Instruct | |
| python inference.py | |
| ``` | |
| ### Start API server | |
| ```bash | |
| python app.py | |
| # or: uvicorn api.app:app --host 0.0.0.0 --port 7860 | |
| ``` | |
| ### Verify environment | |
| ```bash | |
| python scripts/verify_openenv.py | |
| ``` | |
| ### Reproduce benchmark scores | |
| ```bash | |
| python scripts/evaluate.py | |
| ``` | |
| --- | |
| ## Docker | |
| ```bash | |
| # Build | |
| docker build -t synapse-x . | |
| # Run (HF_TOKEN supplied at runtime — never bake it into the image) | |
| docker run -p 7860:7860 \ | |
| -e HF_TOKEN=your_hf_token \ | |
| -e API_BASE_URL=https://router.huggingface.co/v1 \ | |
| -e MODEL_NAME=meta-llama/Meta-Llama-3-8B-Instruct \ | |
| synapse-x | |
| ``` | |
| --- | |
| ## API | |
| | Endpoint | Method | Purpose | | |
| |----------|--------|---------| | |
| | `/health` | GET | Liveness check | | |
| | `/reset` | GET, POST | Start new episode (select task and seed) | | |
| | `/step` | POST | Apply one action | | |
| | `/state` | GET | Inspect current environment state | | |
| | `/tasks` | GET | List all task presets | | |
| | `/validate` | GET | Submission readiness check | | |
| | `/grade` | POST | Score an action trace | | |
| Grade smoke test: | |
| ```bash | |
| curl -X POST http://127.0.0.1:7860/grade \ | |
| -H "Content-Type: application/json" \ | |
| -d '{"task_name":"easy","actions":[]}' | |
| ``` | |
| --- | |
| ## Hugging Face Spaces Deployment | |
| | Setting | Type | Value | | |
| |---------|------|-------| | |
| | `HF_TOKEN` | Secret | Your token (required for router inference) | | |
| | `API_BASE_URL` | Variable | `https://router.huggingface.co/v1` | | |
| | `MODEL_NAME` | Variable | `meta-llama/Meta-Llama-3-8B-Instruct` | | |
| | `LOCAL_IMAGE_NAME` | Variable | `synapse-x` | | |
| | `PORT` | Variable | `7860` (default, can leave blank) | | |
| Deployment checks: | |
| - Space status reaches `Running` | |
| - Logs show Uvicorn loading `api.app:app` on `0.0.0.0:7860` | |
| - `POST /reset` returns JSON with `tasks` list | |
| - `/validate` returns task names and endpoint metadata | |
| --- | |
| ## Observation Space | |
| Each `reset()` returns: | |
| | Field | Type | Range | Description | | |
| |-------|------|-------|-------------| | |
| | `tasks` | `list[Task]` | — | Sorted task list | | |
| | `time` | `int` | 0–30 | Current timestep | | |
| | `resources` | `float` | [0.0, 1.0] | Available resource pool | | |
| | `episode_done` | `bool` | — | Terminal flag | | |
| Each `Task`: | |
| | Field | Type | Description | | |
| |-------|------|-------------| | |
| | `id` | int | Stable task identifier | | |
| | `name` | str | Human-readable label | | |
| | `dependencies` | list[int] | Prerequisite task IDs | | |
| | `release_time` | int | Wave-arrival step (triage mode) | | |
| | `released` | bool | Whether task is actionable now | | |
| | `priority` | float [0,1] | Reward importance | | |
| | `risk` | float [0,1] | Current failure likelihood | | |
| | `uncertainty` | float [0,1] | PRISM uncertainty level | | |
| | `deadline` | float | Steps remaining until auto-fail | | |
| | `completed` | bool | Completion flag | | |
| | `failed` | bool | Failure flag | | |
| | `future_risk` | float [0,1] | ECHO forecasted risk | | |
| | `deadline_pressure` | float [0,1] | Normalized urgency signal | | |
| | `resources_required` | float [0,1] | Resource cost to execute | | |
| | `delay_count` | int | Number of delays applied | | |
| --- | |
| ## Action Space | |
| ```json | |
| {"action_type": "execute"|"delay"|"reallocate", "task_id": <int>} | |
| ``` | |
| | Action | Effect | | |
| |--------|--------| | |
| | `execute` | Attempt completion via PRISM, spend `resources_required` | | |
| | `delay` | Defer work, reduce deadline, accumulate delay penalty | | |
| | `reallocate` | Recover up to 0.2 resources before next execution | | |
| Invalid actions return a small negative reward instead of crashing. | |
| --- | |
| ## Reward Function | |
| Dense, continuous reward clamped to **[-2.0, 2.0]** per step. Terminal completion bonus: **+2.0**. | |
| Shared shaping (all actions): | |
| - `+0.1 × resources` | |
| - `-0.03 × current_time` | |
| | Action | Reward formula | | |
| |--------|---------------| | |
| | `execute` success | `1.2×priority + 0.5×deadline + 0.9×deadline_pressure² - 0.3×uncertainty - 0.4×future_risk + shared` | | |
| | `execute` failure | `-0.7×risk - 0.3×uncertainty - 0.4×future_risk + shared` | | |
| | `delay` | `-0.15 - 0.2×delay_count + shared` | | |
| | `reallocate` | `0.05 + 0.1×resources + shared` | | |
| Grader score: `0.5×completion_rate + 0.3×efficiency + 0.2×reward_score` — always in **[0.0, 1.0]**. | |
| --- | |
| ## ECHO / PRISM / CASCADE-X Math | |
| ``` | |
| ECHO future_risk = clamp(risk × (1 + 0.1 × time), 0, 1) | |
| ECHO deadline_pressure = 1 − clamp(deadline / max_deadline, 0, 1) | |
| PRISM effective_risk = clamp(risk + Normal(0, uncertainty × 0.3), 0, 1) | |
| PRISM success = rng.random() > effective_risk | |
| ``` | |
| CASCADE-X (hard mode): | |
| ``` | |
| delay amplification : risk += 0.03 × delay_count | |
| failed dependency : risk += 0.12 × failed_parents | |
| pressure feedback : risk += 0.05 × normalized_pressure² | |
| phase transition : if normalized_pressure > 0.75 → uncertainty ×= 1.2 | |
| ``` | |
| --- | |
| ## Repository Structure | |
| ``` | |
| SYNAPSE-X/ | |
| ├── inference.py # Official submission entrypoint (runs all 4 tasks) | |
| ├── app.py # Local API server launcher | |
| ├── openenv.yaml # OpenEnv spec metadata | |
| ├── requirements.txt | |
| ├── Dockerfile | |
| ├── README.md | |
| ├── env/ | |
| │ ├── environment.py # Core SynapseXEnvironment | |
| │ ├── grader.py # Task configs + deterministic grader | |
| │ ├── models.py # Pydantic models | |
| │ ├── echo.py # ECHO predictor | |
| │ ├── prism.py # PRISM uncertainty engine | |
| │ └── reward.py # Reward shaping | |
| ├── agents/ | |
| │ ├── baseline.py # SYNAPSE-CORE-X deterministic agent | |
| │ └── random_agent.py # Random comparison agent | |
| ├── api/ | |
| │ └── app.py # FastAPI application | |
| ├── scripts/ | |
| │ ├── inference.py # Multi-task benchmarking script | |
| │ ├── evaluate.py # Full benchmark runner | |
| │ ├── verify_openenv.py # Pre-submission verification | |
| │ ├── check_hard_mode.py# Hard-mode strategy comparison | |
| │ └── report.py # Multi-run variance report | |
| ├── configs/ | |
| │ └── config.yaml | |
| └── docs/ | |
| ├── ARCHITECTURE.md | |
| └── RESEARCH.md | |
| ``` | |
| --- | |
| ## Medium Difficulty Note | |
| Medium is designed so one of the three tasks typically fails under canonical seed. This is intentional — it models realistic partial-completion scenarios where an agent should secure the safe win first rather than chase a brittle perfect run. A judge seeing 67% completion here should know this is by design, not a bug. | |
| ## Why Hard Challenges Frontier Models | |
| Hard mode combines four adversarial properties: | |
| 1. **Dependency chain** — wrong execution order immediately blocks downstream tasks | |
| 2. **CASCADE-X** — upstream failure amplifies risk across all dependents | |
| 3. **PRISM uncertainty** — even correct choices can still fail stochastically | |
| 4. **Phase transitions** — when normalized pressure exceeds 0.75, uncertainty multiplies by 1.2 | |
| This makes hard mode a planning problem, not greedy ranking. | |