--- title: SYNAPSE-X emoji: 🧠 colorFrom: blue colorTo: purple sdk: docker pinned: false app_port: 7860 --- # SYNAPSE-X **SYNAPSE-X** is a reproducible benchmark for real-world operational decision-making under uncertainty, deadlines, and constrained resources. It models the choices an autonomous assistant or ops agent faces in incident response, escalation handling, and moderation triage — not just "what has highest value," but *what should be done now*, with limited time, imperfect information, and failure spillover across dependent tasks. --- ## What Makes SYNAPSE-X Distinct Three coupled mechanisms produce benchmark-relevant, non-linear behaviour: | Mechanism | Role | |-----------|------| | **ECHO** | Forecasts how task risk evolves over time, exposing `future_risk` and `deadline_pressure` so agents can reason ahead | | **PRISM** | Injects seeded Gaussian noise into execution outcomes, preserving reproducibility while preventing trivial deterministic policies | | **CASCADE-X** | Links failures and delays across dependency chains, turning hard-mode into a true order-sensitive planning problem | --- ## Tasks Four structurally different task presets, each with a deterministic grader returning scores in **[0.0, 1.0]**: | Task | Difficulty | Tasks | Description | |------|-----------|-------|-------------| | `easy` | Easy | 3 | Low-risk, generous deadlines — basic execution ordering | | `medium` | Medium | 3 | Moderate risk; partial completion is expected and intentional | | `hard` | Hard | 3 | Dependency chain + CASCADE-X dynamics | | `triage` | Wave-based | 5 | Content moderation; tasks arrive in release waves | --- ## Baseline Scores Measured with `python scripts/evaluate.py` (canonical seed): | Task | Baseline | Random | Difficulty | |------|:--------:|:------:|------------| | Easy | 0.9778 | 0.6779 | Easy | | Medium | 0.6786 | 0.0712 | Medium | | Hard | 0.9882 | 0.6905 | Hard | | Triage | 1.0000 | 0.8400 | Wave-based | --- ## Quick Start ### Install ```bash pip install -r requirements.txt ``` ### Run inference (all tasks) ```bash python inference.py ``` Expected output (baseline-fallback mode): ``` [START] task=easy env=synapse-x model=baseline-fallback [STEP] step=1 action={"action_type":"execute","task_id":0} reward=2.00 done=false error=null [STEP] step=2 action={"action_type":"execute","task_id":1} reward=2.00 done=false error=null [STEP] step=3 action={"action_type":"execute","task_id":2} reward=2.00 done=true error=null [END] success=true steps=3 score=0.978 rewards=2.00,2.00,2.00 [START] task=medium env=synapse-x model=baseline-fallback ... [START] task=hard env=synapse-x model=baseline-fallback ... [START] task=triage env=synapse-x model=baseline-fallback ... ``` ### Run with LLM ```bash export HF_TOKEN=your_token export API_BASE_URL=https://router.huggingface.co/v1 export MODEL_NAME=meta-llama/Meta-Llama-3-8B-Instruct python inference.py ``` ### Start API server ```bash python app.py # or: uvicorn api.app:app --host 0.0.0.0 --port 7860 ``` ### Verify environment ```bash python scripts/verify_openenv.py ``` ### Reproduce benchmark scores ```bash python scripts/evaluate.py ``` --- ## Docker ```bash # Build docker build -t synapse-x . # Run (HF_TOKEN supplied at runtime — never bake it into the image) docker run -p 7860:7860 \ -e HF_TOKEN=your_hf_token \ -e API_BASE_URL=https://router.huggingface.co/v1 \ -e MODEL_NAME=meta-llama/Meta-Llama-3-8B-Instruct \ synapse-x ``` --- ## API | Endpoint | Method | Purpose | |----------|--------|---------| | `/health` | GET | Liveness check | | `/reset` | GET, POST | Start new episode (select task and seed) | | `/step` | POST | Apply one action | | `/state` | GET | Inspect current environment state | | `/tasks` | GET | List all task presets | | `/validate` | GET | Submission readiness check | | `/grade` | POST | Score an action trace | Grade smoke test: ```bash curl -X POST http://127.0.0.1:7860/grade \ -H "Content-Type: application/json" \ -d '{"task_name":"easy","actions":[]}' ``` --- ## Hugging Face Spaces Deployment | Setting | Type | Value | |---------|------|-------| | `HF_TOKEN` | Secret | Your token (required for router inference) | | `API_BASE_URL` | Variable | `https://router.huggingface.co/v1` | | `MODEL_NAME` | Variable | `meta-llama/Meta-Llama-3-8B-Instruct` | | `LOCAL_IMAGE_NAME` | Variable | `synapse-x` | | `PORT` | Variable | `7860` (default, can leave blank) | Deployment checks: - Space status reaches `Running` - Logs show Uvicorn loading `api.app:app` on `0.0.0.0:7860` - `POST /reset` returns JSON with `tasks` list - `/validate` returns task names and endpoint metadata --- ## Observation Space Each `reset()` returns: | Field | Type | Range | Description | |-------|------|-------|-------------| | `tasks` | `list[Task]` | — | Sorted task list | | `time` | `int` | 0–30 | Current timestep | | `resources` | `float` | [0.0, 1.0] | Available resource pool | | `episode_done` | `bool` | — | Terminal flag | Each `Task`: | Field | Type | Description | |-------|------|-------------| | `id` | int | Stable task identifier | | `name` | str | Human-readable label | | `dependencies` | list[int] | Prerequisite task IDs | | `release_time` | int | Wave-arrival step (triage mode) | | `released` | bool | Whether task is actionable now | | `priority` | float [0,1] | Reward importance | | `risk` | float [0,1] | Current failure likelihood | | `uncertainty` | float [0,1] | PRISM uncertainty level | | `deadline` | float | Steps remaining until auto-fail | | `completed` | bool | Completion flag | | `failed` | bool | Failure flag | | `future_risk` | float [0,1] | ECHO forecasted risk | | `deadline_pressure` | float [0,1] | Normalized urgency signal | | `resources_required` | float [0,1] | Resource cost to execute | | `delay_count` | int | Number of delays applied | --- ## Action Space ```json {"action_type": "execute"|"delay"|"reallocate", "task_id": } ``` | Action | Effect | |--------|--------| | `execute` | Attempt completion via PRISM, spend `resources_required` | | `delay` | Defer work, reduce deadline, accumulate delay penalty | | `reallocate` | Recover up to 0.2 resources before next execution | Invalid actions return a small negative reward instead of crashing. --- ## Reward Function Dense, continuous reward clamped to **[-2.0, 2.0]** per step. Terminal completion bonus: **+2.0**. Shared shaping (all actions): - `+0.1 × resources` - `-0.03 × current_time` | Action | Reward formula | |--------|---------------| | `execute` success | `1.2×priority + 0.5×deadline + 0.9×deadline_pressure² - 0.3×uncertainty - 0.4×future_risk + shared` | | `execute` failure | `-0.7×risk - 0.3×uncertainty - 0.4×future_risk + shared` | | `delay` | `-0.15 - 0.2×delay_count + shared` | | `reallocate` | `0.05 + 0.1×resources + shared` | Grader score: `0.5×completion_rate + 0.3×efficiency + 0.2×reward_score` — always in **[0.0, 1.0]**. --- ## ECHO / PRISM / CASCADE-X Math ``` ECHO future_risk = clamp(risk × (1 + 0.1 × time), 0, 1) ECHO deadline_pressure = 1 − clamp(deadline / max_deadline, 0, 1) PRISM effective_risk = clamp(risk + Normal(0, uncertainty × 0.3), 0, 1) PRISM success = rng.random() > effective_risk ``` CASCADE-X (hard mode): ``` delay amplification : risk += 0.03 × delay_count failed dependency : risk += 0.12 × failed_parents pressure feedback : risk += 0.05 × normalized_pressure² phase transition : if normalized_pressure > 0.75 → uncertainty ×= 1.2 ``` --- ## Repository Structure ``` SYNAPSE-X/ ├── inference.py # Official submission entrypoint (runs all 4 tasks) ├── app.py # Local API server launcher ├── openenv.yaml # OpenEnv spec metadata ├── requirements.txt ├── Dockerfile ├── README.md ├── env/ │ ├── environment.py # Core SynapseXEnvironment │ ├── grader.py # Task configs + deterministic grader │ ├── models.py # Pydantic models │ ├── echo.py # ECHO predictor │ ├── prism.py # PRISM uncertainty engine │ └── reward.py # Reward shaping ├── agents/ │ ├── baseline.py # SYNAPSE-CORE-X deterministic agent │ └── random_agent.py # Random comparison agent ├── api/ │ └── app.py # FastAPI application ├── scripts/ │ ├── inference.py # Multi-task benchmarking script │ ├── evaluate.py # Full benchmark runner │ ├── verify_openenv.py # Pre-submission verification │ ├── check_hard_mode.py# Hard-mode strategy comparison │ └── report.py # Multi-run variance report ├── configs/ │ └── config.yaml └── docs/ ├── ARCHITECTURE.md └── RESEARCH.md ``` --- ## Medium Difficulty Note Medium is designed so one of the three tasks typically fails under canonical seed. This is intentional — it models realistic partial-completion scenarios where an agent should secure the safe win first rather than chase a brittle perfect run. A judge seeing 67% completion here should know this is by design, not a bug. ## Why Hard Challenges Frontier Models Hard mode combines four adversarial properties: 1. **Dependency chain** — wrong execution order immediately blocks downstream tasks 2. **CASCADE-X** — upstream failure amplifies risk across all dependents 3. **PRISM uncertainty** — even correct choices can still fail stochastically 4. **Phase transitions** — when normalized pressure exceeds 0.75, uncertainty multiplies by 1.2 This makes hard mode a planning problem, not greedy ranking.