multi-agent-lab / docs /architecture /scenario-authoring.md
agharsallah
feat: Implement early termination for the judge in Twenty Sprouts game and enhance event handling
0460922
|
Raw
History Blame Contribute Delete
28.2 kB
# Adding a Scenario on Top of the Core
A scenario is **data, not code**. The engine โ€” ledger, conductor, projections,
routing, memory, tools โ€” never learns the name of your world. You add a world by
dropping YAML into `config/`, and only write Python for the one thing the generic
turn can't express (usually: calling a tool).
This guide walks one complete scenario from empty files to a running app, naming
exactly which of the four stable contracts each step touches, and flagging the one
gotcha โ€” cross-agent visibility โ€” that decides whether your cast actually
collaborates. See `next-steps/architecture-review-and-next-steps.md` ยง3 for the
why.
---
## The mental model: four contracts, three of them YAML
| Contract | File | You touch it byโ€ฆ |
|---|---|---|
| **Event schema** | `src/core/events.py` | *Nothing.* Kinds are open + format-validated (`<ns>.<name>`). You mint `rumor.whispered` in YAML; no engine edit. |
| **Agent manifest** | `src/core/manifest.py` | Writing `config/agents/<name>.yaml`. |
| **Scenario config** | `src/core/config.py` | Writing `config/scenarios/<name>.yaml`. |
| **Tool contract** | `src/tools/registry.py` | A `tools: [...]` grant in a manifest; Python only to register a *new* tool. |
The rule of thumb: **if you're editing a file under `src/core/`, stop โ€” you've
probably left the configurable surface.** The exceptions are deliberate escape
hatches (a handler in `src/agents/handlers.py`, a custom renderer in
`src/ui/render.py`), covered below.
### Two ways agents act each turn
The conductor schedules agents two ways (`src/core/conductor.py`):
1. **Tick** โ€” `schedule.tick_every: N` fires the agent every N turns (`0` = every
turn). Use for a heartbeat producer or a periodic judge.
2. **Subscription** โ€” `subscribes_to: [kind, โ€ฆ]` queues the agent to react when an
event of that kind is appended. Cascades drain within the same tick, so a single
`user.injected` can ripple through a chain in one turn (the Governor bounds the
cascade).
Most casts mix both: a tick-driven producer plus subscription-driven reactors.
### The three scenario shapes you'll build
- **Divergent** (Thousand Token Wood): a lead emits `world.observed`; everyone riffs
on the shared, evolving scene. Collaboration rides `world.observed` (globally
visible) โ€” works today.
- **Convergent** (Mystery Roots): specialists build toward a verdict โ€” clue โ†’
hypothesis โ†’ rebuttal โ†’ judgment. **Requires real cross-agent reads** (see
โš ๏ธ below).
- **Serial / publish-gate** (the worked example here, and Phase 6): producers feed a
judge that holds a publish gate; an artifact ships on approval.
---
## Worked example: ๐Ÿ“ฐ The Rumour Mill
A quiet town. A `rumor-starter` invents gossip; a `gossip` agent embellishes
whatever it just heard; a `fact-checker` consults the oracle tool and rates
plausibility; a `gazette` judge publishes the day's most believed story every five
turns. It exercises **every** extension point: custom kinds, subscription cascades,
a tool grant + handler, salience memory, and the visibility gotcha.
### Step 1 โ€” Decide the shape and cast
| Agent | Role | Profile | Fires by | May emit | Reads (subscribes) |
|---|---|---|---|---|---|
| `rumor-starter` | worker | `tiny` | tick 1 | `rumor.whispered` | โ€” |
| `gossip` | worker | `fast` | subscription | `rumor.whispered` | `rumor.whispered` |
| `fact-checker` | worker | `balanced` | subscription | `fact.checked` | `rumor.whispered` |
| `gazette` | judge | `strong` | tick 5 | `news.published` | โ€” |
Three custom kinds (`rumor.whispered`, `fact.checked`, `news.published`) โ€” none of
which the engine has ever heard of.
```mermaid
flowchart TD
Seed(["genesis: world.observed (the seed)"]) --> RS["rumor-starter ยท tick 1<br/>rumor.whispered"]
RS --> G["gossip<br/>subscribes rumor.whispered<br/>emits rumor.whispered"]
RS --> FC["fact-checker<br/>subscribes rumor.whispered<br/>oracle tool โ†’ fact.checked"]
G -.->|begets more gossip| G
G --> GZ["gazette ยท tick 5<br/>news.published"]
FC --> GZ
```
### Step 2 โ€” Write the agent manifests
`config/agents/rumor-starter.yaml`:
```yaml
name: rumor-starter
role: worker
persona: >
You are the town's first whisper. Invent ONE fresh, specific, harmless rumour
sparked by the current scene. One sentence. Start with 'Did you hearโ€”'.
subscribes_to: []
may_emit:
- rumor.whispered # a brand-new kind, minted here, zero engine edits
schedule:
tick_every: 1
model_profile: tiny # routed to a <=4B model
memory:
window: 6
tools: []
```
`config/agents/gossip.yaml`:
```yaml
name: gossip
role: worker
persona: >
You are the town gossip. Take the LATEST rumour you heard and embellish it with
one wilder, more specific detail. One sentence. Start with 'Actually, I heardโ€”'.
subscribes_to:
- rumor.whispered # woken whenever any rumour is whispered
may_emit:
- rumor.whispered # gossip begets gossip (a subscription cascade)
model_profile: fast
memory:
window: 8
use_salience: true # rank what to remember, not just the most recent
salience_top_k: 6
tools: []
```
`config/agents/fact-checker.yaml`:
```yaml
name: fact-checker
role: worker
handler: fact-checker # custom behaviour: it calls a tool (Step 5)
persona: >
You are the Fact-Checker. Read the latest rumour and the omen the oracle gives
you. Rate the rumour 'plausible' or 'nonsense' in one sentence, citing the omen.
subscribes_to:
- rumor.whispered
may_emit:
- fact.checked
model_profile: balanced
memory:
window: 8
tools:
- oracle # capability grant โ€” the others have none
```
`config/agents/gazette.yaml`:
```yaml
name: gazette
role: judge
persona: >
You are the Daily Gazette. From the rumours and fact-checks so far, publish the
SINGLE most believable, most interesting story as today's headline. One sentence.
Start with 'TODAY'S NEWS:'.
subscribes_to: []
may_emit:
- news.published
schedule:
tick_every: 5 # press runs every five turns
model_profile: strong
memory:
window: 12
use_salience: true
salience_top_k: 8
tools: []
```
### Step 3 โ€” Write the scenario config
`config/scenarios/rumor-mill.yaml`:
```yaml
name: rumor-mill
title: "๐Ÿ“ฐ The Rumour Mill"
goal: >
Turn one quiet whisper into a cascade of rumours, fact-check them against the
omens, and publish the town's most believable story.
default_seed: The mayor's hat was found on the church steeple, brim full of acorns.
example_seeds:
- The mayor's hat was found on the church steeple, brim full of acorns.
- Every dog in town began walking backwards at noon, and no one will say why.
- The baker swears the new apprentice has no reflection, only excellent bread.
genesis_text: "The town square hums with the day's first whisper: {seed}"
cast:
- rumor-starter
- gossip
- fact-checker
- gazette
governor:
max_turns: 2000
max_calls_per_turn: 16
max_total_calls: 20000
```
That's a runnable world. The registry will discover it on next load; `app.py`
auto-lists any scenario in `config/scenarios/` (preferred ones first, then the rest
alphabetically โ€” see `app.py:21`).
### Step 4 โ€” Custom event kinds: there's nothing to do
`rumor.whispered`, `fact.checked`, and `news.published` are valid the instant you
write them: the schema validates the *shape* (`<namespace>.<name>`, lowercase,
dot-separated โ€” `src/core/events.py:33`), and the *authority* to emit is your
manifest's `may_emit`. No enum to extend, no registration. This is ADR-0009.
### Step 5 โ€” A handler (only because `fact-checker` calls a tool)
Three of the four agents are pure YAML + the generic `ManifestAgent`. Only
`fact-checker` needs Python, because it calls the `oracle` tool and folds the result
into both the prompt and the emitted event. Mirror the shipped `FortuneTeller`
pattern (`src/agents/handlers.py`):
```python
# in src/agents/handlers.py
@register_handler("fact-checker")
class FactChecker(ManifestAgent):
"""Consults the oracle, then rates the latest rumour against the omen."""
def _build_extra_prompt(self, projection, recent_events) -> str:
self._omen = ""
if self.tools is not None and "oracle" in self.manifest.tools:
# Capability-checked: raises if the manifest didn't grant `oracle`.
self._omen = self.call_tool("oracle", seed=projection.current_scene).get("omen", "")
if self._omen:
return f"AN OMEN FROM THE ORACLE\n{self._omen}\nUse it to judge the rumour."
return ""
def act(self, run_id, turn, projection, recent_events):
event = super().act(run_id, turn, projection, recent_events)
if getattr(self, "_omen", ""):
event.payload["omen"] = self._omen # tool output is now on the ledger
return event
```
The decorator registers the class under the key `fact-checker`; the manifest's
`handler: fact-checker` binds them. The handler supplies *behaviour only* โ€” every
declarative field still comes from the YAML.
> **Registering a brand-new tool** (instead of reusing `oracle`): add an in-process
> callable to the registry. A tool is a `(name, description, run)` triple where
> `run(**params)` returns a JSON-serialisable dict.
>
> ```python
> # extend src/tools/builtins.py::default_tool_registry()
> registry.register(
> "almanac",
> description="Look up a fact about the town. Params: {topic: str}.",
> run=lambda topic="", **_: {"fact": _ALMANAC.get(topic, "no record")},
> )
> ```
>
> Grant it with `tools: [almanac]` in a manifest. To back it with an MCP server
> instead of an in-process callable, set the MCP env gate (ADR-0017) โ€” the
> capability check runs *first* either way, so transport is never the security
> boundary.
### Step 6 โ€” โš ๏ธ The visibility gotcha (read this or `gossip` talks to itself)
`gossip` and `fact-checker` both `subscribe_to: rumor.whispered` โ€” so the conductor
*wakes* them when a rumour is whispered. **But subscribing does not grant reading.**
An agent's memory is `own events โˆช _GLOBALLY_VISIBLE` (`src/core/memory.py`);
`subscribes_to` is *not* part of it (the conductor only uses it for triggering). So
`gossip`, woken by a custom `rumor.whispered` kind it cannot see, would embellish a void.
**If your agents just talk, there is no gotcha.** `_GLOBALLY_VISIBLE` includes the public
speech kinds โ€” `agent.spoke` and `oracle.spoke` โ€” alongside `world.observed`,
`judge.verdict`, `user.injected`, `run.started`, `agent.reflected` (ADR-0023 follow-up). So
any agent that **speaks** is heard by every peer (recalled across the whole run) and by the
judge (which gets the complete transcript). The gotcha bites only for a **custom,
non-speech kind** like `rumor.whispered`. Two options:
**(a) No engine edit** โ€” emit the shared content on a globally-visible kind. Easiest: have
the producers `may_emit: [agent.spoke]` (now globally visible) so the rumour is ordinary
public speech the whole cast recalls. (Or `world.observed` for narration โ€” the Thousand
Token Wood trick โ€” at the cost that each line overwrites `current_scene`.)
**(b) Make a custom kind first-class** โ€” fold subscriptions into the visibility filter so a
subscribed custom kind is also *readable*:
```python
# src/core/memory.py โ€” pass `reads: frozenset[str]` (the manifest's subscribes_to) into
# EpisodicMemory / SalienceMemory, then at each visibility filter site:
if e.actor == self.agent_name or e.kind in _GLOBALLY_VISIBLE or e.kind in self.reads:
```
โ€ฆwiring `reads=frozenset(self.manifest.subscribes_to)` where `base.py` builds the memory
layers. Then `gossip` recalls the rumour it was woken by. (A scenario-level
`shared_kinds: [...]` could let a whole cast see a custom kind without subscribing.)
**Within a round, peers' lines are also injected directly.** `ContextBuilder` surfaces the
recent table as a **WHAT'S BEEN SAID** block for workers, and the **complete ordered
transcript** (`THE EXCHANGE TO JUDGE`) for judges (`src/core/context.py`, role-aware). A
private `agent.thought` is never shared either way โ€” it rides the mind-reader alone. So:
make a contributor **speak** (`agent.spoke`) if peers/the judge must build on it; reserve
`agent.thought` for genuinely private reasoning (the mind-reader UI).
### Step 7 โ€” Rendering: free for text, custom for shapes
Any event carrying a `text` payload renders on stage **for free** via the projection
fallback (`src/core/projections.py:36`): `rumor.whispered` and `fact.checked` show up
under *Agent Activity*. You only write rendering code when:
- you want a kind in a specific lane โ€” e.g. route `news.published` to *Judge Notes*
by emitting `judge.verdict` instead, or add a branch in `StageProjection.apply`;
- your payload isn't text โ€” an image, a table, an episode card (Phase 6 adds
`render_image` for `image.generated`). Then add a helper in `src/ui/render.py`,
and for a structurally different world a `render_<scenario>_stage()` like the
serial plan. The Rumour Mill needs none of this.
### Step 8 โ€” Validate and test (the modularity invariant)
Prove your world is config, not a fork. Follow `tests/test_modularity.py`:
```python
def test_rumor_mill_runs(tmp_path):
from src.core.conductor import Conductor
from src.core.registry import Registry
from collections import Counter
# Loads YOUR config/ tree unchanged:
registry = Registry.from_dir() # or Registry.from_dir(tmp_path) for a fixture world
scenario = registry.build_scenario("rumor-mill")
c = Conductor(scenario, governor=registry.governor_for("rumor-mill"))
c.reset(scenario.default_seed)
for _ in range(6):
c.step()
kinds = Counter(e.kind for e in c.ledger.events)
assert kinds["rumor.whispered"] >= 1
assert kinds["news.published"] >= 1 # the gazette fired on tick 5
```
Anything malformed fails loudly at load: an unknown agent in `cast`, a bad kind, an
extra field โ€” `validate_world` / `validate_scenario` raise with a precise message
(`src/core/config.py`). That same validation is what lets a UI form or an LLM *emit*
a world and check it before running (ADR-0011).
### Step 9 โ€” Run it
```bash
uv run pytest tests/ -q # your new test + the 260 that must stay green
uv run app.py # ๐Ÿ“ฐ The Rumour Mill now appears in the dropdown
```
No API key needed โ€” the deterministic stub serves every profile. Set
`MODAL_WORKSPACE` (the Modal binding behind `config/models.yaml`) to go live on
the small models you deploy.
---
## Pattern: hidden-role games (๐Ÿ•ต The Steeped)
A fourth shape worth calling out โ€” **social deduction**, where each mind holds
*secret* information and the drama is who can hide it. `config/scenarios/the-steeped.yaml`
is the shipped example: a word-pair bluff where the herd shares a word (COFFEE) and a
lone spy holds another (TEA). It's built entirely on the patterns above, plus two small
techniques:
- **Secret info rides the persona.** The manifest's `persona` is injected verbatim into
every prompt โ€” and *only* that agent's prompt โ€” so it's the natural home for private
state: `spy-nil`'s persona says "you alone hold TEA โ€” never use a tea-only tell like
*steep*," while the herd's personas hold COFFEE. No engine field, no shared leak: the
secret lives where only that mind can read it. The say-vs-think split
(`output_extra_fields: [thought, mood]`) then lets the mind-reader watch the spy
*think* about hiding while it *says* something bland.
- **The reveal is a structured verdict payload.** The unmasking is just a
`judge.verdict` whose payload carries a `reveal: [{agent, secret, role}]` list โ€” the
exact shape the Fishbowl verdict banner renders (`view_model` โ†’ `render_verdict`).
`spy-host` uses a handler (`@register_handler("spy-host")` in `src/agents/handlers.py`)
that runs the generic judge turn for the verdict *text*, then attaches the `reveal` โ€”
the same "decorate the emitted event" move `FortuneTeller` uses for `omen`. Riding the
real ledger, the reveal scrubs and replays like any other event.
### Variant: code-dealt secret, asymmetric roles (โ“ Twenty Sprouts)
The Steeped puts the secret in the *persona*; Twenty Sprouts shows the other end of the
spectrum โ€” the secret is **dealt by code** and the two seats play *different jobs*. It's
the reference for a hidden-word game with a ground-truth answer (`src/agents/twenty_sprouts.py`):
- **Code owns the word; it rides the ledger privately.** `SecretKeeper` deals a word as a
pure function of the seed and stamps it on every keeper event under a `secret` payload
key. Memory/context only ever surface an event's `text` (`_displayable`), so the word is
ground truth on the shared log that the guesser's prompt never sees โ€” the same discipline
as the spy words, but enforced in code rather than trusted to a persona.
- **Asymmetric handlers shape the play.** The keeper is an *answerer* that never leaks: its
handler hands the model the guesser's exact open question, re-asks once if the reply asks a
question or spells the word, and โ€” as an absolute guarantee โ€” scrubs the word from the
spoken line before it ships (a small model *will* sometimes write "a glowing emberโ€ฆ", so
code, not the prompt, is what keeps the secret). The guesser is an *asker* whose handler
mines the ledger into a CONFIRMED / RULED OUT / ALREADY ASKED dossier so it never repeats โ€”
and it is forced to **commit a guess between questions** (when the facts converge, every few
questions once it knows enough, and as a hard stop near the round's end) so it actually
names the word instead of interrogating forever. Two handlers, one shared ledger โ€” neither
agent calls the other.
- **The audience sees the secret; the cast never does.** `view_model_at` lifts the latest
`secret` off the ledger into a `secret` / `secret_holder` field, which the stage core
renders as a dashed "๐Ÿ”’ only you can see" badge. It's a director's-eye peek for the human
watching โ€” it is never placed in any agent's prompt (verified in
`tests/test_twenty_sprouts.py`). Use this field for any future hidden-information run
where the watcher should be in on the secret.
- **Ground truth gets a code-stamped scoreboard.** The scenario declares a
[`competition:` block](../schema/scenario-config.md#competition-who-can-win-and-who-decides)
(`kind: versus`, ADR-0029) naming the teams โ€” `spy: [spy-nil]` vs the herd. The
judge's `winner` field (a well-known typed extra field, validated with one re-ask)
is its *accusation*; the `SpyHost` handler then scores it in code: the herd wins
iff the accused really is the spy, and the verdict payload gains `accused`,
`correct`, and a team-label `winner`. This is the load-bearing split โ€” the model
provides the judgment drama, code stamps the bookkeeping where truth exists.
(Contrast `mystery-roots`, `kind: judged`: no ground truth, so the model's
validated `winner` *is* the result, no handler needed.) Offline, the handler
recovers the accusation from the verdict text, so the no-API-key demo still
produces a full deterministic scoreboard.
Offline, curated lines + a per-role mood bias in `src/models/provider.py` (`_STUB_*`,
keyed by agent name) give the bluff a coherent arc with no API key; live, a real small
model improvises from the personas. Either way the cast never calls each other โ€” they
only post to the shared log, and the seam shows.
---
## Extension-point cheat sheet
> "I want to add a scenario whereโ€ฆ"
| โ€ฆthe goal is | You write | Engine edit? |
|---|---|---|
| a new cast on existing patterns | agent + scenario YAML | none |
| a hidden-role / secret-info game | secret in each `persona`; reveal via a verdict `handler` | none (handler in `agents/`) |
| a scenario that produces a *winner* | a `competition:` block; `winner` in the judge's `output_extra_fields` (+ a scoring `handler` if ground truth exists) | none (ADR-0029) |
| a new event kind | just use it in `may_emit` | none |
| an agent that calls an existing tool | a `handler` + `tools:` grant | none (handler in `agents/`) |
| an agent that calls a *new* tool | register it in `builtins.py` + grant it | tool registration only |
| agents that read each other's custom kinds | **visibility fix** (review #1) | ~5 lines, once, for all scenarios |
| a non-text artifact (image, card) | a `render_*` helper | `src/ui/render.py` only |
| the world to *end* when X happens | terminal condition (review #8) | conductor โ€” not yet supported |
| a different prompt strategy for one agent | override `_build_extra_prompt` in a handler | none (handler) |
| a different prompt strategy for *all* agents | edit `ContextBuilder.build` | `src/core/context.py` (one file, all agents) |
## Common pitfalls
- **Subscribing โ‰  reading (custom kinds only).** An agent woken by a *custom* kind
(e.g. `rumor.whispered`) cannot read it until you make it visible (Step 6). Public
speech (`agent.spoke` / `oracle.spoke`) is globally visible, so talking agents need no
fix.
- **Make contributors speak, not think.** A worker that emits `agent.thought` is invisible
to peers and the judge (thoughts ride the mind-reader alone). If its contribution should
be built on or judged โ€” a clue, a rebuttal, an argument โ€” it must `agent.spoke`. Reserve
`agent.thought` for genuinely private reasoning. (This is exactly what Mystery Roots'
clue-gatherer / devils-advocate got wrong before the ADR-0023 follow-up.)
- **The first `judge.verdict` ends the show.** The Fishbowl autoplay loop pauses the
moment a verdict lands at the head (`FishbowlSession.has_verdict`). So a judge that
`subscribes_to: [world.observed]` fires on the *genesis* event and resolves the run
on turn 1 โ€” the cast never interacts. To let the show breathe, pace the judge with
`schedule.tick_every: N` (and `subscribes_to: []`): it stays silent for N turns of
interaction, then delivers one closing verdict. `N` is your "verdict after X
iterations" knob; keep it below the scenario's `governor.max_turns` so the verdict
lands before the budget trips. The autoplay backstop is derived from that budget
(`max_total_calls`), so a long show is never cut off early. Thousand Token Wood is
the worked example: judge `tick_every: 16`, `max_turns: 20`.
- **End a game the moment it's won, not at the timeout.** A scheduled-only judge always
runs to its tick, even if the contest was decided ten turns earlier โ€” and because it
reads only the *latest* lines, a win that scrolled off (Twenty Sprouts guessed "compass"
then kept guessing) can be scored a miss. For a scenario with a code-known win condition,
make the judge a `JudgedCompetition` subclass, give it `subscribes_to: [agent.spoke]`,
and override `has_early_winner(recent_events)` to return `True` once the game is decided
(scanning *all* of the run, not just the head). The judge then rules the instant the win
lands and **abstains** (its `act()` returns `None` โ€” a real "not yet," not an error) on
every other wake, while keeping its `tick_every` as the no-win timeout. `SproutJudge` is
the worked example.
- **`tick_every: 0` means *every* turn**, not *never*. Use `subscribes_to: []` +
no `tick_every` for an agent that should fire only on subscription.
- **`max_consecutive` does nothing** today (review #9) โ€” don't rely on it to throttle
a chatty agent; tune `tick_every` and budgets instead.
- **Custom kinds need a `text` payload** to render on stage for free. No `text` โ†’
invisible unless you add a renderer.
- **One shared ledger across scenarios** when `DATABASE_URL` is set (`app.py:34`);
use `scripts/resume_run.py` (one DB per scenario) for isolated durable runs.
---
## Scenario authoring checklist: making it arena-grade
A scenario that *runs* is not yet a scenario that *competes*. The arena (ADR-0029,
[arena roadmap](next-steps/arena-roadmap.md) W2โ€“W3) asks every shipped scenario to
declare how it ends and who โ€” if anyone โ€” wins, so the leaderboard can attribute wins
to models. Run down this list before you ship:
- [ ] **Premise** โ€” `goal`, `default_seed`, `example_seeds`, and `genesis_text` are
all set. The goal is rendered into every prompt; the seeds are what the Summon
dropdown offers; the genesis is the audience's opening line.
- [ ] **Cast** โ€” every name in `cast:` resolves to a manifest in `config/agents/`
(`validate_world` fails loudly if not).
- [ ] **Governor** โ€” explicit budgets (`max_turns`, `max_total_calls`, ideally
`max_total_tokens` and `hourly_budget_usd`). The budget is the show's outer wall.
- [ ] **`competition` block** โ€” explicit, even when it's `kind: none`. The model
defaults to `none` so old YAML still validates, but shipped scenarios must say
what they are out loud (the contract test checks).
- [ ] **End condition** โ€” a judge whose `schedule.tick_every` fires *within* the
governor's budget (the first `judge.verdict` drops the curtain โ€” see the pitfall
above), or `kind: none` with a closing voice like the Wood's reckoning.
- [ ] **Structured verdict** โ€” the judge lists `winner` in `output_extra_fields`
and runs `handler: judged-competition` (or a ground-truth subclass), so the
verdict carries `payload.winner` as data, not prose.
- [ ] **Reveal text** โ€” the verdict reads as an *ending*. Hidden-info games attach a
`reveal: [{agent, secret, role}]` payload (the banner renders it); judged games
make the verdict prose name and justify the winner.
### The `competition` block, one example per kind
**`versus` with teams** โ€” asymmetric sides; ground truth decided in code
(`config/scenarios/the-steeped.yaml`, scored by the `spy-host` handler):
```yaml
competition:
kind: versus
teams:
spy: [spy-nil]
herd: [spy-cara, spy-bex, spy-ovo]
```
**`versus` with symmetric seats** ([ADR-0030](../adr/0030-arena-scenarios-and-offline-winners.md)) โ€”
identical manifests, different models; the cleanest "which model plays better" arena
(`config/scenarios/debate-duel.yaml`):
```yaml
competition:
kind: versus
symmetric_seats:
- debater-a
- debater-b
```
**`judged`** โ€” a judge names the winning *agent* via its verdict's `winner` field
(`config/scenarios/mystery-roots.yaml`, `open-table.yaml`):
```yaml
competition:
kind: judged
```
**`none`** โ€” collaborative or showcase; full session and history, no winner, no
leaderboard rows (`config/scenarios/thousand-token-wood.yaml`, `oracle-grove.yaml`):
```yaml
competition:
kind: none
```
The eight shipped scenarios, for orientation: ๐Ÿ•ต The Steeped (versus, teams,
ground-truth `SpyHost`) ยท โ“ Twenty Sprouts (versus, teams, ground-truth
`SproutJudge`) ยท โš”๏ธ Debate Duel (versus, symmetric seats) ยท ๐ŸŽญ Beat Battle (versus,
symmetric seats) ยท ๐Ÿ” Mystery Roots (judged) ยท ๐Ÿ’ฌ Open Table (judged) ยท ๐Ÿ„ Thousand
Token Wood (none) ยท ๐Ÿ”ฎ Oracle Grove (none, tool-use showcase).
### The checklist is enforced, not aspirational
`tests/test_scenario_contract.py` loads every YAML under `config/scenarios/` and
composes a full `WorldConfig` (this matters โ€” the cross-cast competition rules, e.g.
team/seat members must be in the cast and team labels must not collide with agent
names, run on `WorldConfig`, **not** on the per-file `validate_scenario` path). It
then asserts:
- an **explicit** `competition` block is present on every shipped scenario;
- `judged`/`versus` scenarios include a cast member with `role: judge` that emits
`judge.verdict` (a *test-level* check โ€” the schema stays permissive so a partial
world from the Lab still validates);
- every team member and symmetric seat is in the scenario's `cast`;
- symmetric seats are *actually* symmetric โ€” their manifests are identical except
`name`, `hue`, `archetype`, `model_profile`, and `model_endpoint`, so a seat win
measures the model, not a stacked persona.
```bash
uv run pytest tests/test_scenario_contract.py -q # your scenario is arena-grade when this is green
```
If your scenario fails the judge check but genuinely has no winner, the fix is one
line: say `kind: none` and mean it.