Spaces:
Running
A newer version of the Gradio SDK is available: 6.22.0
Scenario Uniqueness Audit (Phase 2)
Read-only audit of openra_bench/scenarios/packs/*.yaml. Scope: 212 YAML
files (excluding TEMPLATE.yaml); 23 in-flight econ-* packs were
deliberately skipped (other agents own them), leaving 189 packs probed.
The probe was a stall policy โ Command.observe() only โ run on the
hard tier, seed=1, of every pack, through the same code path
(openra_bench.eval_core.run_level) the bench evaluator uses. Each
result is the engine-evaluated outcome (win / loss / draw),
final game tick, and the static profile read straight from the YAML.
Headline numbers
- 181 / 189 packs (95.8%) โ stall LOSSES (healthy: no-cheat bar holds).
- 3 packs โ stall WINS (the no-cheat bar is broken).
- 3 packs โ stall DRAWS (no real LOSS reachable from a stall play).
- 2 packs โ Rust engine panic at
reset(seed=1)(bothadversarial-siege/adversarial-skirmishโ alreadystatus: quarantinein-pack so they are NOT in the default eval set, but they remain on disk and panic when explicitly invoked).
Quarantined packs are still YAML-discoverable and were probed. They are
flagged [Q] in the lists below.
Section 1 โ Defect packs
A "defect" here means stall WINS (the lazy bar fell) or stall DRAWS
(no real reachable LOSS โ the timeout collapses to draw degeneracy
because the fail predicate never trips). Both are explicit violations
of the "no defect, no cheat" bar in CLAUDE.md.
1A. Stall WINS (3 โ bar fell)
| Pack | Outcome | Turn | Tick | Diagnosis |
|---|---|---|---|---|
mid-economy-under-fire |
WIN | 11 | 993 | Stall achieves the win predicate without any agent action. Hard win-clause is economy_value_gte:4000 AND harvโฅ2 AND units_lost_lte:2. With the current engine, the 3 starter harvesters auto-harvest (no harvest command needed โ they begin in harvest mode), the perimeter 1tnks auto-fire and kill the lone raider 1tnk, and EV trips 4000 by tick ~993 with zero losses. The pack header says "Stall (only observe)โฆ harvs never harvest โ EV stays at 0 โ LOSS." That is no longer true under the current engine โ harvs in harvest mode work without re-issued orders. Fix: either remove the harvesters from the starting placement (require the model to issue Command.harvest) or replace the 1tnk raider with a stronger raider wave that out-attritions an idle defense ring. |
combat-naval-shore-strike |
WIN | 3 | 198 | Stall wins inside 3 turns. The hard tier places two destroyers in the water channel on stance:2 (Defend), and the entire shore garrison on stance:0 (HoldFire). The destroyers auto-fire on every garrison unit in range and clear all 5 enemies in <200 ticks; the agent never sends a command. Fix: set the destroyers to stance:0 so the agent must issue an attack_unit order; or place the shore garrison just outside the destroyer's auto-target range so the agent must manually fire. The whole capability (target an across-shore target with a ranged unit) collapses because the engine fires automatically. |
def-with-ambush |
(intended) WIN | 24 | 1233 | Not a defect โ exempt by design. CLAUDE.md ยงTriage coverage: "1 pack (def-with-ambush) is exempt by design (positional-discipline scenario where do-nothing IS the intended policy)." The capability under test is hold the ambush position; the four stance:2 flanker tanks auto-engage the rusher band when it enters weapon range. Listed here only for completeness โ no action required. |
1B. Stall DRAWS (3 โ no LOSS reachable)
A draw means the win predicate was never met AND the fail predicate
never tripped, so the level resolves to draw (outcome score 0.5).
This is a defect because it makes the lazy play indistinguishable
from a partial / borderline play, and inflates the lazy score by
treating it as half a win.
| Pack | Outcome | Turn | Tick | Diagnosis |
|---|---|---|---|---|
economy-time-box [Q] |
DRAW | 80 | 7203 | Already quarantined in-file (reason: "redundant with economy-force-buildup"). The hard tier has no fail_condition at all (fail_condition: None), so a 12000-tick budget that never produces 6 units / 5 buildings simply times out as draw. This is one of the canonical CLAUDE.md defect classes ("no fail_condition, or only triggers on full force-wipe; a stall / preserve / partial outcome silently draws"). Since the pack is quarantined the lazy-bar drift is harmless to the default eval; if the pack is ever un-quarantined, add fail_condition: {after_ticks: 12001} plus {not: own_units_gte: 1}. |
spec-spy-infiltrate |
DRAW | 3 | 273 | Hard tier fail_condition: {after_ticks: 4001} โ i.e. the only failure mode is the deadline. But the agent's only combat-relevant units are 2 spies (spy) and the scenario also places one defender e1. With stall, the spies are passive while the engine auto-dones at turn 3 (likely because the defender e1 falls to auto-fire or because all agent combat units are destroyed). The outcome lands as DRAW because the spy-only force is wiped before the within_ticks ever bites. Fix: add {not: own_units_gte: 1} (or unit_type_count_gte:{spy, 1}) to the fail_condition.any_of so a spy-wipe is a LOSS, not a draw. Same fix likely needed for easy / medium (uninspected; consider auditing). |
def-bridge-chokepoint |
(mixed: usually LOSS at tick 1623, occasionally DRAW in one fresh process) | 18 | 1623 | The first probe (subprocess #1) reported DRAW; three subsequent re-runs from a fresh process consistently reported LOSS (17 losses, fail-clause not has_building:fact trips when the procs/fact fall under the rifle assault). This is borderline โ the pack is probably healthy, but the one-off DRAW observation suggests a possible determinism-edge case (race between the engine auto-done and the predicate evaluator). Recommend a 4-seed re-probe (1โ4) over a clean Python process to confirm. If the LOSS is stable, no fix required; if the DRAW resurfaces, add {not: own_units_gte: 1} to the fail clause. |
1C. Engine panics at reset (2)
| Pack | Status | Error |
|---|---|---|
adversarial-siege |
quarantine (consolidated into adversarial-duel) |
pathfinder.rs:175 index out of bounds: len 5120, index 5382 at env.reset(seed=1). The hard tier places actors outside the playable bounds (CLAUDE.md footgun #6) for the chosen spawn. Pack is excluded from default eval; no eval-time impact, but flag the file is still selectable via explicit --packs. |
adversarial-skirmish |
quarantine (consolidated into adversarial-duel) |
Same panic at env.reset. Same situation as above. |
Recommend either deleting these two YAMLs (the consolidation is the
documented future state) or moving them to packs/_archived/ so
discover_packs skips them (it already skips _* and TEMPLATE).
Section 2 โ Duplicate / near-duplicate clusters
Methodology: each pack was fingerprinted by
(capability, tools, actor_types_sorted, win_predicate_keys_sorted).
Five strict (= all four fields equal) clusters surfaced; a second pass
relaxed tools and actor_types to surface near-duplicates with the
same capability + same win-key set, which surfaced a further ~10
clusters. Each cluster is judged on the per-pack real_world_meaning
and the level_description of the hard tier โ if these tell the same
story, the cluster is genuinely redundant; if not, the predicate-key
match is coincidental and the packs probe distinct skills.
2A. Strict duplicates (action required)
adversarial-siege+adversarial-skirmishโ both already quarantined as "consolidated intoadversarial-duel". Keep:adversarial-duel. Action: physical removal (delete or_archive) โ they are the only two packs that crash the engine at load, so removing them eliminates the crash class entirely.artofwar-decoy-sacrifice+artofwar-lure-the-tigerโ bothcapability: reasoning, both use botguard, identical hard-tier actor list[2tnk,e1,e3,fact,jeep], identical win-keys{units_in_region_gte, units_lost_lte, within_ticks}, both described as "main-force reaches an objective by diverting a leashed defender." Keep:artofwar-lure-the-tiger(the more doctrinally complete framing โ leash mechanic; "the strong defender" is the load-bearing pull). Merge / delete:artofwar-decoy-sacrificeโ its "spend a decoy unit" angle is covered by theunits_lost_lteclause in the lure pack, which already permits a small attrition budget the model can spend on the bait sub-force.scout-cycle-keep-info-fresh+scout-track-enemy-movementโ bothcapability: perception, identical hard-tier actor list, both use thethen-chainedunits_killed_gtewin predicate, both probe "scout multiple times because the world changes between observations." Both use thescheduled_events:hook (cycle = reinforcement waves; track = enemy march legs). These are genuinely distinct skills โ cycle is "re-observe a stale region for new arrivals", track is "follow a moving target across the map." Keep both, but verify the briefing of each clearly distinguishes the cause of staleness so the model is not asked to solve cycle-by-track-strategy or vice-versa. (No action required if the briefings are unambiguous, which they appear to be on inspection.)build-defensive-skirt-corners+build-defensive-tower-line(strict 4-field match) AND the wider quartetbuild-defensive-tower-cluster+def-in-depth-vs-single(same capability + win-key set, near-identical actor list, all userusherbot). These four packs all probe "where do you place a finite pillbox budget to survive an incoming rush?" with topology variants: skirt (one in each map-relative corner of the building), cluster (tight wrap), line (across the choke), depth (two thinner bands). Each topology is genuinely distinct in the optimal answer, so all four packs probe a real choice axis. Keep all four but consider promoting them into a single named "Defense Topology Suite" in the eval-cell catalog so the four cells score together (one "topology-IQ" metric) rather than as four unrelated reasoning packs. (No file action required; this is a catalog / metric grouping recommendation.)
2B. Near-duplicates by win-predicate keys (informational โ same
predicate idiom is used across distinct scenarios; verify each is load-bearing)
These groups all share an identical (capability, win_predicate_keys)
fingerprint. Each group below is not necessarily a duplicate, but
flags that all members rely on the same predicate idiom and the
scenarios should diverge on actor composition / pressure / bot to
remain discriminative.
(reasoning, then+within_ticks)โ N=10:build-power-online-first,build-sequence-tech-cheapest,build-sequence-tech-fastest,build-sequence-tech-most-resilient(not in cluster โ was filtered earlier),lh-econ-army-victory,lh-opening-to-defense-to-counter,lh-opening-to-tech-to-army,lh-progression-stage-locked,rob-objective-change-midway,tech-balanced-econ-then-tech,lh-build-army-coordinate-multifront-attack. These are all ordered-multi-phase scenarios. Distinct intents (build-order vs long-horizon vs objective-shift); thethenpredicate is the common engine machinery, not a duplication signal.(reasoning, building_count_gte+building_in_region+within_ticks)โ N=7:build-sell-and-rebuild-elsewhere,def-in-depth,def-tower-line-vs-cluster,expansion-aggro-3-base-greedy,mcv-deploy-second-base,mcv-deploy-third-base,mfb-mirror-base-east-west. All probe "build / re-build N buildings in a target region." Genuinely distinct intents (expansion-greed, mirror-base symmetry, MCV-redeploy under pressure); keep all.(reasoning, building_count_gte+units_killed_gte+within_ticks)โ N=6:adv-rps-counter-pick,build-tech-skip-decision,def-counter-battery,def-walls-vs-towers,def-while-building,lh-tech-rush-vs-army-rush. All "kill K enemies and have N buildings standing." Reads as honest differentiated tests.(action, own_units_gte+units_killed_gte+within_ticks)โ N=5:combat-flanking-attack,combat-harass-aggro-commit,combat-kite-and-pull,combat-kite-jeep-vs-tank,combat-tanya-vs-rush. Five combat micro packs โ kite, flank, harass. Each probes a different micro idiom; keep all but verify that the win bar (units_killed_gte) is tuned per pack so the intended micro is load-bearing, not just "auto-fire wins." Two are explicitly kite (combat-kite-and-pull,combat-kite-jeep-vs-tank) โ possible duplicate: kite-and-pull uses generic 2tnk/3tnk vs jeep-vs-tank's unit-class asymmetry. Keep both only ifcombat-kite-jeep-vs-tankactually requires the jeep speed advantage thatcombat-kite-and-pulldoesn't (read the briefings; they look distinct).(action, reach_region+units_lost_lte+within_ticks)โ N=4:proc-no-attack-passive-only,proc-tool-use-multi-distractor,proc-tool-use-with-distractor,strict-toolban-fidelity-under-pressure. All procedural-compliance packs โ "reach the goal without using a forbidden tool / without attacking." Theproc-tool-use-*pair is a possible duplicate: "with-distractor" (one distractor tool) vs "multi-distractor" (multiple). Both are the same skill at different scales โ recommend collapsing into a single pack with three tiers (no distractor / one / many) rather than two top-level packs.(perception, buildings_discovered_gte+units_lost_lte+within_ticks)โ N=4:perception-target-vs-fog,scout-detect-enemy-tech,scout-discover-hidden-base,scout-multiple-fog-areas. All "discover K enemy buildings under fog." Each probes a distinct discrimination (target-vs-fog = ignore decoys; detect-enemy-tech = read the tech-tree; hidden-base = single hidden compound; multiple-fog-areas = K disjoint regions). Keep all.
2C. Overlapping prefixes worth a catalog review (not duplicates,
but the area is dense)
combat-*(25 packs),def-*(18),scout-*(14),build-*(14),lh-*(13),proc-*(10),mfb-*(8),rob-*(8),tp-*(7),coord-*(7),mcv-*(6),spec-*(5),economy-*(5),tech-*(4),strategy-*(4),perception-*(4),artofwar-*(4),tempo-*(2),maint-*(2),mid-*(2),expansion-*(3),coordination-*(2),harass-*(1),risk-*(1),power-*(1),navigation-*(1),defense-*(1),longhorizon-*(1),custom-*(1),building-*(1),rush-*(1),reasoning-*(2),action-*(2),adv-*(2),adversarial-*(3),strict-*(3). The combat / def / scout / build axes carry ~70 packs combined; a follow-up audit explicitly checking pairwise overlap inside each prefix (briefing-level read) would surface ~5โ10 more candidate collapses.
Section 3 โ Capability coverage matrix
Cross-referenced against the phase ร decision-type matrix sketched
in PAPER_PLAN.md ยง12.2 and CLAUDE.md. Each cell lists the packs
that primarily measure it (a pack may legitimately serve more than one
cell โ only the dominant assignment is listed; cross-cutting packs are
flagged).
Opening
| Decision | Packs | Density |
|---|---|---|
| MCV deploy site (where to plant) | mcv-deploy-and-build, mcv-deploy-defensible-site, mcv-deploy-near-resource, mcv-deploy-relocate-under-pressure, mcv-deploy-second-base, mcv-deploy-third-base |
6 packs (dense; good) |
| Build-order commit | build-sequence-tech-cheapest, build-sequence-tech-fastest, build-sequence-tech-most-resilient, build-tech-skip-decision, tech-balanced-econ-then-tech, tech-aggro-all-in, tech-turtle-defensive-tech, tech-production-planning, build-power-online-first, power-budget-online, building-and-planning |
11 packs (dense; good) |
| Defense-direction commit (which side to anticipate) | def-position-expected-direction, def-position-revealed-direction, def-multi-direction, def-surprise-flank-react, def-pre-position-mobile-reserve |
5 packs (good) |
| Rush-defense | defense-rush-survive, rush-hour, build-defensive-tower-cluster, build-defensive-tower-line, build-defensive-skirt-corners, def-in-depth, def-in-depth-vs-single, def-walls-vs-towers, def-tower-line-vs-cluster, def-while-building |
10 packs (dense) |
Early-mid
| Decision | Packs | Density |
|---|---|---|
| Harass / harass-preserve | harass-response-preserve, combat-harass-aggro-commit, combat-harass-balanced-hit-and-run, combat-skirmish-then-disengage, combat-bait-counter-attack, combat-kite-and-pull, combat-kite-jeep-vs-tank, combat-retreat-after-engagement, combat-suicide-charge-mission |
9 packs (dense) |
| Exact-count perception | perception-count-the-threat, perception-count-the-threat-small-k, scout-count-defenders, scout-detect-incoming-army |
4 packs (adequate) |
| Scout-direction commit | scout-detect-base-direction, scout-far-frontier, scout-frontier-reading (=perception-frontier-reading), scout-multiple-fog-areas, scout-and-report, scout-and-survive, scout-discover-hidden-base, scout-detect-enemy-tech, scout-map-reveal-percent-target, reasoning-frontier-commit |
10 packs (dense; possible over-coverage) |
Mid
| Decision | Packs | Density |
|---|---|---|
| Live-economy defense | mid-economy-under-fire (DEFECT โ stall wins), econ-harvester-defense-raid [skipped โ econ-*], econ-protect-harvester-route [skipped] |
~3 packs, 1 defective โ the only non-econ exemplar is broken. Gap when the in-flight econ- land โ verify they cover.* |
| Tech-switch on scout | mid-tech-switch-on-scout, lh-scout-react-counter, adv-rps-counter-pick |
3 packs (adequate) |
| Second front | mfb-base-1-defend-base-2-build, mfb-supply-line-link-between-bases, mfb-mirror-base-east-west, mfb-two-base-simultaneous, mfb-third-base-against-clock, expansion-balanced-2-base-defended, expansion-aggro-3-base-greedy, coord-diversionary-attack |
8 packs (dense) |
| Replan after loss | rob-unit-loss-recovery, rob-partial-base-loss-continue, lh-recovery-after-mid-game-loss, build-engineer-rebuild-after-loss, def-retreat-and-rebuild, build-sell-and-rebuild-elsewhere |
6 packs (good) |
Mid-late
| Decision | Packs | Density |
|---|---|---|
| Concede vs hold | mid-concede-vs-hold |
1 pack (THIN โ single exemplar) |
| Isolate vs split | combat-divide-and-conquer, combat-pincer-coordination, combat-prevent-retreat, combat-formation-tank-wedge |
4 packs (adequate) |
| Tempo double-window | tempo-double-window, tempo-strike-window, coordination-staggered-window, tp-survive-and-strike-at-window |
4 packs (adequate) |
| Decoy / lure / feint | artofwar-decoy-sacrifice, artofwar-lure-the-tiger, artofwar-indirect-approach, artofwar-sequenced-citadel, combat-bait-counter-attack, coord-diversionary-attack, def-with-ambush |
7 packs (good) |
| Multi-front | mfb-rotating-production-pressure, mfb-redundant-tech-buildings, mfb-tech-base-vs-economy-base, rob-multiple-simultaneous-pressures |
4 packs (adequate) |
Late
| Decision | Packs | Density |
|---|---|---|
| Sustained multi-front | lh-build-army-coordinate-multifront-attack, mfb-rotating-production-pressure |
2 packs (thin) |
| Base-trade race | None directly | 0 packs (GAP) |
| Counter-strategy read | adv-rps-counter-pick, combat-attack-from-behind-fog, lh-tech-rush-vs-army-rush |
3 packs (adequate) |
| Superweapon timing | spec-nuke-strike, spec-tanya-c4-strike, spec-engineer-capture, spec-spy-infiltrate (DEFECT โ stall draws), spec-thief-steal-cash |
5 packs, 1 defective (good once spec-spy is fixed) |
| Credit-only final phase | lh-credit-only-final-phase |
1 pack (thin) |
Cross-cutting
| Capability | Packs | Density |
|---|---|---|
| Procedural compliance under pressure | proc-checklist-no-deviation, proc-conditional-branch-action, proc-instruction-following-edge-case, proc-no-attack-passive-only, proc-only-build-no-combat, proc-only-defend-no-attack, proc-ordered-action-strict, proc-strict-toolban-fidelity, proc-tool-use-multi-distractor, proc-tool-use-with-distractor, strict-production-bom, strict-sequence, strict-toolban-fidelity-under-pressure, tp-pressure-procedural |
14 packs (very dense) |
| Long-horizon multi-phase | lh-* (13 packs), longhorizon-opening-to-assault, lh-100-turn-marathon-survival |
14 packs (dense) |
| Coordination across squads | coord-converge-on-target, coord-cover-and-move, coord-diversionary-attack, coord-mutual-support, coord-relay-attack, coord-relay-vision-chain, coord-squad-handoff, coordination-ordered-rendezvous, coordination-staggered-window, action-multiunit-coordination, action-sequenced-execution, combat-pincer-coordination, combat-heli-flank |
13 packs (dense) |
| Adversarial 1v1 (full macro) | adversarial-duel. The full 1v1 battleground lives in one_v_one.py (openra_bench/), not as a meta.capability: adversarial pack. |
1 pack (the catalog says this is by design โ full macro is the live ladder) |
Section 4 โ Capability gaps (uncovered or thin cells)
Cells worth new packs, in priority order:
Base-trade race (mid-late / late) โ no current pack tests "your base falls while you race the enemy's; whoever finishes first wins." Sketch: hard tier places both players' bases halfway to killable, no defenders left; the agent has one army worth attacking with, the enemy is doing the same to the agent's base; win by destroying the enemy's
factbefore they destroy yours. Winbuilding_count_gte:{type:fact,n:0,owner:enemy}paired withhas_building:fact(own) and an aggressivewithin_ticks. Pack id:combat-base-trade-race.Concede-vs-hold (mid-late) โ only
mid-concede-vs-holdcurrently. Sketch a second pack:mid-concede-vs-rebuildwhere a forward outpost is lost-cost (committing reinforcements throws good money after bad), and the right call is to abandon it and rebuild on a held line. Win condition: keep the rear-linefactalive AND have โฅ4tent/weapon the rear line at deadline; throwing reinforcements at the forward outpost fails the resource budget. Pack id:mid-concede-forward-rebuild-rear.Live economy defense (mid) โ the only non-econ pack (
mid-economy-under-fire) is currently a stall-WIN defect. After fixing it, add a sibling packmid-economy-rebuild-harvester-linewhere the agent must re-issueharvestcommands to a freshly produced harvester after a raider kills the starter ones (puts the load-bearing capability on theharvestorder, not on the passive auto-harvest).Long-horizon credit-only final phase โ
lh-credit-only-final-phaseis the only exemplar. Sketch sibling:lh-credit-only-bait-windowโ the agent must spend the last credits on a single decisive strike window rather than on attrition. Different decision (one shot vs accumulation).Sustained multi-front (late) โ thin (2 packs). Sketch:
mfb-three-front-rotationโ three simultaneously-attacked bases with the agent's army too small to defend all three at once; the decision is which two to hold and which one to let fall while the army cycles. Cross-pack withcoord-relay-attack.MCV deploy under timer + crossfire โ current MCV packs cover site selection but not "deploy NOW or the MCV is destroyed by an incoming wave that hits in 4 turns." Sketch:
mcv-deploy-emergency-relocationโ the starting MCV stands in the path of an incoming squad; the agent must deploy + re-build OR move + re-deploy further west before the squad arrives.Information freshness across modalities โ
scout-cycle-keep-info-freshandscout-track-enemy-movementcover this for the structured channel, but the perception ablation grid (channel ร fog) doesn't currently have an information-freshness pack that is explicitly easier to solve with the labelled image (theimagechannel is advantaged when the model can spot newly-spawned units inter-tick). Sketch:perception-freshness-image-advantage.Single-pack
meta.capability: adversarialโ onlyadversarial-duelcarries this tag. PAPER_PLAN.md ยง12.2 notes the imbalance. The full 1v1 lives inone_v_one.py(correct), but a second adversarial-tagged pack that tests "reactive opponent selects a counter from a small menu mid-game" (different from RPS pre-game commit) would put real teeth on the tag.Engineer / Tanya / spy under fire (specialist packs) โ only one pack each (
spec-engineer-capture,spec-tanya-c4-strike,spec-spy-infiltrate,spec-thief-steal-cash,combat-tanya-vs-rush). Each specialist has one happy-path scenario. Sketch a "stealth + target priority" pack per specialist that requires the model to pick which of N enemy assets to hit with the one-shot specialist.Naval / amphibious โ
combat-naval-shore-strikeis the only naval pack (and it's currently a stall-WIN defect). Once fixed, addcombat-naval-amphibious-landing(a transport ship lands infantry on a contested shore) andcombat-naval-anti-air-vs-bomberpending the air-unit engine work. These are documented as out-of-scope inPAPER_PLAN.md ยง11.1for the air variant but the naval ones are now feasible thanks towater_rect:.
Section 5 โ Map mis-bindings (rush-hour-arena used where intent
calls for a custom map)
Of 189 probed packs, 183 use rush-hour-arena; 3 use a custom
map (navigation-confined-hard-only, custom-map-no-enemy,
strategy-dilemma/-gauntlet/-twobody); 1 uses a generator
(combat-naval-shore-strike); 2 are quarantined-and-crashing.
Packs whose stated intent calls for a non-rush-hour geometry:
| Pack | Stated intent | Current map | Recommendation |
|---|---|---|---|
combat-heli-flank |
Helicopter assault from a flank; the brief explicitly references aircraft flight over terrain. | rush-hour-arena (no terrain features to flank around) | Author / use a custom map with a forested ridge that channels ground units one way and aircraft another. (Engine air-unit support is out-of-scope per PAPER_PLAN ยง11.1 โ pack is currently a ground-only proxy.) |
def-bridge-chokepoint |
"A water band cuts the map east-to-west with bridges." | rush-hour-arena + water_rect overlay |
The water_rect overlay correctly synthesizes the bridge geometry on top of rush-hour, but the chokepoint-narrowing aspect ("attackers must pass through 3 narrow openings") works against an open-arena map. Consider a custom bridges-arena map authored once โ the visual / minimap reading test is more honest when the bridges are real terrain rather than a YAML overlay on an unrelated map. |
combat-naval-shore-strike |
Naval ship in a water channel. | uses naval-arena generator (correct) โ but the generator spec lives in the pack file rather than a named .oramap. |
When a second naval pack is authored (per gap #10), promote naval-arena to a real data/maps/naval-arena.oramap so the bench has a canonical naval geometry. |
def-with-ambush |
"Concealed flanking defenders catch the band in an L-ambush down a lane toward the construction yard." Doctrinal answer requires real linear terrain. | rush-hour-arena (open) | The pack synthesizes the lane purely with actor placement (a single e1 fixing defender at x=15, flankers at x=40). Visually the geometry is invisible on the minimap โ a custom corridor-arena map with a real lane would make the spatial decision legible from the image. (Capability is sound; this is a perception-channel polish issue.) |
mfb-supply-line-link-between-bases |
"Supply line corridor between two bases." | rush-hour-arena | Similar polish issue โ a custom multi-base map with two separated valleys would make the supply-line decision visible on the minimap rather than implicit in the actor coordinates. |
combat-hold-chokepoint |
"Defend a narrow chokepoint corridor." | rush-hour-arena (no narrow corridor) | Same; either author a chokepoint-arena map or accept that the chokepoint is implied by enemy spawn geometry alone. |
mcv-deploy-second-base, mcv-deploy-third-base |
"Plant a second / third base at a defensible site." | rush-hour-arena (uniform; no defensible sites) | The map is symmetric and uniform โ every cell is roughly as defensible as every other. The capability becomes "pick a cell that satisfies the building_in_region predicate," not "pick a defensible cell." Add a custom map with terrain features (cliff edges, narrow passes) so the defensibility geometry is real. |
expansion-aggro-3-base-greedy, expansion-balanced-2-base-defended, expansion-turtle-1-base-fortified |
Expansion decisions over distinct base sites. | rush-hour-arena (only ~3 sensible base sites given map size) | Authoring a wider arena with 5+ distinguishable expansion sites would put real teeth on the trilemma. |
Summary: roughly 8โ10 packs would benefit from a custom map, but all are currently functional (the no-cheat bar holds). The map mis-binding is a perception-channel polish issue, not a correctness issue. Priority: author 2 custom maps that unlock several packs each:
bridges-arenaโ would replacewater_rectoverlay ondef-bridge-chokepointand could hostcombat-naval-amphibious-landing.chokepoint-arenaโ narrow corridor forcombat-hold-chokepoint,def-with-ambush,mfb-supply-line-link-between-bases.
Section 6 โ Recommended actions (prioritized)
P0 โ Hard defects (gates the no-cheat headline number)
Fix
mid-economy-under-fireโ currently the only stall-WIN defect among non-exempt packs. Either:- (a) Remove the 3 starter harvesters; require the model to issue
Command.harvest(...)on a fresh harvester (the load-bearing capability), OR - (b) Upgrade the raider to a force that out-attritions an idle
defense ring (e.g. 3ร
1tnk+ 2รe3over 60 turns), so a stall play loses harvesters and theharv,2clause fails.
- (a) Remove the 3 starter harvesters; require the model to issue
Fix
combat-naval-shore-strikeโ set the destroyerstanceto0(HoldFire) so the agent must issue an explicitattack_unitorder. Verify the intended capability (manual cross-shore attack) still wins; stall now loses by deadline.Fix
spec-spy-infiltrateโ add{not: own_units_gte: 1}(or{not: unit_type_count_gte: {type: spy, n: 1}}) to thefail_condition.any_ofso a spy-wipe is a real LOSS, not a DRAW. Audit the easy / medium tiers for the same defect.Delete or
_archive/adversarial-siegeandadversarial-skirmishโ both already quarantined; deletion removes the only two load-time-crashing packs.Re-probe
def-bridge-chokepointover seeds 1โ4 in clean processes โ confirm the one-off DRAW observation was a transient (subsequent re-runs gave LOSS consistently); if the DRAW recurs, add{not: own_units_gte: 1}to itsfail_condition.any_of.
P1 โ Coverage gaps (new packs to author)
In priority order from ยง4: combat-base-trade-race,
mid-concede-forward-rebuild-rear, mid-economy-rebuild-harvester-line,
mfb-three-front-rotation, mcv-deploy-emergency-relocation,
lh-credit-only-bait-window, perception-freshness-image-advantage,
combat-naval-amphibious-landing, and per-specialist target-priority
packs (spec-engineer-priority-targets, spec-tanya-priority-targets).
P2 โ Duplicate consolidation
- Merge / delete
artofwar-decoy-sacrificeintoartofwar-lure-the-tiger. - Merge
proc-tool-use-with-distractor+proc-tool-use-multi-distractorinto a single pack with two tiers. - Re-read
combat-kite-and-pullvscombat-kite-jeep-vs-tankand either merge or sharpen the asymmetry-of-units distinction in the second pack's briefing.
P3 โ Map / catalog polish
- Author
data/maps/bridges-arena.oramapand re-binddef-bridge-chokepoint+ (future)combat-naval-amphibious-landing. - Author
data/maps/chokepoint-arena.oramapand re-binddef-with-ambush,combat-hold-chokepoint,mfb-supply-line-link-between-bases. - Promote the four "defensive topology" packs
(
build-defensive-tower-cluster,build-defensive-tower-line,build-defensive-skirt-corners,def-in-depth-vs-single) into a single grouped eval-cell suite so they score one "topology-IQ" metric.
Appendix โ Probe methodology
Each pack was probed by:
import openra_train # builds the engine; required for env.reset
from openra_bench.scenarios.loader import compile_level, load_pack
from openra_bench.eval_core import run_level
pack = load_pack(pack_yaml_path)
c = compile_level(pack, "hard")
ep = run_level(c, lambda rs, Command: [Command.observe()], seed=1)
# ep.outcome โ {"win", "loss", "draw"}; ep.signals.game_tick = final tick.
The driver (/tmp/stall_probe_driver.py) invoked one subprocess per
pack so a Rust engine panic in any one pack did not abort the run.
189 packs probed in 48 seconds wall-time (Apple M-series).
Hard tier seed=1 only. A full audit would extend to seeds 1โ4 (the documented "hard seed" range in CLAUDE.md) and to easy/medium tiers, but seed-1 hard is the highest-pressure cell per pack โ defects visible on any tier are visible here.
Static profile was extracted from the compiled CompiledLevel
(authoritative โ the engine sees the merged scenario, not the raw
YAML).
Raw probe results JSON: /tmp/stall_probe_results.json (regeneratable
by re-running /tmp/stall_probe_driver.py).