Commit History

Adversarial 1v1 spotlight: ladder family + rating + Elo wiring
f5e23f8

yxc20098 commited on

cargo tools: enter_transport + unload, 1:1 congruence
755ab44

yxc20098 commited on

Unified Battle Viewer in app.py + run/model playback identity
0a488d3

yxc20098 commited on

guard tool: 1:1 congruence with engine Command.guard
18d038a

yxc20098 commited on

Wire goal tracker into scoring + leaderboard
eeedfdf

yxc20098 commited on

Playback: capture model reasoning + per-turn goal tracker + viewer
f77eea7

yxc20098 commited on

Scenario-controlled tool allow/deny + default core set
f912cfc

yxc20098 commited on

S7 bench: set_stance + patrol tools (schema 1:1, 17==17)
eca1d53

yxc20098 commited on

S7 bench: surrender tool + loss outcome (tool schema 1:1, 15==15)
09ac234

yxc20098 commited on

Step 4 (bench): interrupt-driven loop + playback/trace capture
03c65ab

yxc20098 commited on

Step 2: concurrent scenario execution (--concurrency N)
3771d77

yxc20098 commited on

Pipeline step 7: per-episode playback persistence
28c736f

yxc20098 commited on

Bench: surface S9 spatial tensor in render_state (multimodal reach)
41a0d2e

yxc20098 commited on

Pairwise adversarial eval + Elo (the user's 'pairwise conditions')
b04adfc

yxc20098 commited on

Generalization-gap metric: held-out split in run_eval + leaderboard
03e4efa

yxc20098 commited on

Catalog C13: harvest-economy packs — closes #14 (user economy families)
7a25eb3

yxc20098 commited on

S1 bench wiring: resources/capacity + economy_value win predicates
1481b7f

yxc20098 commited on

Research-grounded 200-level scenario catalog (12 categories) + integ
aa7cb95

yxc20098 commited on

Scenario brainstorm + 2 runnable economy packs
4b4e8f0

yxc20098 commited on

Single source of truth: port rush-hour + 3 strategy scenarios to packs/
7380c10

yxc20098 commited on

adapter: use engine map_info for true map dims (S9), synthesis fallback
1b62e34

yxc20098 commited on

Custom-map no-enemy scenario pack + robustness regression
a1d3e46

yxc20098 commited on

#13: robustness regression suite (fail-safe + determinism)
59f03db

yxc20098 commited on

Building & Planning scenario family (user-specified first set)
a919131

yxc20098 commited on

#12: real custom-map terrain via dynamic .oramap registry
83c6b8f

yxc20098 commited on

#6 leaderboard: data layer + run_eval publish + Gradio tab
b98ab1a

yxc20098 commited on

Bench: economy scenario pack + full-loop integ test + starting_cash constraint
dc028b6

yxc20098 commited on

tests: reset reveals sight (engine fix); robust in-bounds move targets
ef9ab46

yxc20098 commited on

Bench: consume S9 economy obs + economy win-conditions + full toolset
5a1cf72

yxc20098 commited on

Add scoring + P/R/A diagnostics + run_eval CLI
5b68a55

yxc20098 commited on

Add provider-agnostic model agent (vLLM/OpenRouter/Bedrock)
715cbbc

yxc20098 commited on

Add Rust-backed eval stack: scenario packs, adapter, spine, integration tests
098c6e0

yxc20098 commited on

Add HF identity verification and anonymous submission support
6f326d5

yxc20098 commited on

Add server-side game aggregation with minimum 5-game threshold
e642d5b

yxc20098 commited on

Add minimum game count, difficulty multiplier, and updated scoring formula
65eceb8

yxc20098 commited on

Security hardening: XSS prevention, input validation, rate limiting
9422de7

yxc20098 commited on

Update docs: CLI submission, agent identity, replay downloads, API endpoints
3a2bab2

yxc20098 commited on

Add agent URL hyperlinks, replay downloads, and submit_with_replay endpoint
9fead46

yxc20098 commited on

Merge pull request #3 from yxc20089/feature/upload-form-5tiers
afa3975
unverified

yxc20098 commited on

Fix CI: remove openra_env dependency from get_agent_fn
3c41d18

yxc20098 commited on

Fix CI: add openra-rl-util to requirements.txt
3450be4

yxc20098 commited on

Revert "Fix CI: import scoring from evaluate_runner instead of openra_rl_util"
c07d9e8

yxc20098 commited on

Fix CI: import scoring from evaluate_runner instead of openra_rl_util
50772fa

yxc20098 commited on

Add upload form, API endpoint, 5 difficulty tiers, real game data
45ef63c

yxc20098 commited on

Move Try experience to OpenRA-RL Space, remove Try tab from Bench
8ce66d2

yxc20098 commited on

Rename Evaluate tab to Try, stream LLM agent gameplay
2771ddd

yxc20098 commited on

Add in-browser evaluation via Evaluate tab
824262a

yxc20098 commited on

Connect evaluation harness to HF-hosted OpenRA-RL environment
44493a3

yxc20098 commited on

Fix HF Space build: remove unused deps from requirements.txt
4236023

yxc20098 commited on

Move HF Space sync to openra-rl org
5962a76

yxc20098 commited on