Upload mouse-run-run-1: trained agents, rollouts, evaluation, cross-architecture study
a0ebc13 verified | # MARL Experiment Summary | |
| Generated: `2026-07-02T00:51:16.148597+00:00` | |
| Tables: `runs/tables/mouse-run-run-1` | |
| ## Valid Pairs | |
| | Task | Valid pairs | | |
| | --- | ---: | | |
| | `non_social` | 10 | | |
| | `social` | 10 | | |
| ## Training Metrics (final update, mean ± std across pairs) | |
| | Metric | `non_social` (n=10) | `social` (n=10) | | |
| | --- | ---: | ---: | | |
| | collisions per episode | 2.02 ± 0.745 | 4.758 ± 6.29 | | |
| | chaser return | 6.425 ± 0.352 | -1.839 ± 7.419 | | |
| | explorer return | 61.896 ± 9.459 | 23.844 ± 21.326 | | |
| | chaser partner vision | 0 ± 0 | 0.326 ± 0.228 | | |
| | explorer partner vision | 0 ± 0 | 0.326 ± 0.228 | | |
| | chaser new fields | 82.535 ± 1.845 | 15.063 ± 4.095 | | |
| | explorer new fields | 74.988 ± 6.285 | 51.49 ± 12.231 | | |
| | final distance | 5.044 ± 0.426 | 5.429 ± 1.504 | | |
| | policy loss | -0.000564 ± 0.000452 | 0.000734 ± 0.005 | | |
| | value loss | 5.034 ± 0.193 | 7.117 ± 0.718 | | |
| | entropy | 0.25 ± 0.071 | 0.631 ± 0.166 | | |
| | approx kl | 0.000188 ± 0.000384 | 0.004 ± 0.006 | | |
| | chaser grad norm | 0.296 ± 0.005 | 2.553 ± 1.475 | | |
| | explorer grad norm | 0.313 ± 0.059 | 1.061 ± 0.595 | | |
| | max param abs | 1.782 ± 0.497 | 2.376 ± 0.987 | | |
| ## Random-Opponent Evaluation | |
| ### Opponent: `random_chaser` | |
| | Metric | `non_social` (n=10) | `social` (n=10) | | |
| | --- | ---: | ---: | | |
| | collisions per episode | 2.872 ± 0.316 | 0.572 ± 0.192 | | |
| | chaser return | -3.292 ± 0.037 | -2.605 ± 0.208 | | |
| | explorer return | 61.532 ± 9.081 | 17.34 ± 18.824 | | |
| | chaser partner vision | 0 ± 0 | 0.266 ± 0.031 | | |
| | explorer partner vision | 0 ± 0 | 0.266 ± 0.031 | | |
| | chaser new fields | 33.841 ± 0.19 | 33.939 ± 0.097 | | |
| | explorer new fields | 74.874 ± 6.064 | 45.239 ± 12.602 | | |
| | final distance | 4.862 ± 0.21 | 5.437 ± 0.242 | | |
| | average distance | 4.996 ± 0.075 | 5.584 ± 0.193 | | |
| | degenerate fraction | 0.882 ± 0.019 | 0.925 ± 0.039 | | |
| ### Opponent: `random_explorer` | |
| | Metric | `non_social` (n=10) | `social` (n=10) | | |
| | --- | ---: | ---: | | |
| | collisions per episode | 3.451 ± 0.325 | 8.976 ± 5.699 | | |
| | chaser return | 6.316 ± 0.366 | 3.174 ± 7.155 | | |
| | explorer return | -1.299 ± 0.334 | -10.191 ± 6.452 | | |
| | chaser partner vision | 0 ± 0 | 0.605 ± 0.185 | | |
| | explorer partner vision | 0 ± 0 | 0.605 ± 0.185 | | |
| | chaser new fields | 82.564 ± 1.901 | 17.55 ± 5.312 | | |
| | explorer new fields | 32.97 ± 0.246 | 30.951 ± 1.595 | | |
| | final distance | 5.005 ± 0.188 | 4.094 ± 1.119 | | |
| | average distance | 5.191 ± 0.038 | 3.777 ± 0.992 | | |
| | degenerate fraction | 0.845 ± 0.011 | 0.928 ± 0.059 | | |
| Failed evaluations: 0 | |
| ## Figures | |
| ### training curves | |
|  | |
| Training curves, mean ± std across 20 valid pairs. | |
| ### evaluation comparison | |
|  | |
| Evaluation metrics over 40 successful evaluations, grouped by opponent mode and task. | |
| ## Degenerate Episode Exclusions | |
| | Excluded episodes | Episodes with records | Records | | |
| | ---: | ---: | ---: | | |
| | 82 | 225 | 9 | | |
| ## Failed Attempts | |
| - count: 0 | |