Upload mouse-run-run-1: trained agents, rollouts, evaluation, cross-architecture study
a0ebc13 verified MARL Experiment Summary
Generated: 2026-07-02T00:51:16.148597+00:00
Tables: runs/tables/mouse-run-run-1
Valid Pairs
| Task | Valid pairs |
|---|---|
non_social |
10 |
social |
10 |
Training Metrics (final update, mean ± std across pairs)
| Metric | non_social (n=10) |
social (n=10) |
|---|---|---|
| collisions per episode | 2.02 ± 0.745 | 4.758 ± 6.29 |
| chaser return | 6.425 ± 0.352 | -1.839 ± 7.419 |
| explorer return | 61.896 ± 9.459 | 23.844 ± 21.326 |
| chaser partner vision | 0 ± 0 | 0.326 ± 0.228 |
| explorer partner vision | 0 ± 0 | 0.326 ± 0.228 |
| chaser new fields | 82.535 ± 1.845 | 15.063 ± 4.095 |
| explorer new fields | 74.988 ± 6.285 | 51.49 ± 12.231 |
| final distance | 5.044 ± 0.426 | 5.429 ± 1.504 |
| policy loss | -0.000564 ± 0.000452 | 0.000734 ± 0.005 |
| value loss | 5.034 ± 0.193 | 7.117 ± 0.718 |
| entropy | 0.25 ± 0.071 | 0.631 ± 0.166 |
| approx kl | 0.000188 ± 0.000384 | 0.004 ± 0.006 |
| chaser grad norm | 0.296 ± 0.005 | 2.553 ± 1.475 |
| explorer grad norm | 0.313 ± 0.059 | 1.061 ± 0.595 |
| max param abs | 1.782 ± 0.497 | 2.376 ± 0.987 |
Random-Opponent Evaluation
Opponent: random_chaser
| Metric | non_social (n=10) |
social (n=10) |
|---|---|---|
| collisions per episode | 2.872 ± 0.316 | 0.572 ± 0.192 |
| chaser return | -3.292 ± 0.037 | -2.605 ± 0.208 |
| explorer return | 61.532 ± 9.081 | 17.34 ± 18.824 |
| chaser partner vision | 0 ± 0 | 0.266 ± 0.031 |
| explorer partner vision | 0 ± 0 | 0.266 ± 0.031 |
| chaser new fields | 33.841 ± 0.19 | 33.939 ± 0.097 |
| explorer new fields | 74.874 ± 6.064 | 45.239 ± 12.602 |
| final distance | 4.862 ± 0.21 | 5.437 ± 0.242 |
| average distance | 4.996 ± 0.075 | 5.584 ± 0.193 |
| degenerate fraction | 0.882 ± 0.019 | 0.925 ± 0.039 |
Opponent: random_explorer
| Metric | non_social (n=10) |
social (n=10) |
|---|---|---|
| collisions per episode | 3.451 ± 0.325 | 8.976 ± 5.699 |
| chaser return | 6.316 ± 0.366 | 3.174 ± 7.155 |
| explorer return | -1.299 ± 0.334 | -10.191 ± 6.452 |
| chaser partner vision | 0 ± 0 | 0.605 ± 0.185 |
| explorer partner vision | 0 ± 0 | 0.605 ± 0.185 |
| chaser new fields | 82.564 ± 1.901 | 17.55 ± 5.312 |
| explorer new fields | 32.97 ± 0.246 | 30.951 ± 1.595 |
| final distance | 5.005 ± 0.188 | 4.094 ± 1.119 |
| average distance | 5.191 ± 0.038 | 3.777 ± 0.992 |
| degenerate fraction | 0.845 ± 0.011 | 0.928 ± 0.059 |
Failed evaluations: 0
Figures
training curves
Training curves, mean ± std across 20 valid pairs.
evaluation comparison
Evaluation metrics over 40 successful evaluations, grouped by opponent mode and task.
Degenerate Episode Exclusions
| Excluded episodes | Episodes with records | Records |
|---|---|---|
| 82 | 225 | 9 |
Failed Attempts
- count: 0

