DMM / README.md
tviskaron's picture
Add DMM checkpoints and model card
b8cc825
|
Raw History Blame Contribute Delete
1.2 kB
---
license: mit
tags:
- pytorch
- multi-agent-path-finding
- mapf
- decentralized
---
# DMM: Decentralized Master-Mind
DMM is a decentralized multi-agent pathfinding policy that refines agents' action
intents over several local communication rounds before committing to actions.
The models are pretrained on expert solutions with imitation learning and
optionally fine-tuned with MICPO, a critic-free reinforcement learning method.
[Paper](https://arxiv.org/abs/2609.32019) · [Code](https://github.com/CognitiveAISystems/DMM)
| Checkpoint | Imitation pretraining iterations | MICPO optimizer updates |
| --- | ---: | ---: |
| [DMM-08M.pt](DMM-08M.pt) | 1,000,000 | — |
| [DMM-3M.pt](DMM-3M.pt) | 1,000,000 | — |
| [DMM-MICPO-08M.pt](DMM-MICPO-08M.pt) | 1,000,000 | 96,000 |
| [DMM-MICPO-3M.pt](DMM-MICPO-3M.pt) | 1,000,000 | 96,000 |
MICPO fine-tuning comprises 500 outer iterations (96,000 optimizer updates).
Both model sizes use four communication rounds. Training details are reported in
[the paper](https://arxiv.org/html/2609.32019v1#S5.SS1).
See the [GitHub repository](https://github.com/CognitiveAISystems/DMM) for code, usage instructions, training, and evaluation.