--- license: mit tags: - pytorch - multi-agent-path-finding - mapf - decentralized --- # DMM: Decentralized Master-Mind DMM is a decentralized multi-agent pathfinding policy that refines agents' action intents over several local communication rounds before committing to actions. The models are pretrained on expert solutions with imitation learning and optionally fine-tuned with MICPO, a critic-free reinforcement learning method. [Paper](https://arxiv.org/abs/2609.32019) · [Code](https://github.com/CognitiveAISystems/DMM) | Checkpoint | Imitation pretraining iterations | MICPO optimizer updates | | --- | ---: | ---: | | [DMM-08M.pt](DMM-08M.pt) | 1,000,000 | — | | [DMM-3M.pt](DMM-3M.pt) | 1,000,000 | — | | [DMM-MICPO-08M.pt](DMM-MICPO-08M.pt) | 1,000,000 | 96,000 | | [DMM-MICPO-3M.pt](DMM-MICPO-3M.pt) | 1,000,000 | 96,000 | MICPO fine-tuning comprises 500 outer iterations (96,000 optimizer updates). Both model sizes use four communication rounds. Training details are reported in [the paper](https://arxiv.org/html/2609.32019v1#S5.SS1). See the [GitHub repository](https://github.com/CognitiveAISystems/DMM) for code, usage instructions, training, and evaluation.