File size: 1,196 Bytes
642781c
 
b8cc825
 
 
 
 
642781c
b8cc825
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
---
license: mit
tags:
  - pytorch
  - multi-agent-path-finding
  - mapf
  - decentralized
---

# DMM: Decentralized Master-Mind

DMM is a decentralized multi-agent pathfinding policy that refines agents' action
intents over several local communication rounds before committing to actions.
The models are pretrained on expert solutions with imitation learning and
optionally fine-tuned with MICPO, a critic-free reinforcement learning method.

[Paper](https://arxiv.org/abs/2609.32019) · [Code](https://github.com/CognitiveAISystems/DMM)

| Checkpoint | Imitation pretraining iterations | MICPO optimizer updates |
| --- | ---: | ---: |
| [DMM-08M.pt](DMM-08M.pt) | 1,000,000 | — |
| [DMM-3M.pt](DMM-3M.pt) | 1,000,000 | — |
| [DMM-MICPO-08M.pt](DMM-MICPO-08M.pt) | 1,000,000 | 96,000 |
| [DMM-MICPO-3M.pt](DMM-MICPO-3M.pt) | 1,000,000 | 96,000 |

MICPO fine-tuning comprises 500 outer iterations (96,000 optimizer updates).
Both model sizes use four communication rounds. Training details are reported in
[the paper](https://arxiv.org/html/2609.32019v1#S5.SS1).

See the [GitHub repository](https://github.com/CognitiveAISystems/DMM) for code, usage instructions, training, and evaluation.