Repository navigation
Add MAPPO trainer (multi-agent PPO with a centralized critic) - #6328
MohabYasser2 wants to merge 3 commits into
Conversation
MAPPO (Multi-Agent PPO, https://arxiv.org/abs/2103.01955) trains each agent's policy like PPO on its own observations, with a centralized critic that also sees the observations of the agent's group. The critic reuses the attention-based MultiAgentNetworkBody from MA-POCA, and the advantage is computed with GAE on the centralized value. Individual rewards stay with the agent that earned them; the group reward is shared. Agents outside a group are trained like PPO. Enabled with `trainer_type: mappo`. The trainer is registered next to PPO, SAC and POCA, and MAPPO uses the same on-policy schedule defaults.
Unit tests for the optimizer and trainer (updates, centralized value estimates, critic memories, groupmate observations, curiosity, group rewards), MAPPO in the simple RL tests (multi-agent, visual, memory and self-play), and MAPPO in the trainer normalization and saver tests.
|
|
Example configurations for DungeonEscape, PushBlockCollab, SoccerTwos and StrikersVsGoalie that mirror the MA-POCA ones, MAPPO in the training configuration, overview and agent design docs, and a changelog entry.
c015031 to
1d29c41
Compare
|
Hi @MohabYasser2, Thank you for your contribution. It is nice work, and I appreciate that there are thorough tests, and updated documentation. We found that most of the MAPPO optimizer and trainer duplicate MA-POCA. ML-Agents has a plugin system for trainers, made exactly for this case. You can write a pip package that registers in Please see:
Please note that the critic has a memory issue that you should fix. When an agent is removed before the rest of its group, the corresponding entry in We will close this PR for now. Thank you again. |
Proposed change(s)
This PR adds MAPPO (Yu et al., The Surprising Effectiveness of PPO in Cooperative, Multi-Agent Games) as a new trainer,
trainer_type: mappo.MAPPO is PPO with decentralized actors and a centralized critic. Each agent's policy acts on its own observations, exactly as with PPO, while the value function also sees the observations of the other agents in its
SimpleMultiAgentGroup. Advantages are computed with GAE on that centralized value. Agents that are not in a group get a critic that only sees their own observations, so they are trained like PPO.The change is additive and opt-in:
mlagents/trainers/mappo/(new):MAPPOTrainer,TorchMAPPOOptimizerandMAPPOSettings. They reuse existing building blocks: the PPO loss and GAE, and the attention-basedMultiAgentNetworkBodythat MA-POCA already uses for its critic.mlagents/plugins/trainer_type.py: registers the trainer and its settings (4 lines).mlagents/trainers/settings.py: MAPPO joins PPO and POCA in the on-policy schedule defaults (1 line).trainer_type: mappo.How it differs from MA-POCA
AddGroupReward()is still shared), so agents in the same group can be given different, role-specific rewards, for example a goalkeeper and a striker.Also included
config/mappo/for the same environments asconfig/poca/(DungeonEscape, PushBlockCollab, SoccerTwos, StrikersVsGoalie). They mirror the MA-POCA hyperparameters so the two trainers can be compared on equal terms; they have not been tuned for MAPPO.Training-Configuration-File.md, a paragraph inML-Agents-Overview.md(cooperative multi-agent training) and a note inLearning-Environment-Design-Agents.md(multi-agent groups).Useful links (Github issues, JIRA tickets, ML-Agents forum threads etc.)
Types of change(s)
Checklist
Other comments
Tests performed locally (Python 3.10.11, torch 2.6.0 CPU, Windows 10):
test_mappo.py(28 tests): optimizer updates for discrete and continuous actions, visual and vector observations, with and without LSTM; centralized value estimates, including terminal andignore_donehandling and per-agent critic memories; the critic responding to groupmate observations; curiosity; group rewards and end of episode; trajectory processing.test_simple_rl.py: MAPPO learns the simple multi-agent environment (vector and visual observations), learns the LSTM memory environment, and trains under self-play (test_simple_mappo,test_visual_mappo,test_recurrent_mappo,test_simple_ghost_mappo): 10 passed.test_trainers.py) and to the saver tests (checkpoint save and load of the optimizer and reward providers).pre-commit runon every changed file: black, flake8, mypy, pyupgrade and the repository's local hooks pass. The Ruby-based search-and-replace hook could not run on my machine, so I checked its two patterns on the changed markdown by hand.config/mappo/parses intoRunOptionswithMAPPOSettings.This is my first contribution here, so I'm happy to adjust the structure, naming or documentation to fit the project. The contribution guide recommends opening an issue first; since this change is additive and opt-in I went straight to a PR, but I'm glad to move the design discussion to an issue if you prefer.