Skip to content

Add MAPPO trainer (multi-agent PPO with a centralized critic) - #6328

Closed
MohabYasser2 wants to merge 3 commits into
Unity-Technologies:developfrom
MohabYasser2:mappo-trainer
Closed

MohabYasser2 wants to merge 3 commits into
Unity-Technologies:developfrom
MohabYasser2:mappo-trainer

Conversation

@MohabYasser2

Copy link
Copy Markdown

Proposed change(s)

This PR adds MAPPO (Yu et al., The Surprising Effectiveness of PPO in Cooperative, Multi-Agent Games) as a new trainer, trainer_type: mappo.

MAPPO is PPO with decentralized actors and a centralized critic. Each agent's policy acts on its own observations, exactly as with PPO, while the value function also sees the observations of the other agents in its SimpleMultiAgentGroup. Advantages are computed with GAE on that centralized value. Agents that are not in a group get a critic that only sees their own observations, so they are trained like PPO.

The change is additive and opt-in:

  • mlagents/trainers/mappo/ (new): MAPPOTrainer, TorchMAPPOOptimizer and MAPPOSettings. They reuse existing building blocks: the PPO loss and GAE, and the attention-based MultiAgentNetworkBody that MA-POCA already uses for its critic.
  • mlagents/plugins/trainer_type.py: registers the trainer and its settings (4 lines).
  • mlagents/trainers/settings.py: MAPPO joins PPO and POCA in the on-policy schedule defaults (1 line).
  • No existing trainer, network or buffer is modified. Moving an existing PPO or POCA configuration to MAPPO is a one-line change: trainer_type: mappo.

How it differs from MA-POCA

  • MAPPO has no counterfactual baseline: the advantage comes from GAE on the centralized value, as in the paper.
  • MA-POCA adds the individual rewards of groupmates to each agent's extrinsic reward. MAPPO keeps individual rewards individual (the group reward from AddGroupReward() is still shared), so agents in the same group can be given different, role-specific rewards, for example a goalkeeper and a striker.
  • For agents that are removed from the scene mid-episode, MA-POCA's absorbing-state handling keeps crediting them with the group's later rewards, while MAPPO bootstraps from its critic. The docs recommend MA-POCA for that case.

Also included

  • Example configurations in config/mappo/ for the same environments as config/poca/ (DungeonEscape, PushBlockCollab, SoccerTwos, StrikersVsGoalie). They mirror the MA-POCA hyperparameters so the two trainers can be compared on equal terms; they have not been tuned for MAPPO.
  • Documentation: a MAPPO section in Training-Configuration-File.md, a paragraph in ML-Agents-Overview.md (cooperative multi-agent training) and a note in Learning-Environment-Design-Agents.md (multi-agent groups).
  • A changelog entry under Unreleased.

Useful links (Github issues, JIRA tickets, ML-Agents forum threads etc.)

Types of change(s)

  • Bug fix
  • New feature
  • Code refactor
  • Breaking change
  • Documentation update
  • Other (please describe)

Checklist

  • Added tests that prove my fix is effective or that my feature works
  • Updated the changelog (if applicable)
  • Updated the documentation (if applicable)
  • Updated the migration guide (if applicable): not needed, nothing changes for existing configurations

Other comments

Tests performed locally (Python 3.10.11, torch 2.6.0 CPU, Windows 10):

  • New test_mappo.py (28 tests): optimizer updates for discrete and continuous actions, visual and vector observations, with and without LSTM; centralized value estimates, including terminal and ignore_done handling and per-agent critic memories; the critic responding to groupmate observations; curiosity; group rewards and end of episode; trajectory processing.
  • test_simple_rl.py: MAPPO learns the simple multi-agent environment (vector and visual observations), learns the LSTM memory environment, and trains under self-play (test_simple_mappo, test_visual_mappo, test_recurrent_mappo, test_simple_ghost_mappo): 10 passed.
  • MAPPO added to the trainer normalization test (test_trainers.py) and to the saver tests (checkpoint save and load of the optimizer and reward providers).
  • The full non-slow trainer test suite: 438 passed, 3 skipped.
  • pre-commit run on every changed file: black, flake8, mypy, pyupgrade and the repository's local hooks pass. The Ruby-based search-and-replace hook could not run on my machine, so I checked its two patterns on the changed markdown by hand.
  • Every file in config/mappo/ parses into RunOptions with MAPPOSettings.

This is my first contribution here, so I'm happy to adjust the structure, naming or documentation to fit the project. The contribution guide recommends opening an issue first; since this change is additive and opt-in I went straight to a PR, but I'm glad to move the design discussion to an issue if you prefer.

MAPPO (Multi-Agent PPO, https://arxiv.org/abs/2103.01955) trains each
agent's policy like PPO on its own observations, with a centralized
critic that also sees the observations of the agent's group. The critic
reuses the attention-based MultiAgentNetworkBody from MA-POCA, and the
advantage is computed with GAE on the centralized value.

Individual rewards stay with the agent that earned them; the group
reward is shared. Agents outside a group are trained like PPO.

Enabled with `trainer_type: mappo`. The trainer is registered next to
PPO, SAC and POCA, and MAPPO uses the same on-policy schedule defaults.
Unit tests for the optimizer and trainer (updates, centralized value
estimates, critic memories, groupmate observations, curiosity, group
rewards), MAPPO in the simple RL tests (multi-agent, visual, memory and
self-play), and MAPPO in the trainer normalization and saver tests.
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

Example configurations for DungeonEscape, PushBlockCollab, SoccerTwos
and StrikersVsGoalie that mirror the MA-POCA ones, MAPPO in the
training configuration, overview and agent design docs, and a changelog
entry.
@montplaisir

Copy link
Copy Markdown
Contributor

Hi @MohabYasser2,

Thank you for your contribution. It is nice work, and I appreciate that there are thorough tests, and updated documentation.

We found that most of the MAPPO optimizer and trainer duplicate MA-POCA. ML-Agents has a plugin system for trainers, made exactly for this case. You can write a pip package that registers in mlagents.trainer_type. After that, when you pick trainer_type: mappo your package will be used.

Please see:

Please note that the critic has a memory issue that you should fix. When an agent is removed before the rest of its group, the corresponding entry in value_memory_dict should be removed.

We will close this PR for now. Thank you again.

@montplaisir montplaisir closed this Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants