diff --git a/com.unity.ml-agents/Documentation~/Python-LLAPI.md b/com.unity.ml-agents/Documentation~/Python-LLAPI.md index 1f05a3574b..58a6992dce 100644 --- a/com.unity.ml-agents/Documentation~/Python-LLAPI.md +++ b/com.unity.ml-agents/Documentation~/Python-LLAPI.md @@ -92,14 +92,14 @@ Similarly to `DecisionSteps` and `DecisionStep`, `TerminalSteps` (with `s`) cont A `TerminalSteps` has the following fields : -- `obs` is a list of numpy arrays observations collected by the group of agent. The first dimension of the array corresponds to the batch size of the group (number of agents requesting a decision since the last call to `env.step()`). +- `obs` is a list of numpy arrays observations collected by the group of agent. The first dimension of the array corresponds to the batch size of the group (number of agents whose episodes ended since the last call to `env.step()`, including interrupted episodes). - `reward` is a float vector of length batch size. Corresponds to the rewards collected by each agent since the last simulation step. - `agent_id` is an int vector of length batch size containing unique identifier for the corresponding Agent. This is used to track Agents across simulation steps. - `interrupted` is an array of booleans of length batch size. Is true if the associated Agent was interrupted since the last decision step. For example, if the Agent reached the maximum number of steps for the episode. It also has the two following methods: -- `len(TerminalSteps)` Returns the number of agents requesting a decision since the last call to `env.step()`. +- `len(TerminalSteps)` Returns the number of agents whose episodes ended since the last call to `env.step()`, including interrupted episodes. - `TerminalSteps[agent_id]` Returns a `TerminalStep` for the Agent with the `agent_id` unique identifier. A `TerminalStep` has the following fields: