From f280537bfdc4da1326982658d19a0aaceb031acd Mon Sep 17 00:00:00 2001 From: shabeeth <90326442+shabeeth2@users.noreply.github.com> Date: Tue, 6 Oct 2026 22:23:59 +0530 Subject: [PATCH 1/3] docs: correct TerminalSteps length description --- com.unity.ml-agents/Documentation~/Python-LLAPI.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/com.unity.ml-agents/Documentation~/Python-LLAPI.md b/com.unity.ml-agents/Documentation~/Python-LLAPI.md index 1f05a3574b..4f34eaa08c 100644 --- a/com.unity.ml-agents/Documentation~/Python-LLAPI.md +++ b/com.unity.ml-agents/Documentation~/Python-LLAPI.md @@ -99,7 +99,7 @@ A `TerminalSteps` has the following fields : It also has the two following methods: -- `len(TerminalSteps)` Returns the number of agents requesting a decision since the last call to `env.step()`. +- `len(TerminalSteps)` Returns the number of agents whose episodes ended since the last call to `env.step()`, including interrupted episodes. - `TerminalSteps[agent_id]` Returns a `TerminalStep` for the Agent with the `agent_id` unique identifier. A `TerminalStep` has the following fields: From ce6f53e0a758a80719fa6b71d30591e6f20c71a4 Mon Sep 17 00:00:00 2001 From: shabeeth <90326442+shabeeth2@users.noreply.github.com> Date: Tue, 6 Oct 2026 22:24:01 +0530 Subject: [PATCH 2/3] Clarify 'obs' field description in TerminalSteps Updated the description of 'obs' in TerminalSteps to clarify the batch size context. --- com.unity.ml-agents/Documentation~/Python-LLAPI.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/com.unity.ml-agents/Documentation~/Python-LLAPI.md b/com.unity.ml-agents/Documentation~/Python-LLAPI.md index 4f34eaa08c..ebecf5b31c 100644 --- a/com.unity.ml-agents/Documentation~/Python-LLAPI.md +++ b/com.unity.ml-agents/Documentation~/Python-LLAPI.md @@ -92,7 +92,7 @@ Similarly to `DecisionSteps` and `DecisionStep`, `TerminalSteps` (with `s`) cont A `TerminalSteps` has the following fields : -- `obs` is a list of numpy arrays observations collected by the group of agent. The first dimension of the array corresponds to the batch size of the group (number of agents requesting a decision since the last call to `env.step()`). +- `obs` is a list of numpy arrays observations collected by the group of agent. The first dimension of the array corresponds to the batch size of the group (number of agents whose episodes ended since the last call to `env.step()`, including interrupted episodes`). - `reward` is a float vector of length batch size. Corresponds to the rewards collected by each agent since the last simulation step. - `agent_id` is an int vector of length batch size containing unique identifier for the corresponding Agent. This is used to track Agents across simulation steps. - `interrupted` is an array of booleans of length batch size. Is true if the associated Agent was interrupted since the last decision step. For example, if the Agent reached the maximum number of steps for the episode. From e1c2deea7a5d8832e67c84b24ad2d8d258923446 Mon Sep 17 00:00:00 2001 From: shabeeth <90326442+shabeeth2@users.noreply.github.com> Date: Tue, 6 Oct 2026 22:24:04 +0530 Subject: [PATCH 3/3] Fix formatting in Python-LLAPI.md documentation --- com.unity.ml-agents/Documentation~/Python-LLAPI.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/com.unity.ml-agents/Documentation~/Python-LLAPI.md b/com.unity.ml-agents/Documentation~/Python-LLAPI.md index ebecf5b31c..58a6992dce 100644 --- a/com.unity.ml-agents/Documentation~/Python-LLAPI.md +++ b/com.unity.ml-agents/Documentation~/Python-LLAPI.md @@ -92,7 +92,7 @@ Similarly to `DecisionSteps` and `DecisionStep`, `TerminalSteps` (with `s`) cont A `TerminalSteps` has the following fields : -- `obs` is a list of numpy arrays observations collected by the group of agent. The first dimension of the array corresponds to the batch size of the group (number of agents whose episodes ended since the last call to `env.step()`, including interrupted episodes`). +- `obs` is a list of numpy arrays observations collected by the group of agent. The first dimension of the array corresponds to the batch size of the group (number of agents whose episodes ended since the last call to `env.step()`, including interrupted episodes). - `reward` is a float vector of length batch size. Corresponds to the rewards collected by each agent since the last simulation step. - `agent_id` is an int vector of length batch size containing unique identifier for the corresponding Agent. This is used to track Agents across simulation steps. - `interrupted` is an array of booleans of length batch size. Is true if the associated Agent was interrupted since the last decision step. For example, if the Agent reached the maximum number of steps for the episode.