Skip to content

feat(trajectory): live trajectories for trajectory/v1 - #115

Open
canersoz-scaleai wants to merge 14 commits into
scaleapi:mainfrom
canersoz-scaleai:canersoz/fd-3719-live-trajectory-follower
Open

canersoz-scaleai wants to merge 14 commits into
scaleapi:mainfrom
canersoz-scaleai:canersoz/fd-3719-live-trajectory-follower

Conversation

@canersoz-scaleai

@canersoz-scaleai canersoz-scaleai commented Oct 8, 2026 •

Copy link
Copy Markdown

What and why

Today a trajectory only exists after its task ends, so every runtime that wants a live view builds its own read. This moves live trajectories into the framework, and the live view is guaranteed to match the final trajectory.

Protocol

  • enable(TRAJECTORY_V1, live=True) adds two cursor reads to get: by task_id or by context_id, each with after and an optional limit. The existing reads are unchanged.
  • Agents get a TrajectoryLog on request.trajectory. They call set_format(...) once, then append(event) as events happen; the format can't change after the first event.
  • The framework keeps each task's log and state, plus each context's turns in order. The store is bounded to 1,024 contexts, and running ones are never evicted.
  • If an agent appends and returns no native_trajectory, the log is its final trajectory. native_trajectory still works for agents that only have one at the end.

Kernel

  • follow_trajectory: a reusable follower that polls one task's cursor and hands each batch to a sink until the task ends.
  • prompt_agent gets live_trajectory: bool = False. When it's on and the agent supports it, each turn's events are written as chunks under the turn's trajectory prefix while the turn runs, and end.json records how the turn ended, including when the step is cancelled or the follower stops early. When it's off, nothing changes.
  • Each turn's next.json names the turn that followed it, or null after the step's last, so a reader can follow a multi-turn step in order and knows when it is over.

How it was tested

  • CI's unit, integration-local and integration-local-slow jobs pass. New unit tests cover the cursor reads, states, eviction, cancel, the follower, the chunk, end and link writes, and a check that the log's final trajectory is byte-identical to the equivalent native_trajectory.
  • A new integration test runs a two-turn prompt_agent conversation with a deployed agent on the local defaults: each turn is stored under its own prefix while it runs, its chunks join to its final trajectory, and the turns link in order.
  • End to end with a CLI agent on a deployed stack, single-turn and two-turn: chunks appeared during each turn, the joined chunks matched each turn's final trajectory byte for byte, and the two turns linked in order.

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with only the previously reported oversized echo-agent prompt issue remaining.

Fix All in CursorFindings

  1. P2 Long step count fails turn ▶
Fix with agent prompt
### Issue 1
tst/data/a2a_agent/agent.py:undefined-84
A prompt containing `live-steps ` followed by more than 4,300 digits makes `int()` raise before `MAX_LIVE_STEPS` can cap the count. The echo agent then fails the turn instead of replying. Check the digit count before converting it.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR adds live trajectory logs and cursor reads to the agent framework, then lets prompt_agent save supported agents’ events while each turn runs.

  • Agents can append trajectory events while tasks run.
  • Prompt runs can save each supported turn’s trajectory as it happens.

Diagram

sequenceDiagram
  participant Step as prompt_agent
  participant Agent as A2A agent
  participant Log as TrajectoryLog
  participant Store as Object store
  Step->>Agent: Send turn
  Agent->>Log: Append events
  loop While the turn runs
    Step->>Agent: Read trajectory cursor
    Agent->>Log: Read next events
    Step->>Store: Write event chunk
  end
  Agent-->>Step: Final task result
  Step->>Store: Write end and next markers
Loading

Reviews (7) · Last reviewed commit: "test(a2a_agent): bound the echo agent's ..." · Reviewed by Greptile

canersoz-scaleai and others added 6 commits October 7, 2026 14:39
Agents can now push trajectory events into a framework-owned, append-only log
(`request.trajectory.append(event)`), and clients can read it while the task
runs. `enable(TRAJECTORY_V1, live=True)` advertises two framework-owned cursor
reads on the existing `get`: `{task_id, after, limit?}` and
`{context_id, after, limit?}`, answered with
`{context_id, task_id, state, format, events, next, has_more}`.

- Additive: the `{task_id}`, `{task_id, objects}`, `{context_id}` and
  `{context_id, objects}` reads keep their request and response shapes. A
  `native_trajectory` is still the final record and is encoded as before, so
  the final object is unchanged for every existing agent.
- When an agent appends and returns no `native_trajectory`, the log is the
  task's final trajectory.
- States: pending, running, completed, failed, canceled. A terminal state is
  reported only once the log is sealed, and the log is sealed before the
  task's terminal A2A event.
- Tasks are registered before the executor's first await, so an id returned by
  `message/send` is never 404. Context reads follow execution order.
- With `live=True` the framework answers the `{context_id}` reads from its log
  unless the agent binds its own handlers, which still take precedence.
- The completed-trajectory cache is replaced by a store bounded at 256 MiB and
  1,024 contexts. Whole contexts are evicted, least recently used first, and
  never one with a running task.

FD-3719

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gh eviction

- `limit` on the cursor reads is validated as 1 to 1,000 (HTTP 400 outside),
  instead of being quietly lowered to the cap. The cap moves to `extensions.py`
  as `MAX_TRAJECTORY_READ_EVENTS`, next to the request models it bounds.
- The trajectory store evicts by count only (1,024 contexts, least recently
  used first), so finished trajectories are kept at least as long as the
  previous 1,024-task cache kept them; the 256 MiB byte cap is removed.
- Eviction never drops the context whose task has just ended, so its final
  record is still there when the client fetches it. Previously a context over
  the byte cap could be evicted at the moment its task was sealed.

FD-3719

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`TrajectoryLog.append` raises until `set_format` has named the events'
format, so every live trajectory carries one for readers to parse it by.
A context read reports the most recent format any of its tasks set, so a
turn that has not started yet no longer hides the earlier turns' format.

FD-3719

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hunks

`prompt_agent` gains `live_trajectory` (default off; also settable per run
through `step_params`). When on and the agent's card offers the cursor read,
each turn's trajectory is followed while the agent works and stored under
the turn's trajectory prefix:

  <prefix><target_a2a_task_id>/live/<after:08d>-<next:08d>.jsonl  per page
  <prefix><target_a2a_task_id>/live/meta.json                     {"format"}
  <prefix><target_a2a_task_id>/live/end.json                      {"state", "next"}

- `agent_env.a2a_agent.trajectory_follower`: `live_trajectory_endpoint`
  reads the card; `follow_trajectory` reads one task from its first event,
  hands each page to a sink, retries transport errors and 5xx, stops on 404,
  and returns once the task has ended and every event is read.
- `snapshot_utils.live_trajectory`: the chunk sink (write-once objects, so a
  repeated write is a no-op) and `following_live_trajectory`, which drains
  after a turn that ended normally (at most 60 s) and stops at once after
  one that failed. A follower's failure never fails the turn.
- `send_and_wait` takes `on_sent`, called with the peer's task id.
- With the flag off, an agent without the cursor read, or a prefix outside
  the configured store, the step runs exactly as before; `to_dict` records
  `live_trajectory` only when it is on.

FD-3719

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When the step stops following before the agent's task ends (a cancelled
step, or a turn whose send/wait raised), the follower is stopped at once and
never reads a terminal state, so `end.json` was never written and readers
could not tell the turn was over: AMV reported a cancelled run's turn as live.

`following_live_trajectory` now writes `end.json` itself in that case,
`{"state": "canceled" | "failed", "next": <events stored>}`, bounded to 5 s
so the cancel is not held up. `end.json` is write-once, so a state the
follower already read (e.g. `completed`) stays.

FD-3719

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ive one

The same events, returned as a `native_trajectory` by one agent and appended
to the log by another, give the same `{task_id}` answer and the same uploaded
bytes (non-ASCII text, U+2028, nested objects, floats and nulls included), so
porting an agent to the log does not change its final trajectory.

FD-3719

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@canersoz-scaleai
canersoz-scaleai marked this pull request as ready for review October 8, 2026 17:54
@canersoz-scaleai
canersoz-scaleai requested a review from a team as a code owner October 8, 2026 17:54
Comment thread src/agent_env/task_step/snapshot_utils/live_trajectory.py Outdated
canersoz-scaleai and others added 3 commits October 8, 2026 15:25
…wing stops

A turn whose step returned normally got no end.json when the follower died on
a client error, the agent no longer knew the task, or the drain timed out, so
readers kept waiting on it. The step now reports the state its task ended in,
and the turn is marked ended in it whenever the follower did not read the end.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…o turns

The echo agent records its trajectory through the framework's log, with the
same final bytes, and a live-steps prompt has it append steps a second apart.
A two-turn conversation on the local defaults then checks each turn is stored
under its own prefix while it runs and joins to its final trajectory.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Readers parse every event by the one format a read reports, so set_format
now refuses a different format after the first append; naming the same
format again is still allowed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread tst/data/a2a_agent/agent.py
Each turn's end.json ends that turn only, so a reader of a multi-turn
step's stored live trajectory could not tell which turn came next or when
the step was over. Each sent turn's next.json now names the turn sent after
it, written as that turn's message goes out, and the step's last turn gets
{"turn": null} once the step ends, on failure too.

The turn loop is re-indented under the try that marks the last turn; the
rest of its diff is whitespace only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
canersoz-scaleai and others added 3 commits October 8, 2026 17:29
…object-url names

main renamed PromptResponse's trajectory fields to object-url names, so the
live trajectory tests read agent_trajectory_object_url and
target_agent_per_turn_trajectory_object_urls.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…z/live-trajectory-turn-links

# Conflicts:
#	src/agent_env/task_step/task_steps/prompt_agent.py
Comment thread src/agent_env/task_step/snapshot_utils/live_trajectory.py
A prompt that happened to contain a large live-steps count held its turn
for that many seconds; the echo agent now appends at most ten.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
request.trajectory.set_format(TRAJECTORY_FORMAT)
prompt = "\n".join(part.text for part in request.parts if isinstance(part, TextPart))
if steps := LIVE_STEPS.search(prompt):
for step in range(min(int(steps.group(1)), MAX_LIVE_STEPS)):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Long step count fails turn

A prompt containing live-steps followed by more than 4,300 digits makes int() raise before MAX_LIVE_STEPS can cap the count. The echo agent then fails the turn instead of replying. Check the digit count before converting it.

Prompt To Fix With AI
This is a comment left during a code review.
Path: tst/data/a2a_agent/agent.py
Line: 84

Comment:
**Long step count fails turn**

A prompt containing `live-steps ` followed by more than 4,300 digits makes `int()` raise before `MAX_LIVE_STEPS` can cap the count. The echo agent then fails the turn instead of replying. Check the digit count before converting it.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Cursor Fix in Claude Code Fix in Codex

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not changing this one: the echo agent is test-only, and the only prompt that sends live-steps is our integration test, with a single digit. MAX_LIVE_STEPS already bounds any realistic count.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That context makes this non-actionable. Since this is a test-only echo agent and the integration input uses a single-digit count, MAX_LIVE_STEPS bounds the exercised behavior as intended. I’m withdrawing the finding; no change is needed for this PR.

@canersoz-scaleai canersoz-scaleai changed the title Canersoz/fd 3719 live trajectory follower feat(trajectory): live trajectories for trajectory/v1 Oct 9, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants