Skip to content

fix(a2a): give up on an agent whose sandbox died, and cancel a prompt that timed out - #82

Merged
earakely-scale merged 5 commits into
mainfrom
edgararakelyan/prompt-agent-dead-agent-and-cancel
Oct 6, 2026
Merged

earakely-scale merged 5 commits into
mainfrom
edgararakelyan/prompt-agent-dead-agent-and-cancel

Conversation

@earakely-scale

@earakely-scale earakely-scale commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

What and why

Two gaps in how core waits on an A2A agent:

  • A dead agent was waited out in full. If an agent's sandbox died mid-prompt, poll_a2a_task kept retrying until the prompt's own timeout. On Modal, every poll of a terminated sandbox fails with Server disconnected without sending a response, so a 1200 s prompt failed 20 minutes after the agent was gone.
  • A prompt that timed out left the agent working. poll_a2a_task raised TimeoutError without telling the agent, which kept running and calling its model until its sandbox went away. The agent framework couldn't have stopped it anyway: its executor answered tasks/cancel with UnsupportedOperation. The abandoned turn also kept holding its context's lock, so the next message in that context waited behind it.

Changes:

  • poll_a2a_task(..., sandbox_id=...), the sandbox the agent runs on:
    • The agent is given up on once it has gone 60 s without answering, over at least 3 polls in a row. The poll raises AgentUnreachableError, naming the sandbox.
    • Transport errors and 502/503/504 count as no answer. Any other reply, errors included, counts as the agent answering. A tunnel blip drops a poll or two and is ridden out.
    • AgentUnreachableError subclasses TimeoutError, so each caller handles a dead agent as it did before, only sooner. For example, a user-sim that dies still ends the conversation rather than failing the run.
  • Whenever poll_a2a_task stops waiting, it sends tasks/cancel first (new cancel_a2a_task): out of time, and also when it gives up on an unreachable agent, in case that agent was only cut off. It is best effort: a cancel that fails is logged, and the TimeoutError or AgentUnreachableError is still what's raised.
  • Callers:
    • prompt_agent passes the solver's sandbox, and a deployed user-sim's. A human, reached through the human A2A URL, keeps the full timeout.
    • rubrics_verifier passes its judge's sandbox.
    • A caller that passes nothing keeps today's waiting, but now also sends the cancel on timeout.
  • Agent framework (agentenv-framework-protocol): _StandardExecutor.cancel marks the task canceled. The A2A SDK then cancels the coroutine running run(), which sees CancelledError, and the context lock is released. The README tells agent authors to stop any process run() started before re-raising.

prompt_agent already forwards timeout_seconds to agents that accept it, so a healthy agent usually times itself out first. The client-side cancel covers an agent that hangs or ignores that setting, and a user-sim or human turn.

Not in this PR: what a cancel does to the processes an agent's run() started is up to run(). A handler that kills its CLI only on its own timeout leaves the CLI running after a cancel (see the local results below), so such agents need a one-line change on their side once they take this release.

How it was tested

  • make unit-test: 6274 passed. New tests:

    • tst/unit/a2a_agent/protocol_test.py, on a fake clock:
      • a dead agent is given up on within the window, naming its sandbox. Covered for disconnects, refused connections, read timeouts and 502s;
      • a 40 s blip is ridden out;
      • three unanswered polls are needed, however far apart;
      • an agent answering with 500s or JSON-RPC errors, and an agent with no sandbox named, are waited on until the timeout;
      • a timeout sends tasks/cancel, and so does giving up on an unreachable agent;
      • a cancel that is unsupported, gets a 500 or can't connect still reports the timeout.
    • tst/unit/task_step/test_prompt_agent_unreachable.py:
      • which sandbox each poll watches: the solver's, a deployed user-sim's, and none for a human;
      • a dead solver fails the step, naming its sandbox;
      • a dead user-sim ends the conversation, not the run.
    • packages/agentenv-protocol/tests/test_a2a_agent.py: through the A2A SDK's request handler, tasks/cancel marks a running task canceled, run() sees CancelledError, and the next task in the same context runs. It fails on main, where the cancel is answered with an error.
  • Local chaos with real processes. The agent is a real agent-framework server whose run() starts a sleep 900 child the way a CLI agent does, with a TCP proxy in front standing in for the sandbox tunnel:

    Scenario Result
    prompt times out (10 s) TimeoutError at 10.4 s; the task is canceled; the next turn in the same context completes in 1.0 s
    the same, run() kills its child on cancel the child is gone. A run() that kills it only on its own timeout leaves the child running after the cancel
    agent killed behind the proxy AgentUnreachableError 58.5 s after the kill; the cancel it then sends fails at once and is logged
    agent killed, polled on its own port AgentUnreachableError 58.5 s after the kill (ConnectError)
    agent frozen with SIGSTOP AgentUnreachableError after 119.4 s (three read timeouts, then the 10 s cancel attempt)
    agent killed, no sandbox named TimeoutError at the 90 s deadline, with the failed cancel logged
    proxy down for 40 s mid-prompt ridden out; the task completes
  • A real Claude Code agent on Modal:

    Scenario Result
    sandbox terminated 45 s into a prompt AgentUnreachableError naming the sandbox, 60.4 s after the kill (20 min on main). Re-run on the final head: 59.0 s, and the cancel it then sends fails at once and is logged
    healthy prompt running sleep 150 completed in 159.5 s with no failed poll
    the poll's 30 s timeout expires before the agent's own TimeoutError at 30 s. The cancel was sent; the agent, still on the previous protocol, answered UnsupportedOperation and was still working 5 s later. That is the gap the framework change closes once agents upgrade

    On another remote provider, a gone sandbox's gateway answers 404, which already failed the step at once (2.1 s after the kill). A healthy sleep 150 prompt there completed in 156.0 s with no failed poll.

    Every sandbox was torn down afterwards.

🤖 Generated with Claude Code

RetriggerConfidence Score: 4/5

The PR is not yet safe to merge because a malformed cancel reply can still hide a timeout.

Fix All in CursorFindings

  1. P1 Bad cancel reply hides timeout ▶
Fix with agent prompt
### Issue 1
src/agent_env/a2a_agent/protocol.py:undefined-217
If the cancel endpoint returns HTTP 200 with valid JSON such as `null` or a list, calling `.get("error")` raises `AttributeError`. This hides the original `TimeoutError`, so a user-sim timeout fails the run instead of ending the conversation. Treat an unexpected cancel reply as a failed best-effort cancel.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

A2A polling now gives up sooner when a sandbox agent stops answering, and asks agents to cancel tasks when polling times out.

  • Sandbox-backed agent polls give up after prolonged silence.
  • Timed-out A2A tasks ask agents to stop working.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Poll task] --> B{Task finished?}
  B -->|Yes| C[Return task]
  B -->|No| D{Sandbox stopped answering?}
  D -->|Yes| E[Ask agent to cancel]
  D -->|No| F{Time ran out?}
  F -->|Yes| E
  F -->|No| A
  E --> G[Raise timeout error]
Loading

Reviews (2) · Last reviewed commit: "fix(a2a): ask an agent given up on as un..."

earakely-scale and others added 4 commits October 6, 2026 14:09
… that timed out

A prompt to an agent whose sandbox died waited out its whole timeout (20
minutes in practice) before failing, and a prompt that timed out left the
agent working on it, spending tokens, until its sandbox went away.

- poll_a2a_task takes the agent's sandbox_id. With it, once the agent has
  gone 60 s without answering, over at least 3 polls in a row, it raises
  AgentUnreachableError naming the sandbox. Transport errors and 502/503/504
  count as no answer; any other reply, errors included, is the agent
  answering. A tunnel blip drops a poll or two and is ridden out.
  AgentUnreachableError is a TimeoutError, so callers handle it as they
  handled the same dead agent before, only sooner.
- Out of time, poll_a2a_task sends tasks/cancel before raising. Best effort:
  a cancel that fails is logged and the timeout is still what's raised.
- prompt_agent passes the solver's sandbox and a deployed user-sim's; a
  human keeps the full timeout. rubrics_verifier passes its judge's.
- The agent framework implements tasks/cancel: it marks the task canceled
  and the SDK cancels run(), which sees CancelledError. It answered
  UnsupportedOperation before, so no agent stopped on a cancel.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…mments

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@earakely-scale
earakely-scale requested a review from a team as a code owner October 6, 2026 21:20
Comment thread src/agent_env/a2a_agent/protocol.py
"params": {"id": task_id},
}, timeout=_CANCEL_TIMEOUT_SECONDS)
resp.raise_for_status()
error = resp.json().get("error")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Bad cancel reply hides timeout

If the cancel endpoint returns HTTP 200 with valid JSON such as null or a list, calling .get("error") raises AttributeError. This hides the original TimeoutError, so a user-sim timeout fails the run instead of ending the conversation. Treat an unexpected cancel reply as a failed best-effort cancel.

Prompt To Fix With AI
This is a comment left during a code review.
Path: src/agent_env/a2a_agent/protocol.py
Line: 215

Comment:
**Bad cancel reply hides timeout**

If the cancel endpoint returns HTTP 200 with valid JSON such as `null` or a list, calling `.get("error")` raises `AttributeError`. This hides the original `TimeoutError`, so a user-sim timeout fails the run instead of ending the conversation. Treat an unexpected cancel reply as a failed best-effort cancel.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Cursor Fix in Claude Code Fix in Codex

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Leaving this as is. An A2A endpoint answers with a JSON-RPC object; an error status or a non-JSON body (say, a proxy's HTML page) raises HTTPStatusError or ValueError, which are caught. A 200 whose body is JSON null or a list is not something an agent or the proxies in front of one produce, and the repo conventions ask for no defensive code for a case no caller produces.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That’s fair. The existing handling covers the failure modes produced by an A2A endpoint or an intervening proxy, while a successful 200 response with non-object JSON is outside the protocol contract and not produced by the callers here. Given the repository’s preference against speculative defensive code, I withdraw this finding; leaving it as-is is reasonable.

An agent cut off for over a minute, rather than dead, would otherwise keep
working once it can be reached again. poll_a2a_task now sends tasks/cancel
whenever it stops waiting on a task.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@earakely-scale
earakely-scale merged commit 6b76fa7 into main Oct 6, 2026
16 checks passed
@earakely-scale
earakely-scale deleted the edgararakelyan/prompt-agent-dead-agent-and-cancel branch October 6, 2026 21:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant