Repository navigation
fix(a2a): give up on an agent whose sandbox died, and cancel a prompt that timed out - #82
Conversation
… that timed out A prompt to an agent whose sandbox died waited out its whole timeout (20 minutes in practice) before failing, and a prompt that timed out left the agent working on it, spending tokens, until its sandbox went away. - poll_a2a_task takes the agent's sandbox_id. With it, once the agent has gone 60 s without answering, over at least 3 polls in a row, it raises AgentUnreachableError naming the sandbox. Transport errors and 502/503/504 count as no answer; any other reply, errors included, is the agent answering. A tunnel blip drops a poll or two and is ridden out. AgentUnreachableError is a TimeoutError, so callers handle it as they handled the same dead agent before, only sooner. - Out of time, poll_a2a_task sends tasks/cancel before raising. Best effort: a cancel that fails is logged and the timeout is still what's raised. - prompt_agent passes the solver's sandbox and a deployed user-sim's; a human keeps the full timeout. rubrics_verifier passes its judge's. - The agent framework implements tasks/cancel: it marks the task canceled and the SDK cancels run(), which sees CancelledError. It answered UnsupportedOperation before, so no agent stopped on a cancel. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…mments Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-agent-dead-agent-and-cancel
| "params": {"id": task_id}, | ||
| }, timeout=_CANCEL_TIMEOUT_SECONDS) | ||
| resp.raise_for_status() | ||
| error = resp.json().get("error") |
There was a problem hiding this comment.
Bad cancel reply hides timeout
If the cancel endpoint returns HTTP 200 with valid JSON such as null or a list, calling .get("error") raises AttributeError. This hides the original TimeoutError, so a user-sim timeout fails the run instead of ending the conversation. Treat an unexpected cancel reply as a failed best-effort cancel.
Prompt To Fix With AI
This is a comment left during a code review.
Path: src/agent_env/a2a_agent/protocol.py
Line: 215
Comment:
**Bad cancel reply hides timeout**
If the cancel endpoint returns HTTP 200 with valid JSON such as `null` or a list, calling `.get("error")` raises `AttributeError`. This hides the original `TimeoutError`, so a user-sim timeout fails the run instead of ending the conversation. Treat an unexpected cancel reply as a failed best-effort cancel.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.There was a problem hiding this comment.
Leaving this as is. An A2A endpoint answers with a JSON-RPC object; an error status or a non-JSON body (say, a proxy's HTML page) raises HTTPStatusError or ValueError, which are caught. A 200 whose body is JSON null or a list is not something an agent or the proxies in front of one produce, and the repo conventions ask for no defensive code for a case no caller produces.
There was a problem hiding this comment.
That’s fair. The existing handling covers the failure modes produced by an A2A endpoint or an intervening proxy, while a successful 200 response with non-object JSON is outside the protocol contract and not produced by the callers here. Given the repository’s preference against speculative defensive code, I withdraw this finding; leaving it as-is is reasonable.
An agent cut off for over a minute, rather than dead, would otherwise keep working once it can be reached again. poll_a2a_task now sends tasks/cancel whenever it stops waiting on a task. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
What and why
Two gaps in how core waits on an A2A agent:
poll_a2a_taskkept retrying until the prompt's own timeout. On Modal, every poll of a terminated sandbox fails withServer disconnected without sending a response, so a 1200 s prompt failed 20 minutes after the agent was gone.poll_a2a_taskraisedTimeoutErrorwithout telling the agent, which kept running and calling its model until its sandbox went away. The agent framework couldn't have stopped it anyway: its executor answeredtasks/cancelwithUnsupportedOperation. The abandoned turn also kept holding its context's lock, so the next message in that context waited behind it.Changes:
poll_a2a_task(..., sandbox_id=...), the sandbox the agent runs on:AgentUnreachableError, naming the sandbox.AgentUnreachableErrorsubclassesTimeoutError, so each caller handles a dead agent as it did before, only sooner. For example, a user-sim that dies still ends the conversation rather than failing the run.poll_a2a_taskstops waiting, it sendstasks/cancelfirst (newcancel_a2a_task): out of time, and also when it gives up on an unreachable agent, in case that agent was only cut off. It is best effort: a cancel that fails is logged, and theTimeoutErrororAgentUnreachableErroris still what's raised.prompt_agentpasses the solver's sandbox, and a deployed user-sim's. A human, reached through the human A2A URL, keeps the full timeout.rubrics_verifierpasses its judge's sandbox.agentenv-framework-protocol):_StandardExecutor.cancelmarks the task canceled. The A2A SDK then cancels the coroutine runningrun(), which seesCancelledError, and the context lock is released. The README tells agent authors to stop any processrun()started before re-raising.prompt_agentalready forwardstimeout_secondsto agents that accept it, so a healthy agent usually times itself out first. The client-side cancel covers an agent that hangs or ignores that setting, and a user-sim or human turn.Not in this PR: what a cancel does to the processes an agent's
run()started is up torun(). A handler that kills its CLI only on its own timeout leaves the CLI running after a cancel (see the local results below), so such agents need a one-line change on their side once they take this release.How it was tested
make unit-test: 6274 passed. New tests:tst/unit/a2a_agent/protocol_test.py, on a fake clock:tasks/cancel, and so does giving up on an unreachable agent;tst/unit/task_step/test_prompt_agent_unreachable.py:packages/agentenv-protocol/tests/test_a2a_agent.py: through the A2A SDK's request handler,tasks/cancelmarks a running task canceled,run()seesCancelledError, and the next task in the same context runs. It fails onmain, where the cancel is answered with an error.Local chaos with real processes. The agent is a real agent-framework server whose
run()starts asleep 900child the way a CLI agent does, with a TCP proxy in front standing in for the sandbox tunnel:TimeoutErrorat 10.4 s; the task iscanceled; the next turn in the same context completes in 1.0 srun()kills its child on cancelrun()that kills it only on its own timeout leaves the child running after the cancelAgentUnreachableError58.5 s after the kill; the cancel it then sends fails at once and is loggedAgentUnreachableError58.5 s after the kill (ConnectError)SIGSTOPAgentUnreachableErrorafter 119.4 s (three read timeouts, then the 10 s cancel attempt)TimeoutErrorat the 90 s deadline, with the failed cancel loggedA real Claude Code agent on Modal:
AgentUnreachableErrornaming the sandbox, 60.4 s after the kill (20 min onmain). Re-run on the final head: 59.0 s, and the cancel it then sends fails at once and is loggedsleep 150TimeoutErrorat 30 s. The cancel was sent; the agent, still on the previous protocol, answeredUnsupportedOperationand was stillworking5 s later. That is the gap the framework change closes once agents upgradeOn another remote provider, a gone sandbox's gateway answers 404, which already failed the step at once (2.1 s after the kill). A healthy
sleep 150prompt there completed in 156.0 s with no failed poll.Every sandbox was torn down afterwards.
🤖 Generated with Claude Code
The PR is not yet safe to merge because a malformed cancel reply can still hide a timeout.
Fix with agent prompt
Summary
A2A polling now gives up sooner when a sandbox agent stops answering, and asks agents to cancel tasks when polling times out.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart TD A[Poll task] --> B{Task finished?} B -->|Yes| C[Return task] B -->|No| D{Sandbox stopped answering?} D -->|Yes| E[Ask agent to cancel] D -->|No| F{Time ran out?} F -->|Yes| E F -->|No| A E --> G[Raise timeout error]Reviews (2) · Last reviewed commit: "fix(a2a): ask an agent given up on as un..."