Skip to content

fix(task): task run, run-batch and the local runner tear down what a run deployed - #67

Merged
earakely-scale merged 8 commits into
mainfrom
edgararakelyan/task-run-teardown
Oct 6, 2026
Merged

earakely-scale merged 8 commits into
mainfrom
edgararakelyan/task-run-teardown

Conversation

@earakely-scale

@earakely-scale earakely-scale commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

What and why

Only agent-env run and agent-env eval run tore a run down. A task started any other way left every env stack, agent and sandbox it deployed running, whether it passed or failed:

  • agent-env task run
  • task run-batch
  • the explorer's LocalRunner

Remote sandboxes stayed up until their TTL; local ones stayed until someone removed them by hand.

Changes:

  • task run and run-batch:
    • Each run is torn down as it ends, passed or failed. The context JSON is still written.
    • Ctrl-C or SIGTERM mid-run cancels the runs, tears them down and exits 130 (143 for SIGTERM). A second signal stops the teardown. This uses the same Interrupts handling as eval run.
    • --keep holds what the runs left up until Ctrl-C, then tears it down, as agent-env run --keep does. It lists each sandbox, plus the folder for a local one. While it holds, another terminal can resume a run with --start-step and --context-json.
    • A failed single run is still raised to the CLI, so it's reported as before.
  • LocalRunner tears a run down however it ends: completed, failed, cancelled, or interrupted by shutdown. The explorer can't resume from an earlier run's context (it returns 501 for that), so nothing needs those sandboxes afterwards.
  • eval run's teardown report printer moves to cli/teardown_output.py, so both commands share it.

Task.run() itself stays teardown-free. The hub and the worker keep their own cleanup.

How it was tested

  • make unit-test: 6195 passed. New tests:

    • tst/unit/cli/task_run_teardown_test.py covers a run that passes, one that fails, parallel runs, --keep held then released by Ctrl-C, --keep with Ctrl-C mid-run, Ctrl-C mid-run with --k 2 (exits 130), and run-batch.
    • local_runner_test.py covers a run that completes, raises or is cancelled, each torn down once.
  • Real runs with local stores and local sandboxes. The task is deploy_sandbox then run_docker_container, plus a 60-second step for the signal case:

    Run Result
    task run on main the container, its image and the sandbox folder stay
    task run "Tore down 1 sandbox"; nothing left
    task run --keep, then SIGTERM listed the kept sandbox and its folder, held it up, then tore it down and exited 0; nothing left
    SIGTERM during the 60 s step "Cancelling and tearing down", "Tore down 1 sandbox", exit 143; nothing left

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

No new finding was established in the changes since the previous review.

Summary

Task runs and batches now tear down deployed sandboxes, and the local runner cleans up after its runs too. Both task commands add signal handling and a --keep option for runs that need to stay up temporarily.

  • Task runs and batches clean up what they deploy.
  • The local runner cleans up resources after every run.

Reviews (8) · Last reviewed commit: "Let an abandoned run tear down what stop..."

…run deployed

Only `agent-env run` and `agent-env eval run` tore a run down. A task
started by `task run`, `task run-batch` or the explorer's LocalRunner
left every env stack, agent and sandbox it deployed running, passed or
failed; a local one stayed until removed by hand.

Each now tears its runs down as they end. `--keep` on `task run` and
`run-batch` holds what they left up until Ctrl-C, as `agent-env run
--keep` does; while it holds, another terminal can resume a run with
--start-step and --context-json. Ctrl-C or SIGTERM mid-run cancels the
runs, tears them down and exits 130 (143); a second one stops the
teardown. A failed single run is still raised to the CLI as before.

The teardown report printer moves out of `eval run` so both commands
share it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@earakely-scale
earakely-scale requested a review from a team as a code owner October 6, 2026 16:58
Comment thread src/agent_env/cli/task/run.py Outdated
Comment thread src/agent_env/runner/local_runner.py
Comment thread src/agent_env/cli/task/run.py
…ardown outlast stop(), and write a kept failed run's context

With --keep, a run that finished was held while another still ran. A
Ctrl-C then cancelled the other, skipped the hold and exited, leaving
the finished run up. Ctrl-C now tears down what the finished runs kept
too.

LocalRunner.stop() cancelled every run in flight, including one tearing
down after it had ended, so its sandboxes could stay up. stop() now
cancels only the runs still working and waits for the teardowns.

A failed run kept with --keep wrote no context, so --context-json had
nothing to resume it from. Its context is now written.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@earakely-scale

Copy link
Copy Markdown
Collaborator Author

@greptileai review

Comment thread src/agent_env/runner/local_runner.py Outdated
Comment thread src/agent_env/cli/task/run.py
…rror when its context can't be written

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread src/agent_env/runner/local_runner.py Outdated
Comment thread src/agent_env/runner/local_runner.py Outdated
stop()'s wait began when shutdown did, so a run that took most of it to
stop its steps had its teardown cancelled before it could finish. Each
run now bounds its own teardown by TEARDOWN_WAIT_SECONDS, counted from
when the teardown begins, and stop() waits for every run as it did
before. The hold's poll interval is named too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread src/agent_env/runner/local_runner.py Outdated
A step that ignored its cancel kept a run from ever starting its teardown
timeout, so stop() could wait forever. stop() now waits at most
STOP_WAIT_SECONDS plus TEARDOWN_WAIT_SECONDS, enough for a run to unwind
its steps and then tear down in full, and then shuts down without
whatever is still going.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread src/agent_env/runner/local_runner.py
A run whose step outlasted both of stop()'s waits never reached its own
teardown, so what it had deployed stayed up after shutdown. stop() now
tears down the sandboxes such a run has recorded before leaving it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread src/agent_env/runner/local_runner.py Outdated
Comment thread src/agent_env/runner/local_runner.py Outdated
… a cleanup that fails

Once stop() tears down a run it gave up on, the run's own teardown skips
if its step ever ends, so the two never remove the same sandbox. A
shutdown cleanup that runs out of time or raises is logged with the run
it left.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread src/agent_env/runner/local_runner.py
…nup is over

Skipping an abandoned run's own teardown for good left anything stop()'s
cleanup missed, or the run recorded later, up. The skip now lasts only
while stop()'s cleanup runs; a run that ends after it tears down
whatever is still recorded, and teardown_run passes over what is already
down.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@earakely-scale
earakely-scale merged commit d42dbc6 into main Oct 6, 2026
13 checks passed
@earakely-scale
earakely-scale deleted the edgararakelyan/task-run-teardown branch October 6, 2026 20:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant