Skip to content

🤖 bug: a migration whose exit is never confirmed blocks its workspace's teardown until restart #5589

Description

@ThomasK33

Problem

Since #5579 (#5522), a refused or failed background migration stays pending until its command's exit is confirmed. If that confirmation never comes, the pending entry stays for the rest of the backend process. That happens when the command is stuck in uninterruptible I/O or when a remote runtime's exitCode rejects. This fails closed: no checkout is deleted. But it hurts liveness until the next restart:

  • Removal and archive throw at their 60 s drain deadline with "Retry once the runtime responds". A retry cannot succeed before a restart.
  • Session disposal calls backgroundProcessManager.cleanup(workspaceId) with no deadline (agentSession.ts ~:1816). That call waits forever and keeps its admission seal. So every later move to the background in that workspace is refused (and its command is killed), and the session's final teardown (listener removal, session.end) never runs.

Fix direction (not decided)

Bound the session-disposal drain, or give a migration that is stuck pending a recovery path (for example, settle it once a probe confirms that its process group is gone). The bound must keep the rule that an unconfirmed stop blocks destructive cleanup.

Found by the final check of #5579.


Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $43.48

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions