Skip to content

fix(artifact): bound image saves and recover failed publication - #89

Open
morluto wants to merge 4 commits into
scaleapi:mainfrom
morluto:codex/fix-artifact-publication
Open

morluto wants to merge 4 commits into
scaleapi:mainfrom
morluto:codex/fix-artifact-publication

Conversation

@morluto

@morluto morluto commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Closes #88

Apply one deadline to nonblocking reads from both Docker-save pipes, stream gzip output, and retain a bounded 16 KiB stderr tail. On failure, terminate the owned process group, close the parent pipe ends, and reap the save process. File and byte uploads now reuse the existing attempt-isolated locator scheme so a failed document insert cannot block the next publication.

Reproduced failure Before After
Stalled save with saturated stderr; 0.2 s deadline Still blocked at 2 s; external watchdog required Fails in 0.206 s
Retry after document insert fails Three retries collide with orphan object Next retry publishes version 1
Temporary archive after failed save Cleanup unreachable while blocked Removed
Retained stderr Unbounded Last 16 KiB; no stderr staging file
Detached child retains stdout/stderr Hang or cleanup failure Returns in 1.501 s at a 1.5 s deadline; archive removed

Measured with local subprocesses and real SQLite/filesystem stores. Tests cover stalled stdout, saturated stderr, surviving descendants, compression errors, valid 16 MiB gzip output, concurrent writes, and both file-upload APIs.

Only newly published object locators change. Existing saved locators still load. Failed publication can leave an isolated object; this change makes retry safe rather than adding garbage collection.

Validation: 6,332 unit/protocol tests passed, 13 existing capability skips; plugin API check passed. Independent correctness and maintainability reviews are clean. Docker/cloud integration remains for CI.

RetriggerConfidence Score: 5/5

The PR appears safe to merge; the earlier publication hang and cleanup findings are addressed.

What we checked:

  • Docker save runs without a shell: The save passes the image name as one argument to Docker. The timeout lines only raise an exception.

Summary

Docker image saves now use one deadline and clean up failed archives. File uploads use a fresh object location for each attempt, so a failed document write does not block a retry.

  • Docker image saves finish or clean up under one deadline.
  • File uploads get a fresh object location for every try.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Start Docker save] --> B[Read stdout and stderr]
  B --> C[Write gzip archive]
  C --> D{Save finished in time?}
  D -- Yes --> E[Publish archive]
  D -- No --> F[Kill and reap save]
  F --> G[Remove temporary archive]
Loading

Reviews (3) · Last reviewed commit: "fix(artifact): enforce deadlines for inh..." · Reviewed by Greptile

@morluto
morluto requested a review from a team as a code owner October 7, 2026 01:36
Comment thread src/agent_env/artifact/artifacts/docker_image.py Outdated
Comment thread src/agent_env/artifact/artifacts/docker_image.py Outdated
Comment thread src/agent_env/artifact/artifacts/docker_image.py Outdated

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(artifact): bound image saves and recover interrupted file publication

1 participant