Repository navigation
feat(bundle): check where a run deploys before writing anything, and build the infra envs it needs - #59
Merged
Conversation
…build the infra envs it needs A plan-time walk, shared by the run and the dry run, resolves each deploy's sandbox provider as its step does and refuses, before the lock or any write: an image only this machine has on a provider that isn't the local one (every provider of a chain must pass), a VM sandbox on a provider that can't create one, containers on the local provider when docker doesn't answer, and a step or judge that names no agent when the default agent isn't in the store. The infra envs a gateway deploy on the local provider needs (the gateway, the service-db for local Postgres state, the website browser for websites) are built once the writes are done and before any task starts: when missing, or built by another agent-env release, for this host's platform, into local stores only, each id locked so concurrent runs build it once. `agent-env up` uses the same in-process bootstrap. The infra put commands' bodies become library functions in `agent_env.env.bootstrap`, the per-id locks move to `store.local_state`, and the build-metadata helpers to `utils.build_metadata`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…oy, and say why Modal can't serve websites From review: - An env type with a deploy() of its own (a plugin's env, which may set up a gateway itself) crashed the walk; it's now left alone, as deploy_env's own preflight leaves it. So is an env the server provider refuses (a multi or website env), which deploy_env's preflight reports. - A website env on Modal is refused because Modal's gateway runs in containers and can't serve websites, rather than asking for a website browser env that wouldn't help. - An agent whose image is another of the bundle's writes gets that write's own refusal instead of a lookup in the store. - A run that needs no infra no longer reads the configured stores. - A version rebuild only replaces an infra env agent-env built from the Dockerfile it ships; one built from someone's own Dockerfile is left alone. - The provider-resolution test also pins which default each deploy falls back to. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ds for amd64 and arm64 Chrome has no Linux arm64 build, so a website browser built for an Apple Silicon host, as the bootstrap builds it, failed at `playwright install chrome`. The image now installs Playwright's Chromium and the MCP server starts it with `--browser chromium`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lth check
The gateway's compose health-checks every MCP server with a python3 socket
probe and starts only once each one passes, but the website browser image
(node:22-slim) has no python3. So its check never passed and a gateway deploy
with websites failed at `docker compose up` ("dependency website-browser failed
to start"). The image now installs python3-minimal.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
earakely-scale
marked this pull request as ready for review
October 6, 2026 05:15
…ice-db env too A deploy_env step attaching an existing state instance (env_state_instance_id) was taken to need no service-db, but the gateway reads it whenever the store is local Postgres, the attached instance's type included. The walk now reads the attached instance's own type. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t on every release infra_to_build rebuilt any infra env another agent-env release had built, so the first gateway run after every upgrade rebuilt the gateway, the service-db's three images and the browser, though they rarely change; two installs sharing a state root rebuilt on every switch; and `up` needed docker after each upgrade. A put now records a digest of what agent-env builds the env from (the shipped Dockerfiles, the files they copy and the build args) as build_inputs_sha256, and an env built from the shipped Dockerfiles is rebuilt only when its recorded digest isn't this release's. Envs put before this have no digest and are rebuilt once. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t it records Passing build_inputs_sha256 through --metadata replaced the digest the put computed, so the next local run would rebuild an unchanged image. The computed digest now comes last. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…first Every gateway deploy on a remote provider runs the same infra images, so a refused bundle repeated each infra problem once per task: 12 lines for 2 tasks, about 250 for 50. Each problem is now reported once, at the first deploy that has it, with how many more do. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Collaborator
Author
Verification on 4eedc13Real runs on the local sandbox provider, each from a fresh state root where it matters: all of them on macOS arm64, and 1, 2, 4 and 7 again on Ubuntu x86_64. The bundle deploys the integration suites' Slack MCP server and Slack website as store envs; a second bundle deploys a VM sandbox.
No sandbox folder or container was left after any of them. Found and fixed in 4eedc13: a refused run repeated each shared infra problem once per task, 12 lines for the 2-task bundle on |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
agent-env runnow checks where each task deploys before it writes anything, and builds the infra envs a gateway deploy on the local sandbox provider needs, so a clean install can run a bundle that deploys an MCP server env.Until now, both kinds of problem surfaced only mid-run, after the bundle's writes and often after a sandbox had started:
Modal sandbox create failed, or a tarball crawling through exec chunks);hello --sandbox modal);Script failed (exit 1)atdocker load);deploy_agentstep that names no agent, on a store without the default agent;Env default not found), which onlyagent-env upcould build, and only with its extra and a config file.What it does
Preflight (
bundle/preflight.py). One walk over the tasks a run will run, in the run and the dry run alike, before the bundle's lock and writes.deploy_env,deploy_agent,deploy_sandbox, and a rubrics judge it deploys itself) resolves its sandbox provider as its step does when it runs:--sandbox, else the step's field, else the config default. A test pins the walk to each step's own resolution.file://tarball. That covers a bundle agent built from its Dockerfile, an agent's or env's store image, the infra envs' images, and adeploy_sandboximage.docker infodoesn't answer. A VM sandbox on the local provider is a work folder, sohellostill needs no Docker.sandbox_typethat names no provider.Infra envs (
env/bootstrap.py). The gateway, the service-db that local Postgres state runs from, and the website browser a gateway adds for websites.build_inputs_sha256, and the run compares it with this release's, so an upgrade that doesn't touch the infra rebuilds nothing. Envs put before this have no digest and are rebuilt once. One built from someone's own Dockerfile is left alone.up.agent-env upuses the same in-process bootstrap instead of a subprocess per put command.The website browser image.
--browser chromium. Chrome has no Linux arm64 build, so on Apple Silicon the image, built for the host, failed atplaywright install chrome.docker compose up(dependency website-browser failed to start). The image now installspython3-minimalitself. (The commit message for this says the image never had python3; it did, through Chrome.)Moved. The behaviour is the same, apart from the service-db put's progress lines, which are worded a little differently.
store/local_state.holding_locks, so the bootstrap can use them.utils/build_metadata.py;agent_env.cli.utilsre-exports them.Behaviour changes
agent-env upwith stores that aren't local refuses a missing gateway or service-db env, naming its put command. Before, it built them into those stores.agent-env upbuilds for this host's platform. Before, it built linux/amd64, which runs emulated on Apple Silicon. Nothing outside this machine can use these images either way.agent-env uprebuilds infra whose build inputs changed (and, once, infra put before agent-env recorded them).--metadatabefore building rather than after.Limits
file://tarball. A configuration that mixes a local registry with a remote object store is refused on remote providers even where a VM provider could load the tarball.task runandeval rundon't bootstrap; onlyagent-env runandupdo.Tests
modaland onlocal,modal; a remote image passes;deploy_sandboxloopback image refused;modaland chains refused, but not onmodal_vmor local;sandbox_type;--sandbox);up's output and its one-line error.End to end
From a clean state with the default local stores, I put the integration suites' Slack MCP server as a store env. The bundle's task deploys it with
sandbox_type: local.agent-env run <bundle> --dry-runWould build first: service-db env 'default-db' (missing),gateway env 'default' (missing); nothing written--dry-run --sandbox modal, no infra yetagent-env run <bundle>--dry-run --sandbox modal, infra presentA website env, on Linux amd64 and on macOS arm64. From a clean state, I put the integration suites' Slack website. A driver then bootstraps the infra through
ensure_default_envs, which builds the gateway, the service-db and the Chromium browser for the host. It deploys the website env on the local sandbox and drives the site through the gateway's tools, as the gateway integration test does.list_website_urlsbrowser_navigate/browser_snapshotbrowser_closeEach took about 2 minutes. Without
python3-minimal, the Chromium image failed the gateway's health check on Linux, so the deploy failed atdocker compose up.🤖 Generated with Claude Code
The PR appears safe to merge; no outstanding issue was found.
Summary
Bundle runs now check whether selected tasks can deploy before writing bundle data, then build any missing local gateway infrastructure before tasks start.
agent-env upshares that bootstrap, which also refreshes agent-env-built images when their build inputs change.Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR A[Plan selected tasks] --> B[Check deployment providers and images] B --> C{Checks pass?} C -->|No| D[Stop before writes] C -->|Yes| E[Write bundle] E --> F[Build needed local infra] F --> G[Run tasks]Reviews (6) · Last reviewed commit: "fix(bundle): report a problem several de..."