Skip to content

feat(bundle): check where a run deploys before writing anything, and build the infra envs it needs - #59

Merged
earakely-scale merged 9 commits into
mainfrom
edgararakelyan/run-preflight-and-infra
Oct 6, 2026
Merged

earakely-scale merged 9 commits into
mainfrom
edgararakelyan/run-preflight-and-infra

Conversation

@earakely-scale

@earakely-scale earakely-scale commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

agent-env run now checks where each task deploys before it writes anything, and builds the infra envs a gateway deploy on the local sandbox provider needs, so a clean install can run a bundle that deploys an MCP server env.

Until now, both kinds of problem surfaced only mid-run, after the bundle's writes and often after a sandbox had started:

  • an image only this machine has, deployed on another provider (Modal sandbox create failed, or a tarball crawling through exec chunks);
  • a VM asked of a provider that can't create one (hello --sandbox modal);
  • containers on the local provider with Docker stopped (Script failed (exit 1) at docker load);
  • a deploy_agent step that names no agent, on a store without the default agent;
  • a missing gateway or service-db env (Env default not found), which only agent-env up could build, and only with its extra and a config file.

What it does

Preflight (bundle/preflight.py). One walk over the tasks a run will run, in the run and the dry run alike, before the bundle's lock and writes.

  • Each deploy (deploy_env, deploy_agent, deploy_sandbox, and a rubrics judge it deploys itself) resolves its sandbox provider as its step does when it runs: --sandbox, else the step's field, else the config default. A test pins the walk to each step's own resolution.
  • A comma-separated provider is a chain whose deploys fall back from one provider to the next, so every provider in it must pass.
  • Refused:
    • An image only this machine has on a provider that isn't the local one. "Only this machine" means a reference to a loopback registry, or a file:// tarball. That covers a bundle agent built from its Dockerfile, an agent's or env's store image, the infra envs' images, and a deploy_sandbox image.
    • A VM sandbox on a provider that can't create one (Modal containers, and any chain).
    • Containers on the local provider when docker info doesn't answer. A VM sandbox on the local provider is a work folder, so hello still needs no Docker.
    • A deploy step or judge that names no agent, when the default agent isn't in the store. The step deploys the configured default, which no ref names, so nothing checked it before.
    • A website env on Modal, whose gateway runs each server in a container of its own and can't serve websites.
    • A step's sandbox_type that names no provider.
  • A problem several deploys share, such as an infra image every gateway deploy runs, is reported once, at the first deploy that has it, with how many more do.
  • Nothing deployed through a plugin's env provider, an env type with a deploy() of its own, or a plugin step is checked.

Infra envs (env/bootstrap.py). The gateway, the service-db that local Postgres state runs from, and the website browser a gateway adds for websites.

  • Which ones. The walk names the infra each gateway deploy on the local provider needs:
    • the service-db only for local Postgres state;
    • the website browser only when the env has websites.
  • When they're built.
    • One is built when the store doesn't hold it, or when agent-env built it from the Dockerfiles it ships and what it was built from has changed. A put records a digest of its build inputs (the shipped Dockerfiles, the files they copy, the build args) as build_inputs_sha256, and the run compares it with this release's, so an upgrade that doesn't touch the infra rebuilds nothing. Envs put before this have no digest and are rebuilt once. One built from someone's own Dockerfile is left alone.
    • The build happens once the run's writes are done and before any task starts, for this host's platform.
    • Each id is locked while it's checked and built, so concurrent runs build it once.
  • Stores that aren't local. agent-env never builds infra into them. One missing there is refused, naming its put command; one another release built there is used as it is.
  • Remote providers. A deploy on another provider needs infra that provider can reach. Missing or local infra is refused there, since a local build wouldn't help.
  • Dry run. It lists what the run would build first.
  • up. agent-env up uses the same in-process bootstrap instead of a subprocess per put command.

The website browser image.

  • Chromium. It installs Playwright's Chromium and starts the MCP server with --browser chromium. Chrome has no Linux arm64 build, so on Apple Silicon the image, built for the host, failed at playwright install chrome.
  • python3. The gateway's compose health-checks every MCP server with a python3 socket probe and starts only once each passes. Chrome's package pulled python3 into the image; Playwright's Chromium doesn't, so the Chromium image failed that check and a gateway deploy with websites failed at docker compose up (dependency website-browser failed to start). The image now installs python3-minimal itself. (The commit message for this says the image never had python3; it did, through Chrome.)

Moved. The behaviour is the same, apart from the service-db put's progress lines, which are worded a little differently.

  • The three infra put commands' bodies are now library functions the commands call.
  • The per-id file locks move from the bundle ledger to store/local_state.holding_locks, so the bootstrap can use them.
  • The build-metadata helpers move to utils/build_metadata.py; agent_env.cli.utils re-exports them.

Behaviour changes

  • agent-env up with stores that aren't local refuses a missing gateway or service-db env, naming its put command. Before, it built them into those stores.
  • agent-env up builds for this host's platform. Before, it built linux/amd64, which runs emulated on Apple Silicon. Nothing outside this machine can use these images either way.
  • agent-env up rebuilds infra whose build inputs changed (and, once, infra put before agent-env recorded them).
  • The infra put commands check --metadata before building rather than after.
  • The website browser is Chromium instead of Chrome, for envs built from this release on.

Limits

  • The image check reads references only: a loopback registry host, or a file:// tarball. A configuration that mixes a local registry with a remote object store is refused on remote providers even where a VM provider could load the tarball.
  • Platforms aren't recorded or checked. Remote providers can't reach locally built images yet, so a platform mismatch can't arise.
  • task run and eval run don't bootstrap; only agent-env run and up do.

Tests

  • Walk:
    • local images refused on modal and on local,modal; a remote image passes;
    • a bundle-built agent refused remotely, and a deploy_sandbox loopback image refused;
    • VM sandboxes on modal and chains refused, but not on modal_vm or local;
    • an unknown sandbox_type;
    • the Docker probe, skipped for a VM sandbox;
    • the default agent for a deploy step and for a judge, and judges that need none;
    • infra named for local gateway deploys (service-db only with local state, the browser only with websites), rebuilt when their build inputs changed, refused in stores that aren't local, and missing or unreachable for remote providers;
    • the dry-run listing, and that a run builds its infra after its writes and before its tasks;
    • each deploy resolves its provider as its step does (4 steps × default, step field, --sandbox);
    • a problem several deploys share, reported once.
  • Bootstrap:
    • missing infra built in order, host-native, once;
    • rebuilt once when their build inputs change or were never recorded, and left alone when built from someone else's Dockerfile;
    • a put records its build inputs, and the digest covers the build args but not Python caches;
    • nothing built without Docker, or into stores that aren't local;
    • up's output and its one-line error.
  • The Docker probe's answers.
  • Unit tier: 5,835 passed.

End to end

From a clean state with the default local stores, I put the integration suites' Slack MCP server as a store env. The bundle's task deploys it with sandbox_type: local.

command result
agent-env run <bundle> --dry-run lists Would build first: service-db env 'default-db' (missing), gateway env 'default' (missing); nothing written
--dry-run --sandbox modal, no infra yet refused: the env's image is in this machine's registry, and Modal needs the gateway and service-db envs, which the store doesn't hold
agent-env run <bundle> builds the service-db and gateway envs for this host, then deploys the env on the local sandbox through the gateway: passed (unscored, no verifier), torn down, 80 s in all
the same run again builds nothing; deploys in 24 s
--dry-run --sandbox modal, infra present refused, naming the env's image and each of the four infra images as in this machine's registry

A website env, on Linux amd64 and on macOS arm64. From a clean state, I put the integration suites' Slack website. A driver then bootstraps the infra through ensure_default_envs, which builds the gateway, the service-db and the Chromium browser for the host. It deploys the website env on the local sandbox and drives the site through the gateway's tools, as the gateway integration test does.

host browser image list_website_urls browser_navigate / browser_snapshot after browser_close
Ubuntu 22.04, x86_64 (an EC2 devbox) amd64 the frontend's URL the "Slack Workspace" page gone
macOS, Apple Silicon arm64 the frontend's URL the "Slack Workspace" page gone

Each took about 2 minutes. Without python3-minimal, the Chromium image failed the gateway's health check on Linux, so the deploy failed at docker compose up.

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no outstanding issue was found.

Summary

Bundle runs now check whether selected tasks can deploy before writing bundle data, then build any missing local gateway infrastructure before tasks start. agent-env up shares that bootstrap, which also refreshes agent-env-built images when their build inputs change.

  • Bundle runs check deploy requirements before writing anything.
  • Local gateway deploys build the infrastructure they need before tasks start.
  • The website browser runs Chromium with the gateway's health-check tool.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Plan selected tasks] --> B[Check deployment providers and images]
  B --> C{Checks pass?}
  C -->|No| D[Stop before writes]
  C -->|Yes| E[Write bundle]
  E --> F[Build needed local infra]
  F --> G[Run tasks]
Loading

Reviews (6) · Last reviewed commit: "fix(bundle): report a problem several de..."

earakely-scale and others added 4 commits October 5, 2026 20:33
…build the infra envs it needs

A plan-time walk, shared by the run and the dry run, resolves each deploy's
sandbox provider as its step does and refuses, before the lock or any write:
an image only this machine has on a provider that isn't the local one (every
provider of a chain must pass), a VM sandbox on a provider that can't create
one, containers on the local provider when docker doesn't answer, and a step or
judge that names no agent when the default agent isn't in the store.

The infra envs a gateway deploy on the local provider needs (the gateway, the
service-db for local Postgres state, the website browser for websites) are
built once the writes are done and before any task starts: when missing, or
built by another agent-env release, for this host's platform, into local stores
only, each id locked so concurrent runs build it once. `agent-env up` uses the
same in-process bootstrap. The infra put commands' bodies become library
functions in `agent_env.env.bootstrap`, the per-id locks move to
`store.local_state`, and the build-metadata helpers to `utils.build_metadata`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…oy, and say why Modal can't serve websites

From review:
- An env type with a deploy() of its own (a plugin's env, which may set up a
  gateway itself) crashed the walk; it's now left alone, as deploy_env's own
  preflight leaves it. So is an env the server provider refuses (a multi or
  website env), which deploy_env's preflight reports.
- A website env on Modal is refused because Modal's gateway runs in containers
  and can't serve websites, rather than asking for a website browser env that
  wouldn't help.
- An agent whose image is another of the bundle's writes gets that write's own
  refusal instead of a lookup in the store.
- A run that needs no infra no longer reads the configured stores.
- A version rebuild only replaces an infra env agent-env built from the
  Dockerfile it ships; one built from someone's own Dockerfile is left alone.
- The provider-resolution test also pins which default each deploy falls back to.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ds for amd64 and arm64

Chrome has no Linux arm64 build, so a website browser built for an Apple
Silicon host, as the bootstrap builds it, failed at `playwright install chrome`.
The image now installs Playwright's Chromium and the MCP server starts it with
`--browser chromium`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lth check

The gateway's compose health-checks every MCP server with a python3 socket
probe and starts only once each one passes, but the website browser image
(node:22-slim) has no python3. So its check never passed and a gateway deploy
with websites failed at `docker compose up` ("dependency website-browser failed
to start"). The image now installs python3-minimal.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@earakely-scale
earakely-scale marked this pull request as ready for review October 6, 2026 05:15
@earakely-scale
earakely-scale requested a review from a team as a code owner October 6, 2026 05:15
Comment thread src/agent_env/bundle/preflight.py Outdated
earakely-scale and others added 3 commits October 5, 2026 22:21
…ice-db env too

A deploy_env step attaching an existing state instance (env_state_instance_id)
was taken to need no service-db, but the gateway reads it whenever the store is
local Postgres, the attached instance's type included. The walk now reads the
attached instance's own type.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t on every release

infra_to_build rebuilt any infra env another agent-env release had built, so
the first gateway run after every upgrade rebuilt the gateway, the service-db's
three images and the browser, though they rarely change; two installs sharing
a state root rebuilt on every switch; and `up` needed docker after each upgrade.
A put now records a digest of what agent-env builds the env from (the shipped
Dockerfiles, the files they copy and the build args) as build_inputs_sha256,
and an env built from the shipped Dockerfiles is rebuilt only when its recorded
digest isn't this release's. Envs put before this have no digest and are
rebuilt once.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread src/agent_env/env/bootstrap.py Outdated
earakely-scale and others added 2 commits October 6, 2026 08:20
…t it records

Passing build_inputs_sha256 through --metadata replaced the digest the put
computed, so the next local run would rebuild an unchanged image. The computed
digest now comes last.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…first

Every gateway deploy on a remote provider runs the same infra images, so a refused bundle repeated each infra
problem once per task: 12 lines for 2 tasks, about 250 for 50. Each problem is now reported once, at the first
deploy that has it, with how many more do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@earakely-scale

Copy link
Copy Markdown
Collaborator Author

Verification on 4eedc13

Real runs on the local sandbox provider, each from a fresh state root where it matters: all of them on macOS arm64, and 1, 2, 4 and 7 again on Ubuntu x86_64. The bundle deploys the integration suites' Slack MCP server and Slack website as store envs; a second bundle deploys a VM sandbox.

# check result
1 first run, empty store builds the service-db, the gateway and the website browser, then deploys both envs; each records a digest matching this release's
2 the same run again builds nothing
3 infra put by v0.9.1270, so no digest each rebuilt once (built before agent-env recorded its build inputs), then nothing
4 a gateway source file changed only the gateway rebuilt (its build inputs changed), and the new image has the change
5 two first runs at once, needing different infra each infra env built once: one run printed waiting for another agent-env run… and built only what was left
6 Ctrl-C during the service-db build, and as the gateway build starts Aborted!, no build left running, no half-written env, locks released; the next run built what was missing and passed
7 docker info failing refused before any write, run and dry run, with infra present or missing; the VM sandbox task still ran
8 --sandbox modal, modal_vm, local,modal, e2b, an unknown name each refused before any write as intended; the VM task passes on modal_vm and local and is refused on modal and the chain
9 agent-env up already registered when current; builds from empty; with Docker down, a one-line error naming --no-bootstrap, which then starts
10 a user's own env under an infra id (their Dockerfile, or no build metadata) left alone

No sandbox folder or container was left after any of them.

Found and fixed in 4eedc13: a refused run repeated each shared infra problem once per task, 12 lines for the 2-task bundle on modal. Each problem is now reported once, with how many more deploys have it: 8 lines there, and 8 for a 5-task bundle, which printed 27 before.

@earakely-scale
earakely-scale merged commit 5279acb into main Oct 6, 2026
14 checks passed
@earakely-scale
earakely-scale deleted the edgararakelyan/run-preflight-and-infra branch October 6, 2026 16:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant