Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
43 commits
Select commit Hold shift + click to select a range
988137f
Add production pre-merge preparation
mchwang Oct 7, 2026
9d7351b
Harden pre-merge preparation lifecycle
mchwang Oct 7, 2026
5b1d4b3
Close pre-merge review lifecycle gaps
mchwang Oct 7, 2026
b11a8ee
Close pre-merge authorization and cleanup gaps
mchwang Oct 7, 2026
715ab39
Guard final pre-merge readiness races
mchwang Oct 7, 2026
c016a54
Close pre-merge settlement safety gaps
mchwang Oct 7, 2026
d2c5065
Bound pre-merge shutdown settlement
mchwang Oct 7, 2026
a258492
Bind merge admission to preparation readiness
mchwang Oct 7, 2026
c6fd5c3
Reauthorize merge and bound rewrite scans
mchwang Oct 7, 2026
7b5c112
Close pre-merge timeout and retry gaps
mchwang Oct 7, 2026
2c5c910
Bound pre-merge reload and retry settlement
mchwang Oct 7, 2026
88200c6
Reject normalized command arguments
mchwang Oct 7, 2026
9068862
Keep legacy command reviews repairable
mchwang Oct 8, 2026
a0acf18
Seed fresh runner repositories for review
mchwang Oct 8, 2026
49beda1
Merge current main into F5
mchwang Oct 8, 2026
3d6f0ff
Preserve pre-merge lifecycle history
mchwang Oct 8, 2026
d062643
Merge remote-tracking branch 'origin/main' into codex/f5-rebase-admis…
mchwang Oct 8, 2026
cee9e01
Close final rebase admission races
mchwang Oct 8, 2026
fb2187d
Seal final admission validation windows
mchwang Oct 8, 2026
f4ab96b
Keep merge refresh authorization-scoped
mchwang Oct 8, 2026
d212471
Address F5 review findings
mchwang Oct 8, 2026
c0262e6
Revalidate trust after pre-merge inspection
mchwang Oct 8, 2026
11c3853
Keep settled rebase failures current
mchwang Oct 8, 2026
f94501b
Bind queue checks after final authorization
mchwang Oct 8, 2026
3458a84
Bound final merge and rebase admission
mchwang Oct 8, 2026
f1eaea1
Share preparation validation deadlines
mchwang Oct 8, 2026
7780422
Bind preparation policy and cleanup budget
mchwang Oct 8, 2026
1be0f9c
Keep unsettled rebase failures current
mchwang Oct 8, 2026
81c9c1d
Retain refreshed remote commits
mchwang Oct 8, 2026
3f3fb43
Preserve command-check exit semantics
mchwang Oct 8, 2026
1e6cb57
Merge remote-tracking branch 'origin/main' into codex/f5-rebase-admis…
mchwang Oct 8, 2026
629a933
Retain startup snapshots and cleanup budget
mchwang Oct 8, 2026
a857be3
Merge remote-tracking branch 'origin/main' into codex/f5-rebase-admis…
mchwang Oct 8, 2026
9a00e47
Keep cleanup failures current for recovery
mchwang Oct 8, 2026
3adc132
Fail closed on pre-merge lifecycle gaps
mchwang Oct 9, 2026
d7109d8
Bound multibyte command diagnostics
mchwang Oct 9, 2026
e713ed7
Merge remote-tracking branch 'origin/main' into codex/f5-rebase-admis…
mchwang Oct 9, 2026
232102b
Close pre-merge recovery races
mchwang Oct 9, 2026
966c38c
Reserve rebase cleanup handoff
mchwang Oct 9, 2026
265d4cd
Arm conflict deadline guard before starting bounded work
mchwang Oct 9, 2026
d5628a9
Keep later attribution choices binding across head moves
mchwang Oct 9, 2026
111a0fe
Count every attribution choice against execution approvals
mchwang Oct 9, 2026
d638eb9
Bind pre-execution approvals to the current snapshot
mchwang Oct 9, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 7 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

Review agent-made Git changes one plan item at a time. The approved plan lists each item's files and acceptance checks; the review engine shows which item produced each change and flags foreign or overlapping work.

**Status:** the plan/linking library, SQLite store, and local review screen are implemented. Run `npm run demo` and open its private local URL. Ask runs Claude Code inside the locked-down agent container for read-only answers; choose it in Settings. Codex is refused in every phase for now, including Ask, because it reads code only by running commands ([why](docs/implementation/agent-isolation.md#codex-is-refused-in-every-phase)). Ask needs Docker, plus `CLAUDE_CODE_OAUTH_TOKEN` (from `claude setup-token`). The first question builds the agent image, which can take a few minutes. A configured GitHub review can merge only after the guarded exact-head gate passes. The agent container, vendor-only network and Claude/Codex adapters are implemented ([agent isolation](docs/implementation/agent-isolation.md)); only Ask uses them so far. Automated rebasing, plan command execution, and code-writing agents are not implemented. The paired human review experiment was cancelled before results were recorded and no longer blocks roadmap work; optional future validation is tracked in [#19](https://github.com/codeabovelab/codeboost/issues/19).
**Status:** the plan/linking library, SQLite store, local review screen, opt-in code-writing runner, durable local rebase, and exact-head plan command checks are implemented. Run `npm run demo` and open its private local URL. Ask and the runner use Claude Code inside locked-down containers; Codex is refused because it reads code only by running commands ([why](docs/implementation/agent-isolation.md#codex-is-refused-in-every-phase)). They need Docker, plus `CLAUDE_CODE_OAUTH_TOKEN` (from `claude setup-token`). Startup builds the agent image, which can take a few minutes. A configured GitHub review can merge only after the guarded exact-head gate passes. The agent container, vendor-only network, credential-free command-check profile, and Claude/Codex adapters are described in [agent isolation](docs/implementation/agent-isolation.md). Pushing a rebased head, waiting for its required checks, and handing that exact pair to guarded merge remain follow-up work. The paired human review experiment was cancelled before results were recorded and no longer blocks roadmap work; optional future validation is tracked in [#19](https://github.com/codeabovelab/codeboost/issues/19).

## Development

Expand All @@ -14,7 +14,7 @@ npm run typecheck
npm test
```

Tests create disposable local repositories. They do not invoke agents, access GitHub, or execute plan acceptance commands.
Tests create disposable local repositories and exercise command checks in disposable credential-free containers. They do not invoke model providers or access GitHub.

## Guarded GitHub merge

Expand All @@ -31,7 +31,7 @@ An existing-store configuration may add a trusted GitHub binding:
}
```

The issue must match the stored plan. `pullRequest` is required unless the configuration has a `runner` block (see below). The authenticated `gh` account must be able to read the pull request, issue timeline, applicable rulesets, and classic branch protection, and to merge the PR. Codeboost unions required checks from both rule sources, requires strict server-enforced current-base checks, rechecks the base and head immediately before merging, and passes the reviewed head to `gh pr merge --match-head-commit`. Missing permissions and ambiguous rule responses block the merge. When the branch uses a merge queue, codeboost adds the exact reviewed head to the queue and treats the merge as done only when GitHub confirms it merged; a queued pull request is not merged. If GitHub removes it from the queue while the plan, snapshot, review and head are unchanged, Merge offers a retry after the status refresh; if the head changed, it goes back to review first. A moved base and any unexecuted `cmd:` acceptance check remain blocked until [#22](https://github.com/codeabovelab/codeboost/issues/22) adds the runner path.
The issue must match the stored plan. `pullRequest` is required unless the configuration has a `runner` block (see below). The authenticated `gh` account must be able to read the pull request, issue timeline, applicable rulesets, and classic branch protection, and to merge the PR. Codeboost unions required checks from both rule sources, requires strict server-enforced current-base checks, rechecks the base and head immediately before merging, and passes the reviewed head to `gh pr merge --match-head-commit`. Missing permissions and ambiguous rule responses block the merge. When the branch uses a merge queue, codeboost adds the exact reviewed head to the queue and treats the merge as done only when GitHub confirms it merged; a queued pull request is not merged. If GitHub removes it from the queue while the plan, snapshot, review and head are unchanged, Merge offers a retry after the status refresh; if the head changed, it goes back to review first. With the opt-in runner, `prepare-merge` handles a moved base and runs every allowed `cmd:` acceptance check on the resulting exact head; merge remains blocked until that evidence is current.

## Runner (opt-in)

Expand All @@ -51,7 +51,7 @@ A configuration with a `github` block may also add a `runner` block. The runner
- `diagnosticsCapBytes`: the size retention trims that directory back to. It is a target, not a hard limit: the file just saved is always kept, even when it alone passes the cap. The default is 256 MiB.
- `limits`: task storage limits (`workBytes`, `workInodes`, `metadataBytes`, `metadataInodes`).

The runner needs Docker and `CLAUDE_CODE_OAUTH_TOKEN`. At start, before the server opens, codeboost recovers what an earlier run left and builds the agent image; a recovery it cannot finish safely stops startup with what to do. `POST /api/runner` with `start` or `resume` runs the plan; the review screen has no runner buttons yet. A task paused for an out-of-scope change stays paused until the plan is amended, every changed path is declared, the remaining items validate against the runner's audited tree, and a person sends `approve-continuation` with the current `expectedStateVersion`, `expectedReviewVersion`, and an action ID. Then `resume` starts at the first unfinished item. `GET /api/runner` reports the checkpoint and whether the continuation is approved. When a run has completed every item, codeboost pushes the task head to a `codeboost/…` branch with your `gh` credentials and opens a ready pull request into `baseBranch`; a task that needs a person gets a draft pull request with its problems. A publish that was refused (for example, the branch holds a commit codeboost did not make) is retried with the `publish` action, and `GET /api/runner` shows the last outcome under `publish`. When a task is cancelled, codeboost closes the open pull requests it opened for it (the branch is kept). A pull request a person moved is left open and reported, and `close-pull-requests` retries a close that failed. Merge PR then merges that pull request, so `github.pullRequest` can be left out; if it is set, it must name the task's pull request, or merging is blocked. Demos never publish.
The runner needs Docker and `CLAUDE_CODE_OAUTH_TOKEN`. At start, before the server opens, codeboost recovers what an earlier run left and builds the agent image; a recovery it cannot finish safely stops startup with what to do. `POST /api/runner` with `start` or `resume` runs the plan; the review screen has no runner buttons yet. A task paused for an out-of-scope change stays paused until the plan is amended, every changed path is declared, the remaining items validate against the runner's audited tree, and a person sends `approve-continuation` with the current `expectedStateVersion`, `expectedReviewVersion`, and an action ID. Then `resume` starts at the first unfinished item. `GET /api/runner` reports the checkpoint and whether the continuation is approved. When a run has completed every item, codeboost pushes the task head to a `codeboost/…` branch with your `gh` credentials and opens a ready pull request into `baseBranch`; a task that needs a person gets a draft pull request with its problems. Before merge, `prepare-merge` refreshes the published PR's exact base/head, rebases onto a moved base, refreshes attribution and approvals, and runs allowed `cmd:` checks in a credential-free read-only container. The check result counts only for that exact head and command list. A publish that was refused (for example, the branch holds a commit codeboost did not make) is retried with the `publish` action, and `GET /api/runner` shows the last outcome under `publish`. When a task is cancelled, codeboost closes the open pull requests it opened for it (the branch is kept). A pull request a person moved is left open and reported, and `close-pull-requests` retries a close that failed. Merge PR then merges that pull request, so `github.pullRequest` can be left out; if it is set, it must name the task's pull request, or merging is blocked. Demos never publish.

## Library

Expand Down Expand Up @@ -87,8 +87,8 @@ Inputs such as `planText` and the ledger must come from the trusted runner. `run
- The caller selects and trusts the repository and its Git administrative directory. Normal Git discovery, linked-worktree gitfiles, and symlinked gitdirs are supported; object-storage links and alternates inside that selected gitdir are rejected. This adapter is not a filesystem-containment boundary for untrusted repository roots.
- Git administrative metadata and object storage must remain unchanged during a read; these library checks do not isolate a concurrently hostile filesystem.
- Ownership uses line diffs, not semantic inference. Within one replacement block, new lines inherit all affected owners conservatively. Function context comes from Git hunk headers, not an AST.
- The importer requires accurate typed base entries, stable plan identity, a selected issue, and a trusted checkout path-identity function. It rejects path traversal, Git metadata paths, and traversal through a listed file/symlink/submodule. Runtime symlink and write-scope enforcement belong to the future container/runner; plan validation alone is not a sandbox.
- Allowed commands restrict accidents, not hostile programs or changed scripts. Parsing returns argv and never executes it. An unlisted valid command is a warning and must not run until allowed.
- Container isolation, vendor-only network access and credential handling are implemented by the lane D boundary (`agents/`), not by this library. Ask runs in that boundary in the read-only "questions" phase: it sees a clone of the reviewed head, supplied review context, and nothing else from your computer. Safe dependency installation is not implemented.
- The importer requires accurate typed base entries, stable plan identity, a selected issue, and a trusted checkout path-identity function. It rejects path traversal, Git metadata paths, and traversal through a listed file/symlink/submodule. The opt-in runner adds runtime symlink and write-scope enforcement; plan validation alone is not a sandbox.
- Allowed commands restrict accidents, not hostile programs or changed scripts. The plan library only parses them into argv. The opt-in runner executes an exact allowed argv only in its credential-free command-check container; an unlisted command never runs.
- Container isolation, vendor-only network access and credential handling are implemented by the lane D boundary (`agents/`), not by this library. Ask, code-writing runs, conflict resolution, and credential-free command checks use that boundary with phase-specific mounts and tools. Safe dependency installation is not implemented.

See [implementation decisions and evidence](docs/implementation/build-step-1.md) and the [plan format](docs/plan-format.md).
27 changes: 27 additions & 0 deletions agents/adapters/runner.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
import type { InvocationHandle } from '../contract.ts';
import { readCapturedFile } from '../container/profile.ts';
import { createPhasePolicy, createRunnerCommand, MAX_COMMAND_SCHEMA_BYTES } from '../policy.ts';
import { launchInvocation } from './supervisor.ts';
import { setUpProfile } from './setup.ts';
import { assertAdapterRequest, capAdapterInvocationBudget, createAdapterInvocationBudget,
type AgentAdapterOptions, type AgentAdapterRequest } from './types.ts';
import { join } from 'node:path';

/** Run exact approved argv arrays in the read-only review container, without a provider credential or external host. */
export function startRunnerCommandInvocation(request: AgentAdapterRequest,
options: AgentAdapterOptions = {}): InvocationHandle {
if (request.invocation.vendor !== 'runner' || request.invocation.phase !== 'review')
throw new Error('Command checks require a runner-owned review invocation.');
const policy = createPhasePolicy(request.invocation);
const remaining = options.invocationBudget
? capAdapterInvocationBudget(options.invocationBudget, options.timeoutMs)
: createAdapterInvocationBudget(request.invocation, options.timeoutMs);
const raw = readCapturedFile(join(request.inputDirectory, 'schema.json'), 'Runner command input', MAX_COMMAND_SCHEMA_BYTES).content;
const commands = new TextDecoder('utf-8', { fatal: true }).decode(raw);
const command = createRunnerCommand(policy, commands);
assertAdapterRequest(request);
return launchInvocation(request.invocation, remaining, (signal, start) => setUpProfile(request, remaining, signal,
network => ({ ...request, policy, network, command }),
profile => start(profile, { ...options, processLifecycle: request.processLifecycle, invocationBudget: remaining,
diagnosticOutput: true })));
}
43 changes: 34 additions & 9 deletions agents/adapters/supervisor.ts
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,8 @@ export interface SupervisorOptions {
readonly secrets?: Readonly<Record<string, string>>;
readonly timeoutMs?: number;
readonly limits?: Partial<CaptureLimits>;
/** Keep bounded, lossy diagnostics without letting command output change the process exit result. */
readonly diagnosticOutput?: boolean;
/** Trusted monotonic budget carried from adapter setup. */
readonly invocationBudget?: () => number;
readonly decode?: (profile: ContainerProfile, rawStdout: Buffer, maximumBytes: number,
Expand Down Expand Up @@ -480,6 +482,9 @@ export function startProfileInvocation(profile: ContainerProfile, options: Super
throw new Error('invocationBudget cannot exceed the production ten-minute ceiling.');
if (options.processLifecycle && profile.deferredOutput)
throw new Error('Durably tracked invocations cannot use concurrent deferred-output controls.');
if (options.diagnosticOutput && (invocation.vendor !== 'runner' || invocation.phase !== 'review'
|| options.decode || profile.deferredOutput))
throw new Error('Lossy diagnostic output is reserved for runner command checks.');
} catch (error) {
return rejectProfile(profile, error, true, disposeContainerProfile,
isDeadlineError(error) ? 'timeout' : undefined, options.processLifecycle);
Expand All @@ -498,6 +503,7 @@ export function startProfileInvocation(profile: ContainerProfile, options: Super

const stdoutChunks = new ByteCollector(), stderrChunks = new ByteCollector();
let stdoutBytes = 0, stderrBytes = 0, combinedBytes = 0;
let outputTruncated = false;
let stopReason: StopReason | undefined, failureDetail: string | undefined;
let closed = false, terminating = false, settlementComplete = false;
let decodedOutput: DecodedOutput | undefined, decodePromise: Promise<void> | undefined;
Expand Down Expand Up @@ -595,7 +601,8 @@ export function startProfileInvocation(profile: ContainerProfile, options: Super
combinedBytes += retained.length;
}
if (chunk.length > available) {
if (final) stopReason ??= 'output-limit';
if (options.diagnosticOutput) outputTruncated = true;
else if (final) stopReason ??= 'output-limit';
else stop('output-limit');
}
};
Expand Down Expand Up @@ -777,20 +784,38 @@ export function startProfileInvocation(profile: ContainerProfile, options: Super
});
}
}
// Publish only strictly valid UTF-8: replacement characters would grow the result past the byte ceilings.
// Only output cut at a capture limit may end in an incomplete character, which the streaming decode then drops;
// otherwise the decode flushes, so a trailing lone lead byte fails closed.
// Agent answers publish only strictly valid UTF-8: replacement characters would grow the result past the byte
// ceilings. Runner command output is diagnostic rather than an answer, so it is decoded lossily and bounded again.
// Only strict output cut at a capture limit may end in an incomplete character, which streaming decode drops;
// otherwise strict decode flushes, so a trailing lone lead byte fails closed.
const truncated = stopReason === 'output-limit';
const strictText = (value: Buffer) => {
try { return new TextDecoder('utf-8', { fatal: true }).decode(value, truncated ? { stream: true } : undefined); }
catch { return undefined; }
};
const stdoutText = strictText(finalStdout), stderrText = strictText(finalStderr);
if (stdoutText === undefined || stderrText === undefined) {
stopReason ??= 'capture-failure';
failureDetail ??= 'Captured output is not valid UTF-8.';
const boundedDiagnostic = (value: Buffer, maximum: number) => {
const encoded = Buffer.from(new TextDecoder('utf-8').decode(value));
if (encoded.length <= maximum) return encoded;
return Buffer.from(new TextDecoder('utf-8', { fatal: true }).decode(encoded.subarray(0, maximum), { stream: true }));
};
if (options.diagnosticOutput) {
finalStdout = boundedDiagnostic(finalStdout, Math.min(limits.stdoutBytes, limits.combinedBytes));
finalStderr = boundedDiagnostic(finalStderr,
Math.max(0, Math.min(limits.stderrBytes, limits.combinedBytes - finalStdout.length)));
if (outputTruncated) {
const notice = Buffer.from('[codeboost: command output truncated]\n');
const maximum = Math.max(0, Math.min(limits.stderrBytes, limits.combinedBytes - finalStdout.length));
finalStderr = maximum <= notice.length ? notice.subarray(0, maximum)
: Buffer.concat([boundedDiagnostic(finalStderr, maximum - notice.length), notice]);
}
} else {
const stdoutText = strictText(finalStdout), stderrText = strictText(finalStderr);
if (stdoutText === undefined || stderrText === undefined) {
stopReason ??= 'capture-failure';
failureDetail ??= 'Captured output is not valid UTF-8.';
}
finalStdout = Buffer.from(stdoutText ?? ''); finalStderr = Buffer.from(stderrText ?? '');
}
finalStdout = Buffer.from(stdoutText ?? ''); finalStderr = Buffer.from(stderrText ?? '');
if (stopReason) finalStderr = withDiagnostic(finalStderr, finalStdout.length, stopReason, limits, failureDetail);
const result = Object.freeze({ attemptId: invocation.attemptId, context: invocation.context,
exitCode, signal: finalSignal, ...(stopReason ? { stopReason } : {}),
Expand Down
1 change: 1 addition & 0 deletions agents/container/Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ RUN npm install --global --allow-scripts=@anthropic-ai/claude-code \
&& install --directory --owner=10001 --group=10001 --mode=0755 /work /work/.git

COPY --chmod=0555 container/probe.sh /usr/local/bin/codeboost-container-probe
COPY --chmod=0555 container/command-check.mjs /usr/local/bin/codeboost-command-check
COPY --chmod=0444 network/proxy.mjs /usr/local/lib/codeboost-egress-proxy.mjs

LABEL org.opencontainers.image.base.name="docker.io/library/node:26.7.0-bookworm@sha256:e929171d35b9df7773a3ec5b068e387fa109441dc90f91e6560af5d39b7e9bf1" \
Expand Down
Loading
Loading