Follow the repository conventions in CLAUDE.md. Read DESIGN.md before making visual or interaction changes.
For features with background jobs, polling, retries, cancellation, or shutdown:
- Define the lifecycle states and ownership before implementation: pending, running, completed, failed, cancelled, stale, and closing.
- Treat persisted state, in-memory jobs, subprocesses, HTTP requests, and rendered UI as separate state holders. Define how each transitions and settles.
- Never apply a background response without proving it is still current. Use a generation, attempt ID, version, or guarded merge so older polling responses cannot overwrite newer actions.
- When the local state has more than one version counter (for example a task's state version and its plan's review version), an in-flight action owns all of them, and applying its response requires every one to be unchanged. One counter does not cover writes that advance only another.
- A current response may still move a record only along a transition allowed from the state the action was guarded for. Derive the new status from that state as well as the response; a response alone must never move a record out of a state that waits for a person.
- Do not release a concurrency slot when cancellation is requested. Keep the job tracked until its underlying invocation or subprocess has terminated.
- Do not let a retry replace a locally active job, even when its persisted lease has expired or wall-clock time changes.
- Validate retry context against the current snapshot, plan revision, assignment, and referenced code. If any context is stale, disable retry and require a new request.
- Preserve the original timeout, cancellation, and shutdown reason through every layer. Do not replace actionable errors with generic cancellation text.
- Begin shutdown by rejecting new work at the outer admission boundary. Drain already-admitted HTTP requests, then cancel and await jobs, then close storage.
- Bound the HTTP drain during shutdown. After its grace period, abort and await owned work before awaiting server closure so an admitted poll cannot deadlock teardown.
- Polling endpoints should read only the state they need. Do not rebuild Git history or the full review merely to retrieve background-job status.
- Back off recurring external-status polling to a bounded cap. Reset the interval only after a meaningful lifecycle change or explicit user action.
- A background response must not erase text, selections, attachments, navigation changes, or other input made after the request started.
- Clear a submitted draft only if its current value and attachment still match what was submitted. Treat this as compare-and-swap behavior.
- Preserve completed historical results, but visibly mark them stale when their snapshot, plan revision, assignment, or referenced code no longer matches.
- When polling updates one part of the screen, update only that state. Preserve scroll position unless the user was already following the bottom.
- While a request is in flight, do not disable the control that has keyboard focus; disabling it drops focus to the page. Mark it
aria-disabled, ignore repeat activation with an in-flight guard, and test that focus stays on the control after the response. - When a row or control's visual selection determines the current content or input, expose the same state with the appropriate accessibility attribute, such as
aria-currentoraria-selected, and test it across navigation.
Before opening or updating a PR for asynchronous behavior, test every applicable interleaving with controllable promises, clocks, and partial requests:
- Poll starts, then a user action completes, then the old poll returns.
- A job lease expires, then retry is attempted while the original job still runs.
- Timeout fires, then the provider remains unsettled temporarily, then retry is attempted.
- Shutdown starts, then a new request arrives.
- A request is partially received, then shutdown starts, then the request completes.
- An abort error fires, then subprocess close arrives later.
- Submit starts, then the user edits the composer or switches items, then the response returns.
- Referenced code is reassigned or the snapshot changes, then retry or rendering occurs.
- A synchronous caller hook inside a lifecycle (such as a callback that records a spawned process group) throws, blocks past the deadline, or aborts the signal, then the lifecycle continues. Timers and listeners that bound the lifecycle must already be armed when the hook runs, and a signal already aborted when its listener is attached must still take effect.
- A child process exits, but a descendant outside its process group still holds its output pipes.
Every reproduced race requires a failing-before and passing-after regression. Assert both the visible result and the durable state when they can diverge.
- Before requesting or re-requesting an automated Copilot review, self-review the full current diff, fix every issue found, and repeat the self-review and fix cycle until a complete pass finds no new issues. Re-run the relevant validation after fixes, then complete the mandatory independent-agent review loop below before requesting Copilot review.
- Each self-review and independent-agent review pass examines the full current PR diff against the base, including changed tests, configuration and documentation, and rereads every changed function in full. Do not review only the fixes or lines changed since the previous round. Code unchanged since the first commit of the PR still gets reviewed in every pass.
- Treat every behavioural claim the change makes, in code comments, the PR body or docs (for example "pauses every 1,000 entries", "settles only after exit", "never throws", "bounded by N seconds"), as something to verify. Trace each claim through every path that can break it, including nested loops, callbacks, error paths and early returns, and give it a test that fails if the claim is false.
- A test's setup must leave the state the production path would: if production never runs a step (such as a commit that refreshes Git's index), the test must not run it before the behaviour under test either.
- Independent-agent review is mandatory before every Copilot review request, including the first request and every re-request, for all changes, including documentation-only changes. The author's self-review is not a substitute. Use a reviewer that did not implement the change and does not share the author's working context:
/codex review, a separate review agent, or/code-reviewathigheffort or above. - Repeat independent-agent review -> address findings -> re-run relevant validation -> independent-agent review until a complete pass on the latest head reports no new issues and no earlier valid findings remain unresolved. A pass that found issues is not clean merely because they were fixed afterward; the fixes require another independent full-diff pass. Do not stop after a fixed number of rounds.
- Only after that independent loop is clean may Copilot review be requested on the same reviewed and validated head. If Copilot finds issues, including summary-only concerns, address them and re-run relevant validation, then repeat the independent-agent review/fix loop until clean before requesting another Copilot review. Repeat this sequence until Copilot also completes a review with no new issues and no earlier valid findings remain unresolved. Never request Copilot in parallel with an unfinished independent-review cycle.
- A clean review applies only to its exact base/head pair. Any subsequent change, including documentation or review-lesson updates, rebases or base integration, invalidates that result. Complete the independent-agent review/fix loop and relevant validation again for the new pair before requesting Copilot.
- If an independent-agent or Copilot review cannot complete because of quota, timeout, tool failure or unavailability, report the blocked gate. A missing or incomplete review is not a clean result; do not bypass the independent-review prerequisite.
- Run final validation against the exact pushed head after the last change.
- Report current test counts separately from historical milestone counts.
- Before requesting automated Copilot review, report the current head, CI state, mergeability, unresolved threads, deferred follow-up issues, and the latest clean independent-agent review with its exact base/head pair.
- In evidence records, label cited commits as baselines, intermediate checkpoints, or validated heads. Keep final exact-head results in a place that can name the resulting commit, such as the PR body or CI record.
- Reproduce summary-only review concerns or turn them into a concrete follow-up issue. Do not repeatedly patch vague wording without a failure case.
- A validation fixture for a summary-only concern must assert the disputed intermediate representation or state before using a downstream outcome as evidence that the concern was exercised.
- For each self-review, independent-agent and Copilot round, record the review type and reviewer, exact base/head pair, findings, what changed, what was declined and why, and the regression evidence. Record the clean independent-agent pass that authorizes each Copilot request; re-requests must follow the same prerequisite.
- Treat review-lesson extraction as a merge gate. Before invoking merge, classify every review finding in the PR body as: covered by an existing rule (cite it), captured by a new rule in this branch (cite it), or one-off (record why). Do not merge until this audit is complete and every required
AGENTS.mdupdate is included in the reviewed head. Omit rules that merely repeat existing guidance.
- When asked to start, continue or resume an issue or pull request, first locate any existing PR, branch, worktree and active task for it. Fetch current remote refs and inspect the exact base/head pair, working-tree status and untracked files before creating another branch or implementation.
- Do not duplicate work merely because it is absent from the current checkout. Inspect open PRs, issues and worktrees, and identify what remains uncovered.
- Before starting parallel work, compare touched files and shared dependencies. Prefer non-overlapping work; when overlap is necessary, name one integration owner and define ownership of shared files.
- Report focused tests, affected-suite tests, the full suite, type checking, CI and real-Docker validation as separate gates. Passing one does not imply that another passed, and a focused suite never substitutes for a required full-suite, CI or real-Docker gate.
- When validation fails because of machine load, Docker exhaustion, timeout or infrastructure instability, record the observed failure separately from product failures. Targeted checks may diagnose the change, but the required gate must eventually pass or remain explicitly blocked.
- Give heavyweight Docker and integration validation one active owner per repository. Do not launch competing runs that make their results unreliable.
- Distinguish requested, started, pushed, review-pending, checks-pending, mergeable, merged and deployed states. Never collapse an intermediate state into completion.
- After opening, updating, retargeting, marking ready, closing or merging a pull request, read the remote record back and confirm the exact head and resulting state before reporting success.
- A successful command, a merge button click, a green local run or another agent's status summary is evidence of that step only; it is not proof of the downstream result.
- Before changing a stacked PR, record its current parent, base and head, plus any overlapping files with adjacent PRs.
- When resolving overlap, preserve the intended behavior, documentation and regression coverage from both sides. Do not resolve shared-file conflicts by blindly choosing one branch.
- A parent merge, rebase, restack or base change invalidates prior review and validation evidence. Re-run the applicable validation and independent-review loop on the new base/head pair.
- Do not restack dependent PRs until the parent merge is confirmed remotely.
- Any operating rule duplicated between
AGENTS.mdandCLAUDE.mdmust be changed in both files in the same commit. Validate that mirrored sections remain textually identical.
- A bounded safety scan must fail closed when its limit is exceeded. Never truncate evidence and report the result as clear.
- Align subprocess output limits with every payload the schema accepts, or tighten the upstream page and field bounds; valid bounded input must not fail only because the transport budget is smaller.
- Pass text whose size follows user or agent input to a subprocess on stdin, never as an argument: the OS limits one argument's size and cannot pass a NUL, so valid bounded input can fail to spawn.
- Exclude the subject of a duplicate or supersession check by stable identity only. A shared branch name or other mutable attribute does not prove two records are the same subject.
- Exclude the subject's own records only in the states the exclusion is for. A check that skips its own open PR must still count its own merged PR, or its own close of the issue, as done.
- Preserve repository identity with pull request numbers in cross-reference scans. Never resolve or exclude a repository-qualified reference by number alone.
- When a relation can be added and removed (a manually linked PR, a label, an assignment), replay its add and remove events in order and count only its latest state. An add event alone does not prove the relation still holds.
- After the final asynchronous external validation, re-read the local generation immediately before an irreversible action. A generation check performed before that await is insufficient.
- Check an operation's source-state preconditions before any shortcut or early return that writes state or reports success, not only on the main path.
- Honour a cancellation signal that is already aborted before the first durable write, not only after awaits: a path with no await otherwise writes after the caller cancelled.
- Batch and briefly cache read-only status probes, and give the combined operation an overall deadline below the serving request timeout.
- Count the subprocess shutdown grace periods (SIGTERM-to-SIGKILL wait, pipe drain) inside that overall deadline: the operation ends when its processes have settled, not when the abort fires.
- After an external state change (ready, draft, close), read the record back and require the new state before recording success; a command that exits 0 does not prove the change applied.
- Budget a multi-stage validation across all sequential stages; giving each stage the full request allowance does not create an overall deadline.
- Preserve the distinction between an explicit unbound identity and missing or malformed authorization metadata. Missing or malformed identities must fail closed.
- Validate every field used to classify an external record as clear, including enum values and required nullable fields. Partial records and malformed policy objects must fail closed.
- Select every member the API documents for a union or enum you classify (for example every GraphQL
Closertype). Fail closed only on types outside that list; a missing known member turns every result into a false unknown. - Validate coupled lifecycle fields as allowed combinations. A terminal-looking conclusion must not override an active or unknown status.
- Treat a successful external command as the transition it actually performed. If it can enqueue or schedule work, model and verify that lifecycle before reporting the final action as complete.
- For safety-critical API responses, require and validate every requested field before any early return, including terminal-success paths. Treat an omitted field differently from an explicit
nullallowed by the API contract. - Once an irreversible external command succeeds, do not convert later refresh or rendering failures into action failure. Return the committed result, keep repeat controls disabled, and require fresh confirmation of terminal state.
- Track an in-flight irreversible subprocess as part of server shutdown. Abort it, await its settlement, and only then close the state it depends on.
- Set the shutdown admission flag before snapshotting active work, and enforce it again at the irreversible action boundary for requests admitted before shutdown began.
- An approval must not override server-computed plan-scope violations. Irreversible gates must block attributed out-of-scope files until the plan is amended.
- When startup acquires a store, process, listener, or other resource before later dependency construction, close that resource on every construction failure. Prefer validating dependencies before acquisition when possible.
- When deriving a review configuration for a clone, experiment, fork, or new identity, clear external action bindings unless they are re-established and validated for the derived target.
- A deadline must abort and await the underlying operation before releasing its in-flight ownership; rejecting only the caller can leave untracked work running.
- Invalidate pre-action status caches after both successful and refused external mutations before rendering or fetching status again. Use a generation guard so reads started before or during the mutation cannot repopulate the cache afterward.
- Keep irreversible integrations disabled in demo mode even when configuration or an injected dependency is present. After a stale or refused irreversible action, keep its control disabled until fresh state is loaded.
- When a durable external-action attempt is bound to an older snapshot, require approvals or evidence recorded against the replacement generation before another action, whether or not the prior outcome explicitly requested fresh review. A mismatch with the old context is not itself fresh review.
- If the external lifecycle mechanism or mode changes between validation passes, abort before the irreversible command. Create durable lifecycle ownership from the final stable mode, never from an earlier observation.
- When an irreversible command has an ambiguous timeout, cancellation, transport, or unknown outcome, retain durable in-flight ownership and reconcile external state before enabling retry. Only a confirmed refusal may become retryable failure.
- Record in-flight ownership before the first external write of an operation, not before a later step: when a push moves an open PR's head, the refresh it belongs to is already in flight.
- Order an operation's external writes so that any step the remote can definitely refuse comes before the writes that cannot be taken back. A refusal after a push or a description change leaves a half-updated record.
- Correlate retry observations to the current attempt with an immutable external identity or event boundary, and fail closed when multiple post-boundary action sequences appear. Matching only the resource or commit identity can replay another attempt's terminal event.
- When recovery looks for the result of an earlier attempt, recognise the identities of all earlier attempts, abandoned ones included. A result that appears after its attempt was abandoned must be adopted, not treated as foreign, or recovery can never settle.
- Read an identity marker only from its reserved position (for example the first line of a description). Text around it may be agent-controlled, so a marker-shaped string elsewhere must never identify or disqualify a record.
- When the external record is observed, correct the local copy of any field the operation owns (such as a draft flag) from it. A change that landed but whose confirmation was lost is otherwise never repaired, because the next run sees nothing left to change.
- Make an idempotency key required at the API boundary for every replayable action, and look up its saved outcome before any other guard, including in-flight, validation and coordinator shutdown guards. The one exception is the server's HTTP 503 during shutdown, which applies nothing; the client must keep the key and resend it. Save every definite outcome under the key (refusals before admission too), keep the saved response current with the durable outcome it reports, and replay failures as failures with complete result fields. Only a passing, nothing-applied outcome (shutdown, abort, deadline, storage error) stays resendable.
- When a refusal must also change durable state (for example, handing an expired task to a person), commit that change outside the refused transaction, together with the saved refusal. Never write it inside the transaction the refusal rolls back.
- Treat the cleanup handle of an external resource (container, volume, network, temporary directory) as owned state. If removal fails, keep the handle, record it durably before its in-memory owner can be dropped (shutdown, crash, abandon, restart), and fail closed until removal is confirmed. Never delete the durable evidence before the final release report has been saved.
- Give every subprocess an explicit allowlisted environment. Pass credentials only to the component that needs them, through a separate channel. Name-based scrubbing of an inherited environment is not isolation. Run Git with the repository's hardened invocation: no user or system config, no hooks, no lazy fetch, no network protocols.
- Build that allowlist from each tool's documented credential and configuration channels on every supported platform (for example, the D-Bus session bus that a Linux keyring uses), and test that each is passed.
- Treat paths read from a durable record or discovered on disk as untrusted. Before deleting, opening or probing one, validate its exact location and name, not only its basename, and never follow a link to it. Keep files that other local users must not plant or swap, such as lock files, in a directory only the current user can write. Write durable records through a unique temporary file opened exclusively, and delete it if the write fails.
- Exclude other processes with an OS-level lock held for the owner's lifetime, keyed by the resource's stable identity rather than a path spelling. A PID liveness check never authorizes taking over a lock. Run shared one-time startup work single-flight under that lock, and keep the lock until the work has finished.
- When a subprocess reports a problem only as a warning and carries on, decide pass or fail by what each message means for the result, not by whether anything was printed: fail on messages that mean it did less than it should (for example could not read a path), and let through messages about harmless input the agent controls. Test both a benign case and a failing case, and filter the output as it arrives so that no volume of benign messages can push a failure out of a bounded buffer.
- Quote or escape agent-controlled text (file names, paths, branch names) wherever it lands in output that people or tools parse, such as diffs, notices, logs or reports, so it cannot forge that output's structure.
- Neutralise issue references (
#N,GH-N,owner/repo#N, issue URLs) in any text codeboost writes that can become a commit message, such as a PR title or description, and neutralise @-mentions in any of that text that is not fenced (a title). Code fences do not protect commit messages, and GitHub closes issues from closing keywords in default-branch commits. - Never spread a collection whose size follows unbounded input into function arguments (
Math.max(...runs)); engines limit the argument count, so use a loop. - A hardened Git invocation must also keep Git out of nested repositories and populated submodules, whose own config and hooks are the agent's: pass
--ignore-submoduleson the command line (the config default does not bind plumbing or override.gitmodules), and never run Git with a nested repository as its working directory. - Decide a filesystem fact (whether a name exists, where a link resolves, what a path's entries are) on the file system that will serve it, not on a copy. A host that folds case or Unicode normalization, or decodes names as text, answers differently from the container's; check links and names inside the container, as its kernel resolves them, by raw bytes.
- Never infer that agent-written content is unchanged from timestamps, sizes or an index's stat cache: not every write moves them (tmpfs, for one). Compare the content itself.
- Hand agent-chosen names to tools as literal data: disable or sidestep pathspec magic (a leading
:), quote them in the tool's own input format (a leading"in--stdin-paths), pass them as bytes, and keep empty fields when splitting NUL-separated output. Check a name's safety only where it is reported, never refusing unchanged input for a name the report never carries. - Check a trust root before anything reads it. When a tool reads config or rules from storage an agent could have reached (Git and the metadata volume), compare that storage with its trusted baseline first, and run nothing if it changed.
- When one run reports a result and then acts on the same data, serialize the report before the action reads it: a read that changes data (Perl autovivification, a lazy default) must not change what the caller checks.
- When the runner predicts what an external tool will produce (a commit ID from its inputs), refuse up front any input the tool would normalize (Git writes a
-0000offset as+0000and re-encodes Unicode noncharacters), so the prediction cannot fail later on valid-looking input.
- Keep experimental PRs as drafts with automated review disabled until the assigned human decision is recorded. An automated review invalidates reviewer blindness; replace the affected package rather than reusing it.
- If an experiment is cancelled, record it as cancelled rather than passed, remove it from roadmap prerequisites, and track any future validation as explicitly non-blocking. Do not infer product claims from preparation work or incomplete trials.