Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
f91089a
fix(check): canonicalise threshold aliases at the parse boundary
dekobon Aug 1, 2026
ca1ce91
fix(check): scale the soft tier by metric direction
dekobon Aug 1, 2026
6ec2b96
fix(check): send offender rows to stdout
dekobon Aug 1, 2026
9035491
fix(check): accept a rationale on a suppression marker
dekobon Aug 1, 2026
845c3db
feat(check): preview a candidate limit at both tiers
dekobon Aug 2, 2026
ecebdaa
chore(baseline): stop git text-merging the baseline file
dekobon Aug 2, 2026
6245f57
fix(baseline): record a start_line only where it disambiguates
dekobon Aug 2, 2026
7b9a5ed
chore(dev): add an idempotent worktree bootstrap
dekobon Aug 2, 2026
19f53d8
chore(make): end both gates with one verdict line
dekobon Aug 2, 2026
fa79ede
docs: record the batch-session operational traps
dekobon Aug 2, 2026
be45b88
docs(changelog): consolidate entries from the #1165-#1174 batch
dekobon Aug 2, 2026
5f3e6f1
refactor(baseline): split path and body-hash units out
dekobon Aug 2, 2026
3f04efb
refactor: fold repeated pipelines and orderings
dekobon Aug 2, 2026
880115d
refactor(check): reuse TierSpec::ratio in the preview
dekobon Aug 2, 2026
8811246
fix: resolve twelve findings from two branch reviews
dekobon Aug 2, 2026
abdcc71
test: make six batch assertions able to fail
dekobon Aug 2, 2026
9ecd36d
fix(check): snap a scaled floor already on the sig-fig grid
dekobon Aug 2, 2026
2298588
fix(suppression): drop rationale after a bare verb
dekobon Aug 2, 2026
9aec606
test: derive the corpus fixture suffix instead of spelling it
dekobon Aug 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
195 changes: 7 additions & 188 deletions .bca-baseline.toml

Large diffs are not rendered by default.

11 changes: 6 additions & 5 deletions .claude/hooks/bca-guidance.txt
Original file line number Diff line number Diff line change
Expand Up @@ -17,11 +17,12 @@ simpler -- not the number smaller.
- When the complexity is essential, suppress with a reason. Some
functions are irreducibly complex and clearest left whole -- a
dispatch match, a hand-rolled parser table, an exhaustive state
machine. For these, add a suppression marker with a one-line
rationale: // bca: suppress(<metrics>) inside the function (canonical
names; it is nexits, never exit -- an unknown metric voids the whole
marker). A clear function with an honest suppression beats a
"compliant" tangle.
machine. For these, add a suppression marker with the rationale on the
same line: // bca: suppress(<metrics>) -- why, inside the function.
Anything after the metric list is free text. Use canonical names; it
is nexits, never exit -- an unknown name warns and is skipped, and the
names beside it still suppress. A clear function with an honest
suppression beats a "compliant" tangle.
- Keep the fix where the violation is. The flag is scoped to the
function you just edited. Fix it there, mention anything larger you
noticed, and do not widen the change into a module rewrite to bring
Expand Down
50 changes: 50 additions & 0 deletions .claude/rules/tool-output.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,50 @@
# Tool Output Rule

A truncated tool result is not the result. When output is too large to
inline, the harness persists it and shows a fragment headed
`Preview (first 2KB)` alongside the path to the full text. That fragment
is shaped exactly like a finished answer: no marker at the cut, no total
row count, nothing at the end saying more was dropped.

## Why it yields a wrong answer rather than an error

A preview is a prefix, so what survives is decided by output order, and
output order is rarely importance order. The commands used here most
often put the rows you need past the cut at least as reliably as before
it:

- `sort | uniq -c` — the counts you are hunting are the largest, and
they sort **last** unless you passed `-rn`.
- `rg` over a tree — hits arrive in path order, so a sweep across
`src/languages/` or `src/getter/` shows the alphabetically early
languages and hides every one after them.
- `cargo test` / `make pre-commit` — the failure summary is at the end,
behind all the passing output.
- `git log`, `git diff --stat` — the last-listed file is the one cut.

Reading a prefix of any of these produces a coherent, plausible, partial
answer. Nothing in it looks incomplete, which is the whole problem: the
tell that normally prompts a second look is absent.

## What it cost here

During #1127 an agent audited a `sort | uniq -c` tally from the preview
alone, missed two rows, and nearly shipped two tests that could no longer
fail.

## How to apply

- Read the persisted file in full before drawing a conclusion from it. A
preview establishes *that* there is output. It never establishes what
the output says.
- Better, do not generate output you will have to re-read. Aggregate in
the command instead: `| wc -l` for a count, `sort -rn | head` to bring
the interesting rows to the front, `rg -c` in place of `rg`,
`--name-only` in place of a full diff.
- Treat every "there are no other X" claim as requiring the full text.
Absence is precisely what a prefix cannot establish, and it is the
conclusion these sweeps are usually run to reach.
- The same discipline covers filters you added yourself. A `| head -20`
written to keep the output small is a truncation you must account for
when reading the result, and unlike the harness preview it leaves no
trace at all.
111 changes: 104 additions & 7 deletions .claude/skills/batch-fix/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -347,7 +347,7 @@ Use the issue data (title, body, comments) cached from Step 0a to populate
each agent's prompt.

Pass each agent the full agent prompt (see below) with `<ISSUE_NUMBER>`,
`<ISSUE_TITLE>`, and `<ISSUE_BODY>` substituted.
`<ISSUE_TITLE>`, `<ISSUE_BODY>`, and `<INTEGRATION_BRANCH>` substituted.

#### Worktree mode (`ISOLATION_MODE=worktree`)

Expand All @@ -365,12 +365,23 @@ For a **multi-issue wave**: launch ALL agents in a single message block
`model: "opus"`. Do NOT use `run_in_background` -- wait for all agents in
the wave to complete before proceeding.

**Known limitation**: Worktree agents fork from `INTEGRATION_BRANCH` at the
moment they are spawned. Within a single multi-issue wave, agents do not
see each other's in-flight work — they only see the integration-branch tip
that existed when the wave started. The merge in Step 4b reconciles their
results mechanically. For tightly coupled issues that must build on each
other, use `--sequential` or run them as a single `/fix-issue`.
**Worktree agents fork from the base, not from the integration-branch
tip.** Each worktree starts at the commit `INTEGRATION_BRANCH` was created
from in Step 1 — `main` — not at wherever the integration branch has since
advanced to. Every wave therefore begins from the same pristine pre-batch
tree, and a Wave-3 agent cannot see what Waves 1 and 2 already merged.
Serializing same-crate issues into separate waves does not, on its own,
prevent them from colliding: both agents edit the pre-batch file. Disjoint
hunks still merge cleanly in Step 4b. Semantic coupling does not, and it
fails quietly — the later agent writes against a helper the earlier one
replaced, or re-does a fix that already landed.

**Remedy — substitute `<INTEGRATION_BRANCH>` into every worktree-mode
agent prompt.** The prompt's setup step then merges it, which works
because worktrees share the object store (see "Setup — Environment
Verification"). That covers the cross-wave case only. Agents within one
wave run concurrently and still cannot see each other; for tightly coupled
issues use `--sequential`, or run them as a single `/fix-issue`.

#### Branch mode (`ISOLATION_MODE=branch`)

Expand Down Expand Up @@ -429,6 +440,19 @@ improvement over worktree mode where agents fork independently from `main`.

> **Branch mode**: Skip this step — results are processed inline in Step 4a.

**Merge only once every agent in the wave has returned, and only against
a clean integration worktree.** Serialisation is the safeguard here; git
is not. Two things follow:

- `git checkout <INTEGRATION_BRANCH>` fails outright while another
worktree holds that branch (`fatal: '<branch>' is already used by
worktree at …`). That refusal means an agent is still running. Wait for
it and retry — nothing is broken, and there is nothing to repair.
- `git merge` refuses only when the merge would overwrite an uncommitted
change to a file it actually touches. Uncommitted work in any other
file is carried through silently and ends up inside the merge commit.
Do not treat a successful merge as evidence that the tree was idle.

For each agent result in the wave:

The worktree agent returns one of:
Expand Down Expand Up @@ -550,6 +574,27 @@ git checkout <INTEGRATION_BRANCH>
Then run `make pre-commit` — the canonical gate per "Validation gates"
in `AGENTS.md` (it adds udeps, doc warnings, the lint families, the
self-scan gates, and `make snapshot-anchors` on top of the cargo trio).

**Capture it to a per-invocation log and read the verdict from the log**
— never from a reported exit status, which in this harness is the status
of the trailing command on the line, not of `make`:

```bash
log=$(mktemp /tmp/bca-pre-commit.XXXXXX.log)
make pre-commit >"$log" 2>&1
grep '^BCA_GATE:' "$log"
```

`BCA_GATE: pass (gate=pre-commit)` or
`BCA_GATE: fail (gate=pre-commit, exit=2, stage=_pc-fmt)` — exactly one
line, and it is the last thing the gate writes. Neither a `make[2]: ***
… Error 2` line nor a green-looking tail is a verdict: the gate is a
parallel DAG, so a failing stage is reported the instant it fails and
other stages keep producing successful output after it. No `BCA_GATE:`
line at all means the run never finished. A fixed path such as
`/tmp/pc.log` is shared with every concurrent agent on the host and has
already caused one agent to diagnose another's failure — hence `mktemp`.

If `make` is unavailable, fall back to:

```bash
Expand Down Expand Up @@ -735,6 +780,37 @@ your worktree.

In branch mode: the orchestrator has verified a clean repo before launching you.

**In worktree mode, run `make worktree-setup` before your first
`make pre-commit`.** A fresh worktree has neither the integration corpora
under `tests/repositories/` (24 tests fail without them: 5 corpus tests and
19 CLI tests analysing a real `DeepSpeech` source file) nor
`big-code-analysis-py/.venv` (`py-typecheck` reports ~33 mypy errors,
`py-test` dies with "Couldn't find a virtualenv"). Both are bootstrap
artifacts, not regressions in your change. The target is idempotent and a
~100 ms no-op afterwards.

If a corpus checkout is interrupted — it is long enough to hit a command
timeout — the submodule is left with its files deleted but its HEAD already
at the recorded SHA, so **a plain `git submodule update --init` is a silent
no-op**. Re-run `make worktree-setup`; it detects that state and escalates
to `--force`. Do not conclude the corpus is fine because a re-run exited 0.

**In worktree mode, merge the integration branch before you investigate:**

```bash
git merge <INTEGRATION_BRANCH> --no-edit
```

Your worktree forked from the commit the integration branch was created
from, not from its current tip, so without this you are reading the
pre-batch tree and cannot see fixes that earlier waves already merged.
On a fresh worktree the merge is a fast-forward. It is allowed even
though `<INTEGRATION_BRANCH>` is checked out elsewhere: git refuses to
*check out* a branch another worktree holds, not to merge from one.

In branch mode, skip this — your branch was created from the integration
branch's current tip in Step 4a and already contains it.

Try to activate Serena:

```
Expand Down Expand Up @@ -972,6 +1048,27 @@ cargo test --workspace --all-features
pre-commit run --all-files
```

**When you redirect a gate to a log file, use a path that is yours alone
— put your issue number in it.** Every agent in a wave runs on the same
host and shares `/tmp`, so an obvious fixed name such as `/tmp/pc.log`
gets written by all of them at once. That has already
produced a wrong diagnosis here: one agent paired its own exit status
with a neighbour's log, read a `fmt-check` failure out of it, and spent
the attempt on code it had never touched.

```bash
make pre-commit > /tmp/pc-<ISSUE_NUMBER>.log 2>&1
grep '^BCA_GATE:' /tmp/pc-<ISSUE_NUMBER>.log
```

Step 6a explains how to read that verdict line and why an exit status
reported by the harness is not one. The requirement here is narrower: the
path must be unique to you. Before believing a failure you did not
expect, confirm the log is describing your tree — a passing `make
pre-commit` run names its project root well over a hundred times, so
`grep -c "$PROJECT_ROOT" /tmp/pc-<ISSUE_NUMBER>.log` returning 0 means
you are reading someone else's run.

If any check fails on code you changed, fix and retry (one attempt).
If it fails again or fails on code you did not change, report as FAILED.

Expand Down
11 changes: 11 additions & 0 deletions .gitattributes
Original file line number Diff line number Diff line change
Expand Up @@ -11,3 +11,14 @@ tree-sitter-preproc/** linguist-vendored
# the byte-exact comparison. Covering all `*.py` forecloses the same trap for
# any future generated-Python gate.
*.py text eol=lf

# The self-scan baseline is generated wholesale by
# `make self-scan-write-baseline-headroom`; nothing in it is
# hand-authored. A textual merge of two branches that both touched a
# baselined file yields hunks that are plausible-looking and wrong on
# *both* sides, because neither side's measurement describes the merged
# tree. `-merge` makes git refuse the text merge and leave the whole
# file conflicted, which says "regenerate" rather than "pick a side".
# Resolve with:
# make self-scan-write-baseline-headroom
.bca-baseline.toml -merge
20 changes: 20 additions & 0 deletions .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -128,6 +128,26 @@ repos:
entry: python3 -m unittest -q utils/check-excluded-manifests-test.py
pass_filenames: false

# Self-tests for the worktree-setup submodule classifier. It is
# what decides when to run a destructive `git submodule update
# --force`, so a logic change there re-runs them (#1171).
- id: worktree-setup-test
name: worktree-setup-test
language: system
files: '^utils/worktree-setup(-test)?\.py$'
entry: python3 -m unittest -q utils/worktree-setup-test.py
pass_filenames: false

# Self-tests for the pre-commit/ci verdict line. gate-status.sh is
# what stands between a red gate and a log that reads as green, so
# a logic change there re-runs them (#1172).
- id: gate-status-test
name: gate-status-test
language: system
files: '^utils/gate-status(-test)?\.sh$'
entry: bash utils/gate-status-test.sh
pass_filenames: false

# Sync-test for check-grammar-crate.py's EXTENSIONS table against
# src/langs.rs (#869). Re-runs when the script, its test, or the
# source-of-truth language table changes.
Expand Down
80 changes: 74 additions & 6 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -139,6 +139,12 @@ and `cargo run -p big-code-analysis-web --`.
fabricated measurement tables here. Build argument lists as arrays and
expand them `"${ARR[@]}"`. See
[`.claude/rules/shell.md`](.claude/rules/shell.md).
- **Tool output**: a truncated result is not a result. Persisted output
shown as a `Preview (first 2KB)` fragment carries no marker at the cut,
and the orderings in routine use here — `sort | uniq -c`, `rg` over a
tree, a test run's trailing summary — put the rows that matter past it.
Read the persisted file, or aggregate in the command. See
[`.claude/rules/tool-output.md`](.claude/rules/tool-output.md).
- **Code search**: `rg` (ripgrep). Never `grep` via Bash.
- **File search**: `fd` (or `fdfind` on Debian/Ubuntu). Never `find` via
Bash.
Expand Down Expand Up @@ -212,6 +218,15 @@ and `cargo run -p big-code-analysis-web --`.

## Validation gates

In a checkout you have not run the gate in before, run
`make worktree-setup` first. It checks out the integration corpora and
the Python-bindings venv; without them `make pre-commit` reports 24 test
failures and ~33 mypy errors that are bootstrap artifacts rather than
regressions. It is idempotent and a ~100 ms no-op afterwards, and it
repairs an interrupted corpus checkout — a state a plain
`git submodule update --init` cannot fix, because the recorded SHA
already matches and the re-run is a silent no-op (#1171).

Before considering a change done, run `make pre-commit` from the repo
root. It is the canonical entry point for the full validation gate
and runs, in one parallel pass: the cargo trio (`cargo fmt --check`,
Expand Down Expand Up @@ -242,6 +257,30 @@ Python stage is skipped with a clear "X not found" message when the
corresponding tool is absent). `make ci` runs the same checks without
auto-fix, mirroring CI behaviour.

**Read the outcome from the `BCA_GATE:` line, nothing else.** Both gates
end with exactly one of `BCA_GATE: pass (gate=pre-commit)` or
`BCA_GATE: fail (gate=pre-commit, exit=2, stage=_pc-fmt)` on stdout, and
no other line of either gate's output starts with that token. Do not
infer an outcome from make's `Error N` lines: the gate is a parallel
DAG, so a stage's `make[2]: *** [… _pc-fmt] Error 2` is printed the
instant it fails and is routinely followed by a hundred lines of other
stages'
*successful* output. Capture each run to its own log path — a fixed
`/tmp/pc.log` is shared with every other checkout and agent on the host
— and grep that log rather than trusting a reported exit status, which
in some tooling is the status of the trailing `echo` rather than of
make:

```bash
log=$(mktemp /tmp/bca-pre-commit.XXXXXX.log)
make pre-commit >"$log" 2>&1
grep '^BCA_GATE:' "$log"
```

No `BCA_GATE:` line at all is a third state, not a pass: the run
crashed, was killed, or was interrupted. See
[`CONTRIBUTING.md`](CONTRIBUTING.md), "Reading the verdict".

**`_native.pyi` is stubtest-gated (#673).** The hand-written PyO3 stub
`big-code-analysis-py/python/big_code_analysis/_native.pyi` is no
longer "kept in lockstep by hand" on trust alone: `make py-stubtest`
Expand Down Expand Up @@ -285,6 +324,29 @@ purely procedural: do not bypass pre-commit, and refresh the baseline
with `make self-scan-write-baseline-headroom` in the commit that moved
the metric. A red gate on `main` traces directly to skipping this step.

The same rule governs **merges**. `.bca-baseline.toml` is marked
`-merge` in `.gitattributes`, so git leaves it wholly conflicted rather
than splicing two branches' entries together. That is deliberate:
neither side's recorded values describe the merged tree, so hand
resolution is always wrong here, not merely tedious. Regenerate with
`make self-scan-write-baseline-headroom` and stage the result.

**Price a candidate limit at both tiers before calling it free.**
Converging a limit onto a cluster of existing values is never free while
a proportional soft tier is active — the soft tier measures *distance to
the limit*, so a limit chosen to sit exactly on a population's value
maximises soft-tier breach by construction, and none of those functions
can ever clear the band because they *are* the limit. The natural
measurement is the misleading one: `bca check --threshold <m>=<limit>`
is applied last and absolutely, never scaled, so it has no soft tier to
report. Use `bca check --explain-threshold <metric>=<limit>`, which
reports both tiers plus how many offenders each already has in the
baseline, and weigh the *new-entry* count. This repo's `nargs 7 → 6`
was approved on a hard-tier zero and would have bought 74 permanent
baseline entries (#1143, #1169); the same trap applies to a
`[thresholds.lang.<slug>]` override (#1141). See
[Tightening a limit onto a cluster](big-code-analysis-book/src/recipes/thresholds.md#converging-onto-a-cluster).

If GNU Make 4 or any of the optional tools (`taplo`, `rumdl`,
`shellcheck`, `shfmt`, `checkmake`, `actionlint`, `cargo-nextest`,
`ruff`, `mypy`, `pyright`, `maturin`) are unavailable, fall back to the
Expand Down Expand Up @@ -474,12 +536,18 @@ smaller.
- **When the complexity is essential, suppress with a reason.** Some
functions are irreducibly complex *and clearest left whole* — a
dispatch `match`, a hand-rolled parser table, an exhaustive state
machine. For these, add an in-source marker with a one-line rationale
rather than contorting the code:
`// bca: suppress(<metrics>)` inside the function (per-file:
`// bca: suppress-file(<metrics>)`). Use canonical metric names — it
is `nexits`, **never** `exit`; an unknown identifier warns *and voids
the entire marker*. `tokens` is not suppressible. See
machine. For these, add an in-source marker rather than contorting
the code, with the rationale on the same line:
`// bca: suppress(<metrics>) — <why>` inside the function (per-file:
`// bca: suppress-file(<metrics>) — <why>`). Anything after the metric
list is free text; no separator is required. **Name the metrics** if
you want to write a reason: a *bare* verb (`// bca: suppress`, no
list) takes no trailing text at all, because nothing distinguishes a
rationale from prose *about* the marker — put the reason on the line
above if the `All` scope is really what you want.
Use canonical metric names — it is `nexits`, **never** `exit`; an
unknown identifier warns and is skipped, while the recognised names in
the same marker still suppress. `tokens` is not suppressible. See
[Suppression markers](big-code-analysis-book/src/commands/suppression.md)
and the full recipe at
[`recipes/agent-feedback.md`](big-code-analysis-book/src/recipes/agent-feedback.md).
Expand Down
Loading