Skip to content

feat(eval): calibrated matcher, fail-closed control verdicts, and a negative result - #68

Merged
ManSio merged 14 commits into
mainfrom
exp/prompt-robustness-v3
Oct 4, 2026
Merged

ManSio merged 14 commits into
mainfrom
exp/prompt-robustness-v3

Conversation

@ManSio

@ManSio ManSio commented Oct 3, 2026

Copy link
Copy Markdown
Owner

Stacked on #67 (guards). Rebased onto chore/bump-urllib3-pyjwt-cve — merge that first.

This PR is, in the end, mostly a negative result with numbers, plus the harness fixes needed to reach it honestly.

The finding

A string matcher is not a measurement instrument for paraphrase invariance on this corpus.

Hand-adjudicating all 17 rows the v3 matcher rejected out of 360 live calls:

count share
Live calls 360 1.000
Rejected by matcher 17 0.047
...of those, answer factually correct 16 false negatives
...of those, answer factually wrong 1 true error caught
  • Signal (real model errors): 1/360 = 0.28%
  • Noise (matcher false negatives): 16/360 = 4.44%
  • Noise-to-signal: 16 : 1

Detecting a 0.28% event requires an instrument below 0.28% error. The best string matcher implemented here sits at 4.44% — 16 times too coarse. Tuning does not change this.

Why calibrating on live answers cannot work

All 120 answers in the frozen live sample are factually correct. There are no negatives. So FP rate = 0/120 is trivially true — the denominator is empty — and any matcher that catches one form scores 1.0/0. That is a tautology, not a test. Model errors live in the non-first repeats, and a "take the first row" selection rule excludes exactly those.

The implication is broader than this corpus: you cannot calibrate a matcher on model output, because the class you must calibrate for (wrong answers) is nearly absent from it. Negatives have to be invented, and then you are calibrating against your own imagination — which is what v2 did, at FP 0.212.

Second defect, unrelated to the matcher

invariance_given_base_correct is conditioned on base_ok, so its denominator is a function of the model: n = 30/21/27/27/24/30. At a fixed 0.90 threshold the gate needs 4 disagreements at n=30 but only 3 at n<=27 — verified arithmetically. Verdict columns are therefore not comparable across models. Worse, the conditioning selects exactly the rows where the model was right, which is where stability is easiest, so weaker models score better: longcat/en has the worst accuracy (0.7) and the best invariance (1.0) with STRICT_PASS.

Harness fixes

  • Control arms returned N/A and exited 0, so a planted 1/30 break (invariance 0.967) scored as PASS — the control could not report its own blindness. Now CONTROL_OK/CONTROL_FAIL with exit 1; injected_hits is a per-run delta because the counter accumulates across models.
  • Live arms now distinguish STRICT_PASS (1.0) from CONDITIONAL_PASS (>=0.90 with the disagreeing rows named for manual review).
  • calibrate_matcher.py is a gate that fails on FN, on FP and on an empty population (exit 0/1/2), with a --selftest proving all three can fail.
  • The v3 matcher replaces loose any_of substrings with three declarative rule kinds. Calibration on 156 labelled items: recall 0.867 → 1.0, FP 0.212 → 0.0.

What this PR does NOT establish

  • Not that string-based matching is impossible in principle — only that it does not work on this corpus with this instrument.
  • Does not evaluate LLM-as-judge or metamorphic rules.
  • Labelling is by a single annotator, with no inter-annotator agreement. My own phrasing was already wrong once (Water boils at 212 degrees Fahrenheit was marked no_match until corrected, recorded as revision 1.1).
  • The train/heldout split in live_sample.json is perfectly confounded with row kind (even indices are all base), so it is systematic, not random. Noted in the frozen hypotheses file.
  • Of the heldout half, one item is no longer blind evidence: it exposed the ё-normalization inconsistency between matcher paths, and the fix changed its result.

Reproduce

python -m pytest experiments/prompt_robustness/test_harness_sanity.py -q
python experiments/prompt_robustness/calibrate_matcher.py --dataset experiments/prompt_robustness/frozen/dataset_v3.json --split all
python experiments/prompt_robustness/calibrate_matcher.py --selftest

MSCodeBase Agent added 14 commits October 3, 2026 06:07
KNOWN_ISSUES.md held 448 lines against a 300-line limit (rule R1), which blocked
every commit on main. 224 of those lines were not board entries: 29 blocks
injected verbatim from AGENT_DIARY.md by AutoDocUpdater during a full reindex,
each carrying "Источник: AGENT_DIARY.md" and "Статус: автоматически
синхронизировано".

A board that mirrors the diary holds every issue twice and cannot be trimmed
without losing information. The diary stays the source; the board keeps the 24
hand-authored sections. Live board is now 224 lines.

scripts/archive_autosync_entries.py does this by marker rather than by date, so
the classification does not rot as the file ages, and a second run finds nothing
to move and changes nothing. Nothing is deleted — the blocks are appended to
docs/archive/KNOWN_ISSUES_2026_10_AUTO_SYNC.md, per §4.8 R4.
pip-audit failed on both test jobs with vulnerabilities that appeared after the
last green main run on 2026-09-29, so this is a time-dependent failure rather
than a regression from any recent change:

  urllib3 2.7.0   PYSEC-2026-4175 / -4176 / -4177   -> 2.8.0
  PyJWT  2.13.0   13 advisories, highest fix 2.15.0 -> 2.15.0

Both are transitive: neither is imported as an API. PyJWT appears in the codebase
only as a token pattern in src/core/redact.py, and the line above the pin claimed
a verified RS256 roundtrip against pyjwt 2.13.0 — that claim is re-verified here
rather than assumed, and the comment is updated to the version actually tested.

Verified after the bump, not before:
  - pip-audit -r requirements-lock.txt --no-deps --disable-pip -> "No known
    vulnerabilities found", rc=0
  - RS256 roundtrip with cryptography 50.0.0, and `import mcp` -> ok
  - pytest tests/ -> 1989 passed, 6 skipped, 98 deselected
  - scripts/smoke_e2e.py -> SMOKE E2E: PASSED, which exercises the real embed
    (llama.cpp), the real rerank (BGE-M3) and a real index search, i.e. the
    network stack urllib3 actually backs
…y guard as UNPROVEN

CI reported `NEGATIVE CONTROLS: FAILED (broken=0, unproven=1)` with
`dead_guard_classifier` marked UNPROVEN — green locally, red on every runner.

Root cause: tests/test_negative_controls_runner.py proved the digest pin by
editing the REAL fixture
scripts/negative_controls/fixtures/dead_guard.py and restoring it in `finally`.
Under `pytest -n auto` another worker read that fixture's digest inside the
window between the write and the restore, computed a different digest, and
classified a healthy guard as UNPROVEN. The digest guard was never wrong; the
race was between two tests, and it presented as a guard defect.

Every digest matched locally and in a clean checkout, which is why this survived
local verification — the trigger is worker count, not content.

The test now copies the fixture into a scratch directory it creates and removes,
and never touches the tracked file. It asserts a PROVEN control before mutating,
so the property under test is still "a byte change flips PROVEN to UNPROVEN".

Adds tests/test_no_tracked_file_mutation.py so the class cannot come back. It is
a test rather than a script because a script nobody runs is not a gate, and its
first version collected zero tests under pytest while reading as coverage.

That guard's selftest earned its place immediately: it caught a real second
instance in tests/test_planted_break_gate.py, which wrote
experiments/planted_break/results.json from two xdist workers with no atomicity.
That write is now temp-file + os.replace, the artifact is gitignored, and the
one exemption is recorded with its reason.

Verified: 1991 passed, 6 skipped, two consecutive runs of the exact CI command
with -n auto.
check_third_party_data: R1 personal email, R2 exported-profile signature, R3 bulk
dump (advisory). Vendored dependency metadata is allowlisted as legitimate
attribution. --selftest proves positive 2/2 and negative 2/2 so a matcher that
cannot fail cannot pass; --all over 1897 tracked files reports blocking=0.

check_parallel_sessions: lists foreign worktrees with branches and uncommitted
files, blocks when a staged file is also modified in another tree, advisory
otherwise. Wired into .githooks/pre-commit; AGENTS.md documents both.

Three defects were found while verifying them. Each is reproduced by a test.

1. GIT_DIR leak. git exports GIT_DIR into the hook environment and _git()
inherited it, so 'git -C <other worktree> status' kept reading the CURRENT
repository. The gate inspected the wrong tree: with GIT_DIR set it reported
three pre-commit-stage files as modified in D:/Project/wt-pr-v3 and exited 1
while that worktree was clean. Bidirectional - it also let real cross-tree
conflicts through. Fix drops GIT_DIR/GIT_WORK_TREE/GIT_INDEX_FILE/
GIT_COMMON_DIR from the child environment.

2. The third-party gate could not be tested at all. Its R1 fixture held a real
person's address (pascal.cescato@gmail.com), so committing the detector would
have published a third party's email - exactly what the rule forbids. Test
fixtures used gmail.com addresses and tripped R1 on synthetic data. Both now
use the RFC 2606 reserved domain personal.invalid, which can never be
registered, and personal_domains is injectable so R1 is still proven to fire.

3. R2 counted 11 profile markers in the detector itself, because PROFILE_MARKERS
lives in that file. R2 is skipped for exactly two paths; R1 still applies to
them, so a dump cannot be hidden there.
gate_zero_full_suite() picked pythonw.exe on Windows to avoid a console flash.
pythonw has no console, so every test that spawns a subprocess (git, lock
helpers, discriminators) dies inside subprocess with OSError [WinError 50].

Measured on the same tree and commit: 'python -m pytest tests/' gives
2006 passed and 0 failed, 'pythonw.exe -m pytest tests/' gives 72 failed and
6 errors, every failure WinError 50 from subprocess.

CREATE_NO_WINDOW already suppresses the console window, so pythonw bought
nothing and cost the entire gate.
…gates

Adds the session's own notes on the parallel-session incidents, the third-party
data audit and the verification gate work.

KNOWN_ISSUES.md was 366 lines against a 300-line limit, so check_known_issues
failed and blocked every commit. Rotated 79 lines of closed entries per the
monthly-archival rule: only sections whose own header carries Fixed/Closed/
REFUTED, no Open and no P1 moved. Closed entries from 2026-09 went to
docs/archive/KNOWN_ISSUES_2026_09.md, the 2026-10-03 one to a new
KNOWN_ISSUES_2026_10.md. Live file is 289 lines; open items untouched.

While rotating, found two pre-existing defects NOT caused by this change and
deliberately left alone: docs/archive/KNOWN_ISSUES_2026_09.md holds the
'Exp E13' section four times over roughly 700 lines in two encodings, and the
live file carries CP1251-read-as-UTF-8 mojibake in the 2026-09-19 onward
headers. Both belong to the registry owner.
run_model() defaulted to base_repeats=2 while the CLI default and the frozen
manifest specify 3, so any caller omitting the argument silently produced 100
rows instead of the declared 120 per model.

manifest.json also now records the harness SHA it actually describes plus the
v2 findings (matcher false negatives, verdict sensitivity, abstention).
The transport read the CLI response from the wrong stream, so every long
output was parsed as empty and rows were silently marked invalid. Tests now
cover banner-on-stderr/body-on-stdout, both RU and EN arms, and the 120-row
per-model expectation implied by base_repeats=3.
360 live calls, 0 invalid, all three models PASS. FINDINGS.md documents the two
meta-results: q_http_404 alias coverage caused 8/360 false negatives, so every
sub-1.0 value measures the matcher rather than the model, and the 0.90 gate
does not see a single 1/30 planted break (0.9667). Values stay quarantined
until dataset v3 recalibrates the matcher.
Matcher: three declarative rule kinds replace loose any_of substrings.
value_context separates boiling from freezing and 100C from 212F;
subject_value requires the year next to its subject, so '1991' in an
unrelated sentence stops counting as an answer about Python; last_mention
resolves hedged answers by the final mention, not the first. Numeric
anchors keep word boundaries, stems stay substrings. yo normalization is
now identical on both matcher paths.

Calibration on 156 labelled items: recall 0.867 -> 1.0, FP 0.212 -> 0.0.
Rules were fitted on the tune half only; the heldout half is reported
separately in results/calibration_v3_heldout.json.

Verdicts: live arms report STRICT_PASS (invariance 1.0) or
CONDITIONAL_PASS (>= threshold, with the disagreeing rows named for
manual review) instead of a single PASS. Control arms report
CONTROL_OK/CONTROL_FAIL and exit 1 on failure; they previously returned
N/A, so a planted 1/30 break scored 0.967 and exited 0 - the control
could not report its own blindness. injected_hits is a per-run delta
because the transport counter accumulates across models.
Adds findings 4-6: substring matching is not calibratable in principle
(recall 0.867, FP 0.212 both directions, hence neither v2 metric class is
publishable), the three declarative rules that replace it, and the control
arm that previously returned N/A and therefore could not report its own
blindness. Records that the heldout half is no longer fully blind for the
one item that exposed the yo-normalization inconsistency.
…N class

360 live calls, invalid=0. All metrics recomputed from raw rows independently,
0 mismatches against the harness summary.

Finding 7: the invariance denominator is a function of the model itself
(n = 30/21/27/27/24/30), because it is conditioned on base_ok which the matcher
defines. At a fixed 0.90 threshold the gate needs 4 disagreements at n=30 but
only 3 at n<=27, so verdict columns cannot be compared across models. The
conditioning also selects exactly the rows where the model was right, which is
where stability is easiest, so weaker models score better: longcat/en has the
worst accuracy (0.7) and the best invariance (1.0) with STRICT_PASS.

Finding 8: v3 fixed the q_http_404 false negatives it was built for, but its
bare rule only fires on exact equality with one anchor, so a correct answer
carrying a unit, a symbol or a case ending (100C, 100C (212F), '1912 godu.',
'Kislorod (O2)') fails the require check. Four case-cells regressed.

Finding 9: recall 1.0 on 156 self-authored items was real and not predictive.
Every positive I wrote carried the predicate; the models answer with the bare
value. Calibrating against one's own phrasing distribution calibrates against
a distribution that does not exist.
…to signal

Cheap experiment on 120 already-collected live answers (no new model calls).
Sample frozen before labelling with matched/agrees_with_base stripped.

Finding 10: all 120 live answers are factually correct, so FP is not measurable
- its denominator is empty. H1 and H2 are therefore degenerate rather than
confirmed: any matcher reaches recall 1.0 on positives and FP 0 because there
is nothing to catch. Model errors live in non-first repeats, and the
first-row selection rule excluded exactly those.

Finding 11: hand-adjudicating all 17 rows the v3 matcher rejected gives 16
false negatives against 1 genuine model error. Signal 1/360 = 0.28 percent,
noise 16/360 = 4.44 percent, noise-to-signal 16:1. Catching a 0.28 percent
event needs an instrument below 0.28 percent error; the best string matcher
here sits at 4.44 percent, 16 times too coarse. All 16 failures are one
class: value plus decoration.

H5 confirmed: 64/120 answers are three words or fewer. H4 (no rule exists in
principle) is NOT established - 120 rows cannot prove non-existence, and they
contain no negatives to test against.
…gistry owner

The parallel session on chore/bump-urllib3-pyjwt-cve currently holds uncommitted
edits to AGENT_DIARY.md, KNOWN_ISSUES.md and WISDOM.md, so per the registry
ownership rule this is a proposal rather than an edit. Carries the post-mortem,
five distilled facts for WISDOM, two open issues, and an explicit list of what
these findings do NOT establish.
@coderabbitai

coderabbitai Bot commented Oct 3, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: a0bfdb74-6ef4-4877-8bcf-713db4a0e14b
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ManSio
ManSio merged commit 5983cda into main Oct 4, 2026
6 of 7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant