feat(eval): calibrated matcher, fail-closed control verdicts, and a negative result - #68
Merged
Merged
Conversation
added 14 commits
October 3, 2026 06:07
KNOWN_ISSUES.md held 448 lines against a 300-line limit (rule R1), which blocked every commit on main. 224 of those lines were not board entries: 29 blocks injected verbatim from AGENT_DIARY.md by AutoDocUpdater during a full reindex, each carrying "Источник: AGENT_DIARY.md" and "Статус: автоматически синхронизировано". A board that mirrors the diary holds every issue twice and cannot be trimmed without losing information. The diary stays the source; the board keeps the 24 hand-authored sections. Live board is now 224 lines. scripts/archive_autosync_entries.py does this by marker rather than by date, so the classification does not rot as the file ages, and a second run finds nothing to move and changes nothing. Nothing is deleted — the blocks are appended to docs/archive/KNOWN_ISSUES_2026_10_AUTO_SYNC.md, per §4.8 R4.
pip-audit failed on both test jobs with vulnerabilities that appeared after the
last green main run on 2026-09-29, so this is a time-dependent failure rather
than a regression from any recent change:
urllib3 2.7.0 PYSEC-2026-4175 / -4176 / -4177 -> 2.8.0
PyJWT 2.13.0 13 advisories, highest fix 2.15.0 -> 2.15.0
Both are transitive: neither is imported as an API. PyJWT appears in the codebase
only as a token pattern in src/core/redact.py, and the line above the pin claimed
a verified RS256 roundtrip against pyjwt 2.13.0 — that claim is re-verified here
rather than assumed, and the comment is updated to the version actually tested.
Verified after the bump, not before:
- pip-audit -r requirements-lock.txt --no-deps --disable-pip -> "No known
vulnerabilities found", rc=0
- RS256 roundtrip with cryptography 50.0.0, and `import mcp` -> ok
- pytest tests/ -> 1989 passed, 6 skipped, 98 deselected
- scripts/smoke_e2e.py -> SMOKE E2E: PASSED, which exercises the real embed
(llama.cpp), the real rerank (BGE-M3) and a real index search, i.e. the
network stack urllib3 actually backs
…y guard as UNPROVEN CI reported `NEGATIVE CONTROLS: FAILED (broken=0, unproven=1)` with `dead_guard_classifier` marked UNPROVEN — green locally, red on every runner. Root cause: tests/test_negative_controls_runner.py proved the digest pin by editing the REAL fixture scripts/negative_controls/fixtures/dead_guard.py and restoring it in `finally`. Under `pytest -n auto` another worker read that fixture's digest inside the window between the write and the restore, computed a different digest, and classified a healthy guard as UNPROVEN. The digest guard was never wrong; the race was between two tests, and it presented as a guard defect. Every digest matched locally and in a clean checkout, which is why this survived local verification — the trigger is worker count, not content. The test now copies the fixture into a scratch directory it creates and removes, and never touches the tracked file. It asserts a PROVEN control before mutating, so the property under test is still "a byte change flips PROVEN to UNPROVEN". Adds tests/test_no_tracked_file_mutation.py so the class cannot come back. It is a test rather than a script because a script nobody runs is not a gate, and its first version collected zero tests under pytest while reading as coverage. That guard's selftest earned its place immediately: it caught a real second instance in tests/test_planted_break_gate.py, which wrote experiments/planted_break/results.json from two xdist workers with no atomicity. That write is now temp-file + os.replace, the artifact is gitignored, and the one exemption is recorded with its reason. Verified: 1991 passed, 6 skipped, two consecutive runs of the exact CI command with -n auto.
check_third_party_data: R1 personal email, R2 exported-profile signature, R3 bulk dump (advisory). Vendored dependency metadata is allowlisted as legitimate attribution. --selftest proves positive 2/2 and negative 2/2 so a matcher that cannot fail cannot pass; --all over 1897 tracked files reports blocking=0. check_parallel_sessions: lists foreign worktrees with branches and uncommitted files, blocks when a staged file is also modified in another tree, advisory otherwise. Wired into .githooks/pre-commit; AGENTS.md documents both. Three defects were found while verifying them. Each is reproduced by a test. 1. GIT_DIR leak. git exports GIT_DIR into the hook environment and _git() inherited it, so 'git -C <other worktree> status' kept reading the CURRENT repository. The gate inspected the wrong tree: with GIT_DIR set it reported three pre-commit-stage files as modified in D:/Project/wt-pr-v3 and exited 1 while that worktree was clean. Bidirectional - it also let real cross-tree conflicts through. Fix drops GIT_DIR/GIT_WORK_TREE/GIT_INDEX_FILE/ GIT_COMMON_DIR from the child environment. 2. The third-party gate could not be tested at all. Its R1 fixture held a real person's address (pascal.cescato@gmail.com), so committing the detector would have published a third party's email - exactly what the rule forbids. Test fixtures used gmail.com addresses and tripped R1 on synthetic data. Both now use the RFC 2606 reserved domain personal.invalid, which can never be registered, and personal_domains is injectable so R1 is still proven to fire. 3. R2 counted 11 profile markers in the detector itself, because PROFILE_MARKERS lives in that file. R2 is skipped for exactly two paths; R1 still applies to them, so a dump cannot be hidden there.
gate_zero_full_suite() picked pythonw.exe on Windows to avoid a console flash. pythonw has no console, so every test that spawns a subprocess (git, lock helpers, discriminators) dies inside subprocess with OSError [WinError 50]. Measured on the same tree and commit: 'python -m pytest tests/' gives 2006 passed and 0 failed, 'pythonw.exe -m pytest tests/' gives 72 failed and 6 errors, every failure WinError 50 from subprocess. CREATE_NO_WINDOW already suppresses the console window, so pythonw bought nothing and cost the entire gate.
…gates Adds the session's own notes on the parallel-session incidents, the third-party data audit and the verification gate work. KNOWN_ISSUES.md was 366 lines against a 300-line limit, so check_known_issues failed and blocked every commit. Rotated 79 lines of closed entries per the monthly-archival rule: only sections whose own header carries Fixed/Closed/ REFUTED, no Open and no P1 moved. Closed entries from 2026-09 went to docs/archive/KNOWN_ISSUES_2026_09.md, the 2026-10-03 one to a new KNOWN_ISSUES_2026_10.md. Live file is 289 lines; open items untouched. While rotating, found two pre-existing defects NOT caused by this change and deliberately left alone: docs/archive/KNOWN_ISSUES_2026_09.md holds the 'Exp E13' section four times over roughly 700 lines in two encodings, and the live file carries CP1251-read-as-UTF-8 mojibake in the 2026-09-19 onward headers. Both belong to the registry owner.
run_model() defaulted to base_repeats=2 while the CLI default and the frozen manifest specify 3, so any caller omitting the argument silently produced 100 rows instead of the declared 120 per model. manifest.json also now records the harness SHA it actually describes plus the v2 findings (matcher false negatives, verdict sensitivity, abstention).
The transport read the CLI response from the wrong stream, so every long output was parsed as empty and rows were silently marked invalid. Tests now cover banner-on-stderr/body-on-stdout, both RU and EN arms, and the 120-row per-model expectation implied by base_repeats=3.
360 live calls, 0 invalid, all three models PASS. FINDINGS.md documents the two meta-results: q_http_404 alias coverage caused 8/360 false negatives, so every sub-1.0 value measures the matcher rather than the model, and the 0.90 gate does not see a single 1/30 planted break (0.9667). Values stay quarantined until dataset v3 recalibrates the matcher.
Matcher: three declarative rule kinds replace loose any_of substrings. value_context separates boiling from freezing and 100C from 212F; subject_value requires the year next to its subject, so '1991' in an unrelated sentence stops counting as an answer about Python; last_mention resolves hedged answers by the final mention, not the first. Numeric anchors keep word boundaries, stems stay substrings. yo normalization is now identical on both matcher paths. Calibration on 156 labelled items: recall 0.867 -> 1.0, FP 0.212 -> 0.0. Rules were fitted on the tune half only; the heldout half is reported separately in results/calibration_v3_heldout.json. Verdicts: live arms report STRICT_PASS (invariance 1.0) or CONDITIONAL_PASS (>= threshold, with the disagreeing rows named for manual review) instead of a single PASS. Control arms report CONTROL_OK/CONTROL_FAIL and exit 1 on failure; they previously returned N/A, so a planted 1/30 break scored 0.967 and exited 0 - the control could not report its own blindness. injected_hits is a per-run delta because the transport counter accumulates across models.
Adds findings 4-6: substring matching is not calibratable in principle (recall 0.867, FP 0.212 both directions, hence neither v2 metric class is publishable), the three declarative rules that replace it, and the control arm that previously returned N/A and therefore could not report its own blindness. Records that the heldout half is no longer fully blind for the one item that exposed the yo-normalization inconsistency.
…N class 360 live calls, invalid=0. All metrics recomputed from raw rows independently, 0 mismatches against the harness summary. Finding 7: the invariance denominator is a function of the model itself (n = 30/21/27/27/24/30), because it is conditioned on base_ok which the matcher defines. At a fixed 0.90 threshold the gate needs 4 disagreements at n=30 but only 3 at n<=27, so verdict columns cannot be compared across models. The conditioning also selects exactly the rows where the model was right, which is where stability is easiest, so weaker models score better: longcat/en has the worst accuracy (0.7) and the best invariance (1.0) with STRICT_PASS. Finding 8: v3 fixed the q_http_404 false negatives it was built for, but its bare rule only fires on exact equality with one anchor, so a correct answer carrying a unit, a symbol or a case ending (100C, 100C (212F), '1912 godu.', 'Kislorod (O2)') fails the require check. Four case-cells regressed. Finding 9: recall 1.0 on 156 self-authored items was real and not predictive. Every positive I wrote carried the predicate; the models answer with the bare value. Calibrating against one's own phrasing distribution calibrates against a distribution that does not exist.
…to signal Cheap experiment on 120 already-collected live answers (no new model calls). Sample frozen before labelling with matched/agrees_with_base stripped. Finding 10: all 120 live answers are factually correct, so FP is not measurable - its denominator is empty. H1 and H2 are therefore degenerate rather than confirmed: any matcher reaches recall 1.0 on positives and FP 0 because there is nothing to catch. Model errors live in non-first repeats, and the first-row selection rule excluded exactly those. Finding 11: hand-adjudicating all 17 rows the v3 matcher rejected gives 16 false negatives against 1 genuine model error. Signal 1/360 = 0.28 percent, noise 16/360 = 4.44 percent, noise-to-signal 16:1. Catching a 0.28 percent event needs an instrument below 0.28 percent error; the best string matcher here sits at 4.44 percent, 16 times too coarse. All 16 failures are one class: value plus decoration. H5 confirmed: 64/120 answers are three words or fewer. H4 (no rule exists in principle) is NOT established - 120 rows cannot prove non-existence, and they contain no negatives to test against.
…gistry owner The parallel session on chore/bump-urllib3-pyjwt-cve currently holds uncommitted edits to AGENT_DIARY.md, KNOWN_ISSUES.md and WISDOM.md, so per the registry ownership rule this is a proposal rather than an edit. Carries the post-mortem, five distilled facts for WISDOM, two open issues, and an explicit list of what these findings do NOT establish.
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configuration
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #67 (guards). Rebased onto
chore/bump-urllib3-pyjwt-cve— merge that first.This PR is, in the end, mostly a negative result with numbers, plus the harness fixes needed to reach it honestly.
The finding
A string matcher is not a measurement instrument for paraphrase invariance on this corpus.
Hand-adjudicating all 17 rows the v3 matcher rejected out of 360 live calls:
1/360 = 0.28%16/360 = 4.44%Detecting a 0.28% event requires an instrument below 0.28% error. The best string matcher implemented here sits at 4.44% — 16 times too coarse. Tuning does not change this.
Why calibrating on live answers cannot work
All 120 answers in the frozen live sample are factually correct. There are no negatives. So
FP rate = 0/120is trivially true — the denominator is empty — and any matcher that catches one form scores 1.0/0. That is a tautology, not a test. Model errors live in the non-first repeats, and a "take the first row" selection rule excludes exactly those.The implication is broader than this corpus: you cannot calibrate a matcher on model output, because the class you must calibrate for (wrong answers) is nearly absent from it. Negatives have to be invented, and then you are calibrating against your own imagination — which is what v2 did, at FP 0.212.
Second defect, unrelated to the matcher
invariance_given_base_correctis conditioned onbase_ok, so its denominator is a function of the model: n = 30/21/27/27/24/30. At a fixed 0.90 threshold the gate needs 4 disagreements at n=30 but only 3 at n<=27 — verified arithmetically. Verdict columns are therefore not comparable across models. Worse, the conditioning selects exactly the rows where the model was right, which is where stability is easiest, so weaker models score better: longcat/en has the worst accuracy (0.7) and the best invariance (1.0) withSTRICT_PASS.Harness fixes
N/Aand exited 0, so a planted 1/30 break (invariance 0.967) scored asPASS— the control could not report its own blindness. NowCONTROL_OK/CONTROL_FAILwith exit 1;injected_hitsis a per-run delta because the counter accumulates across models.STRICT_PASS(1.0) fromCONDITIONAL_PASS(>=0.90 with the disagreeing rows named for manual review).calibrate_matcher.pyis a gate that fails on FN, on FP and on an empty population (exit 0/1/2), with a--selftestproving all three can fail.any_ofsubstrings with three declarative rule kinds. Calibration on 156 labelled items: recall 0.867 → 1.0, FP 0.212 → 0.0.What this PR does NOT establish
Water boils at 212 degrees Fahrenheitwas markedno_matchuntil corrected, recorded as revision 1.1).live_sample.jsonis perfectly confounded with row kind (even indices are allbase), so it is systematic, not random. Noted in the frozen hypotheses file.ё-normalization inconsistency between matcher paths, and the fix changed its result.Reproduce