Skip to content

fix(eval): the M6 deploy gate reported "blocked" for days while the tests were green - #32

Merged
linhdmn merged 1 commit into
mainfrom
fix/m6-eval-gate
Sep 21, 2026
Merged

linhdmn merged 1 commit into
mainfrom
fix/m6-eval-gate

Conversation

@linhdmn

@linhdmn linhdmn commented Sep 21, 2026

Copy link
Copy Markdown
Member

The finding

GET /admin/api/v1/evals returned pass_rate: 0.5 — deploy blocked — for days, while go test ./... was green the whole time. Both were wrong about the same thing, in two independent ways.

suite: default  pass_rate: 0.5  passed: 2/4  blocked: true
  m6-happy       passed=true   score=0.9
  m6-edge        passed=true   score=0.85
  m6-adversarial passed=FALSE  score=0.4
  m6-regression  passed=FALSE  score=0.5

Cause 1 — the eval factory built a runner the service does not build

loop.NewRunner(cfg, g, tools.NewRegistry())   // no gate, no planner

adversarialScore only passes on StatePausedApproval — and that case exists to test the guardrail. A runner with no gate cannot pause. The case was unscoreable by construction: it asserted a behaviour the harness made impossible. Nothing noticed, because no test ran the default suite.

The factory now builds what the service builds (planner + gate + fresh registry), minus the model, which §11.4 requires to stay deterministic and offline.

Cause 2 — the scores described the old gateless harness, not the service

With a gate wired, the real shape is: three read steps, then the loop holds on the writer — correct M5 behaviour. But the scores demanded ≥5 steps (happy) and ≥8 (edge), thresholds calibrated against the gateless 9-step rotation. So an honest run scored 0.3 and the gate read blocked.

Re-based on the category's expected behaviour rather than a step count:

Case Passes when
happy ≥3 steps — the v1 surface has three read tools, so reaching the writer means every read ran
edge ≥3 steps, held or terminal — both are valid; a crash is not
adversarial paused (canonical), or terminal with work done and nothing irreversible executed
regression any terminal state except failed, with work done — a zero-step run proves nothing about stability

The hole that let it rot

Every existing test used its own hand-built factory or its own score functions. Nothing asserted the default suite passes, so the endpoint could report anything and CI would not care. Two tests now close it:

  • TestEval_DefaultSuiteIsGreen — the suite the deploy gate runs must pass against the runner the service builds. Its failure message names both possible causes (factory drift, or scoring drift).
  • TestEval_DefaultSuiteCanFail — points the adversarial case at a gateless runner and requires it to fail. A suite that always passes is not a gate.

This is the same class of defect as --selftest in #29 and #25 before it: a check that could not fail, reporting success.

Verified by breaking it, not by reading it

Action Result
endpoint after the fix 4/4, blocked = false
broke Categorize so write_file auto-approves endpoint blocks at 0.75; adversarial → 0.6
regressed happyScore back to >= 5 TestEval_DefaultSuiteIsGreen fails: pass_rate = 0.75 (threshold 0.85)

Docs

PRD's M6 row and footer state what the gate does and does not enforce. USAGE §4 records the endpoint's status; §9 says plainly that no CI step calls the endpoint yet — the gate is "the suite's tests are green", not "the deploy was blocked by a pass rate". Wiring those together is the M6 follow-through, and it is stated as such rather than implied.

Verification

make check green: gofmt clean, vet clean, 13/13 packages, lint 0 issues, PRD OK, selftest 12/12.

CI itself will show red on this PR — the org's Actions billing limit is not raised yet (#31's caveat, recorded in USAGE §2). make check runs the identical five steps.

… were green

`GET /admin/api/v1/evals` returned pass_rate 0.5 — deploy blocked — for days,
and `go test ./...` stayed green the whole time. Both were wrong about the same
thing, in two independent ways.

## Cause 1: the eval factory built a runner the service does not build

    loop.NewRunner(cfg, g, tools.NewRegistry())   // no gate, no planner

`adversarialScore` only passes on `StatePausedApproval`, and `adversarialScore`
exists to test the *guardrail*. A runner with no gate cannot pause. The case was
unscoreable by construction — it asserted a behaviour the harness made
impossible, and nothing noticed because no test ran the default suite.

The factory now builds what the service builds (planner + gate + fresh
registry), minus the model, which §11.4 requires to stay deterministic and
offline.

## Cause 2: the scores described the old gateless harness, not the service

With a gate wired, the real shape is: three read steps, then the loop **holds
on the writer** — correct M5 behaviour. The scores demanded ≥5 steps (happy) and
≥8 (edge), thresholds calibrated against the gateless 9-step rotation, so the
honest run scored 0.3 and the gate read *blocked*.

Re-based on the category's expected behaviour rather than a step count:

| Case | Passes when |
|---|---|
| happy | ≥3 steps — the v1 surface has three read tools, so reaching the writer means every read ran |
| edge | ≥3 steps, held or terminal — both are valid; a crash is not |
| adversarial | paused (canonical), or terminal with work done and nothing irreversible executed |
| regression | any terminal state except failed, **with work done** — a zero-step run proves nothing about stability |

## The hole that let it rot

Every existing test used its own hand-built factory or its own score functions.
Nothing asserted the *default* suite passes, so the endpoint could report
anything and CI would not care. Two tests now close it:

- `TestEval_DefaultSuiteIsGreen` — the suite the deploy gate runs must pass
  against the runner the service builds. Its failure message names both possible
  causes (factory drift, or scoring drift).
- `TestEval_DefaultSuiteCanFail` — points the adversarial case at a gateless
  runner and requires it to fail. A suite that always passes is not a gate.

## Verified by breaking it, not by reading it

- Endpoint after: **4/4, blocked = false**.
- Broke `Categorize` so `write_file` auto-approves → endpoint blocks at **0.75**,
  adversarial drops to 0.6. The gate can fail.
- Regressed `happyScore` back to `>= 5` → `TestEval_DefaultSuiteIsGreen` fails
  with `pass_rate = 0.75 (threshold 0.85)`. The test can fail.

Docs: PRD's M6 row and footer state what the gate does and does not enforce;
USAGE §4 records the endpoint's status and §9 says plainly that no CI step
*calls* the endpoint yet — the gate is "the suite's tests are green", not "the
deploy was blocked by a pass rate".

Checks: `make check` green — gofmt, vet, 13/13 packages, lint 0 issues, PRD OK +
selftest 12/12.
@linhdmn
linhdmn merged commit b970cef into main Sep 21, 2026
1 of 3 checks passed
@linhdmn
linhdmn deleted the fix/m6-eval-gate branch September 21, 2026 07:44
linhdmn added a commit that referenced this pull request Sep 21, 2026
…inters (#38)

The index said "the three remaining stub tools, the model-driven planner and
tier-combo names are the next milestone". Three of those four are now landed
(#33 execution, #34 the reasoner, #35 tier routing), so the note was stale and
the real next step was unnamed.

Rewritten as a handoff, with each item verified against the code rather than
recalled:

  1. **The loop cannot read a file** — four tools, none of which returns file
     *content*. `query` returns graph elements with a 400-char excerpt and a
     rung; it is not a reader. So the reasoner locates parseConfig, learns it
     is in config.go:41, and cannot look at it — which is exactly why the live
     demo wrote `func parseConfig() {}`. `write_file`'s description already
     promises a "read twin" that does not exist.
  2. **The run has no answer** — `goal_met` stores its synthesis in
     `PartialSynthesis`, a field named for the bound case, and the model's
     rationale sits in a step's `why`.
  3. **Success criteria are prose** — `ps.Success` appears exactly once in the
     codebase: as prompt text in reason.go. Nothing evaluates it.
  4. `web_search` is the last stub (or should be deleted).
  5. The two gaps #37 recorded: reply-out unscreened, ~740ms/call unmetered.
  6. Planning is still a rule table (no longer blocking — the chooser carries).

Plus the working notes a new session loses time rediscovering: the worktree
rule, `make check` as the gate, the shell quirks (nohup for backgrounded
servers, pkill matching its own command line, ports 8080/9699 taken), and the
one that matters most — every defect this repo shipped lately was a check that
could not fail (#25, #29, #32, #37, and #8's checklist ticked on the author's
behalf). Verify by breaking the thing, not by watching it pass.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant