Skip to content

docs(prd): the framework question was never a decision — D9, and the parity test that makes it falsifiable - #40

Open
linhdmn wants to merge 1 commit into
mainfrom
docs/framework-decision
Open

linhdmn wants to merge 1 commit into
mainfrom
docs/framework-decision

Conversation

@linhdmn

@linhdmn linhdmn commented Sep 21, 2026

Copy link
Copy Markdown
Member

design.md §18 carries the book's framework chooser, and it names LangGraph for exactly this shape: branching/state/HITL-persistence. agentloop is that case. The choice never reached §15 — D1 asks "Runtime: Go vs Python", a language question, and it closed 2026-09-19, after M1's loop had shipped. The book's escape hatch is the 3% parity suite, and no parity harness existed in any .go file — so custom won by default rather than by measurement.

This PR records that, and builds the missing half.

Recorded in §3.1 (not repaired quietly)

  1. The decision was unfalsifiable. Appendix F's F1 said the same thing about M3 from the other direction.
  2. AgentBase does not exist as code. §3.1, §4.3 and DUPLICATION-AUDIT.md's inheritance manifest all call it "the abstraction layer"; the only occurrence in any .go file is a comment at internal/tools/registry_impl.go:3. The seams that do exist are loop.ModelClient, tools.ToolRegistry and xdev.Client.
  3. The reversal path was never costed — reversing now means re-implementing the loop, budget.Guard, the tracer and the approval gate behind a framework to compare them.

What the finding does not say: the book's headline advice (Ch.7, start custom for production) was followed, and its landscape chapter (Ch.2) agrees past a 30% workaround share, which the portfolio is. The narrower, more uncomfortable point is that a recorded decision was never made and the test that would have made it reversible was never built.

The missing half, in code

internal/eval/parity.go — CompareParity pairs two reports by CaseID and returns Ship only when no case lost more than ParityTolerance (0.03) and the candidate is cheaper. That is the book's App. C bar, and the acceptance for §11.4's requirement and M3's cost claim.

Pairing is by id, not index: a case that vanished is reported Unpaired and fails the comparison, because a dropped case is not a case that did not regress. Four tests; all four fail when the tolerance is removed.

Tracker

§15 gains D9 — framework vs custom loop, carrying the question as open with that comparison as its acceptance. docs/check-prd.py gains two properties (the parity harness must be named where §11.4 requires it; D9 must exist and cite its chooser) plus two selftest mutations.

The first mutation found a real blind spot: the filename also appears in §3.1, D9 and the footer, so an unscoped search passed while §11.4 itself named no implementation. The check is now scoped to the section.

Verification

go test ./... -count=1 green · gofmt clean · golangci-lint 0 issues · check-prd.py OK · check-prd.py --selftest — all 14 deliberate breakages caught. PRD footer stamped in the same commit.

…parity test that makes it falsifiable

`design.md` §18 carries the book's framework chooser, and it names LangGraph
for exactly this shape: branching, state, HITL persistence. agentloop is that
case. The choice never reached §15 — D1 asks "Runtime: Go vs Python", a
language question, and it closed 2026-09-19, after M1's loop had shipped. The
book's escape hatch is the 3% parity suite; no parity harness existed in any
`.go` file, so custom won by default rather than by measurement.

Recorded, not repaired quietly, in §3.1:

  - the decision was unfalsifiable, and Appendix F's F1 said the same thing
    about M3 from the other direction;
  - `AgentBase` — called "the abstraction layer" by §3.1, §4.3 and
    DUPLICATION-AUDIT's inheritance manifest — exists only as a comment at
    `internal/tools/registry_impl.go:3`. The seams that do exist are
    `loop.ModelClient`, `tools.ToolRegistry` and `xdev.Client`;
  - the reversal path was never costed.

Then the missing half, in code: `internal/eval/parity.go` (`CompareParity`)
pairs two reports by `CaseID` and returns `Ship` only when no case lost more
than `ParityTolerance` (0.03) *and* the candidate is cheaper — the book's
App. C bar. Pairing is by id, not index: a case that vanished is reported
`Unpaired` and fails, because a dropped case is not a case that did not
regress. Four tests, all failing when the tolerance is removed.

§15 gains D9 carrying the question as **open** with that comparison as its
acceptance; `docs/check-prd.py` gains two properties (the parity harness must
be named where §11.4 requires it, and D9 must exist and cite its chooser) plus
two selftest mutations. The first mutation found a real blind spot: the
filename also appears in §3.1, D9 and the footer, so an unscoped search passed
while §11.4 itself named no implementation — the check is now scoped to the
section, and `--selftest` reports all 14 breakages caught.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant