Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 26 additions & 1 deletion cmd/agentloop/main.go
Original file line number Diff line number Diff line change
Expand Up @@ -72,7 +72,7 @@ func NewServer() *Server {
envOr("AGENTLOOP_ONEGW_URL", "http://127.0.0.1:8080"),
os.Getenv("AGENTLOOP_ONEGW_KEY"),
envOr("AGENTLOOP_ONEGW_COMBO", "dev"),
),
).WithTiers(tiersFromEnv()),
evalRunner: eval.NewRunner(func(cfg loop.RunnerConfig) (*loop.LoopRunner, *budget.Guard, tools.ToolRegistry, error) {
// Same runner the service builds, minus the model: the gate and
// the planner are part of what the cases exercise, so building a
Expand Down Expand Up @@ -105,6 +105,31 @@ func envOr(key, def string) string {
//
// No LeanKG API key: the service is local and its REST surface is
// unauthenticated by design for a single-tenant deployment (PRD §7.4).
// tiersFromEnv maps this loop's three routing tiers onto onegw combos.
//
// The gateway ships whatever combos an operator configured — in this
// portfolio, exactly one (`dev`). Naming a combo here that does not exist
// upstream is how "tiered routing" becomes an error instead of a saving,
// so an unset tier simply falls through to AGENTLOOP_ONEGW_COMBO and the
// loop still runs:
//
// AGENTLOOP_ONEGW_COMBO_PLANNING combo for plan/replan steps
// AGENTLOOP_ONEGW_COMBO_EXECUTION combo for action steps (the hot path)
// AGENTLOOP_ONEGW_COMBO_SYNTHESIS combo for the bound-exit answer
func tiersFromEnv() map[string]string {
out := map[string]string{}
for tier, key := range map[string]string{
loop.TierPlanning: "AGENTLOOP_ONEGW_COMBO_PLANNING",
loop.TierExecution: "AGENTLOOP_ONEGW_COMBO_EXECUTION",
loop.TierSynthesis: "AGENTLOOP_ONEGW_COMBO_SYNTHESIS",
} {
if v := os.Getenv(key); v != "" {
out[tier] = v
}
}
return out
}

func leankgFromEnv() *leankg.Client {
if os.Getenv("AGENTLOOP_LEANKG_OFF") != "" {
return nil
Expand Down
23 changes: 15 additions & 8 deletions docs/PRD.md
Original file line number Diff line number Diff line change
Expand Up @@ -441,6 +441,12 @@ Three of the twelve are the ones that decide whether this is a plan or a wish. *

**Definition of done for M1–M6:** the milestone's acceptance passes in CI, the status column here is updated in the same commit as the work, and each acceptance becomes a named eval case — never a prose claim.

**Tier routing is wired, and the combo names were fiction (2026-09-21).** §13.1's move 3 says "route models by step type (40–70%)", and M3's row claims `tierCombo` does it. It did not: the runner computed a tier per step, stored it on the run record, and **sent one fixed combo forever** — `onegw.Client.Chat` took no model parameter and the client was constructed with a single combo. Worse, `tierCombo` emitted `planning`/`execution`/`tiny`, and this portfolio's onegw ships **one** combo (`dev`), so two of the three names were unrouteable even if they had been sent.

Fixed: `ModelClient.ChatTier(ctx, tier, msgs…)` carries the tier to the wire, `onegw.Client.WithTiers` maps a tier to a combo, and the three tiers are now the three decisions the loop actually makes — **planning**, **execution**, **synthesis** — renamed from the old trio because a tier name that is not a combo upstream is a decorative constant. An unmapped tier falls through to `AGENTLOOP_ONEGW_COMBO`, so a one-combo deployment still runs. `Reply` now records both the combo asked for and the leg that answered, because a fallback is the event tiering has to be judged on.

What this does **not** do: prove the 40–70% figure. That is §11.4's paired parity suite, and the acceptance has not been run.

**Task record:** this table is the status summary, not the task list. Tasks live as **GitHub issues in [`FreePeak/agentloop`](https://github.com/FreePeak/agentloop/issues)** and `todo.md` is the short index of the open ones; no `TASKS.md` is ever created. Two tracker conventions apply to this repo and are *not* yet followed: issues are now banded **P0**–**P3** (labels defined 2026-09-21; the band meanings are one line each in `todo.md`), and the open set is: **P1** [#27](https://github.com/FreePeak/agentloop/issues/27) CI + [#21](https://github.com/FreePeak/agentloop/issues/21) onegw combo reorder; **P2** [#20](https://github.com/FreePeak/agentloop/issues/20) HTMX console + [#8](https://github.com/FreePeak/agentloop/issues/8) TypeSafe checklist; **P3** [#19](https://github.com/FreePeak/agentloop/issues/19), [#15](https://github.com/FreePeak/agentloop/issues/15), [#28](https://github.com/FreePeak/agentloop/issues/28). What is still not tied together is closure: a merged PR does not close its issue, which is why [#8](https://github.com/FreePeak/agentloop/issues/8)'s implemented half still reads as open — that is the remaining half of [#28](https://github.com/FreePeak/agentloop/issues/28).

**Schedule overlay** (from the book's 30-day plan, adapted — the book's week 1 "from-scratch ReAct with no framework" is subsumed by M1, and its weeks 5–8 become M4–M6):
Expand Down Expand Up @@ -736,13 +742,14 @@ Written the way an unfriendly reviewer would write it, then answered. Every find

**Read next.** §13.1 (scope → milestones), §17 (defaults), §18 (where to discount the source), §22 (this document's own weaknesses).

* Last updated: 2026-09-21 (The loop **decides its own steps**. `internal/loop/reason.go` makes one model call per step with the previous step's verbatim result, and the answer (`{tool,args,why,done}`) is what runs — so the loop observes before it reasons, which is the half it was missing. `done` exits `success`/`goal_met`, a new exit reason for the goal predicate firing rather than a bound. A resume replays the approved decision instead of re-asking (a second call can choose a different tool, so the operator would have approved one action and a different one would run). `StepRecord` now carries its `Args` and `Why`. No model client still means the deterministic rotation, unchanged.)
* Last updated: 2026-09-21 (The loop can finally **write and verify**. `internal/xdev` speaks xdev's `rpc` JSONL protocol (ready-frame version gate, event-before-response interleaving, one turn at a time, child killed when the step's budget expires) and `run_tests`/`write_file` run as one xdev turn each, in a sandboxed workspace from `AGENTLOOP_XDEV_DIR`. With no sandbox the two report *no executor configured* and `written`/`ran` stay false — "no sandbox" can never read as "the tests passed". `nextToolDefault` now gives each tool the arguments it needs to be a real call, so a write has a target instead of failing closed on a missing path. `web_search` is the last stub.)
* Last updated: 2026-09-21 (The M6 eval gate was reporting **0.5 — deploy blocked — for days** while `go test ./...` was green. Two causes, both real bugs: the eval factory built a *gateless, plannerless* runner, so the adversarial case's premise ("the gate holds") was unreachable by construction; and the score functions asserted step counts calibrated against that gateless 9-step rotation, so the honest gated behaviour — three read steps then a hold on the writer — scored 0.3. The factory now builds the service's runner (minus the model, as §11.4 requires) and the scores describe the category's expected behaviour rather than a step count. Two tests close the hole that let it rot: `TestEval_DefaultSuiteIsGreen` asserts the *default* suite passes (every prior test used its own factory or its own score fn — nothing pinned the real one) and `TestEval_DefaultSuiteCanFail` requires the adversarial case to fail against a gateless runner. Verified by breaking `Categorize` and watching the endpoint block at 0.75.)
* Last updated: 2026-09-21 (CI exists: `.github/workflows/ci.yml` runs gofmt, `go vet`, `go test`, `golangci-lint` (config pinned in `.golangci.yml`) and `docs/check-prd.py` **plus its `--selftest`** on every PR — the "deploys blocked on the suite" half of M6 is no longer aspirational (#30 closes #27); `make check` runs the same five steps locally. **Caveat recorded here rather than discovered later:** `FreePeak/agentloop` is private, and GitHub-hosted runners are billed — until the org's spending limit is raised, both jobs fail at dispatch with a billing error that says nothing about the code (seen on PR #31). `make check` is the fallback that keeps the gate honest in the meantime. Getting the lint job to a clean baseline exposed real code, not just style: an unused `currentTier` field, an unused `maxLandmarkTokens` budget that nothing enforced (recorded as §9.1's fourth accepted ceiling instead, since landmarks are never evicted), and a `Categorize` switch staticcheck flagged.)
* Last updated: 2026-09-21 (Docs synced to `b492cc2`. §13's M5 row recorded the `Categorize` fix (#25) and its stale test count; the M6 row now says out loud that the "deploys blocked on the full-suite gate" half is **not** enforced — there is no CI in this repo (#27); the §13 task-record line stopped claiming the repo has no commits/remote and now records the tracker state: priority bands P0–P3 defined and every open issue labelled (#28, half done), with issue closure still not PR-linked. `docs/USAGE.md` lost a duplicated line and its "does nothing" summary was split into what reads today vs what still does not. `docs/check-prd.py` stops false-failing on §-refs into other docs — the cause of the long-standing dual-§-ref FAIL, whose refs pointed at `docs/JEV-INTEGRATION.md`, not this PRD.)
* Last updated: 2026-09-21 (**The `query` tool is real**: it reaches LeanKG `POST /api/v1/query` through `internal/leankg`, wired from `AGENTLOOP_LEANKG_URL`, and reports the answering retrieval rung; a LeanKG outage is a recorded failed step, and the `stepError` panic on a `Success=false`/nil-error tool result is fixed. **The approval table now names the real tools**: `query`/`web_search`/`run_tests` are read-category and `write_file` holds on every call — before this all four fell to the fail-closed default, so every gated run paused on step 1 and M5's <10% interruption ceiling was unreachable.) — §4 and §5.1 FR-7, `docs/USAGE.md`, README.*
* Last updated: 2026-09-21 (TypeSafe coupling risk re-verified: onegw `systemone` Kind is **merged** into master — `9faea01`; the open blocker is PR #110 verdict-driven combo reorder, not the merge. `docs/JEV-INTEGRATION.md` §3 rewritten to reflect the split between the merged Kind and the unmerged routing.) — §14 risk row corrected; `docs/JEV-INTEGRATION.md` §3.1/§3.4/§3.5 marked done, §7 checklist itemised.*
* Last updated: 2026-09-21 (Tier routing reaches the wire. `ChatTier(ctx, tier, msgs…)` replaced the tier-less `Chat` on the model interface, `onegw.Client.WithTiers` maps tier→combo, and the three tiers are now planning/execution/synthesis instead of planning/execution/tiny — the old middle names were combos onegw does not ship. `Reply` records the combo asked for *and* the answering leg, so a fallback is visible. Verified against a stub gateway: three calls to `exec-combo`, one to `synth-combo`.)
* 2026-09-21 — (The loop **decides its own steps**. `internal/loop/reason.go` makes one model call per step with the previous step's verbatim result, and the answer (`{tool,args,why,done}`) is what runs — so the loop observes before it reasons, which is the half it was missing. `done` exits `success`/`goal_met`, a new exit reason for the goal predicate firing rather than a bound. A resume replays the approved decision instead of re-asking (a second call can choose a different tool, so the operator would have approved one action and a different one would run). `StepRecord` now carries its `Args` and `Why`. No model client still means the deterministic rotation, unchanged.)
* 2026-09-21 — (The loop can finally **write and verify**. `internal/xdev` speaks xdev's `rpc` JSONL protocol (ready-frame version gate, event-before-response interleaving, one turn at a time, child killed when the step's budget expires) and `run_tests`/`write_file` run as one xdev turn each, in a sandboxed workspace from `AGENTLOOP_XDEV_DIR`. With no sandbox the two report *no executor configured* and `written`/`ran` stay false — "no sandbox" can never read as "the tests passed". `nextToolDefault` now gives each tool the arguments it needs to be a real call, so a write has a target instead of failing closed on a missing path. `web_search` is the last stub.)
* 2026-09-21 — (The M6 eval gate was reporting **0.5 — deploy blocked — for days** while `go test ./...` was green. Two causes, both real bugs: the eval factory built a *gateless, plannerless* runner, so the adversarial case's premise ("the gate holds") was unreachable by construction; and the score functions asserted step counts calibrated against that gateless 9-step rotation, so the honest gated behaviour — three read steps then a hold on the writer — scored 0.3. The factory now builds the service's runner (minus the model, as §11.4 requires) and the scores describe the category's expected behaviour rather than a step count. Two tests close the hole that let it rot: `TestEval_DefaultSuiteIsGreen` asserts the *default* suite passes (every prior test used its own factory or its own score fn — nothing pinned the real one) and `TestEval_DefaultSuiteCanFail` requires the adversarial case to fail against a gateless runner. Verified by breaking `Categorize` and watching the endpoint block at 0.75.)
* 2026-09-21 — (CI exists: `.github/workflows/ci.yml` runs gofmt, `go vet`, `go test`, `golangci-lint` (config pinned in `.golangci.yml`) and `docs/check-prd.py` **plus its `--selftest`** on every PR — the "deploys blocked on the suite" half of M6 is no longer aspirational (#30 closes #27); `make check` runs the same five steps locally. **Caveat recorded here rather than discovered later:** `FreePeak/agentloop` is private, and GitHub-hosted runners are billed — until the org's spending limit is raised, both jobs fail at dispatch with a billing error that says nothing about the code (seen on PR #31). `make check` is the fallback that keeps the gate honest in the meantime. Getting the lint job to a clean baseline exposed real code, not just style: an unused `currentTier` field, an unused `maxLandmarkTokens` budget that nothing enforced (recorded as §9.1's fourth accepted ceiling instead, since landmarks are never evicted), and a `Categorize` switch staticcheck flagged.)
* 2026-09-21 — (Docs synced to `b492cc2`. §13's M5 row recorded the `Categorize` fix (#25) and its stale test count; the M6 row now says out loud that the "deploys blocked on the full-suite gate" half is **not** enforced — there is no CI in this repo (#27); the §13 task-record line stopped claiming the repo has no commits/remote and now records the tracker state: priority bands P0–P3 defined and every open issue labelled (#28, half done), with issue closure still not PR-linked. `docs/USAGE.md` lost a duplicated line and its "does nothing" summary was split into what reads today vs what still does not. `docs/check-prd.py` stops false-failing on §-refs into other docs — the cause of the long-standing dual-§-ref FAIL, whose refs pointed at `docs/JEV-INTEGRATION.md`, not this PRD.)
* 2026-09-21 — (**The `query` tool is real**: it reaches LeanKG `POST /api/v1/query` through `internal/leankg`, wired from `AGENTLOOP_LEANKG_URL`, and reports the answering retrieval rung; a LeanKG outage is a recorded failed step, and the `stepError` panic on a `Success=false`/nil-error tool result is fixed. **The approval table now names the real tools**: `query`/`web_search`/`run_tests` are read-category and `write_file` holds on every call — before this all four fell to the fail-closed default, so every gated run paused on step 1 and M5's <10% interruption ceiling was unreachable.) — §4 and §5.1 FR-7, `docs/USAGE.md`, README.*
* 2026-09-21 — (TypeSafe coupling risk re-verified: onegw `systemone` Kind is **merged** into master — `9faea01`; the open blocker is PR #110 verdict-driven combo reorder, not the merge. `docs/JEV-INTEGRATION.md` §3 rewritten to reflect the split between the merged Kind and the unmerged routing.) — §14 risk row corrected; `docs/JEV-INTEGRATION.md` §3.1/§3.4/§3.5 marked done, §7 checklist itemised.*
* 2026-09-21 — (System One abstraction: Laya as local Jev-compatible backend; `docs/JEV-INTEGRATION.md` rewritten; §3.1/§4.3/§14/§17/§16 updated; issues #8/#15 extended.) — cookbook deep-dive (function_calling, llm_guardrails, intent-routing) mapped to agentloop UCs.
* v1.2.0 — M5 HITL (ApprovalGate, runner pause, timeout denies) + M6 (EvalRunner, 4-category suite, live HTTP API) shipped; §6 API contract reconciled to main.go routes; SSE event vocabulary narrowed to step + done; §13 milestones updated: M0–M4 closed, M5 closed on merge, M6 closed on merge, M7 conditional.*
* v1.1.5 — UI design added (`docs/UI-DESIGN.md`): operator console design document from UI/UX Pro Max skill, one page per §12 surface, chart-type mapping per surface, cross-cutting UX rules, HTMX pattern table. No new PRD prose; M0 was documentation-only.*
* Last updated: 2026-09-21 (System One abstraction: Laya as local Jev-compatible backend; `docs/JEV-INTEGRATION.md` rewritten; §3.1/§4.3/§14/§17/§16 updated; issues #8/#15 extended.) — cookbook deep-dive (function_calling, llm_guardrails, intent-routing) mapped to agentloop UCs.
13 changes: 11 additions & 2 deletions docs/USAGE.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,7 +107,10 @@ Deployment facts, not compiled defaults:
| `AGENTLOOP_XDEV_DIR` | a fresh temp dir | the workspace `write_file`/`run_tests` turns run in |
| `AGENTLOOP_XDEV_OFF` | *(unset)* | any value disables the sandbox; those two tools then report no executor |
| `AGENTLOOP_ONEGW_URL` | `http://127.0.0.1:8080` | gateway; when reachable, the model **chooses each step** |
| `AGENTLOOP_ONEGW_COMBO` | `dev` | combo used as the wire `model` for both step choice and synthesis |
| `AGENTLOOP_ONEGW_COMBO` | `dev` | default combo — the wire `model` when no tier-specific one is set |
| `AGENTLOOP_ONEGW_COMBO_PLANNING` | *(falls back to `COMBO`)* | combo for plan/replan steps |
| `AGENTLOOP_ONEGW_COMBO_EXECUTION` | *(falls back to `COMBO`)* | combo for action steps (the hot path) |
| `AGENTLOOP_ONEGW_COMBO_SYNTHESIS` | *(falls back to `COMBO`)* | combo for a bound-exit answer |

Take the onegw key from `onegw.toml`'s `[auth] [[auth.keys]]`; the combo must
exist there too, since the client sends whatever name you give it and onegw
Expand Down Expand Up @@ -328,7 +331,13 @@ Stated plainly, so nobody discovers it the hard way:
phases and instructions come from a rule table (`planStepCount` on the goal's
word count), so the loop decides **what to do next** but not **how to break
the goal up**. In practice the chooser carries the run, and the plan is a hint.
- **Tier routing is half-wired** — see §6.
- **Tier routing reaches the wire, but the 40–70% number is unproven.** The
loop sends a tier per call (`planning`/`execution`/`synthesis`) and each maps
to a combo; unmapped tiers fall through to `AGENTLOOP_ONEGW_COMBO`, so a
one-combo deployment runs unchanged. What has **not** been run is §11.4's
paired parity suite, which is the only thing that can say whether the routing
actually saves money without costing quality. Until then the saving is a
hypothesis with a working mechanism behind it.
- **M7 is gated shut**, correctly: the gate is a measurement, not a milestone,
and it opens only when a [PRD §10](PRD.md#10-multi-agent-stance) condition is
actually met.
Expand Down
12 changes: 9 additions & 3 deletions internal/loop/m3_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -123,8 +123,14 @@ func TestM3_PlannerDrivenRun(t *testing.T) {
if len(result.Steps) == 0 {
t.Fatal("run has no steps")
}
// Tier should be set from the plan (planning or tiny default)
if result.CurrentTier != "planning" && result.CurrentTier != "tiny" {
t.Errorf("CurrentTier = %q, want planning or tiny", result.CurrentTier)
// The run reports a routing tier, and it is one of the three the loop
// actually routes on. The old assertion named `tiny`, a combo that
// does not exist in onegw and was therefore a tier nothing could route
// to — the exact drift this test now prevents.
switch result.CurrentTier {
case TierPlanning, TierExecution, TierSynthesis:
default:
t.Errorf("CurrentTier = %q, want one of %s/%s/%s",
result.CurrentTier, TierPlanning, TierExecution, TierSynthesis)
}
}
Loading
Loading