Skip to content

perf(task): prune validation and implicit scheduler graphs - #79

Open
morluto wants to merge 3 commits into
scaleapi:mainfrom
morluto:codex/perf-task-graphs
Open

morluto wants to merge 3 commits into
scaleapi:mainfrom
morluto:codex/perf-task-graphs

Conversation

@morluto

@morluto morluto commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Closes #76.

Prune unused task validation work and replace dense graphs with predecessor chains for wholly implicit sequential tasks. Mixed/explicit graphs retain their existing behavior. Both initial scheduling and retry rearming use the reduced representation.

Workload Before After Speedup
Validate 3,000 implicit steps 97.93 ms 0.243 ms 403×
Validate 500 self-retrying steps 737.17 ms 0.049 ms ~15,000×
Build 3,000-step scheduler graph 220.57 ms 0.621 ms 355×

The graph uses 2,999 edges instead of 4,498,500. Peak traced allocations: 39.1 MB → 0.61 MB. Measurements are validation/graph microbenchmarks on macOS arm64, Python 3.12.13, against 602d730, with five timing samples. They do not measure complete task execution; earlier-target retries and mixed graphs retain their existing validation costs.

Reproduce:

PYTHONPATH=src:packages/agentenv-protocol/src .venv/bin/python tst/benchmarks/task_graphs.py --base 602d730 --steps 1000 3000 --retry-steps 500 --repeats 5

Validation:

  • 157 focused task/retry tests passed, including actual execution, persisted retries, resume/end boundaries, callbacks and cancellation.
  • Full unit and protocol suites: 6,237 passed, 13 skipped.
  • Deterministic regressions check graph edge counts and absence of unused prefix slicing/self-retry ancestry.
  • Independent review completed with no actionable findings; plugin API check and diff checks passed.

Two commits separate validation pruning from scheduler graph reduction. This PR is independent of the journal, Explorer and existing context-diff PRs.

RetriggerConfidence Score: 4/5

The PR is not ready to merge because the documented graph benchmark fails against its baseline.

Fix All in CursorFindings

  1. P1 Benchmark graph comparison crashes ▶
  2. P2 Baseline uses newer helpers ▶
Fix with agent prompt
### Issue 1
src/agent_env/task/task.py:289
The new required `dependencies` field breaks the documented comparison with `main`. The benchmark gives the old `_build_scheduler_state` method the current `_SchedulerState` class, but the old method does not pass `dependencies`. The “scheduler graph” row raises `TypeError` instead of printing a measurement. Load the baseline state class with the baseline method, or adapt its return value.

### Issue 2
tst/benchmarks/task_graphs.py:51-56
If `--base` points to a revision with different dependency helpers, the benchmark runs that revision’s methods with helpers from the working tree. The printed speedup may then compare a mix of versions. Load the helpers from the selected revision or reject the comparison when they differ.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

Task validation skips work it does not need, and wholly implicit task graphs use a predecessor chain with the same sequential run order. Mixed and explicit graphs keep their existing dependency behavior.

  • Creating implicit tasks skips unused validation work.
  • Implicit task steps follow a smaller predecessor chain.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Implicit step A] --> B[Implicit step B] --> C[Implicit step C]
  C --> D[Retry rearm counts stored predecessors]
Loading

Reviews (2) · Last reviewed commit: "refactor(task): build the scheduler's de..."

@morluto
morluto marked this pull request as ready for review October 6, 2026 19:54
@morluto
morluto requested a review from a team as a code owner October 6, 2026 19:54
Comment on lines +51 to +56
namespace = {
"_SchedulerState": _SchedulerState,
"_ancestor_ids": _ancestor_ids,
"_dependency_ids": _dependency_ids,
}
exec(compile(extracted, f"{base_ref}:src/agent_env/task/task.py", "exec"), namespace)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Baseline uses newer helpers

If --base points to a revision with different dependency helpers, the benchmark runs that revision’s methods with helpers from the working tree. The printed speedup may then compare a mix of versions. Load the helpers from the selected revision or reject the comparison when they differ.

Prompt To Fix With AI
This is a comment left during a code review.
Path: tst/benchmarks/task_graphs.py
Line: 51-56

Comment:
**Baseline uses newer helpers**

If `--base` points to a revision with different dependency helpers, the benchmark runs that revision’s methods with helpers from the working tree. The printed speedup may then compare a mix of versions. Load the helpers from the selected revision or reject the comparison when they differ.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Cursor Fix in Claude Code Fix in Codex

_drive_dag rebuilt each active step's dependency ids to re-arm after a
retry's rollback, mirroring _build_scheduler_state's chain-or-dense
choice. _build_scheduler_state now records them on _SchedulerState as
`dependencies`, and _rearm reads that, so the choice lives in one place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- `ready`: steps whose deps are all satisfied and can launch immediately.
"""
step_by_id: dict[str, "TaskStep"]
dependencies: dict[str, set[str]]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Benchmark graph comparison crashes

The new required dependencies field breaks the documented comparison with main. The benchmark gives the old _build_scheduler_state method the current _SchedulerState class, but the old method does not pass dependencies. The “scheduler graph” row raises TypeError instead of printing a measurement. Load the baseline state class with the baseline method, or adapt its return value.

Prompt To Fix With AI
This is a comment left during a code review.
Path: src/agent_env/task/task.py
Line: 289

Comment:
**Benchmark graph comparison crashes**

The new required `dependencies` field breaks the documented comparison with `main`. The benchmark gives the old `_build_scheduler_state` method the current `_SchedulerState` class, but the old method does not pass `dependencies`. The “scheduler graph” row raises `TypeError` instead of printing a measurement. Load the baseline state class with the baseline method, or adapt its return value.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Cursor Fix in Claude Code Fix in Codex

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Reduce task validation work and implicit scheduler graph size

2 participants