Skip to content

perf(task): avoid quadratic response history diffs - #75

Merged
earakely-scale merged 1 commit into
scaleapi:mainfrom
morluto:codex/optimize-context-list-diffs
Oct 6, 2026
Merged

earakely-scale merged 1 commit into
scaleapi:mainfrom
morluto:codex/optimize-context-list-diffs

Conversation

@morluto

@morluto morluto commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

What and why

Closes #74.

Every completed task step diffs its context against a snapshot. The top-level list diff scans every stored response against the entire prior response list, so an unchanged history of 100 responses costs 5,050 equality comparisons; appending one response costs 5,150. This makes context bookkeeping quadratic as a run accumulates responses.

Check whether the old list is an unchanged prefix, then scan only the appended tail. Tail entries still undergo the existing membership check so duplicate appends remain no-ops. Changed or reordered prefixes retain the original behavior, including the handling of mutated response records. Deep-copy snapshots and nested metadata diffs keep their existing semantics.

Synthetic Python 3.12.12 / arm64 benchmark, measured without profiler instrumentation as the median of five batches. Each cycle includes a full context snapshot and diff, with distinct stored responses containing 512-character strings and a small structured output:

Stored responses Before After
100 2.9 ms 0.64 ms
500 59 ms 3.0 ms
1,000 273 ms 6.3 ms

These timings measure context bookkeeping rather than complete agent runs. An approximately 246 KiB nested-metadata workload remains about 5.9 ms per cycle. The issue includes a reproduction script.

How it was tested

  • Added a comparison-count regression for unchanged response histories and histories with one appended response. It fails on the original implementation and passes with the optimization, without relying on wall-clock thresholds.
  • Ran the actual context test file with isolated production-module loading, package initialization and repository conftest fixtures bypassed, and five environment-dependent tests excluded: 34 passed. This isolation bypasses unavailable cloud SDK imports and does not replace the normal CI suite.
  • Compared journal output with the original implementation across 58,572 cases, including appends, duplicates, deletions, reordering, and mutated response contents: identical outputs.
  • git diff --check and Python syntax checks passed.

RetriggerConfidence Score: 5/5

The PR appears safe to merge based on the reviewed changes.

What we checked:

  • Changed responses still reach the full scan: A changed response makes the prefix comparison fail. The code then checks the full list, as it did before.

Summary

Context updates now scan only appended items when a saved list stays an unchanged prefix, avoiding repeated comparisons as histories grow. A comparison-count test covers unchanged and appended response histories.

  • Context updates scan only new items when a list keeps its old prefix.

Reviews (1) · Last reviewed commit: "perf(task): avoid quadratic response his..."

@morluto
morluto marked this pull request as ready for review October 6, 2026 19:20
@morluto
morluto requested a review from a team as a code owner October 6, 2026 19:20
@earakely-scale

Copy link
Copy Markdown
Collaborator

Nice PR, approved. Will help get merged assuming integ tests pass

@earakely-scale

Copy link
Copy Markdown
Collaborator

/codebuild_run(310b0ee)

@earakely-scale

Copy link
Copy Markdown
Collaborator

Merging, thanks!

@earakely-scale
earakely-scale merged commit 48de90f into scaleapi:main Oct 6, 2026
11 of 13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf(task): avoid quadratic comparisons when diffing response histories

2 participants