Skip to content

feat: keep retrying cancels until the chain confirms them - #716

Merged
MoonBoi9001 merged 24 commits into
mb9/reviewfrom
mb9/keep-retrying-cancels-until-the-chain-confirms
Oct 2, 2026
Merged

MoonBoi9001 merged 24 commits into
mb9/reviewfrom
mb9/keep-retrying-cancels-until-the-chain-confirms

Conversation

@MoonBoi9001

@MoonBoi9001 MoonBoi9001 commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Dipper now marks an agreement it wants ended as Cancelling before sending its on-chain cancel, and the chain listener retries the cancel until the chain shows it ended, alerting after 10 failed attempts. Only then is it marked cancelled and its end announced. It also fixes the remaining findings from the review of #715.

Warning

Before rolling back below this version, move every Cancelling agreement (status 9) to another status: older builds can't read it.

Offers and cancels share 1 wallet, so marking an agreement Cancelling before its cancel goes
out lets a late offer withdraw itself. The chain listener retries the cancel until the chain
shows the end, with an ERROR after 10 failed attempts; only then is the end announced.
The offer job's withdraw and the on-chain cancel job each read the chain and cancelled only
a live agreement, with their own copy of that logic; both now use the helper the retry uses.
After its offer lands, the offer job rereads the agreement to withdraw the offer if it was
cancelled meanwhile. A failed reread finished the job, leaving such an offer open; it retries.
Accepts dipper missed were announced only if it cancelled within 1 hour of the offer deadline,
so a lagging listener dropped them. Only agreements created before events existed are skipped.
6 test modules each spelled out the whole agreement config, so every new setting meant
editing 6 copies. They now start from a shared test helper and override what they need.
The subgraph reports a withdrawn offer as cancelled by the payer with an accept time of 0.
Recording that as an accept announced an agreement that was never live.
The cancel retry reads the chain ahead of the listener, so marking an ended agreement there
lost who ended it and when; it now waits for the listener, which records both from the chain.
Each sweep took the 100 oldest cancelling agreements, so a backlog stalled the listener and hid
newer ones, and could resend a cancel still being mined. It now takes 10, least recently
checked first, and skips any marked in the last 2 minutes.
A misconfigured or paused manager failed every agreement's cancel, using up their attempts so
none was retried once fixed. Only a cancel mined without ending the agreement now counts, and
an agreement with no stored terms hash, which can never be cancelled, is given up at once.
Retrying it sent the offer again. The chain listener's cancel retry already withdraws the offer
of an agreement left cancelling, so the job now logs the failed read and finishes.
Agreements still cancelling kept the listener at its fast poll rate so their retry ran often,
and one given up on kept it there for good. The retry now runs every 5 minutes at any rate.
An agreement dipper is cancelling is paid until the cancel lands, so it now stays in the fee
estimates the indexer selection service uses to compare indexers.
A lagging listener can mark an agreement expired that was in fact accepted. Such an agreement
couldn't be marked cancelling, so its replacement's acceptance no longer cancelled it.
The sweep for accepted agreements left behind when a request was cancelled resent their cancel
on every run with no limit. It now marks each one cancelling, and the cancel retry finishes it.
The production registry trait gave the retry's 2 queries do-nothing defaults, so a registry
that forgot them would silently never retry a cancel; only the test stub keeps defaults now.
The cancelling agreements query returned each row's failed cancel count, which no caller reads.
A failed read was treated like a clean check, so an offer past its deadline was marked ended
without knowing it wasn't live, as an accept the listener hadn't yet recorded would be.
Agreements created before dipper announced lifecycle events are never announced, but one being
cancelled had its accept recorded from the chain, which then announced it.
Every replaced agreement marked expired was relabelled cancelled, losing its expired event and
the record that the indexer let the offer lapse. Only one the chain shows live is cancelled.
A mined cancel that reverted was retried every 5 minutes without limit, costing gas each time.
It now counts towards the limit; one that reverts before sending is retried, logged as an ERROR.
The retry left an accepted agreement that already ended for the chain listener to confirm, so
one whose end the listener never read stayed cancelling for good. After an hour it marks it.
Each retried cancel can wait 15 seconds to be mined, and the sweep holds up the chain listener,
so 10 slow ones stalled it for minutes. A sweep now leaves what it hasn't reached to the next.
A reassessment no longer queues it when its own cancel fails; the cancel retry handles that.
@MoonBoi9001
MoonBoi9001 added this pull request to stack #717 October 2, 2026 17:05
@MoonBoi9001
MoonBoi9001 marked this pull request as ready for review October 2, 2026 17:56
@MoonBoi9001
MoonBoi9001 removed this pull request from stack #717 October 2, 2026 17:57
@MoonBoi9001
MoonBoi9001 merged commit 884e33d into mb9/review Oct 2, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant