Skip to content

Finalisation and attestation-target selection share one candidate ladder, turning brief justification pauses into hour-long finality stalls #1211

Description

@dimka90

Summary

On a 14-node mixed-client devnet, finalisation froze four times, for up to 68 minutes, while justification kept advancing throughout. Head never stalled and nodes never disagreed on head. The clients were implementing the spec as written.

The cause is a coupling between two rules that are governed by one candidate set. The finalisation gap check in state_transition.py and the attestation-target walk in validator_duties.py both use Slot.is_justifiable_after against the same finalized_slot. Once finalisation slips past δ = 6 the two reinforce each other, and the stall sustains itself until two justifications happen to land on adjacent rungs.

This is a question about intended behaviour, not a proposed fix. The gap check is a safety condition and should not be loosened without analysis.

Setup

14 nodes (4 gean, 6 ethlambda, 4 lantern), 14 validators, 4 committees, 4 s slots. Genesis 2026-09-22 14:35:05 UTC. All clients on leanVM 48a90420.

The four episodes

freeze finalized held at duration peak head − finalized trigger escape
09-22 21:17 6012 ~6 min ~140 340 s justification pause δ ≈ 132
09-23 04:11 12227 68 min ~1030 160 s pause δ 961 → 992
09-24 06:07 35569 38 min ~580 140 s pause δ 529 → 552
09-24 10:56 39905 25 min ~390 140 s pause δ 342 → 361

Healthy baseline between episodes: head − finalized median ~10.

In the 68-minute episode, latest_finalized_slot held at 12227 while latest_justified_slot advanced fourteen times, from 12233 to 13188. It is a finalisation stall, not a justification stall. A check of "is the chain justifying?" stays green throughout.

Every justification lands on a rung

In all four episodes, every justification while finalisation was frozen landed on a slot justifiable after the frozen finalised slot: 42 of 42. The 68-minute episode, as δ from 12227:

6  9  20  36  169  240  342  400  484  529  600  729  841  900  961
   3² 4·5 6²  13²  15·16 18·19 20² 22²  23²  24·25 27²  29²  30²  31²

At each step the finalisation check refused, because a justifiable slot sat strictly between source and target:

from δ   to δ    rungs strictly between           state_transition.py
9        20      12, 16                           refuse
20       36      25, 30                           refuse
36       169     42, 49, 56, 64, 72 ...           refuse
169      240     182, 196, 210, 225               refuse
240      342     256, 272, 289, 306, 324          refuse
342      400     361, 380                         refuse
400      484     420, 441, 462                    refuse
484      529     506                              refuse
529      600     552, 576                         refuse
600      729     625, 650, 676, 702               refuse
729      841     756, 784, 812                    refuse
841      900     870                              refuse
900      961     930                              refuse
961      992     (none)                           finalizes

Fourteen refusals, then one success, matching the observed recovery. The other three episodes escaped the same way, on an adjacent pair: 529 → 552 (23², 23·24) and 342 → 361 (18·19, 19²).

Mechanism

Three rules, one candidate set:

  1. slot.py is_justifiable_after(finalized_slot): δ ≤ 5, a perfect square, or pronic. Consecutive rungs are n², n²+n, (n+1)², so spacing grows as ≈ √δ.
  2. state_transition.py: finalise source only when no justifiable slot lies strictly between source and target. A justification that jumps further than the rung spacing cannot finalise.
  3. validator_duties.py get_attestation_target: every validator walks its target back to a slot justifiable after the same finalized_slot.
  finalised frozen at F
        │
        ├─► rungs sparsen as √δ
        │        │
        │        ▼
        │   targets walk back onto sparse rungs; validators at different
        │   heads land on different rungs and votes scatter
        │        │
        │        ▼
        │   when justification lands it has jumped past a rung
        │        │
        └──  finalisation refuses ──┘

The frozen boundary shows up directly in aggregation. Votes whose target has already justified were skipped 52× the healthy rate during the 68-minute episode (277,128/h against ~5,350/h) and 20× during the 38-minute one, returning to baseline on recovery.

Escape is stochastic. It needs two consecutive justifications on adjacent rungs. Wider spacing at larger δ makes that more likely, but it doesn't set a threshold: the four episodes escaped at δ = 132, 961, 529 and 361. Duration is not predictable from the trigger.

The trigger is ordinary

Each episode began with a justification pause of 140–340 s, during which head ran 32–86 slots ahead and justification resumed past at least one rung.

Client aggregation was healthy through the stalls. During the 68-minute episode the aggregating client had zero empty sessions, an above-baseline success rate and a faster p90 than its own healthy windows, and it produced about one aggregate per slot. The stall doesn't need a degraded client.

One client-side contributor to the pauses has been found and fixed since: a duty clock not anchored to genesis (geanlabs/gean#451). It left that client's four validators voting on a later head, so the honest aggregate sat at 10 of 14, exactly the supermajority, with no margin. With it fixed, the main aggregate is usually 14 of 14 and justification pauses fell by about 60% over 4 h windows. Whether stall frequency drops too is still being measured. Either way, the amplifier described here is independent of what triggers it.

The same ladder at the healthy end

ethlambda documented the ladder's effect at δ = 6, in normal operation (8d90f97, on an unmerged branch):

on a chain whose justified - finalized sits at 6, the justifiable rungs are 3 slots apart (above delta 5 only squares and pronics qualify), so three consecutive slots of validators all vote for the same rung. Once it is justified, every pooled entry hit this filter, select_attestations returned an empty list on round 0, and every aggregator built no candidate body at all. Measured on devnet-5: 48% of slots had zero candidates built fleet-wide, and ~50% of blocks were empty, in a clean 3-on/3-off cycle.

The same holds here. Over one hour of healthy operation, for 318 of 325 targets exactly one slot's votes reached the chain. The first aggregate justifies the target and every later slot voting for it is dropped as already justified. 62.8% of slots put no vote on chain at all.

Those votes cannot help justify. But in the clients checked (gean, ethlambda), block import records a block's attestations as fork-choice votes whatever the justification verdict, so each dropped vote is also a head vote that never reaches other nodes through a block.

The spec already accepts that principle in one place. block_production_tiered.py (#1149) exempts the genesis self-vote from this filter:

Genesis self-votes are exempt from the target-after-source and target-already-justified checks. The state transition drops them, but they carry fork-choice signal.

That reasoning isn't specific to genesis. #1149 keeps the general filter; the unmerged ethlambda branch replaces it with scoring that zeroes a settled target's justification value and keeps its head-vote value.

Questions

  1. Should the finalisation gap check and the attestation-target walk be governed by the same finalized_slot-relative candidate set? Sharing it closes the loop: a frozen finalised slot sparsens the ladder, and the sparse ladder keeps it frozen.
  2. Is an unbounded stall acceptable? Recovery depends on two justifications landing on adjacent rungs, which nothing in the protocol steers toward. A builder that preferred the nearest rung above the source, rather than the furthest, might finalise sooner; feat(lstar): add tiered block-production strategy as a selectable alternative #1149's ordering prefers the furthest.
  3. Should the already-justified filter stay general, or follow the genesis exemption's own reasoning? As it stands about two thirds of cast votes, and their head votes, never reach the chain.
  4. Would a spec test for this be worth adding? Present justifiability fixtures test is_justifiable_after alone; none exercise it together with the finalisation gap check across non-adjacent rungs.

Full Prometheus data and block contents for all four episodes are available.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions