You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
On a 14-node mixed-client devnet, finalisation froze four times, for up to 68 minutes, while justification kept advancing throughout. Head never stalled and nodes never disagreed on head. The clients were implementing the spec as written.
The cause is a coupling between two rules that are governed by one candidate set. The finalisation gap check in state_transition.py and the attestation-target walk in validator_duties.py both use Slot.is_justifiable_after against the same finalized_slot. Once finalisation slips past δ = 6 the two reinforce each other, and the stall sustains itself until two justifications happen to land on adjacent rungs.
This is a question about intended behaviour, not a proposed fix. The gap check is a safety condition and should not be loosened without analysis.
Setup
14 nodes (4 gean, 6 ethlambda, 4 lantern), 14 validators, 4 committees, 4 s slots. Genesis 2026-09-22 14:35:05 UTC. All clients on leanVM 48a90420.
The four episodes
freeze
finalized held at
duration
peak head − finalized
trigger
escape
09-22 21:17
6012
~6 min
~140
340 s justification pause
δ ≈ 132
09-23 04:11
12227
68 min
~1030
160 s pause
δ 961 → 992
09-24 06:07
35569
38 min
~580
140 s pause
δ 529 → 552
09-24 10:56
39905
25 min
~390
140 s pause
δ 342 → 361
Healthy baseline between episodes: head − finalized median ~10.
In the 68-minute episode, latest_finalized_slot held at 12227 while latest_justified_slot advanced fourteen times, from 12233 to 13188. It is a finalisation stall, not a justification stall. A check of "is the chain justifying?" stays green throughout.
Every justification lands on a rung
In all four episodes, every justification while finalisation was frozen landed on a slot justifiable after the frozen finalised slot: 42 of 42. The 68-minute episode, as δ from 12227:
Fourteen refusals, then one success, matching the observed recovery. The other three episodes escaped the same way, on an adjacent pair: 529 → 552 (23², 23·24) and 342 → 361 (18·19, 19²).
Mechanism
Three rules, one candidate set:
slot.pyis_justifiable_after(finalized_slot): δ ≤ 5, a perfect square, or pronic. Consecutive rungs are n², n²+n, (n+1)², so spacing grows as ≈ √δ.
state_transition.py: finalise source only when no justifiable slot lies strictly between source and target. A justification that jumps further than the rung spacing cannot finalise.
validator_duties.pyget_attestation_target: every validator walks its target back to a slot justifiable after the same finalized_slot.
finalised frozen at F
│
├─► rungs sparsen as √δ
│ │
│ ▼
│ targets walk back onto sparse rungs; validators at different
│ heads land on different rungs and votes scatter
│ │
│ ▼
│ when justification lands it has jumped past a rung
│ │
└── finalisation refuses ──┘
The frozen boundary shows up directly in aggregation. Votes whose target has already justified were skipped 52× the healthy rate during the 68-minute episode (277,128/h against ~5,350/h) and 20× during the 38-minute one, returning to baseline on recovery.
Escape is stochastic. It needs two consecutive justifications on adjacent rungs. Wider spacing at larger δ makes that more likely, but it doesn't set a threshold: the four episodes escaped at δ = 132, 961, 529 and 361. Duration is not predictable from the trigger.
The trigger is ordinary
Each episode began with a justification pause of 140–340 s, during which head ran 32–86 slots ahead and justification resumed past at least one rung.
Client aggregation was healthy through the stalls. During the 68-minute episode the aggregating client had zero empty sessions, an above-baseline success rate and a faster p90 than its own healthy windows, and it produced about one aggregate per slot. The stall doesn't need a degraded client.
One client-side contributor to the pauses has been found and fixed since: a duty clock not anchored to genesis (geanlabs/gean#451). It left that client's four validators voting on a later head, so the honest aggregate sat at 10 of 14, exactly the supermajority, with no margin. With it fixed, the main aggregate is usually 14 of 14 and justification pauses fell by about 60% over 4 h windows. Whether stall frequency drops too is still being measured. Either way, the amplifier described here is independent of what triggers it.
The same ladder at the healthy end
ethlambda documented the ladder's effect at δ = 6, in normal operation (8d90f97, on an unmerged branch):
on a chain whose justified - finalized sits at 6, the justifiable rungs are 3 slots apart (above delta 5 only squares and pronics qualify), so three consecutive slots of validators all vote for the same rung. Once it is justified, every pooled entry hit this filter, select_attestations returned an empty list on round 0, and every aggregator built no candidate body at all. Measured on devnet-5: 48% of slots had zero candidates built fleet-wide, and ~50% of blocks were empty, in a clean 3-on/3-off cycle.
The same holds here. Over one hour of healthy operation, for 318 of 325 targets exactly one slot's votes reached the chain. The first aggregate justifies the target and every later slot voting for it is dropped as already justified. 62.8% of slots put no vote on chain at all.
Those votes cannot help justify. But in the clients checked (gean, ethlambda), block import records a block's attestations as fork-choice votes whatever the justification verdict, so each dropped vote is also a head vote that never reaches other nodes through a block.
The spec already accepts that principle in one place. block_production_tiered.py (#1149) exempts the genesis self-vote from this filter:
Genesis self-votes are exempt from the target-after-source and target-already-justified checks. The state transition drops them, but they carry fork-choice signal.
That reasoning isn't specific to genesis. #1149 keeps the general filter; the unmerged ethlambda branch replaces it with scoring that zeroes a settled target's justification value and keeps its head-vote value.
Questions
Should the finalisation gap check and the attestation-target walk be governed by the same finalized_slot-relative candidate set? Sharing it closes the loop: a frozen finalised slot sparsens the ladder, and the sparse ladder keeps it frozen.
Is an unbounded stall acceptable? Recovery depends on two justifications landing on adjacent rungs, which nothing in the protocol steers toward. A builder that preferred the nearest rung above the source, rather than the furthest, might finalise sooner; feat(lstar): add tiered block-production strategy as a selectable alternative #1149's ordering prefers the furthest.
Should the already-justified filter stay general, or follow the genesis exemption's own reasoning? As it stands about two thirds of cast votes, and their head votes, never reach the chain.
Would a spec test for this be worth adding? Present justifiability fixtures test is_justifiable_after alone; none exercise it together with the finalisation gap check across non-adjacent rungs.
Full Prometheus data and block contents for all four episodes are available.
Summary
On a 14-node mixed-client devnet, finalisation froze four times, for up to 68 minutes, while justification kept advancing throughout. Head never stalled and nodes never disagreed on head. The clients were implementing the spec as written.
The cause is a coupling between two rules that are governed by one candidate set. The finalisation gap check in
state_transition.pyand the attestation-target walk invalidator_duties.pyboth useSlot.is_justifiable_afteragainst the samefinalized_slot. Once finalisation slips past δ = 6 the two reinforce each other, and the stall sustains itself until two justifications happen to land on adjacent rungs.This is a question about intended behaviour, not a proposed fix. The gap check is a safety condition and should not be loosened without analysis.
Setup
14 nodes (4 gean, 6 ethlambda, 4 lantern), 14 validators, 4 committees, 4 s slots. Genesis 2026-09-22 14:35:05 UTC. All clients on leanVM
48a90420.The four episodes
Healthy baseline between episodes: head − finalized median ~10.
In the 68-minute episode,
latest_finalized_slotheld at 12227 whilelatest_justified_slotadvanced fourteen times, from 12233 to 13188. It is a finalisation stall, not a justification stall. A check of "is the chain justifying?" stays green throughout.Every justification lands on a rung
In all four episodes, every justification while finalisation was frozen landed on a slot justifiable after the frozen finalised slot: 42 of 42. The 68-minute episode, as δ from 12227:
At each step the finalisation check refused, because a justifiable slot sat strictly between source and target:
Fourteen refusals, then one success, matching the observed recovery. The other three episodes escaped the same way, on an adjacent pair: 529 → 552 (23², 23·24) and 342 → 361 (18·19, 19²).
Mechanism
Three rules, one candidate set:
slot.pyis_justifiable_after(finalized_slot): δ ≤ 5, a perfect square, or pronic. Consecutive rungs are n², n²+n, (n+1)², so spacing grows as ≈ √δ.state_transition.py: finalisesourceonly when no justifiable slot lies strictly betweensourceandtarget. A justification that jumps further than the rung spacing cannot finalise.validator_duties.pyget_attestation_target: every validator walks its target back to a slot justifiable after the samefinalized_slot.The frozen boundary shows up directly in aggregation. Votes whose target has already justified were skipped 52× the healthy rate during the 68-minute episode (277,128/h against ~5,350/h) and 20× during the 38-minute one, returning to baseline on recovery.
Escape is stochastic. It needs two consecutive justifications on adjacent rungs. Wider spacing at larger δ makes that more likely, but it doesn't set a threshold: the four episodes escaped at δ = 132, 961, 529 and 361. Duration is not predictable from the trigger.
The trigger is ordinary
Each episode began with a justification pause of 140–340 s, during which head ran 32–86 slots ahead and justification resumed past at least one rung.
Client aggregation was healthy through the stalls. During the 68-minute episode the aggregating client had zero empty sessions, an above-baseline success rate and a faster p90 than its own healthy windows, and it produced about one aggregate per slot. The stall doesn't need a degraded client.
One client-side contributor to the pauses has been found and fixed since: a duty clock not anchored to genesis (geanlabs/gean#451). It left that client's four validators voting on a later head, so the honest aggregate sat at 10 of 14, exactly the supermajority, with no margin. With it fixed, the main aggregate is usually 14 of 14 and justification pauses fell by about 60% over 4 h windows. Whether stall frequency drops too is still being measured. Either way, the amplifier described here is independent of what triggers it.
The same ladder at the healthy end
ethlambda documented the ladder's effect at δ = 6, in normal operation (
8d90f97, on an unmerged branch):The same holds here. Over one hour of healthy operation, for 318 of 325 targets exactly one slot's votes reached the chain. The first aggregate justifies the target and every later slot voting for it is dropped as already justified. 62.8% of slots put no vote on chain at all.
Those votes cannot help justify. But in the clients checked (gean, ethlambda), block import records a block's attestations as fork-choice votes whatever the justification verdict, so each dropped vote is also a head vote that never reaches other nodes through a block.
The spec already accepts that principle in one place.
block_production_tiered.py(#1149) exempts the genesis self-vote from this filter:That reasoning isn't specific to genesis. #1149 keeps the general filter; the unmerged ethlambda branch replaces it with scoring that zeroes a settled target's justification value and keeps its head-vote value.
Questions
finalized_slot-relative candidate set? Sharing it closes the loop: a frozen finalised slot sparsens the ladder, and the sparse ladder keeps it frozen.is_justifiable_afteralone; none exercise it together with the finalisation gap check across non-adjacent rungs.Full Prometheus data and block contents for all four episodes are available.