Skip to content

[Epic] Webhook retry system #172

Description

@Benzilla

Webhook retries

Goal

Currently webhooks on Zoneless are fire and forget, if the webhook fails to deliver, then we don't retry it. This is bad behaviour, as the platform using Zoneless might miss critical events due to being temporarily down.

Proposed Solution

We create a new initial simple system within Zoneless for retrying failed webhooks. Webhooks essentially deliver Event objects, however we would need a new system to track Webhook Delivery Attempts.

The proposed solution would be to create a new WebhookDelivery system, that acts as a parent object for tracking WebhookDeliveryAttempts. WebhookDelivery objects are linked to the Event they are trying to send and the WebhookEndpoint they are being sent to.

The objects would roughly look like this - however it is okay for the implemetation to differ based on the findings by the implementation:

export interface WebhookDelivery {
  id: string;
  event_id: string;
  webhook_endpoint_id: string;
  status: 'pending' | 'retrying' | 'succeeded' | 'failed';
  next_attempt_at: number | null;
  delivered_at: number | null;
  claim_until: number | null;
  attempts: WebhookDeliveryAttempt[];
}
export interface WebhookDeliveryAttempt {
  attempt_number: number;
  attempted_at: number;
  completed_at: number;
  result: 'succeeded' | 'http_error' | 'timed_out' | 'network_error';
  http_status: number | null;
  duration_ms: number;
  error: string | null;
  url: string;
}

A recovered delivery attempt might look like this:

{
  "id": "whd_z_123",
  "event_id": "evt_z_123",
  "webhook_endpoint_id": "we_z_123",
  "status": "succeeded",
  "next_attempt_at": null,
  "delivered_at": 1789240915,
  "claim_until": null,
  "attempts": [
    {
      "attempt_number": 1,
      "attempted_at": 1789237274,
      "completed_at": 1789237278,
      "result": "http_error",
      "http_status": 500,
      "duration_ms": 3693,
      "error": "HTTP 500",
      "url": "https://example.com/webhooks"
    },
    {
      "attempt_number": 2,
      "attempted_at": 1789240915,
      "completed_at": 1789240918,
      "result": "succeeded",
      "http_status": 200,
      "duration_ms": 3206,
      "error": null,
      "url": "https://example.com/webhooks"
    }
  ]
}

Required changes

  • Add WebhookDelivery and its nested WebhookDeliveryAttempt type in
    shared-types.
  • Add a WebhookDelivery module responsible for creating deliveries,
    recording attempts and controlling delivery state transitions.
  • Update EventService to create one delivery for every subscribed endpoint
    before making the first delivery attempt.
  • Keep WebhookDispatcher responsible for making one HTTP request and
    returning its result. It should not decide whether or when to retry.
  • Update Event.pending_webhooks exactly once when a delivery first succeeds.
  • Add a retry worker that atomically claims due deliveries using an expiring
    lock.
  • Add an operator authenticated internal endpoint that processes a bounded
    batch of due deliveries. Self-hosters can invoke it using cron or Cloud
    Scheduler. This would be similar to how checks for Subscription updates
    or checks for incoming top ups are implemented.
  • Make one immediate attempt followed by three retries using gradual backoff.
    Initially retry after 5 minutes, 1 hours, 6 hours. These should be easily
    configurable in the code for future changes to backoff strategies.
  • Mark the delivery as failed after the final unsuccessful attempt.
  • Stop retrying when the associated endpoint has been disabled or deleted.

Behaviour

  • Any HTTP 2xx response succeeds.
  • Network errors, timeouts and non-2xx responses fail.
  • Retries use the same stored Event and Event ID.
  • Every attempt receives a newly generated signature and timestamp.
  • Multiple endpoints for the same Event are tracked independently.

Acceptance criteria

  • A delivery is persisted before its first network request.
  • A successful first attempt marks the delivery as succeeded.
  • A failed attempt records its result and schedules a retry.
  • A later successful attempt prevents further automatic retries.
  • Failed deliveries become terminal after exhausting the retry policy.
  • Two workers cannot successfully claim the same delivery simultaneously.
  • pending_webhooks cannot be decremented twice for one delivery.
  • Attempt history distinguishes HTTP errors, timeouts and network errors.
  • Tests cover immediate success, eventual recovery, exhausted retries,
    multiple endpoints and concurrent workers.

Non-goals

  • Public Webhook Delivery API endpoints
  • Dashboard delivery history
  • Manual resend
  • Failure notification emails

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or requesthelp wantedExtra attention is needed

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions