Webhook retries
Goal
Currently webhooks on Zoneless are fire and forget, if the webhook fails to deliver, then we don't retry it. This is bad behaviour, as the platform using Zoneless might miss critical events due to being temporarily down.
Proposed Solution
We create a new initial simple system within Zoneless for retrying failed webhooks. Webhooks essentially deliver Event objects, however we would need a new system to track Webhook Delivery Attempts.
The proposed solution would be to create a new WebhookDelivery system, that acts as a parent object for tracking WebhookDeliveryAttempts. WebhookDelivery objects are linked to the Event they are trying to send and the WebhookEndpoint they are being sent to.
The objects would roughly look like this - however it is okay for the implemetation to differ based on the findings by the implementation:
export interface WebhookDelivery {
id: string;
event_id: string;
webhook_endpoint_id: string;
status: 'pending' | 'retrying' | 'succeeded' | 'failed';
next_attempt_at: number | null;
delivered_at: number | null;
claim_until: number | null;
attempts: WebhookDeliveryAttempt[];
}
export interface WebhookDeliveryAttempt {
attempt_number: number;
attempted_at: number;
completed_at: number;
result: 'succeeded' | 'http_error' | 'timed_out' | 'network_error';
http_status: number | null;
duration_ms: number;
error: string | null;
url: string;
}
A recovered delivery attempt might look like this:
{
"id": "whd_z_123",
"event_id": "evt_z_123",
"webhook_endpoint_id": "we_z_123",
"status": "succeeded",
"next_attempt_at": null,
"delivered_at": 1789240915,
"claim_until": null,
"attempts": [
{
"attempt_number": 1,
"attempted_at": 1789237274,
"completed_at": 1789237278,
"result": "http_error",
"http_status": 500,
"duration_ms": 3693,
"error": "HTTP 500",
"url": "https://example.com/webhooks"
},
{
"attempt_number": 2,
"attempted_at": 1789240915,
"completed_at": 1789240918,
"result": "succeeded",
"http_status": 200,
"duration_ms": 3206,
"error": null,
"url": "https://example.com/webhooks"
}
]
}
Required changes
- Add
WebhookDelivery and its nested WebhookDeliveryAttempt type in
shared-types.
- Add a
WebhookDelivery module responsible for creating deliveries,
recording attempts and controlling delivery state transitions.
- Update
EventService to create one delivery for every subscribed endpoint
before making the first delivery attempt.
- Keep
WebhookDispatcher responsible for making one HTTP request and
returning its result. It should not decide whether or when to retry.
- Update
Event.pending_webhooks exactly once when a delivery first succeeds.
- Add a retry worker that atomically claims due deliveries using an expiring
lock.
- Add an operator authenticated internal endpoint that processes a bounded
batch of due deliveries. Self-hosters can invoke it using cron or Cloud
Scheduler. This would be similar to how checks for Subscription updates
or checks for incoming top ups are implemented.
- Make one immediate attempt followed by three retries using gradual backoff.
Initially retry after 5 minutes, 1 hours, 6 hours. These should be easily
configurable in the code for future changes to backoff strategies.
- Mark the delivery as failed after the final unsuccessful attempt.
- Stop retrying when the associated endpoint has been disabled or deleted.
Behaviour
- Any HTTP 2xx response succeeds.
- Network errors, timeouts and non-2xx responses fail.
- Retries use the same stored Event and Event ID.
- Every attempt receives a newly generated signature and timestamp.
- Multiple endpoints for the same Event are tracked independently.
Acceptance criteria
- A delivery is persisted before its first network request.
- A successful first attempt marks the delivery as succeeded.
- A failed attempt records its result and schedules a retry.
- A later successful attempt prevents further automatic retries.
- Failed deliveries become terminal after exhausting the retry policy.
- Two workers cannot successfully claim the same delivery simultaneously.
pending_webhooks cannot be decremented twice for one delivery.
- Attempt history distinguishes HTTP errors, timeouts and network errors.
- Tests cover immediate success, eventual recovery, exhausted retries,
multiple endpoints and concurrent workers.
Non-goals
- Public Webhook Delivery API endpoints
- Dashboard delivery history
- Manual resend
- Failure notification emails
Webhook retries
Goal
Currently webhooks on Zoneless are fire and forget, if the webhook fails to deliver, then we don't retry it. This is bad behaviour, as the platform using Zoneless might miss critical events due to being temporarily down.
Proposed Solution
We create a new initial simple system within Zoneless for retrying failed webhooks. Webhooks essentially deliver Event objects, however we would need a new system to track Webhook Delivery Attempts.
The proposed solution would be to create a new
WebhookDeliverysystem, that acts as a parent object for trackingWebhookDeliveryAttempts.WebhookDeliveryobjects are linked to theEventthey are trying to send and theWebhookEndpointthey are being sent to.The objects would roughly look like this - however it is okay for the implemetation to differ based on the findings by the implementation:
A recovered delivery attempt might look like this:
{ "id": "whd_z_123", "event_id": "evt_z_123", "webhook_endpoint_id": "we_z_123", "status": "succeeded", "next_attempt_at": null, "delivered_at": 1789240915, "claim_until": null, "attempts": [ { "attempt_number": 1, "attempted_at": 1789237274, "completed_at": 1789237278, "result": "http_error", "http_status": 500, "duration_ms": 3693, "error": "HTTP 500", "url": "https://example.com/webhooks" }, { "attempt_number": 2, "attempted_at": 1789240915, "completed_at": 1789240918, "result": "succeeded", "http_status": 200, "duration_ms": 3206, "error": null, "url": "https://example.com/webhooks" } ] }Required changes
WebhookDeliveryand its nestedWebhookDeliveryAttempttype inshared-types.WebhookDeliverymodule responsible for creating deliveries,recording attempts and controlling delivery state transitions.
EventServiceto create one delivery for every subscribed endpointbefore making the first delivery attempt.
WebhookDispatcherresponsible for making one HTTP request andreturning its result. It should not decide whether or when to retry.
Event.pending_webhooksexactly once when a delivery first succeeds.lock.
batch of due deliveries. Self-hosters can invoke it using cron or Cloud
Scheduler. This would be similar to how checks for Subscription updates
or checks for incoming top ups are implemented.
Initially retry after 5 minutes, 1 hours, 6 hours. These should be easily
configurable in the code for future changes to backoff strategies.
Behaviour
Acceptance criteria
pending_webhookscannot be decremented twice for one delivery.multiple endpoints and concurrent workers.
Non-goals