Repository navigation
Add heartbeat listener for liveness probes - #60
Merged
Merged
Conversation
HeartbeatListener touches a file when the worker starts and after every iteration and removes it when the worker stops, so a liveness probe can detect a worker that hangs in a job. It can be enabled with the new heartbeatFile option in create().
The test measured the job with the real clock and expected 0ms, which fails whenever the job run crosses a millisecond boundary.
DanielBadura
approved these changes
Oct 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A process manager only sees whether the worker process is alive, not whether it hangs in a job. This adds a HeartbeatListener that touches a file when the worker starts and after every iteration, and removes it when the worker stops (which since #52 also happens when the job throws). A liveness probe can then restart the worker if the file gets too old. It can be enabled with the new heartbeatFile option in create(), and the docs include a Kubernetes probe example.
If the file can't be written, the listener logs a warning and the worker keeps running. The probe will notice it anyway, so there's no point in crashing the worker over it.
The file is only updated between iterations, so the probe threshold has to be larger than the longest job plus the sleep timer. The docs point that out.
This adds an option next to the one from #59, so whichever is merged second needs a small rebase.