Skip to content

fix(zombie): cap the zombie-detection threshold - #71

Merged
zmstone merged 4 commits into
mainfrom
260925-clamp-zombie-timeout
Sep 29, 2026
Merged

zmstone merged 4 commits into
mainfrom
260925-clamp-zombie-timeout

Conversation

@zmstone

@zmstone zmstone commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Refs #70

Why

Zombie-connection detection fires at request timeout + max_inactive after the latest send, so it scales with the caller's timeout. A production deployment runs with a 600 s request timeout. With the default 10 s max_inactive, a dead connection is reconnected 610 s after the send. The operator cannot shorten that without breaking the long requests the 600 s timeout was chosen for.

What

The threshold becomes

max(MaxInactive, min(MaxInactiveCap, MaxTimeout + MaxInactive))

MaxTimeout is the largest request timeout among the in-flight requests. MaxInactiveCap defaults to 60 s. It is a new pool option, max_inactive_cap, next to the existing max_inactive option, so tests can exercise the cap in seconds.

max_inactive request timeout detection before detection after
10s 5s 15s 15s
10s 45s 55s 55s
10s 600s 610s 60s
10s infinity never 60s
180s 600s 610s 180s

The outer max keeps a raised max_inactive authoritative. When max_inactive >= 60s, the threshold is exactly max_inactive.

Commits

  1. refactor: ExpireAt becomes {CalledAt, Timeout} instead of a precomputed deadline, so the worker can read the request timeout. deadline/1, timeout/1 and is_expired/2 interpret it. The pass-through sites are unchanged. No behaviour change.
  2. fix(zombie): the threshold change, the max_inactive_cap option, tests, version bump to 0.7.6 and changelog.

Reference timestamp

The old code stored the latest deadline (Now + Timeout) and compared now - MaxDeadline > MaxInactive. The new threshold adds MaxTimeout itself, so the reference must be a send-side timestamp. Otherwise the timeout is counted twice. max_sent_expire is renamed to max_sent_at, and max_sent_timeout sits next to it. Both reset to 0 when sent empties.

max_sent_at is now_() taken in put_sent_req/3, when the request goes to gun, not the caller's call time. The two differ by the time the request spent in the pending queue. The send time answers "how long since we sent something and heard nothing", and it needs no special case for infinity. With a uniform timeout under the cap, the new rule fires at the same point as the old one, apart from that queue time.

The infinity hole

max_expire/2 skipped infinity deadlines. A pool whose requests all used Timeout = infinity never recorded a deadline, so zombie detection never fired. infinity now wins the timeout maximum, and the threshold for it is the cap.

Incompatible case

With a request timeout above cap - max_inactive (50 s by default), a response that legitimately takes longer than the cap now has its connection killed at the cap. Received data does not count as activity: the reference timestamp is written only at send time, so a response that streams for 90 s looks the same as a jammed connection. Such deployments must raise max_inactive_cap or max_inactive to the slowest expected response. The changelog states this. Counting received data as activity is a possible follow-up in #70 and is out of scope here.

Log

force_reconnecting_zombie_http_connection now reports inactive_duration_threshold as the computed threshold, last_request_sent_at instead of last_request_expire, and max_request_timeout.

Tests

  • inactive_threshold_test_ (in src/ehttpc.erl): table test of the threshold function, covering every row above, infinity, and the max_inactive >= cap floor.
  • zombie_detect_long_timeout_test: max_inactive = 1s, max_inactive_cap = 2s, request timeout 30 s, server never answers. reconnect fires between 2 s and 3.5 s.
  • zombie_detect_infinity_timeout_test: same, with request timeout infinity.
  • The existing zombie_detect_inflight_* tests pass with only the field rename.

Both new tests fail against the main worker: reconnect does not arrive within 3.5 s.

Run locally on OTP 28: make fmt-check, ./check-style.sh, make compile, make xref, make dialyzer, make eunit. make eunit has 9 failures, all in ehttpc_google_tests:proxy_test_, because tinyproxy is not installed on the host. main has the same 9 failures. The first commit alone passes rebar3 eunit --module=ehttpc,ehttpc_tests. The proxy tests and OTP 25/26/27 need CI.

ExpireAt was an absolute deadline computed in the caller. It becomes
{CalledAt, Timeout}, so the worker can read the request timeout
itself. The zombie-detection threshold needs that value to clamp
detection time for long timeouts.

deadline/1, timeout/1 and is_expired/2 interpret the tuple. The sites
that only pass ExpireAt through are unchanged. No behaviour change.
Zombie detection fired at request timeout + max_inactive after the
newest send, so it scaled with the caller's timeout. With a 600 s
request timeout a dead connection was reconnected only after 610 s,
and lowering the timeout would break the long requests it was chosen
for.

The threshold is now

    max(MaxInactive, min(MaxInactiveCap, MaxTimeout + MaxInactive))

MaxTimeout is the largest request timeout among the in-flight
requests. MaxInactiveCap defaults to 60 s and is a new pool option,
max_inactive_cap. For timeouts up to 50 s with the default
max_inactive the result is unchanged. A raised max_inactive stays
authoritative through the outer max.

The reference timestamp is now the latest send time (max_sent_at),
not the latest deadline. The timeout is added by the threshold, so a
deadline reference would count it twice. The send time is taken with
now_() in put_sent_req/3, when the request goes to gun, rather than
the caller's call time. The two differ by the time spent in the
pending queue, and the send time is what "silent since the last send"
means. It also needs no special case for infinity.

max_expire/2 skipped infinity deadlines, so a pool whose requests all
used an infinity timeout never recorded one and zombie detection never
fired. infinity now wins the timeout maximum, and the threshold for it
is the cap.

max_sent_expire is renamed to max_sent_at, and max_sent_timeout is
tracked next to it. Both reset to 0 when no request is in flight.

The force_reconnecting_zombie_http_connection log reports the computed
threshold, the latest send time and the largest request timeout.

Incompatible case: with a request timeout above the cap minus
max_inactive, a response that takes longer than the cap now gets its
connection killed at the cap. Received data does not count as
activity. Raise max_inactive to the slowest expected response.

Refs #70
Comment thread src/ehttpc.erl Outdated
reset_sent(R).

reset_sent(Requests) ->
Requests#{sent => #{}, max_sent_at => 0, max_sent_timeout => 0}.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Requests#{sent => #{}, max_sent_at => 0, max_sent_timeout => 0}.
Requests#{sent := #{}, max_sent_at := 0, max_sent_timeout := 0}.

Comment thread changelog.md Outdated
detection entirely.
- Behaviour change: with a request timeout above about 50 seconds and the default
`max_inactive`, a connection whose response takes longer than 60 seconds is now killed at
60 seconds. Set the `max_inactive` pool option to the slowest expected response time to

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

max_inactive_cap?

The cap is a pool option, so the changelog names it and lists it as a
way to raise the bound, next to max_inactive.
@zmstone
zmstone merged commit 547cb99 into main Sep 29, 2026
4 checks passed
@zmstone
zmstone deleted the 260925-clamp-zombie-timeout branch October 1, 2026 14:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants