Skip to content

backport: Reconnect instead of reusing idle connections the server closed - #862

Draft
ilyazub wants to merge 1 commit into
httprb:5-x-stablefrom
serpapi:backport-5x/liveness-check
Draft

ilyazub wants to merge 1 commit into
httprb:5-x-stablefrom
serpapi:backport-5x/liveness-check

Conversation

@ilyazub

@ilyazub ilyazub commented Oct 3, 2026 •

Copy link
Copy Markdown

Backport of #859 to 5-x-stable, like the 5.3.0 backports (#775, #783, #789, #792, #800). 6.x requires Ruby 3.2, so 5.x is the newest line apps on Ruby 2.6–3.1 can use, and they hit the same couldn't read response headers when the server closed the connection while it sat idle (#420, #459).

Same check as main. Before reusing a connection, verify_connection! asks Connection#stale?, which polls the socket without blocking, and reconnects if anything is readable.

One difference from main: 5.x doesn't flush an unread body before the next request, it raises StateError. stale? returns false while a response is pending, so that stays as is. I checked with a 2 MiB body left unread and then a second request: the same StateError on 1 connection, on v5.3.1 and on this branch.

Tests:

  • 7 #stale? examples on a UNIXSocket.pair
  • a TLS example with its own server, because "working with SSL" is disabled (OpenSSL::SSL::SSLErrors in test suite #627)
  • "transparently reopens" asserted that the first request after the server closed the socket raised ConnectionError. It now expects the reopen, and so do new examples for a POST, Timeout::PerOperation, Timeout::Global, and a server that writes a 408 on the idle connection
  • "raises" when the server closes after reading the request, so nothing gets resent

Results:

  • MRI 2.7.8: 934 examples, 0 failures, 30 pending on seeds 4242, 777 and 31337. v5.3.1 has 916 and 25 pending; the 5 extra pending are the new shared examples under the disabled SSL context
  • MRI 3.4.8: 2 failures, the same 2 #inspect examples that fail on v5.3.1 (Ruby 3.4 changed Hash#inspect)
  • rubocop clean, yardstick 58.4%
  • Mutations, all killed: dropping the stale? call (6 failures), the pending-response guard (1), the to_io guard (2), the rescue returning false (1), always stale (2), never stale (9)
  • Not run locally: MRI 2.6 and 3.0–3.3

JRuby 9.3.15.0 doesn't load the suite as is: rspec-memory 1.0.4 calls ObjectSpace.memsize_of at load, and JRuby 9.3 raises NoMethodError for it. With a shim that returns 0 and CI=true like the CI job (so the existing :flaky examples were skipped, and the new ones, not tagged yet, ran):

v5.3.1 this branch
full suite, seeds 11/22/33/44 1, 2, 1, 1 failures 2, 1, 1, 2
client_spec, seeds 1–10 2 failures in 10 runs 1 in 10

Every failure is a request WEBrick never answered on a fresh connection: couldn't read response headers, sometimes after its 30 s RequestTimeout, or a timeout. It hits untouched examples too ("is easy for 302", "with query string parameters", the large-body writes, the non-ASCII URLs). "when reading a cached body succeeds" failed 1 of 14 runs on each tree. The 5 new examples that close or write to server sockets out of band failed 2 of 70, so they're tagged :flaky like the original "transparently reopens"; MRI still runs them. "raises" passed 14 of 14. I didn't dig into why WEBrick stalls on JRuby.

Limits are the same as main's. A close that lands after the check still fails; #861 resends idempotent requests on main, and it's a feature, so it isn't backported. And the poll can't see bytes already in OpenSSL's buffer.

A server, or a proxy in front of it, can close a persistent connection
while it sits idle in the client. The socket still looks open: closed?
is false and the kernel accepts the next request into the half-closed
socket, so the failure only surfaces when reading the response, as
"couldn't read response headers" (httprb#420, httprb#459). A server that writes on
the idle connection, like the 408 some send before closing, gets that
response read as the answer to the next request.

Before reusing the connection, check its socket without blocking. On an
idle connection any readable data means it can't carry another request,
so close it and connect again. The check is skipped while a response is
pending, because its unread body is expected data, and when the timeout
class doesn't expose an IO.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant