Conversation
A server, or a proxy in front of it, can close a persistent connection while it sits idle in the client. The socket still looks open: closed? is false and the kernel accepts the next request into the half-closed socket, so the failure only surfaces when reading the response, as "couldn't read response headers" (httprb#420, httprb#459). A server that writes on the idle connection, like the 408 some send before closing, gets that response read as the answer to the next request. Before reusing the connection, check its socket without blocking. On an idle connection any readable data means it can't carry another request, so close it and connect again. The check is skipped while a response is pending, because its unread body is expected data, and when the timeout class doesn't expose an IO.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Backport of #859 to 5-x-stable, like the 5.3.0 backports (#775, #783, #789, #792, #800). 6.x requires Ruby 3.2, so 5.x is the newest line apps on Ruby 2.6–3.1 can use, and they hit the same
couldn't read response headerswhen the server closed the connection while it sat idle (#420, #459).Same check as main. Before reusing a connection,
verify_connection!asksConnection#stale?, which polls the socket without blocking, and reconnects if anything is readable.One difference from main: 5.x doesn't flush an unread body before the next request, it raises
StateError.stale?returns false while a response is pending, so that stays as is. I checked with a 2 MiB body left unread and then a second request: the sameStateErroron 1 connection, on v5.3.1 and on this branch.Tests:
#stale?examples on aUNIXSocket.pairConnectionError. It now expects the reopen, and so do new examples for a POST,Timeout::PerOperation,Timeout::Global, and a server that writes a 408 on the idle connectionResults:
#inspectexamples that fail on v5.3.1 (Ruby 3.4 changedHash#inspect)stale?call (6 failures), the pending-response guard (1), theto_ioguard (2), the rescue returning false (1), always stale (2), never stale (9)JRuby 9.3.15.0 doesn't load the suite as is: rspec-memory 1.0.4 calls
ObjectSpace.memsize_ofat load, and JRuby 9.3 raisesNoMethodErrorfor it. With a shim that returns 0 andCI=truelike the CI job (so the existing:flakyexamples were skipped, and the new ones, not tagged yet, ran):Every failure is a request WEBrick never answered on a fresh connection:
couldn't read response headers, sometimes after its 30 s RequestTimeout, or a timeout. It hits untouched examples too ("is easy for 302", "with query string parameters", the large-body writes, the non-ASCII URLs). "when reading a cached body succeeds" failed 1 of 14 runs on each tree. The 5 new examples that close or write to server sockets out of band failed 2 of 70, so they're tagged:flakylike the original "transparently reopens"; MRI still runs them. "raises" passed 14 of 14. I didn't dig into why WEBrick stalls on JRuby.Limits are the same as main's. A close that lands after the check still fails; #861 resends idempotent requests on main, and it's a feature, so it isn't backported. And the poll can't see bytes already in OpenSSL's buffer.