Skip to content

Reconnect instead of reusing idle connections the server closed - #859

Draft
ilyazub wants to merge 1 commit into
httprb:mainfrom
serpapi:liveness-check-standalone
Draft

ilyazub wants to merge 1 commit into
httprb:mainfrom
serpapi:liveness-check-standalone

Conversation

@ilyazub

@ilyazub ilyazub commented Oct 3, 2026 •

Copy link
Copy Markdown

A persistent client writes its next request into a connection the server already closed, then fails with couldn't read response headers (#420, #459). Before reusing a connection, this branch reads off the previous response and polls the socket without blocking. If anything is readable on an idle connection, the client opens a new one. Usually that's an EOF, a reset, a TLS close_notify or an unsolicited 408, and none of those can carry another request.

#459 was closed pointing at .retriable (#775). That helps, but it's opt-in, it still writes the request into the dead socket and pays the failed round trip before retrying, and it retries non-idempotent requests too. The check here runs before any byte goes out, costs about 1 µs, and is safe for every method, so the default client gets it. .retriable still covers the case below where the close crosses the request.

It also happens with the default config. keep_alive_timeout is 5 s, and plenty of servers close idle connections at 5 s too. They start counting when they send the response and the client starts after it reads it, so the client loses that race even with matching timeouts. Second request 4.99 s after the first (3 s for haproxy.org), default options:

host server closed idle connection after main this branch
www.debian.org 4.97–4.98 s ResponseHeaderError 3/3 200 on a new connection 3/3
lwn.net 4.81–4.93 s ResponseHeaderError 3/3 3/3
www.postgresql.org 4.85–4.88 s SSLError unexpected eof 3/3 3/3
www.sqlite.org 4.90–4.92 s SSLError unexpected eof 3/3 3/3
www.haproxy.org 1.94–4.98 s ResponseHeaderError 4/4 4/4

Servers that keep idle connections open still get reused: 14 hosts, 28 of 28 after 4.5 s idle with default options (www.samba.org was down that run), and 15 hosts, 30 of 30 after 6 s with keep_alive_timeout: 300.

Local repro, plain and TLS give the same result, MRI 3.4.8 and JRuby 10.1.2.0 too:

server main this branch
closes the idle connection ResponseHeaderError 200, new connection
sends a 408 on the idle connection, then closes the 408 comes back as the 2nd response 200, new connection
keeps it open, 1st response's 2 MiB body left unread SocketWriteError: closed stream 200, new connection
keeps it open reused reused
reads the 2nd request, then closes ResponseHeaderError ResponseHeaderError

POSTs, 30 runs per case, counted by the server: after an idle FIN, an idle RST, or a 1 MiB body left unread, main fails all 90 and the server receives none of them. This branch delivers each one exactly once on a new connection, 90/90, on MRI 3.4.8 and JRuby 10.1.2.0. The check runs before any byte is written, so it's safe for every method.

The unread-body case doesn't need an idle server. perform marks the client clean once headers are read, so the next request flushes the old body inside Connection#send_request, the flush closes the connection because the body is over 1 MiB, and the write fails. Flushing before the check fixes it.

Other clients make the same check:

  • urllib3 drops a pooled connection when wait_for_read(sock, timeout=0.0) is true
  • Net::HTTP reconnects in begin_transport on @socket.io.to_io.wait_readable(0) && @socket.eof?
  • Go's transport closes an idle connection when its read loop sees an EOF or an unsolicited response, and special-cases a 408
  • libcurl polls an idle connection before reusing it and treats input pending on it as dead ("The input might be a TLS Notify Close")

stale? is one non-blocking poll: 1.0 µs on MRI 3.4.8, about 1% of a 95 µs loopback request, and 1.5–2.5 µs on JRuby, under 0.5%.

Limits:

  • If the server closes after the check, before the request goes out or while it's in flight (the last row above), the socket can't tell the client in time, and the server may have received the request. Only resending helps there, and only for requests that are safe to repeat. Resend idempotent requests when a reused connection was closed #861 does that for idempotent ones.
  • It polls the file descriptor, so it can't see bytes that already reached a userspace buffer (OpenSSL's or the parser's) together with the previous response. A TLS server sending a session ticket or key update while the connection sits idle would cost a reconnect; I didn't see that in the 58 reuses above.
  • Connection#stale? and Connection#flush_pending_response are public because Client calls them. Happy to mark them @api private if you'd rather not commit to them.

test_connection_reuse_enabled_socket_issue_transparently_reopens asserted that the first request after the server closed the socket raised. It now expects the reopen its name describes. test_connection_reuse_enabled_raises_when_server_closes_after_receiving_request pins down the first limit.

  • MRI 3.4.8: 2356 runs, 0 failures on seeds 4242 and 777, 100% line and branch coverage. rubocop and yardstick pass, steep shows the same 4 warnings as main
  • JRuby 10.1.2.0: 2355 runs, 0 failures
  • .mutant.yml ignores HTTP::Connection* and HTTP::Client*, so CI doesn't mutate this code. Run locally with them included, the new lines have no survivors; the rest are equivalent or in lines this doesn't change.

5-x-stable reuses connections the same way (verify_connection!, client.rb:131), and it's the newest line apps on Ruby < 3.2 can use. The backport is #862.

I ran into this through HTTPS proxies, which closed idle CONNECT tunnels after 10 to 120 seconds.

Local repro script
# Real-socket repro for http.rb persistent-connection reuse.
# Usage: cd <http.rb checkout> && bundle exec ruby -Ilib upstream_repro.rb [scenario ...]
require "http"
require "socket"
require "openssl"

REQ_END = "\r\n\r\n"
OK = "HTTP/1.1 200 OK\r\nContent-Length: 2\r\n\r\nok"

def read_request(io)
  buf = +""
  buf << io.readpartial(4096) until buf.include?(REQ_END)
  buf
rescue EOFError, IOError, SystemCallError, OpenSSL::SSL::SSLError
  nil
end

def tls_context
  key = OpenSSL::PKey::RSA.new(2048)
  cert = OpenSSL::X509::Certificate.new
  cert.version = 2
  cert.serial = 1
  cert.subject = cert.issuer = OpenSSL::X509::Name.parse("/CN=127.0.0.1")
  cert.public_key = key.public_key
  cert.not_before = Time.now - 60
  cert.not_after = Time.now + 3600
  cert.sign(key, OpenSSL::Digest.new("SHA256"))
  ctx = OpenSSL::SSL::SSLContext.new
  ctx.cert = cert
  ctx.key = key
  ctx
end

HANDLERS = {
  "alive" => lambda do |io, stats|
    while read_request(io)
      stats[:requests] += 1
      io.write(OK)
    end
  end,
  "idle_fin" => lambda do |io, stats|
    read_request(io) or return
    stats[:requests] += 1
    io.write(OK)
    sleep 0.1
  end,
  "idle_408" => lambda do |io, stats|
    read_request(io) or return
    stats[:requests] += 1
    io.write(OK)
    sleep 0.1
    io.write("HTTP/1.1 408 Request Timeout\r\nConnection: close\r\nContent-Length: 0\r\n\r\n")
  end,
  "in_flight_drop" => lambda do |io, stats|
    while read_request(io)
      stats[:requests] += 1
      return if stats[:requests] == 2

      io.write(OK)
    end
  end,
  "big_unread" => lambda do |io, stats|
    first = true
    while read_request(io)
      stats[:requests] += 1
      if first
        first = false
        body = "x" * (2 * 1024 * 1024)
        io.write("HTTP/1.1 200 OK\r\nContent-Length: #{body.bytesize}\r\n\r\n#{body}")
      else
        io.write(OK)
      end
    end
  end
}.freeze

def start_server(scenario, tls:)
  tcp = TCPServer.new("127.0.0.1", 0)
  server = tls ? OpenSSL::SSL::SSLServer.new(tcp, tls_context) : tcp
  server.start_immediately = true if tls
  stats = { accepts: 0, requests: 0 }
  thread = Thread.new do
    loop do
      io = server.accept
      stats[:accepts] += 1
      Thread.new(io) do |conn|
        HANDLERS.fetch(scenario).call(conn, stats)
      ensure
        conn.close rescue nil
      end
    rescue OpenSSL::SSL::SSLError, IOError, SystemCallError
      break if tcp.closed?
    end
  end
  [tcp, tcp.addr[1], stats, thread]
end

def run(scenario, tls:)
  tcp, port, stats, thread = start_server(scenario, tls: tls)
  base = "#{tls ? 'https' : 'http'}://127.0.0.1:#{port}"
  ssl = OpenSSL::SSL::SSLContext.new
  ssl.verify_mode = OpenSSL::SSL::VERIFY_NONE
  client = HTTP.timeout(5).persistent(base, timeout: 60)
  r1 = client.get("#{base}/1", ssl_context: ssl)
  r1.to_s unless scenario == "big_unread"
  sleep 0.3
  t0 = Process.clock_gettime(Process::CLOCK_MONOTONIC)
  outcome = begin
    r2 = client.get("#{base}/2", ssl_context: ssl)
    "#{r2.status.code} #{r2.to_s.inspect}"
  rescue StandardError => e
    "#{e.class}: #{e.message[0, 60]}"
  end
  ms = ((Process.clock_gettime(Process::CLOCK_MONOTONIC) - t0) * 1000).round(1)
  client.close
  sleep 0.05
  format("%-5s %-15s %-62s %7.1fms accepts=%d server_requests=%d",
         tls ? "tls" : "plain", scenario, outcome, ms, stats[:accepts], stats[:requests])
ensure
  tcp&.close
  thread&.kill
end

scenarios = ARGV.empty? ? HANDLERS.keys : ARGV
puts "http #{HTTP::VERSION} ruby #{RUBY_VERSION} #{OpenSSL::OPENSSL_LIBRARY_VERSION}"
[false, true].each do |tls|
  scenarios.each { |s| puts run(s, tls: tls) }
end
Real-host script

HOSTS=www.debian.org,lwn.net KAT=5 GAP=4.99 ROUNDS=3 bundle exec ruby -Ilib idle_reuse_real.rb

$stdout.sync = true
require "http"

# Persistent client against real hosts: request, idle GAP seconds, request again.
# Usage: bundle exec ruby -I<http.rb lib> idle_reuse_real.rb
HOSTS = (ENV["HOSTS"] || "www.debian.org,www.gnu.org,www.ruby-lang.org").split(",")
GAP = Float(ENV.fetch("GAP", "6"))
ROUNDS = Integer(ENV.fetch("ROUNDS", "3"))
KAT = Float(ENV.fetch("KAT", "300"))
UA = "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36"

# HTTP.persistent returns an HTTP::Session (6.x) that keeps one Client per origin.
def local_port(http)
  clients = http.instance_variable_get(:@clients)
  client = clients ? clients.values.first : http
  conn = client.instance_variable_get(:@connection) or return
  io = conn.instance_variable_get(:@socket).socket.to_io
  io.closed? ? nil : io.local_address.ip_port
end

puts "http #{HTTP::VERSION} ruby #{RUBY_VERSION} gap=#{GAP}s rounds=#{ROUNDS} keep_alive_timeout=#{KAT}s"
HOSTS.each do |host|
  ROUNDS.times do |round|
    client = HTTP.timeout(10).headers("User-Agent" => UA).persistent("https://#{host}", timeout: KAT)
    stage = "r1"
    outcome = begin
      r1 = client.get("https://#{host}/")
      r1.to_s
      port1 = local_port(client)
      sleep GAP
      stage = "r2"
      r2 = client.get("https://#{host}/")
      r2.to_s
      port2 = local_port(client)
      "r1=#{r1.status.code} r2=#{r2.status.code} #{port1.nil? ? 'no_keep_alive' : port1 == port2 ? 'reused' : 'new_connection'}"
    rescue StandardError => e
      "#{stage} #{e.class}: #{e.message[0, 40]}"
    ensure
      client.close
    end
    puts format("%-20s round=%d %s", host, round + 1, outcome)
  end
end
Kernel trace: Linux, bpftrace kprobes and uprobes on libruby and libssl; an HTTPS proxy closing idle tunnels after 10 s 1-trace-master

A server or proxy can close a keep-alive connection while it sits idle
in a persistent client. The client expires idle connections after
keep_alive_timeout, 5 seconds by default, but a server's idle timeout
can be shorter, and the server starts counting when it sends the
response, before the client has read it. The client only checked
whether its own end of the socket was closed, so it wrote the next
request into the closed connection. That request failed with
HTTP::ResponseHeaderError ("couldn't read response headers"), or with
OpenSSL::SSL::SSLError ("unexpected eof while reading") when the server
skipped close_notify. A server that sends a response before closing,
such as a 408 Request Timeout, was worse: the next request read that
408 as its own response.

Before reusing a connection, read off the previous response, then check
the socket without blocking. Once the previous response has been read,
an idle connection has nothing left to read, so any readable data means
it can't carry another request: an EOF, a reset, a TLS close_notify or
an unsolicited response. Close it then and open a new one. urllib3
makes the same check (wait_for_read(sock, timeout=0.0)), Net::HTTP
reconnects on wait_readable(0) && eof?, and Go's transport drops idle
connections that see an EOF or an unsolicited response.

www.debian.org, lwn.net, www.postgresql.org and www.sqlite.org closed
idle connections 4.8 to 4.98 seconds after the client read a response,
and www.haproxy.org as early as 2 seconds. With the default
keep_alive_timeout, a second request 4.99 seconds later (3 seconds for
www.haproxy.org) failed 16 of 16 times and now succeeds 16 of 16 times
on a new connection. 14 servers that keep idle connections open were
still reused 28 of 28 times after 4.5 seconds idle.

The check happens before any byte of the next request is written, so it
applies to every method, POST included, and can't send a request twice.
It runs after the client is marked dirty, so an exception that
interrupts reading off the previous response still makes the next
request reconnect. It can't help when the server closes the connection
after the request was written; only idempotent requests are safe to
retry then.

Reading off a body larger than 1 MiB closes the connection instead, but
the client then wrote the next request to the closed connection and
raised HTTP::SocketWriteError ("closed stream"). The same check now
reconnects.

test_connection_reuse_enabled_socket_issue_transparently_reopens
asserted that the first request after the server closed the socket
raised. It now expects the client to reopen the connection, as its name
says, and a POST variant covers the same case.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant