Skip to content

Opt-in serialization of response.create to prevent conversation_already_has_active_response (follow-up to the fast-fail) #7514

Description

@areebkhan-tech

Feature Type

I cannot use LiveKit without it

Feature Description

In production with the OpenAI Realtime plugin we see conversation_already_has_active_response errors frequently. They fire whenever a client-initiated generate_reply overlaps a response that is already active on the server — most commonly a server-VAD auto-response (turn_detection with create_response=True) that starts as the app is also asking for a reply, or a reply issued right after an interrupt() before the previous response has cleared.

The recently-added fast-fail (a rejected response.create fails its future immediately with the typed code instead of orphaning it until the 10s timeout — _handle_error) handles the collision, but does not prevent it: the error still surfaces to the session (fires user error handlers, pollutes logs/metrics) and the reply must be retried client-side, so under load the noise is constant.

Root cause: RealtimeSession.generate_reply() always sends response.create unconditionally, even when the session already knows a response is active (_current_generation is a live _ResponseGeneration).

Workarounds / Alternatives

An opt-in flag, serialize_response_create: bool = False, on RealtimeModel. When enabled, generate_reply() waits for a locally-known active response to finish (bounded by a timeout) before sending response.create, turning the common collision from handled into prevented. It only waits — it does not cancel the active response (cancellation stays an explicit interrupt() decision). Default False keeps today's behaviour exactly, and the fast-fail remains the fallback.

Known limitation (residual): this cannot cover the in-transit race — a server response that was just created but whose response.created the client hasn't processed yet is invisible (_current_generation is still None), so the send still goes out and collides. That case is still caught by the fast-fail and can be retried. Serialization only removes the locally-known collisions.

Reproduction

1.OpenAI Realtime with server VAD, create_response=True.
2.Call session.generate_reply(...) while the server is auto-creating a response (or issue a reply immediately after an interrupt).
3.Observe frequent conversation_already_has_active_response error events.

Additional Context

No response

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions