Skip to content

RoomIO: no jitter prebuffer for duplex models (GPT-Live) – audio breaks mid-word on calls #7513

Description

@Uhimanshu9

Bug Description

With GPT-Live (gpt-live-1, via livekit-plugins-openai's GPTLiveModel), the agent's voice breaks up mid-word during room/phone calls. Transcripts are complete and nothing is interrupted; only the played audio has gaps. The same agent on Gemini Live doesn't do it.

Cause: GPT-Live is a duplex model that streams audio at real-time pace. RoomIO's _ParticipantAudioOutput creates its AudioSource with a fixed queue_size_ms=200 and starts playback on the first frame, so the queue never builds any lead. Any arrival delay on the OpenAI connection goes straight to the caller as silence.

PR #7238 fixed exactly this for console mode (ConsoleAudioOutput prebuffers 300 ms, and the comment in cli/_legacy.py says it's "the largest arrival gap observed with GPT-Live"), but RoomIO got no equivalent. #7259, #7260 and #7272 fixed a related startup gap in RoomIO but added no jitter buffer.

Expected Behavior

Duplex-model audio should play smoothly through RoomIO, the same as in console mode after #7238. RoomIO should keep enough reserve that ordinary network delays from the model provider aren't heard.

Reproduction Steps

1. Create an AgentSession with llm=GPTLiveModel(model="gpt-live-1", ...) and start it with RoomIO (default RoomOptions).
2. Join the room from a browser client or a SIP call and have a normal multi-turn conversation.
3. Listen to the agent's longer replies: short breaks or clicks mid-word, especially partway through a sentence.
4. For comparison, run the same agent with python agent.py console. It's smooth there because of #7238's prebuffer.

Operating System

Linux (Ubuntu, GCP VM, kernel 7.0)

Models Used

OpenAI GPT-Live (gpt-live-1) via GPTLiveModel, with a backend Responses model for tool delegation. No separate STT/TTS.

Package Versions

livekit==1.1.18
livekit-agents==1.8.2
livekit-api==1.2.1
livekit-plugins-openai==1.8.2
python==3.12

Session/Room/Call IDs

Self-hosted LiveKit server, not LiveKit Cloud, so there are no Cloud IDs.

Proposed Solution

Bring #7238's prebuffer to RoomIO. For example:
 - a prebuffer_ms option on AudioOutputOptions, or
 - automatic prebuffering when the session's model is a DuplexModel.

Our current workaround, which you could use as a reference: after RoomIO.start() we wrap room_io._audio_output._audio_source with a small class that:

 -holds the first 300 ms of each reply before passing it on
 -releases anything held on wait_for_playout(), so short replies still play
 -drops held audio on clear_queue(), so interruptions stay instant
 -builds the 300 ms back up if the queue runs empty mid-reply.

It works, but it relies on private attributes, so we'd like a supported option.

Additional Context

In a ~4-minute test call with the workaround, the queue ran dry only once (a gap longer than 300 ms), and the voice was noticeably smoother. The cost is about 300 ms of extra delay at the start of each reply, the same tradeoff console mode makes.

Screenshots and Recordings

No response

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions