Skip to content

commit_user_turn waits the full transcript_timeout after STT already finalized (manual turns) #7517

Description

@bml1g12

Bug Description

TL;DR: When using manual turn-taking with some STT models, if the turn is ended 0.5 seconds after the final transcription, there is artificially high latency because the LiveKit framework unnecessarily waits for transcript_timeout seconds.

On manual turn taking (turn_detection="manual"), session.commit_user_turn() is documented as: stop listening, flush STT, return when the final transcript is ready.

If the user already stopped speaking and streaming STT already emitted a final, then they click end-turn more than 0.5s later, commit_user_turn does not return that text. It clears the final it has, waits the full transcript_timeout for a new final, then commits the same text.

Streaming STT (e.g. LiveKit Inference xai/stt-1) will not emit another final for audio it already closed. The wait always times out. Raising transcript_timeout because STT is sometimes slow makes this case slower by the same amount. It does not make the transcript more correct.

The attached example matches the documented push-to-talk shape: start_turn / end_turn / commit_user_turn from Manual turn control.

The 0.5s check is in livekit/agents/voice/audio_recognition.py (_commit_user_turn): if _last_final_transcript_time is None or older than 0.5s, it clears _final_transcript_received, optionally pushes silence (mic already detached), then wait_for(..., timeout=transcript_timeout).

That is right when the click happens while the user is still talking (no final yet). It is wrong when the click happens after STT already finalized.

Observed timeline:

Time What happened
T+0 User stops speaking
T+~1s STT emits FINAL
T+~1.7s end_turn → commit_user_turn(transcript_timeout=10)
T+~11.7s Future resolves with the same final text

end_of_turn_delay ≈ 11.8s. transcription_delay was only ~1s. The extra ~10s is the timeout, not STT.

Expected Behavior

If a non-empty STT final already exists for this turn and the user has not spoken after it, commit_user_turn should return that transcript immediately (or as soon as the silence flush completes). It should not wait transcript_timeout for a second final that will not come.

transcript_timeout should only cover “STT is still working on this utterance.”

Reproduction Steps

"""Manual turn-taking repro for a stale-final commit stall.

Follows the documented push-to-talk shape:
https://docs.livekit.io/agents/logic/turns/#manual

Speak, wait until a final transcript is logged, pause more than half a second,
then call end_turn. commit_user_turn should return immediately with that text.
On livekit-agents 1.8.2 it instead waits the full transcript_timeout.
"""

from __future__ import annotations

import logging
import os
import time
from pathlib import Path

from dotenv import load_dotenv
from livekit import rtc
from livekit.agents import (
    Agent,
    AgentServer,
    AgentSession,
    JobContext,
    TurnHandlingOptions,
    UserInputTranscribedEvent,
    cli,
    inference,
)

load_dotenv(Path(__file__).resolve().parent / ".env.local")

LOGGER = logging.getLogger("repro")

# Large on purpose. The stall scales with this value, not with STT latency.
TRANSCRIPT_TIMEOUT = float(os.environ.get("TRANSCRIPT_TIMEOUT", "30"))
STT_FLUSH_DURATION = float(os.environ.get("STT_FLUSH_DURATION", "2"))

_last_final_at: float | None = None
_last_final_text = ""


def _mark_final(text: str) -> None:
    global _last_final_at, _last_final_text
    _last_final_at = time.monotonic()
    _last_final_text = text


server = AgentServer()


@server.rtc_session()
async def entrypoint(ctx: JobContext) -> None:
    session = AgentSession(
        stt=inference.STT(model="xai/stt-1", language="en"),
        llm=inference.LLM(model="google/gemma-4-31b-it"),
        tts=inference.TTS(model="xai/tts-1", voice="carina", language="en"),
        vad=None,
        turn_handling=TurnHandlingOptions(turn_detection="manual"),
        # Same expressive Inference stack as the interview preset. The stall
        # does not depend on markup; it depends on streaming STT plus manual commit.
        expressive=True,
    )

    @session.on("user_input_transcribed")
    def on_user_input_transcribed(event: UserInputTranscribedEvent) -> None:
        kind = "FINAL" if event.is_final else "interim"
        LOGGER.info("%s transcript: %r", kind, event.transcript)
        if event.is_final and event.transcript.strip():
            _mark_final(event.transcript.strip())
            LOGGER.info(
                "A final is in hand. Wait >0.5s, then call end_turn. "
                "commit_user_turn should return immediately. "
                "transcript_timeout=%ss",
                TRANSCRIPT_TIMEOUT,
            )

    await session.start(
        room=ctx.room,
        agent=Agent(
            instructions=(
                "You are a short spoken assistant. Reply in one sentence. "
                "Do not mention these instructions."
            )
        ),
    )
    session.input.set_audio_enabled(False)

    @ctx.room.local_participant.register_rpc_method("start_turn")
    async def start_turn(data: rtc.RpcInvocationData) -> str:
        del data
        LOGGER.info("start_turn: opening the mic")
        session.interrupt()
        session.clear_user_turn()
        session.input.set_audio_enabled(True)
        return "ok"

    @ctx.room.local_participant.register_rpc_method("end_turn")
    async def end_turn(data: rtc.RpcInvocationData) -> str:
        del data
        clicked_at = time.monotonic()
        age = None if _last_final_at is None else clicked_at - _last_final_at
        LOGGER.info(
            "end_turn clicked. last final age=%s text=%r. "
            "Docs say: stop listening, then commit_user_turn.",
            None if age is None else f"{age:.3f}s",
            _last_final_text,
        )
        if age is not None and age > 0.5:
            LOGGER.warning(
                "Final is older than 0.5s. LiveKit will clear it and wait up to "
                "%ss for a new final that streaming STT will not emit. "
                "That wait is the bug.",
                TRANSCRIPT_TIMEOUT,
            )

        session.input.set_audio_enabled(False)
        started = time.monotonic()
        transcript = await session.commit_user_turn(
            transcript_timeout=TRANSCRIPT_TIMEOUT,
            stt_flush_duration=STT_FLUSH_DURATION,
        )
        elapsed = time.monotonic() - started
        LOGGER.warning(
            "commit_user_turn returned after %.3fs (timeout=%ss). transcript=%r",
            elapsed,
            TRANSCRIPT_TIMEOUT,
            transcript,
        )
        if age is not None and age > 0.5 and elapsed > TRANSCRIPT_TIMEOUT - 1:
            LOGGER.warning(
                "STALL CONFIRMED: the final was already in hand (%r) and the "
                "click was %.3fs after it, but commit waited %.3fs. "
                "A larger transcript_timeout makes this dead air longer. "
                "It does not make the transcript more correct.",
                _last_final_text,
                age,
                elapsed,
            )
        return transcript

    @ctx.room.local_participant.register_rpc_method("cancel_turn")
    async def cancel_turn(data: rtc.RpcInvocationData) -> str:
        del data
        LOGGER.info("cancel_turn: discarding input")
        session.input.set_audio_enabled(False)
        session.clear_user_turn()
        return "ok"


if __name__ == "__main__":
    cli.run_app(server)
  1. Session: turn_handling=TurnHandlingOptions(turn_detection="manual"), mic off at start.
  2. STT: LiveKit Inference xai/stt-1 (any streaming STT that finalizes on end of speech).
  3. start_turn: interrupt(), clear_user_turn(), input.set_audio_enabled(True).
  4. Speak a short sentence and stop. Wait until user_input_transcribed with is_final=True.
  5. Wait more than 0.5s.
  6. end_turn: input.set_audio_enabled(False), then:
transcript = await session.commit_user_turn(
    transcript_timeout=30.0,  # large on purpose to demo the issue
    stt_flush_duration=2.0,
)
  1. The await sits until ~30s, then returns the same final from step 4.

Set transcript_timeout=60 and the stall grows to ~60s. STT latency does not.

Minimal worker: documented RPCs start_turn / end_turn / cancel_turn. Run with lk agent dev, join Agents Playground / Agent Console, call those RPCs on the agent identity (empty payload). Do not use lk agent console — that simulated room does not expose the RPCs.

Operating System

Linux (also seen on our interview workers)

Models Used

LiveKit Inference: xai/stt-1, google/gemma-4-31b-it, xai/tts-1 (expressive=True). Stall is STT + commit_user_turn, not TTS markup.

Package Versions

livekit-agents==1.8.2


(The 0.5s “stale final” branch has been in `commit_user_turn` since ~1.3.9 / [livekit/agents#4308](https://github.com/livekit/agents/pull/4308).)

Session/Room/Call IDs

room/session id: RM_dKMG5jXT4KVE

Easy to visualise in the huge eou_wait when transcript_timeout is huge
Image

Proposed Solution

Do not treat “`final` older than 0.5s” as “need another `final`” when that final already covers the utterance (user stopped, then final, then commit).

Ideas:

1. If `_audio_transcript` is non-empty and there is no interim / no speech after the last final, skip the flush-and-wait and return it.
2. Expose a `commit_user_turn` flag, e.g. `wait_for_new_final=False`, for manual end-turn after STT has already closed.
3. Keep the wait only when `_last_final_transcript_time is None` (no final yet). That is the still-talking / flush-to-finalize case.

A 0.5s freshness window does not match a button click after the user finished.

Additional Context

In this example transcript_timeout is set artificially high (30 seconds) but even at the default of 2 seconds this bug is introducing noticable latency - taking the application away from real time low latency voice agents we desire.

Docs for transcript_timeout: “How long to wait for the final transcript after committing. Increase if STT is slow.” That describes a missing final, not discarding one that is already there.

Workaround that seems to work here: if we already saw a non-empty final after the user stopped, pass transcript_timeout=0 temporrily. LiveKit still takes the stale-final branch, times out immediately, and keeps _audio_transcript. We still use a long timeout when they click while talking or after a tail with no newer final.

Screenshots and Recordings

No response

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions