Skip to content

[vllm-omni] - Does MiniCPM-o 4.5 officially support TTS generation beyond 4,096 Talker positions? #1127

Description

@akshatvishu

起始日期 | Start Date

july/31/2026

Issue

Hello,

We're trying matching the MiniCPM-o 4.5 TTS implementation in vLLM-Omni and found one remaining gap in long-form speech generation.

The official Hugging Face streaming_generate() implementation splits text into 10-token conditions. It preserves the Talker KV cache and continuously increases text_start_pos across all conditions and generated codec tokens.

The released checkpoint uses:

{
  "attention_type": "full_attention",
  "max_position_embeddings": 4096
}

However, the official Hugging Face loop does not appear to stop, reset, or truncate the Talker when the combined sequence exceeds 4,096 positions. This allows speech generation to continue beyond one minute, subject to available GPU memory.

vLLM requires an explicit max_model_len and stops the Talker when the combined text conditions and codec tokens reach that limit. Raising the limit to 8,192 fixes some 500-word cases, but only moves the failure point. Longer responses, such as 1,000 or 1,500 words, can still be truncated.

Could you clarify the intended official behavior?

Should long-form TTS continue using full_attention beyond 4,096 positions with an increased context limit, or should the Talker switch to sliding_window, sliding_recompute, or reindex?

We want to understand which behavior was used for the reported long-speech results and which approach vLLM-Omni should implement for correct long-form TTS. Also, it will be great if you can also share the benchmark script for long-tts regarding this

相关Issues | Reference Issues

vllm-project/vllm-omni#5259

Activity

  1. tc-mb commented on Aug 3, 2026

    @tc-mb
    Collaborator

    Thanks for raising this question.

    MiniCPM-o 4.5 is designed for full-duplex spoken interaction with a human, where speech is produced incrementally and the user may interact or interrupt. It is not intended for uninterrupted long-form narration without interaction.

    The audio codec rate is approximately 25 tokens per second. Since the 4,096 Talker positions include both text conditions and generated codec tokens, this corresponds to roughly 2–2.5 minutes of uninterrupted speech in practice, depending on the text length and condition overhead. This is sufficient for the intended interactive scenario.

    For applications that require longer continuous speech, a sliding-window strategy can be used as a graceful fallback: discard the oldest Talker context while retaining the recent context for subsequent generation. This may reduce long-range consistency or speech quality, but it allows generation to continue with bounded memory and context length.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    questionFurther information is requested

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions