Skip to content

/v3/audio/speech ignores response_format and always returns 32-bit float WAV #4613

Description

@AlphaJack

Describe the bug
The OpenAI-compatible text-to-speech endpoint /v3/audio/speech ignores response_format. For wav, pcm, mp3 and flac it returns the same body: a WAV file with 32-bit IEEE float samples (format tag 3). The response is also sent with Content-Type: application/json; charset=utf-8.

The OpenAI API specifies 16-bit PCM for wav and headerless 16-bit little-endian PCM for pcm, so clients written against it cannot play the output. For example, wyoming_openai (the Wyoming bridge used by Home Assistant) cannot parse the float WAV header and plays static: roryeckel/wyoming_openai#72

To Reproduce

  1. Models repository: OpenVINO/Kokoro-82M-int8-ov pulled from Hugging Face into /models/OpenVINO/Kokoro-82M-int8-ov, graph generated with ovms --configure --model_path /models/OpenVINO/Kokoro-82M-int8-ov --target_device CPU:
    input_stream: "HTTP_REQUEST_PAYLOAD:input"
    output_stream: "HTTP_RESPONSE_PAYLOAD:output"
    node {
        name: "T2sExecutor"
        calculator: "T2sCalculator"
        input_side_packet: "TTS_NODE_RESOURCES:t2s_servable"
        input_stream: "HTTP_REQUEST_PAYLOAD:input"
        output_stream: "HTTP_RESPONSE_PAYLOAD:output"
        node_options: {
            [type.googleapis.com / mediapipe.T2sCalculatorOptions]: {
                models_path: "/models/OpenVINO/Kokoro-82M-int8-ov"
                target_device: "CPU"
                plugin_config: '{"NUM_STREAMS":"1"}'
                }
        }
    }
    
  2. OVMS launch command (container openvino/model_server:latest-gpu):
    --rest_port 8080 --config_path /config/config.json --cache_dir /cache --log_level INFO
    
  3. Client command:
    for rf in wav pcm mp3 flac; do
      curl -s http://localhost:8080/v3/audio/speech -H 'Content-Type: application/json' \
        -d "{\"model\":\"OpenVINO/Kokoro-82M-int8-ov\",\"input\":\"Hello there.\",\"voice\":\"am_adam\",\"response_format\":\"$rf\"}" \
        -o "out.$rf"
      file -b "out.$rf"
    done
  4. Every file is identical (146444 bytes):
    RIFF (little-endian) data, WAVE audio, IEEE Float, mono 24000 Hz
    
    Response headers:
    HTTP/1.1 200 OK
    content-length: 146444
    content-type: application/json; charset=utf-8
    

Expected behavior

  • wav (the default) returns a 16-bit PCM WAV, as OpenAI does.
  • pcm returns headerless 16-bit little-endian PCM.
  • Formats OVMS cannot produce (e.g. mp3, flac) are rejected with a 400 error instead of silently returning a WAV.
  • The Content-Type matches the audio format (e.g. audio/wav).

Logs
Nothing is logged for these requests at --log_level INFO; I have not captured DEBUG logs.

Configuration

  1. OVMS version: OpenVINO Model Server 2026.4.0.869b2186a, OpenVINO backend 2026.4.0-22959-99c81491cc3-releases/2026/4, OpenVINO GenAI backend 2026.4.0.0-3407-7ea2546852a
  2. config.json:
    {
      "model_config_list": [
        { "config": { "name": "OpenVINO/Kokoro-82M-int8-ov", "base_path": "/models/Kokoro-82M-int8-ov-cpu" } }
      ]
    }
  3. Intel Core Ultra 5 125H (Meteor Lake), target device CPU
  4. Model repository:
    /models/Kokoro-82M-int8-ov-cpu/graph.pbtxt
    /models/OpenVINO/Kokoro-82M-int8-ov/{config.json, openvino_config.json, openvino_model.xml, openvino_model.bin, voices/, ...}
    
  5. Model: https://huggingface.co/OpenVINO/Kokoro-82M-int8-ov

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions