Skip to content

build: allow embedding voice-detect.cpp in another ggml project - #1

Merged
mudler merged 1 commit into
masterfrom
feat/embedding
Sep 30, 2026
Merged

mudler merged 1 commit into
masterfrom
feat/embedding

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown

parakeet.cpp is adding speaker identification: it names the slots its diarizer finds by matching each slot's voice against enrolled speakers. It folds voice-detect.cpp in as a static dependency, the same way it already folds ced.cpp, so voice-detect has to build inside a host project that owns its own ggml target and its own dr_wav implementation.

What this changes:

  • VOICEDETECT_TOP_LEVEL: the GGML_* forwarding, the CPU variants setup, the ggml patch step, the CLI and the tests only run when voice-detect is the top-level project. An embedding host keeps its own ggml configuration.
  • if(NOT TARGET ggml) guards the ggml add_subdirectory, so a host that already has a ggml target is reused.
  • VOICEDETECT_EXTERNAL_DR_WAV (off by default): skips voice-detect's own DR_WAV_IMPLEMENTATION, so the host's dr_wav is the only definition and the link has no duplicate drwav_* symbols.
  • CMAKE_SOURCE_DIR becomes CMAKE_CURRENT_SOURCE_DIR / PROJECT_SOURCE_DIR where a host would resolve it to the wrong root.
  • New C function voicedetect_capi_embedding_dim(ctx): the embedding size (192 to 512), 0 for a model with no speaker embedding (the age, gender and emotion heads), -1 for a NULL ctx. It lets a caller size a speaker registry before it runs the encoder. Additive, so no ABI version bump.

This mirrors what ced.cpp did for the same purpose (b7dc30d and e1a3cfa in that repo). Nothing changes for a normal top-level build.

How it was checked:

  • Top-level build with tests: test_smoke passes, test_capi_dim passes with the WeSpeaker GGUF (256) and skips without a model.
  • A scratch host project with its own ggml v0.13.0 and its own dr_wav, VOICEDETECT_EXTERNAL_DR_WAV=ON: it links, prints abi 1, exits 0 and has a single drwav_init definition. It compiles against ggml v0.13.0 even though this repo vendors v0.15.2.
  • The only ggml patch this repo carries (CUDA / cuDNN conv) is not applied in an embedded build. That does not matter on CPU. A host that wants cuDNN conv has to apply it itself.

The parakeet.cpp side is a separate PR that pins this commit as a submodule.

🤖 Generated with Claude Code

parakeet.cpp folds a speaker encoder into its scene stream, as it does
for ced.cpp. That needs voice-detect to build inside a host that already
owns a ggml target and a dr_wav implementation.

Guard the GGML_* forwarding, patch step, CLI and tests behind a top-level
check, reuse an existing ggml target, and add
VOICEDETECT_EXTERNAL_DR_WAV so the host's dr_wav is the only definition.
Add voicedetect_capi_embedding_dim so a caller can size a speaker
registry before running the encoder.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@mudler
mudler merged commit b74a896 into master Sep 30, 2026
6 checks passed
mudler added a commit to mudler/parakeet.cpp that referenced this pull request Sep 30, 2026
localai-org/voice-detect.cpp#1 is merged. Move the submodule from the
feature branch commit b44c586 to the merge commit b74a896 on master. The
tree is identical, so nothing else changes.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
mudler added a commit to mudler/parakeet.cpp that referenced this pull request Sep 30, 2026
…v9) (#78)

* feat(speaker): add SpeakerRegistry for enrolled voices

Enrolled speakers are centroids of L2-normalized embeddings. identify()
takes a threshold and a runner-up margin so a voice that is not enrolled
comes out unknown instead of as the nearest name. The registry saves and
loads as a small binary blob and rejects corrupt or mismatched input.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* fix(speaker): reject a registry blob with dim 0 and speakers

deserialize() now validates that dim >= 1 when n > 0 (number of speakers > 0).
A crafted blob with dim 0 and speakers was accepted, creating entries with
empty sum vectors. A later enroll() would then cause an out-of-bounds heap
write when trying to accumulate the new embedding. Empty registries (dim 0,
n 0) still serialize and deserialize correctly.

The fix adds validation to reject implausible headers and includes tests for:
- A hand-built corrupt blob (dim 0, n 1) must throw
- An empty registry round-trips and accepts subsequent enrollment
- Failed enrollments on non-empty registries leave state unchanged
- names() correctly lists remaining speakers after removal

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(speaker): add SpeakerIdentifier for diarization slots

Collects each slot's clean audio (overlap with other speakers is
skipped), embeds it once it has enough and again as it grows, and names
the slot from the registry. A different name replaces the current one
only after winning twice in a row, and an unknown match never removes a
name. identify_offline reuses the same logic for finished recordings.
Driven by an embedding callback, so it is tested with a fake encoder.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* fix(speaker): test overlap against closed segments and document the open contract

Adds tests for overlap with segments closed in the same call and in an
earlier call, which exercise the history scan. Documents that the caller
must list every started, unclosed segment in `open`, and what
identify_offline does to max_voice_sec and ring_sec.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(speaker): carry speaker names through words, utterances and scene JSON

Words and utterances get a name and a match score, SceneUpdate gets the
current name of every slot, and the scene JSON and the text renderer
print them. With no speaker model the JSON is unchanged, byte for byte.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(speaker): fold voice-detect.cpp in behind SpeakerEncoder

voice-detect.cpp is a submodule built as a static target that shares our
ggml and dr_wav, like ced.cpp. pk::SpeakerEncoder is the only code that
touches it, through voicedetect_capi.h, and PARAKEET_WITH_VOICEDETECT=OFF
builds without it. The test checks that two clips of one voice score
above two clips of different voices on two_speakers.wav, and that the
folded encoder matches a standalone voice-detect build.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(speaker): identify speakers inside the scene stream

SceneStream runs an optional SpeakerIdentifier after diarization, so
words, utterances and the per-slot names in each update carry the
enrolled name. It needs diarization and a registry and says so when
they are missing. Names apply to words committed after the slot is
identified; earlier words keep the label they were emitted with.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* fix(speaker): check speaker preconditions first and test order-independent naming

The speaker checks now run before the generic "needs at least one
model" check, so a speaker-only config gets the specific message. The
test asserts each message, enrolls the voices in reverse arrival order
so slot i cannot map to registry entry i, and checks that a voice
missing from the registry stays unnamed.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(capi): speaker identification, ABI v9

A voice-detect GGUF loads into a fourth context kind. A registry handle
enrolls, saves and loads voices; the scene stream takes a speaker
context and registry through a new begin function; speaker-attributed
ASR has a named variant. Existing signatures are unchanged, and the
scene options grow only at the end, read according to the caller's size.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* fix(capi): read scene options only within the caller's size and avoid a stream leak

The speaker option fields were passed as arguments to a helper that
checked the caller's size, so they were read before the check and a
caller built against the v8 header read past its struct. Each field is
now read by offset only after the size covers it. The scene wrapper is
also allocated after the SceneStream is built, so a throwing constructor
no longer leaks it.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(cli): enroll speakers and name them in scene output

parakeet-cli enroll builds a registry from labeled clips, and scene
--speakers --registry names diarized speakers in the transcript and the
JSON. docs/speaker.md covers the models, the C-API, the timing rule and
what has and has not been measured. The speaker test also streams ASR and
checks that utterances carry the right name when PARAKEET_TEST_GGUF is set.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* fix(cli): write the speaker registry atomically and keep the docs consistent

enroll now writes <registry>.tmp, checks every write and the close, then
renames over the registry, so a failed write leaves the old file intact.
docs/speaker.md explains why the sample scores differ from the measured
ones and drops first-person wording.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* fix(speaker): name a slot while its segment is still open

SpeakerIdentifier only took audio from closed segments, so a speaker
stayed "Speaker N" for their whole first turn and got a name only after
the first pause.

Open segments are now consumed as they grow. Each slot keeps a cursor,
and a call adds the clean audio from the cursor to the segment end, for
closed and open segments alike, so audio taken while a segment was open
is not added again when it closes. A short clean tail at the growing end
is held back and joined to the audio that follows it, so small feeds do
not lose it to the 0.2 s minimum piece.

On tests/fixtures/two_speakers.wav with the low latency preset, slot 0
is now named at 3.4 s of stream time instead of 6.2 s.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* fix(scene): keep one JSON shape for a stream with a speaker part

The scene JSON printed "names" and the per-item "name"/"name_score"
fields only once some slot had been seen, so the first documents of a
stream with a speaker part had a different shape than the later ones.

SceneUpdate now has a `named` flag, set on every update when the stream
has a speaker part. The writer then prints "names" (possibly {}) and the
name fields (empty name, score 0.0000) from the first document on.
Without a speaker part the output is byte for byte what it was, and a
golden test made with the old serializer guards that.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* fix(speaker): save registries atomically and do not replace an unreadable one

A new helper, pk::write_file_atomic, writes <path>.tmp in the same
directory, checks every write and the close, and moves the tmp file over
the target: MoveFileExA with MOVEFILE_REPLACE_EXISTING on Windows, where
rename() fails when the target exists, and rename() elsewhere. On any
failure the tmp file is removed and the old file is left as it was.

parakeet-cli enroll and parakeet_capi_speaker_registry_save both use it.
The C-API save used to truncate the existing file before writing, and
enroll could not add to an existing registry on Windows.

enroll also treated any failure to open the registry as "no registry
yet" and replaced it with a new one. Now only a missing file starts a new
registry. Any other failure exits 1 with the path and the reason, and a
directory given as the registry says so instead of reporting a truncated
file.

Also: test_capi_speaker uses std::filesystem for its temp file so it
builds on Windows, and the named SAS entry point keeps the message of a
std::exception instead of "unknown error".

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* docs(speaker): give a starting threshold per encoder and gate the voice-detect OFF build

docs/speaker.md gets a starting threshold per encoder, from the fixture
numbers only: 0.5 for WeSpeaker ResNet34 and CAM++, 0.7 for ECAPA (an
impostor reached 0.566 there), ERes2Net not measured. It recommends
WeSpeaker to start with and says the threshold needs checking on your
own audio. The code default stays 0.5; the scene help text points ECAPA
users to 0.7. The Enroll section now says what happens to one voice
under two names and to near-duplicate names, the clip-to-clip table says
which clips it uses, and the timing section says a slot is named while
it is still talking.

CI: the CED-OFF job now also builds with PARAKEET_WITH_VOICEDETECT=OFF
and checks that `parakeet-cli enroll` exits 2 with "built without
speaker identification".

SpeakerRegistry refuses a NaN or Inf embedding on enroll and treats a
NaN or Inf probe as unknown, like an all-zero one.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* test(speaker): word a comment without an arrow

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* build: point the voice-detect submodule at localai-org

voice-detect.cpp now lives under localai-org, like ced.cpp. The old
mudler URL redirects, but the canonical one is what a fresh clone and the
docs should use. The pinned commit is unchanged and is on the pushed
feat/embedding branch.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* build: pin voice-detect.cpp to its merge commit

localai-org/voice-detect.cpp#1 is merged. Move the submodule from the
feature branch commit b44c586 to the merge commit b74a896 on master. The
tree is identical, so nothing else changes.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants