build: allow embedding voice-detect.cpp in another ggml project - #1
Merged
Merged
Conversation
parakeet.cpp folds a speaker encoder into its scene stream, as it does for ced.cpp. That needs voice-detect to build inside a host that already owns a ggml target and a dr_wav implementation. Guard the GGML_* forwarding, patch step, CLI and tests behind a top-level check, reuse an existing ggml target, and add VOICEDETECT_EXTERNAL_DR_WAV so the host's dr_wav is the only definition. Add voicedetect_capi_embedding_dim so a caller can size a speaker registry before running the encoder. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
mudler
approved these changes
Sep 30, 2026
mudler
added a commit
to mudler/parakeet.cpp
that referenced
this pull request
Sep 30, 2026
localai-org/voice-detect.cpp#1 is merged. Move the submodule from the feature branch commit b44c586 to the merge commit b74a896 on master. The tree is identical, so nothing else changes. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
mudler
added a commit
to mudler/parakeet.cpp
that referenced
this pull request
Sep 30, 2026
…v9) (#78) * feat(speaker): add SpeakerRegistry for enrolled voices Enrolled speakers are centroids of L2-normalized embeddings. identify() takes a threshold and a runner-up margin so a voice that is not enrolled comes out unknown instead of as the nearest name. The registry saves and loads as a small binary blob and rejects corrupt or mismatched input. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(speaker): reject a registry blob with dim 0 and speakers deserialize() now validates that dim >= 1 when n > 0 (number of speakers > 0). A crafted blob with dim 0 and speakers was accepted, creating entries with empty sum vectors. A later enroll() would then cause an out-of-bounds heap write when trying to accumulate the new embedding. Empty registries (dim 0, n 0) still serialize and deserialize correctly. The fix adds validation to reject implausible headers and includes tests for: - A hand-built corrupt blob (dim 0, n 1) must throw - An empty registry round-trips and accepts subsequent enrollment - Failed enrollments on non-empty registries leave state unchanged - names() correctly lists remaining speakers after removal Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(speaker): add SpeakerIdentifier for diarization slots Collects each slot's clean audio (overlap with other speakers is skipped), embeds it once it has enough and again as it grows, and names the slot from the registry. A different name replaces the current one only after winning twice in a row, and an unknown match never removes a name. identify_offline reuses the same logic for finished recordings. Driven by an embedding callback, so it is tested with a fake encoder. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(speaker): test overlap against closed segments and document the open contract Adds tests for overlap with segments closed in the same call and in an earlier call, which exercise the history scan. Documents that the caller must list every started, unclosed segment in `open`, and what identify_offline does to max_voice_sec and ring_sec. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(speaker): carry speaker names through words, utterances and scene JSON Words and utterances get a name and a match score, SceneUpdate gets the current name of every slot, and the scene JSON and the text renderer print them. With no speaker model the JSON is unchanged, byte for byte. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(speaker): fold voice-detect.cpp in behind SpeakerEncoder voice-detect.cpp is a submodule built as a static target that shares our ggml and dr_wav, like ced.cpp. pk::SpeakerEncoder is the only code that touches it, through voicedetect_capi.h, and PARAKEET_WITH_VOICEDETECT=OFF builds without it. The test checks that two clips of one voice score above two clips of different voices on two_speakers.wav, and that the folded encoder matches a standalone voice-detect build. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(speaker): identify speakers inside the scene stream SceneStream runs an optional SpeakerIdentifier after diarization, so words, utterances and the per-slot names in each update carry the enrolled name. It needs diarization and a registry and says so when they are missing. Names apply to words committed after the slot is identified; earlier words keep the label they were emitted with. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(speaker): check speaker preconditions first and test order-independent naming The speaker checks now run before the generic "needs at least one model" check, so a speaker-only config gets the specific message. The test asserts each message, enrolls the voices in reverse arrival order so slot i cannot map to registry entry i, and checks that a voice missing from the registry stays unnamed. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(capi): speaker identification, ABI v9 A voice-detect GGUF loads into a fourth context kind. A registry handle enrolls, saves and loads voices; the scene stream takes a speaker context and registry through a new begin function; speaker-attributed ASR has a named variant. Existing signatures are unchanged, and the scene options grow only at the end, read according to the caller's size. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(capi): read scene options only within the caller's size and avoid a stream leak The speaker option fields were passed as arguments to a helper that checked the caller's size, so they were read before the check and a caller built against the v8 header read past its struct. Each field is now read by offset only after the size covers it. The scene wrapper is also allocated after the SceneStream is built, so a throwing constructor no longer leaks it. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(cli): enroll speakers and name them in scene output parakeet-cli enroll builds a registry from labeled clips, and scene --speakers --registry names diarized speakers in the transcript and the JSON. docs/speaker.md covers the models, the C-API, the timing rule and what has and has not been measured. The speaker test also streams ASR and checks that utterances carry the right name when PARAKEET_TEST_GGUF is set. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(cli): write the speaker registry atomically and keep the docs consistent enroll now writes <registry>.tmp, checks every write and the close, then renames over the registry, so a failed write leaves the old file intact. docs/speaker.md explains why the sample scores differ from the measured ones and drops first-person wording. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(speaker): name a slot while its segment is still open SpeakerIdentifier only took audio from closed segments, so a speaker stayed "Speaker N" for their whole first turn and got a name only after the first pause. Open segments are now consumed as they grow. Each slot keeps a cursor, and a call adds the clean audio from the cursor to the segment end, for closed and open segments alike, so audio taken while a segment was open is not added again when it closes. A short clean tail at the growing end is held back and joined to the audio that follows it, so small feeds do not lose it to the 0.2 s minimum piece. On tests/fixtures/two_speakers.wav with the low latency preset, slot 0 is now named at 3.4 s of stream time instead of 6.2 s. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(scene): keep one JSON shape for a stream with a speaker part The scene JSON printed "names" and the per-item "name"/"name_score" fields only once some slot had been seen, so the first documents of a stream with a speaker part had a different shape than the later ones. SceneUpdate now has a `named` flag, set on every update when the stream has a speaker part. The writer then prints "names" (possibly {}) and the name fields (empty name, score 0.0000) from the first document on. Without a speaker part the output is byte for byte what it was, and a golden test made with the old serializer guards that. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(speaker): save registries atomically and do not replace an unreadable one A new helper, pk::write_file_atomic, writes <path>.tmp in the same directory, checks every write and the close, and moves the tmp file over the target: MoveFileExA with MOVEFILE_REPLACE_EXISTING on Windows, where rename() fails when the target exists, and rename() elsewhere. On any failure the tmp file is removed and the old file is left as it was. parakeet-cli enroll and parakeet_capi_speaker_registry_save both use it. The C-API save used to truncate the existing file before writing, and enroll could not add to an existing registry on Windows. enroll also treated any failure to open the registry as "no registry yet" and replaced it with a new one. Now only a missing file starts a new registry. Any other failure exits 1 with the path and the reason, and a directory given as the registry says so instead of reporting a truncated file. Also: test_capi_speaker uses std::filesystem for its temp file so it builds on Windows, and the named SAS entry point keeps the message of a std::exception instead of "unknown error". Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * docs(speaker): give a starting threshold per encoder and gate the voice-detect OFF build docs/speaker.md gets a starting threshold per encoder, from the fixture numbers only: 0.5 for WeSpeaker ResNet34 and CAM++, 0.7 for ECAPA (an impostor reached 0.566 there), ERes2Net not measured. It recommends WeSpeaker to start with and says the threshold needs checking on your own audio. The code default stays 0.5; the scene help text points ECAPA users to 0.7. The Enroll section now says what happens to one voice under two names and to near-duplicate names, the clip-to-clip table says which clips it uses, and the timing section says a slot is named while it is still talking. CI: the CED-OFF job now also builds with PARAKEET_WITH_VOICEDETECT=OFF and checks that `parakeet-cli enroll` exits 2 with "built without speaker identification". SpeakerRegistry refuses a NaN or Inf embedding on enroll and treats a NaN or Inf probe as unknown, like an all-zero one. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * test(speaker): word a comment without an arrow Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * build: point the voice-detect submodule at localai-org voice-detect.cpp now lives under localai-org, like ced.cpp. The old mudler URL redirects, but the canonical one is what a fresh clone and the docs should use. The pinned commit is unchanged and is on the pushed feat/embedding branch. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * build: pin voice-detect.cpp to its merge commit localai-org/voice-detect.cpp#1 is merged. Move the submodule from the feature branch commit b44c586 to the merge commit b74a896 on master. The tree is identical, so nothing else changes. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
parakeet.cpp is adding speaker identification: it names the slots its diarizer finds by matching each slot's voice against enrolled speakers. It folds voice-detect.cpp in as a static dependency, the same way it already folds ced.cpp, so voice-detect has to build inside a host project that owns its own ggml target and its own dr_wav implementation.
What this changes:
VOICEDETECT_TOP_LEVEL: theGGML_*forwarding, the CPU variants setup, the ggml patch step, the CLI and the tests only run when voice-detect is the top-level project. An embedding host keeps its own ggml configuration.if(NOT TARGET ggml)guards the ggmladd_subdirectory, so a host that already has aggmltarget is reused.VOICEDETECT_EXTERNAL_DR_WAV(off by default): skips voice-detect's ownDR_WAV_IMPLEMENTATION, so the host's dr_wav is the only definition and the link has no duplicatedrwav_*symbols.CMAKE_SOURCE_DIRbecomesCMAKE_CURRENT_SOURCE_DIR/PROJECT_SOURCE_DIRwhere a host would resolve it to the wrong root.voicedetect_capi_embedding_dim(ctx): the embedding size (192 to 512), 0 for a model with no speaker embedding (the age, gender and emotion heads), -1 for a NULL ctx. It lets a caller size a speaker registry before it runs the encoder. Additive, so no ABI version bump.This mirrors what ced.cpp did for the same purpose (
b7dc30dande1a3cfain that repo). Nothing changes for a normal top-level build.How it was checked:
test_smokepasses,test_capi_dimpasses with the WeSpeaker GGUF (256) and skips without a model.VOICEDETECT_EXTERNAL_DR_WAV=ON: it links, printsabi 1, exits 0 and has a singledrwav_initdefinition. It compiles against ggml v0.13.0 even though this repo vendors v0.15.2.The parakeet.cpp side is a separate PR that pins this commit as a submodule.
🤖 Generated with Claude Code