docs(examples): add Voice Agent session recorder - #781
Merged
Conversation
GregHolmes
requested review from
deepgram-kiley and
dg-coreylweathers
as code owners
September 3, 2026 11:10
Contributor
|
dg-coreylweathers
approved these changes
Sep 3, 2026
dg-coreylweathers
left a comment
Contributor
There was a problem hiding this comment.
What this PR does
Adds example 31: a standalone Voice Agent script that injects one text question, collects the transcript, function-call, and latency events the server sends, and writes them to one JSON file when the session ends. Closes #775. No SDK code changes.
What I checked
- Does it run end to end? Yes. In a
python:3.12container with a valid key it finished in about 3 seconds and wrote a record with the user question, the assistant reply, and four latency reports (ttt_token_latency,ttt_text_latency,tts_latency,total_latency). - Is the temp-file claim true? Yes. The fallback temp file is created mode 0600. A file passed via
--outputgets the default umask (0644 in my run). - Can the API key leak? No. The key is absent from stdout, stderr, and the record; a fake-key run exposes nothing either.
- Do the event names and imports match the SDK? Yes. All six imported types exist, and
ConversationText,FunctionCallRequest,FunctionCallResponse, andLatencyReportare all in the server response union. - ruff format, ruff check, mypy, py_compile, and CI on 3.10-3.13 pass. No new runtime dependency.
Should-fix (non-blocking, can follow in a later commit)
- A server
Errorevent is silently ignored.on_message(line ~130) handles onlySettingsAppliedandConversationText, so if the think provider fails after the question is injected, the script waits the full 30 seconds and exits with "Timed out waiting for the agent to respond" while the real cause is lost. Example 30 printsErrorevents. Suggest adding, after theConversationTextbranch:
elif getattr(message, "type", None) == "Error":
print(f"Agent error: {message.code} - {message.description}")
assistant_response_event.set()and raising with the server's message after the wait instead of a timeout.
- A failed or timed-out session writes no record. Both
TimeoutErrorraises (lines ~151 and ~155) exit thewithblock beforerecorder.finish()and the write, so the events already collected are lost andended_atstays null. The docstring says "after a successful session," so this is documented, but the failed session is the one a developer most wants to inspect. Suggest atry/finallyaround thewithblock that always writes the record and lets the exception propagate.
Nits, optional
- Docstring (lines 1-13): the settings define no functions, so
function_callsis always[]in a real run. One sentence saying so would set expectations. - Line ~113: a bad or missing key produces a raw traceback ending in
InvalidStatus: HTTP 401; examples 18 and 30 print one line. Wrappingmain()inexcept Exception as exc: print(f"Session failed: {type(exc).__name__}: {exc}", file=sys.stderr)would match them. The key is not in the traceback. - Line ~57:
--outputfiles get default permissions while the temp fallback is 0600 and the docstring stresses sensitive contents. Either open with0o600or say in the docstring that--outputpermissions are the caller's.
GregHolmes
added a commit
that referenced
this pull request
Sep 3, 2026
🤖 I have created a release *beep* *boop* --- ## [7.8.1](v7.8.0...v7.8.1) (2026-09-03) ### Bug Fixes * **TextBuilder:** `ssml_to_deepgram()` now preserves a `<phoneme>` pronunciation when its valid `ph` and `alphabet` attributes appear in either order. ([#741](#741)) ([7fd4b63](7fd4b63)) * **Credentials:** Explicitly passing `api_key=None` continues to disable ambient `DEEPGRAM_API_KEY` lookup, which is important for multi-tenant and test environments. ([#778](#778)) ([e675990](e675990)) * **Credentials:** `DeepgramClient()` and `AsyncDeepgramClient()` now resolve `DEEPGRAM_API_KEY` when constructed, so `load_dotenv()` can run after importing the SDK. Closes [#734](#734). ([#767](#767)) ([ec362ec](ec362ec)) * **Custom transports:** Speak V2 WebSocket connections now honor `transport_factory`, matching the routing behavior of other WebSocket APIs for proxies, test doubles, and custom-hosted transports. ([#766](#766)) ([0980663](0980663)) ### Documentation * **Transcription:** Clarified that Nova-3 assumes English when `language` is omitted; non-English and multilingual audio require an explicit language such as `fr` or `multi`. ([#771](#771)) ([4574337](4574337)) * **Examples:** Added Listen V1 live microphone transcription with optional `sounddevice`, device selection, bounded audio buffering, transcript output, and clean Ctrl-C shutdown. ([#780](#780)) ([08f0471](08f0471)) * **Examples:** Added a resilient Listen V1 live transcription pattern with exponential backoff, reconnect-aware audio buffering, timestamp continuity, and clean shutdown. ([#776](#776)) ([96b2d11](96b2d11)) * **Examples:** Added an application-owned Voice Agent session recorder that serializes received transcripts, function calls, and latency reports as JSON while leaving consent, redaction, retention, and storage policy to the application. Closes [#775](#775). ([#781](#781)) ([30ad152](30ad152)) * **Text-to-Speech:** Corrected streaming synthesis snippets to iterate the response byte chunks instead of accessing a nonexistent `.stream` attribute. ([#749](#749)) ([178724e](178724e)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Validation
poetry run ruff format --check examples/31-voice-agent-session-recording.pypoetry run ruff check examples/31-voice-agent-session-recording.pypoetry run mypy --ignore-missing-imports examples/31-voice-agent-session-recording.pypoetry run python -m py_compile examples/31-voice-agent-session-recording.pypoetry run mypy src/poetry run mypy tests/typecheckpoetry run pytest -rP --cov=deepgram --cov-branch --cov-report=xml --cov-report=term-missing .Closes #775