feat(transcripts): opt-in speaker diarization for transcripts (#168) - #204
Conversation
Deploying podnotes with
|
| Latest commit: |
c9992e7
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://3b19ee7b.podnotes.pages.dev |
| Branch Preview URL: | https://chhoumann-168-diarization.podnotes.pages.dev |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d120aa3c87
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Verified end-to-end against real podcasts (with live API keys)Drove the real per-episode pipeline (download, chunk, provider call, render, file write) inside a real Obsidian instance with live OpenAI and Deepgram keys. All three scenarios produced coherent speaker-labelled transcripts with no error markers.
Scenario 3 specifically exercises the >25 MB multi-chunk path: the byte-split chunks both transcribed and the segments concatenated start-to-finish. The chunk-local speaker-label caveat documented in Total spend was a few dimes. Test keys were injected only into a throwaway isolated vault and scrubbed afterwards (no key material left on disk). |
…split (#168) Codex review (PR #204): createChunkFiles converted every m4a to WAV before checking size, so an m4a episode that already fits under the upload limit was exploded into many uncompressed WAV chunks. That multiplied requests and, for diarization, reset speaker labels at artificial chunk boundaries. The m4a->WAV path exists only to safely SPLIT an m4a (which can't be byte-split), so it is not needed when the file already fits. Return the original (compressed) bytes as a single file whenever the buffer is within the chunk size, before any conversion. OpenAI accepts m4a/mp3 directly under the limit, so a small m4a is now one intact request with consistent speaker labels; larger m4a still converts to WAV to split safely. Helps the plain-Whisper path too (fewer requests for small m4a). Verified end-to-end with a real two-voice m4a: the whole file uploads, decodes, and transcribes accurately through both OpenAI (single-file path) and Deepgram, no errors. Adds unit tests that a small m4a/mp3 yields one original-format file.
Label transcript segments by speaker. Whisper produces no speaker labels,
so diarization is an opt-in path (default off) that routes the episode audio
to a diarization-capable provider instead of plain Whisper:
- OpenAI gpt-4o-transcribe-diarize, reusing the existing OpenAI key (long
episodes are chunked, so speaker labels can reset across chunk boundaries).
- Deepgram pre-recorded, a single whole-file request via requestUrl that keeps
speaker labels consistent across the whole episode (needs its own API key).
The labelled turns fill the existing {{transcript}} template value; a new
"Speaker label format" setting uses a {{speaker}} token. The Deepgram key is a
top-level secret so settings export redacts it (it is never nested under the
wholesale-copied transcript object). A pure migration backfills the diarization
defaults onto existing transcript settings.
Provider parse/render logic is pure and unit-tested; the Deepgram HTTP wrapper
and the credential gate are tested with stubs. Verified in the isolated
worktree vault: plugin loads, defaults present, legacy settings migrate, and
diarization settings persist across reload.
From an ultracode review plus three opposite-model adversarial reviewers: - Settings export named only the "OpenAI API key" while now also exporting the Deepgram key behind the same opt-in. Make the consent copy honest: rename the toggle to "Include API keys", and have the export notice and import confirmation name the specific keys actually present (new describeSecrets). - Settings import bypassed the diarization clamp/backfill, so an imported bogus provider persisted for one session. Run migrateTranscriptSettings inside mergeImportedSettings so import converges with the load path. - OpenAI diarization that failed on EVERY chunk saved a marker-only transcript and reported success, blocking retry. Throw when all chunks fail so no file is written and the run stays retryable (partial failures still keep a marker). - Note that OpenAI per-chunk timestamps are chunk-relative (latent, unused). - Docs: the default speaker label renders as **A:**, not **Speaker A:**. Adds tests for the import provider clamp, describeSecrets, and the OpenAI all-fail / partial-fail behavior. Gates green: lint, typecheck, build, 521 tests.
…split (#168) Codex review (PR #204): createChunkFiles converted every m4a to WAV before checking size, so an m4a episode that already fits under the upload limit was exploded into many uncompressed WAV chunks. That multiplied requests and, for diarization, reset speaker labels at artificial chunk boundaries. The m4a->WAV path exists only to safely SPLIT an m4a (which can't be byte-split), so it is not needed when the file already fits. Return the original (compressed) bytes as a single file whenever the buffer is within the chunk size, before any conversion. OpenAI accepts m4a/mp3 directly under the limit, so a small m4a is now one intact request with consistent speaker labels; larger m4a still converts to WAV to split safely. Helps the plain-Whisper path too (fewer requests for small m4a). Verified end-to-end with a real two-voice m4a: the whole file uploads, decodes, and transcribes accurately through both OpenAI (single-file path) and Deepgram, no errors. Adds unit tests that a small m4a/mp3 yields one original-format file.
8bc7610 to
c9992e7
Compare
# [2.17.0](2.16.0...2.17.0) (2026-06-22) ### Bug Fixes * behavioral-audit logic and robustness fixes (back-end, 1/2) ([#213](#213)) ([4e2845d](4e2845d)) * **download:** default per-episode download path and migrate empty default ([#183](#183)) ([#186](#186)) ([46a6486](46a6486)) * **download:** prevent Android crash and create missing folders on download ([#178](#178)) ([ecd09d5](ecd09d5)), closes [#113](#113) [#86](#86) [#113](#113) [#86](#86) * **lifecycle:** defer mobile podcast view startup ([#208](#208)) ([8bf9ce4](8bf9ce4)) * **notes:** cap note path length and harden folder creation ([#22](#22), [#87](#87)) ([#192](#192)) ([858e280](858e280)), closes [#87-class](#87) * **playback:** persist listened time during playback ([#33](#33)) ([#190](#190)) ([e3433c2](e3433c2)), closes [#191](#191) [#108](#108) [#163](#163) [#183](#183) * **playback:** play local files and downloads on iOS via resource path ([#100](#100)) ([#184](#184)) ([12c503a](12c503a)) * **player:** clear progress on episode switch to stop end-of-playback glitch ([#94](#94)) ([#194](#194)) ([59dccb3](59dccb3)) * **player:** reveal PodNotes view on Play with PodNotes so local files play ([#84](#84)) ([#198](#198)) ([5953625](5953625)) * **settings:** show labelled Add/Remove buttons in podcast search ([#109](#109)) ([#195](#195)) ([33399f5](33399f5)) * show downloaded episodes in the Local Files playlist ([#176](#176)) ([#177](#177)) ([184188c](184188c)) * **timestamps:** capture into the cursor's table cell without breaking the row ([#165](#165)) ([#203](#203)) ([964e342](964e342)) * **transcription:** always transcribe the currently playing episode ([#182](#182)) ([62b488a](62b488a)), closes [#107](#107) * **uri:** preserve '+' in episode titles and paths for timestamp links ([#181](#181)) ([8ad7aa5](8ad7aa5)), closes [#164](#164) * **view:** reliably reveal PodNotes view via command + ribbon icon ([#55](#55)) ([#199](#199)) ([fa1d708](fa1d708)) ### Features * add podcast segment links ([#205](#205)) ([d97e59e](d97e59e)) * **api:** expose generated episode transcripts ([e264465](e264465)), closes [#105](#105) * behavioral-audit UI and interaction fixes (front-end, 2/2) ([#215](#215)) ([894d93d](894d93d)) * **commands:** add playback rate and media timestamp controls ([#206](#206)) ([a8bb44a](a8bb44a)) * **devx:** isolated per-worktree Obsidian E2E vault wrapper ([#188](#188)) ([a8b7a4a](a8b7a4a)) * **episodes:** add a setting to control the Latest Episodes list length ([#114](#114)) ([#200](#200)) ([7b3e3c6](7b3e3c6)) * **notes:** add {{episodelink}} template tag to resume an episode from its note ([#35](#35)) ([#193](#193)) ([8c1ddd6](8c1ddd6)) * **notes:** add podcast feed-level notes ([#163](#163)) ([#187](#187)) ([db0de47](db0de47)), closes [#161](#161) [#160](#160) * **notes:** ship a Bases-friendly default episode note template ([#160](#160)) ([#201](#201)) ([209431d](209431d)), closes [#163](#163) [#183](#183) * **player:** scale episode title font size to its length ([#81](#81)) ([#202](#202)) ([659b6b8](659b6b8)) * **player:** support video episode playback ([#209](#209)) ([f92a91f](f92a91f)) * **queue:** add setting to disable queue auto-population and auto-advance ([#108](#108)) ([#185](#185)) ([cf3d73c](cf3d73c)) * **queue:** allow reordering the playback queue ([#80](#80)) ([#179](#179)) ([d994d61](d994d61)), closes [#173](#173) * **settings:** import/export settings & templates ([#180](#180)) ([a27d23d](a27d23d)), closes [#162](#162) [#162](#162) * **templates:** add {{currentDate}}, {{episodeNumber}}, {{duration}} template variables ([#189](#189)) ([ec573eb](ec573eb)), closes [#75](#75) [#34](#34) [#88](#88) [163/#186](#186) * **templates:** add episode chapters tag ([#207](#207)) ([9c98863](9c98863)) * **transcripts:** opt-in speaker diarization for transcripts ([#168](#168)) ([#204](#204)) ([a96e12f](a96e12f))
|
🎉 This PR is included in version 2.17.0 🎉 The release is available on GitHub release Your semantic-release bot 📦🚀 |
Closes #168.
Summary
Adds opt-in speaker diarization for episode transcripts: instead of one block of unattributed prose, the transcript is labelled by speaker:
It is off by default, so existing transcripts and the plain-Whisper workflow are completely unchanged unless you turn it on. The labelled turns fill the existing
{{transcript}}template value, so no template change is needed to use it.The product decision in this PR
OpenAI Whisper (
whisper-1, the current backend) produces no speaker labels, so diarization needs a diarization-capable backend. I researched the realistic options (verified against official 2026 docs) and this PR ships two user-selectable providers behind one opt-in toggle:gpt-4o-transcribe-diarize/v1/audio/transcriptionsendpoint,diarized_jsonrequestUrlPOSTOpenAI is the zero-friction default (nothing new to sign up for). Deepgram is the "do it properly for long episodes" option (it ingests the whole file in one request, so speaker identity stays stable end to end). Both work on desktop and mobile (
requestUrlbypasses CORS).Options considered and rejected: MacWhisper CLI (macOS-only, the CLI has no diarize flag, can't ship in a cross-platform community plugin) and in-browser WASM diarization (OOMs on iOS). Happy to drop one provider or change the default if you'd prefer a narrower surface.
What's included
src/services/diarization/— a small provider seam:types, puresegments(parse OpenAIdiarized_json+ Deepgram utterances/words into normalized speaker turns and render them),openaiProvider,deepgramProvider(injectablerequestUrl), and a purerequiredTranscriptionKeyPresentcredential gate.TranscriptionServiceroutes to the chosen provider when enabled; the Whisper path is untouched when off.transcript.diarization { enabled, provider, speakerTemplate }plus a top-level secretdiarizationApiKey(kept top-level so the settings export can redact it; it is never nested inside the wholesale-copiedtranscriptobject). A puremigrateTranscriptSettingsbackfills the new defaults onto existing users' settings.{{speaker}}token.docs/docs/transcripts.md) and the e2e seed-parity fixture.Speaker token
The
{{speaker}}token lives in the new Speaker label format setting (default**{{speaker}}:**), e.g. set it to**Speaker {{speaker}}:**or> {{speaker}}:. OpenAI labels speakersA/B/…, Deepgram1/2/… (presented 1-based).Verification
lint,format:check,typecheck,build,test(521 tests).docs:buildcould not run in my worktree (no localmkdocs); the change is plain markdown appended to an already-in-nav page.data.json(nodiarization) migrates correctly, and diarization settings persist across reload.Review
Ran an ultracode multi-dimension review (correctness, security, migration, API-correctness, maintainability; each finding adversarially re-verified against the code) plus three opposite-model (Codex) adversarial reviewers (Skeptic / Architect / Minimalist). No plain-Whisper regression found. Fixes applied from the reviews (second commit):
Notes for the maintainer
src/TemplateEngine.ts: not touched (the speaker token is handled inside the diarization service), to avoid colliding with Buggy Behavior when Capturing Timestamp into Markdown Table #165.