Skip to content

feat(transcripts): opt-in speaker diarization for transcripts (#168) - #204

Merged
chhoumann merged 3 commits into
masterfrom
chhoumann/168-diarization
Jun 17, 2026
Merged

chhoumann merged 3 commits into
masterfrom
chhoumann/168-diarization

Conversation

@chhoumann

Copy link
Copy Markdown
Owner

Closes #168.

Summary

Adds opt-in speaker diarization for episode transcripts: instead of one block of unattributed prose, the transcript is labelled by speaker:

**A:** Welcome to the show.

**B:** Thanks for having me.

It is off by default, so existing transcripts and the plain-Whisper workflow are completely unchanged unless you turn it on. The labelled turns fill the existing {{transcript}} template value, so no template change is needed to use it.

The product decision in this PR

OpenAI Whisper (whisper-1, the current backend) produces no speaker labels, so diarization needs a diarization-capable backend. I researched the realistic options (verified against official 2026 docs) and this PR ships two user-selectable providers behind one opt-in toggle:

Provider Key How Speaker-label consistency
OpenAI gpt-4o-transcribe-diarize reuses your existing OpenAI key same /v1/audio/transcriptions endpoint, diarized_json per-request: a long episode is chunked at OpenAI's 25 MB cap, so labels can reset across chunks
Deepgram a new, separate Deepgram key one whole-file requestUrl POST consistent across the whole episode

OpenAI is the zero-friction default (nothing new to sign up for). Deepgram is the "do it properly for long episodes" option (it ingests the whole file in one request, so speaker identity stays stable end to end). Both work on desktop and mobile (requestUrl bypasses CORS).

Options considered and rejected: MacWhisper CLI (macOS-only, the CLI has no diarize flag, can't ship in a cross-platform community plugin) and in-browser WASM diarization (OOMs on iOS). Happy to drop one provider or change the default if you'd prefer a narrower surface.

What's included

  • src/services/diarization/ — a small provider seam: types, pure segments (parse OpenAI diarized_json + Deepgram utterances/words into normalized speaker turns and render them), openaiProvider, deepgramProvider (injectable requestUrl), and a pure requiredTranscriptionKeyPresent credential gate.
  • TranscriptionService routes to the chosen provider when enabled; the Whisper path is untouched when off.
  • Settings: transcript.diarization { enabled, provider, speakerTemplate } plus a top-level secret diarizationApiKey (kept top-level so the settings export can redact it; it is never nested inside the wholesale-copied transcript object). A pure migrateTranscriptSettings backfills the new defaults onto existing users' settings.
  • Settings UI: a toggle, a provider dropdown, a conditional Deepgram key field, and a "Speaker label format" field using a {{speaker}} token.
  • Docs (docs/docs/transcripts.md) and the e2e seed-parity fixture.

Speaker token

The {{speaker}} token lives in the new Speaker label format setting (default **{{speaker}}:** ), e.g. set it to **Speaker {{speaker}}:** or > {{speaker}}: . OpenAI labels speakers A/B/…, Deepgram 1/2/… (presented 1-based).

Verification

  • Gates green: lint, format:check, typecheck, build, test (521 tests). docs:build could not run in my worktree (no local mkdocs); the change is plain markdown appended to an already-in-nav page.
  • I don't have provider API keys, so the network calls ship instrumented (console logging on failure) and the provider parsing / HTTP wrapper / credential gate are unit-tested with stubs.
  • Verified in a real (isolated) Obsidian vault: the plugin loads with no console errors, defaults are correct, a legacy data.json (no diarization) migrates correctly, and diarization settings persist across reload.

Review

Ran an ultracode multi-dimension review (correctness, security, migration, API-correctness, maintainability; each finding adversarially re-verified against the code) plus three opposite-model (Codex) adversarial reviewers (Skeptic / Architect / Minimalist). No plain-Whisper regression found. Fixes applied from the reviews (second commit):

  • Export consent copy now honestly names both keys (was "OpenAI API key" only while exporting both).
  • Settings import now runs the same diarization clamp/backfill as load (an imported bogus provider no longer persists for a session).
  • OpenAI diarization that fails on every chunk now throws (no marker-only "completed" transcript that blocks retry); partial failures still keep a marker.

Notes for the maintainer

  • New optional paid third-party dependency (Deepgram) if a user chooses that provider. Strictly opt-in, default off, separate key. Flagging since it is a product call.
  • Localized src/TemplateEngine.ts: not touched (the speaker token is handled inside the diarization service), to avoid colliding with Buggy Behavior when Capturing Timestamp into Markdown Table #165.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Jun 16, 2026 •

Copy link
Copy Markdown

Deploying podnotes with  Cloudflare Pages  Cloudflare Pages

Latest commit: c9992e7
Status: ✅  Deploy successful!
Preview URL: https://3b19ee7b.podnotes.pages.dev
Branch Preview URL: https://chhoumann-168-diarization.podnotes.pages.dev

View logs

@chhoumann
chhoumann marked this pull request as ready for review June 16, 2026 19:56

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d120aa3c87

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/services/TranscriptionService.ts
@chhoumann

Copy link
Copy Markdown
Owner Author

Verified end-to-end against real podcasts (with live API keys)

Drove the real per-episode pipeline (download, chunk, provider call, render, file write) inside a real Obsidian instance with live OpenAI and Deepgram keys. All three scenarios produced coherent speaker-labelled transcripts with no error markers.

# Episode Size Provider Result
1 NPR Up First (~12 min) ~11.8 MB (1 chunk) OpenAI gpt-4o-transcribe-diarize 52 turns, labels **A:** / **B:** / **C:**…, hosts cleanly separated from ad-read voices. ~5 min.
2 NPR Up First (~12 min) whole file Deepgram 21 turns, labels **1:** / **2:**…, single whole-file request, consistent across the episode. ~39 s.
3 The Daily (~28 min) ~27.6 MB (2 chunks) OpenAI gpt-4o-transcribe-diarize 188 turns, A–M labels, transcript spans start to closing credits, so both chunks concatenated correctly. No [Error diarizing chunk N] markers. ~14 min.

Scenario 3 specifically exercises the >25 MB multi-chunk path: the byte-split chunks both transcribed and the segments concatenated start-to-finish. The chunk-local speaker-label caveat documented in docs/docs/transcripts.md held (labels are assigned per request), and the output stayed readable.

Total spend was a few dimes. Test keys were injected only into a throwaway isolated vault and scrubbed afterwards (no key material left on disk).

chhoumann added a commit that referenced this pull request Jun 17, 2026
…split (#168)

Codex review (PR #204): createChunkFiles converted every m4a to WAV before
checking size, so an m4a episode that already fits under the upload limit was
exploded into many uncompressed WAV chunks. That multiplied requests and, for
diarization, reset speaker labels at artificial chunk boundaries. The m4a->WAV
path exists only to safely SPLIT an m4a (which can't be byte-split), so it is
not needed when the file already fits.

Return the original (compressed) bytes as a single file whenever the buffer is
within the chunk size, before any conversion. OpenAI accepts m4a/mp3 directly
under the limit, so a small m4a is now one intact request with consistent
speaker labels; larger m4a still converts to WAV to split safely. Helps the
plain-Whisper path too (fewer requests for small m4a).

Verified end-to-end with a real two-voice m4a: the whole file uploads, decodes,
and transcribes accurately through both OpenAI (single-file path) and Deepgram,
no errors. Adds unit tests that a small m4a/mp3 yields one original-format file.
Label transcript segments by speaker. Whisper produces no speaker labels,
so diarization is an opt-in path (default off) that routes the episode audio
to a diarization-capable provider instead of plain Whisper:

- OpenAI gpt-4o-transcribe-diarize, reusing the existing OpenAI key (long
  episodes are chunked, so speaker labels can reset across chunk boundaries).
- Deepgram pre-recorded, a single whole-file request via requestUrl that keeps
  speaker labels consistent across the whole episode (needs its own API key).

The labelled turns fill the existing {{transcript}} template value; a new
"Speaker label format" setting uses a {{speaker}} token. The Deepgram key is a
top-level secret so settings export redacts it (it is never nested under the
wholesale-copied transcript object). A pure migration backfills the diarization
defaults onto existing transcript settings.

Provider parse/render logic is pure and unit-tested; the Deepgram HTTP wrapper
and the credential gate are tested with stubs. Verified in the isolated
worktree vault: plugin loads, defaults present, legacy settings migrate, and
diarization settings persist across reload.
From an ultracode review plus three opposite-model adversarial reviewers:

- Settings export named only the "OpenAI API key" while now also exporting the
  Deepgram key behind the same opt-in. Make the consent copy honest: rename the
  toggle to "Include API keys", and have the export notice and import
  confirmation name the specific keys actually present (new describeSecrets).
- Settings import bypassed the diarization clamp/backfill, so an imported bogus
  provider persisted for one session. Run migrateTranscriptSettings inside
  mergeImportedSettings so import converges with the load path.
- OpenAI diarization that failed on EVERY chunk saved a marker-only transcript
  and reported success, blocking retry. Throw when all chunks fail so no file is
  written and the run stays retryable (partial failures still keep a marker).
- Note that OpenAI per-chunk timestamps are chunk-relative (latent, unused).
- Docs: the default speaker label renders as **A:**, not **Speaker A:**.

Adds tests for the import provider clamp, describeSecrets, and the OpenAI
all-fail / partial-fail behavior. Gates green: lint, typecheck, build, 521 tests.
…split (#168)

Codex review (PR #204): createChunkFiles converted every m4a to WAV before
checking size, so an m4a episode that already fits under the upload limit was
exploded into many uncompressed WAV chunks. That multiplied requests and, for
diarization, reset speaker labels at artificial chunk boundaries. The m4a->WAV
path exists only to safely SPLIT an m4a (which can't be byte-split), so it is
not needed when the file already fits.

Return the original (compressed) bytes as a single file whenever the buffer is
within the chunk size, before any conversion. OpenAI accepts m4a/mp3 directly
under the limit, so a small m4a is now one intact request with consistent
speaker labels; larger m4a still converts to WAV to split safely. Helps the
plain-Whisper path too (fewer requests for small m4a).

Verified end-to-end with a real two-voice m4a: the whole file uploads, decodes,
and transcribes accurately through both OpenAI (single-file path) and Deepgram,
no errors. Adds unit tests that a small m4a/mp3 yields one original-format file.
@chhoumann
chhoumann force-pushed the chhoumann/168-diarization branch from 8bc7610 to c9992e7 Compare June 17, 2026 06:18
@chhoumann
chhoumann merged commit a96e12f into master Jun 17, 2026
2 checks passed
github-actions Bot pushed a commit that referenced this pull request Jun 22, 2026
# [2.17.0](2.16.0...2.17.0) (2026-06-22)

### Bug Fixes

* behavioral-audit logic and robustness fixes (back-end, 1/2) ([#213](#213)) ([4e2845d](4e2845d))
* **download:** default per-episode download path and migrate empty default ([#183](#183)) ([#186](#186)) ([46a6486](46a6486))
* **download:** prevent Android crash and create missing folders on download ([#178](#178)) ([ecd09d5](ecd09d5)), closes [#113](#113) [#86](#86) [#113](#113) [#86](#86)
* **lifecycle:** defer mobile podcast view startup ([#208](#208)) ([8bf9ce4](8bf9ce4))
* **notes:** cap note path length and harden folder creation ([#22](#22), [#87](#87)) ([#192](#192)) ([858e280](858e280)), closes [#87-class](#87)
* **playback:** persist listened time during playback ([#33](#33)) ([#190](#190)) ([e3433c2](e3433c2)), closes [#191](#191) [#108](#108) [#163](#163) [#183](#183)
* **playback:** play local files and downloads on iOS via resource path ([#100](#100)) ([#184](#184)) ([12c503a](12c503a))
* **player:** clear progress on episode switch to stop end-of-playback glitch ([#94](#94)) ([#194](#194)) ([59dccb3](59dccb3))
* **player:** reveal PodNotes view on Play with PodNotes so local files play ([#84](#84)) ([#198](#198)) ([5953625](5953625))
* **settings:** show labelled Add/Remove buttons in podcast search ([#109](#109)) ([#195](#195)) ([33399f5](33399f5))
* show downloaded episodes in the Local Files playlist ([#176](#176)) ([#177](#177)) ([184188c](184188c))
* **timestamps:** capture into the cursor's table cell without breaking the row ([#165](#165)) ([#203](#203)) ([964e342](964e342))
* **transcription:** always transcribe the currently playing episode ([#182](#182)) ([62b488a](62b488a)), closes [#107](#107)
* **uri:** preserve '+' in episode titles and paths for timestamp links ([#181](#181)) ([8ad7aa5](8ad7aa5)), closes [#164](#164)
* **view:** reliably reveal PodNotes view via command + ribbon icon ([#55](#55)) ([#199](#199)) ([fa1d708](fa1d708))

### Features

* add podcast segment links ([#205](#205)) ([d97e59e](d97e59e))
* **api:** expose generated episode transcripts ([e264465](e264465)), closes [#105](#105)
* behavioral-audit UI and interaction fixes (front-end, 2/2) ([#215](#215)) ([894d93d](894d93d))
* **commands:** add playback rate and media timestamp controls ([#206](#206)) ([a8bb44a](a8bb44a))
* **devx:** isolated per-worktree Obsidian E2E vault wrapper ([#188](#188)) ([a8b7a4a](a8b7a4a))
* **episodes:** add a setting to control the Latest Episodes list length ([#114](#114)) ([#200](#200)) ([7b3e3c6](7b3e3c6))
* **notes:** add {{episodelink}} template tag to resume an episode from its note ([#35](#35)) ([#193](#193)) ([8c1ddd6](8c1ddd6))
* **notes:** add podcast feed-level notes ([#163](#163)) ([#187](#187)) ([db0de47](db0de47)), closes [#161](#161) [#160](#160)
* **notes:** ship a Bases-friendly default episode note template ([#160](#160)) ([#201](#201)) ([209431d](209431d)), closes [#163](#163) [#183](#183)
* **player:** scale episode title font size to its length ([#81](#81)) ([#202](#202)) ([659b6b8](659b6b8))
* **player:** support video episode playback ([#209](#209)) ([f92a91f](f92a91f))
* **queue:** add setting to disable queue auto-population and auto-advance ([#108](#108)) ([#185](#185)) ([cf3d73c](cf3d73c))
* **queue:** allow reordering the playback queue ([#80](#80)) ([#179](#179)) ([d994d61](d994d61)), closes [#173](#173)
* **settings:** import/export settings & templates ([#180](#180)) ([a27d23d](a27d23d)), closes [#162](#162) [#162](#162)
* **templates:** add {{currentDate}}, {{episodeNumber}}, {{duration}} template variables ([#189](#189)) ([ec573eb](ec573eb)), closes [#75](#75) [#34](#34) [#88](#88) [163/#186](#186)
* **templates:** add episode chapters tag ([#207](#207)) ([9c98863](9c98863))
* **transcripts:** opt-in speaker diarization for transcripts ([#168](#168)) ([#204](#204)) ([a96e12f](a96e12f))
@github-actions

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 2.17.0 🎉

The release is available on GitHub release

Your semantic-release bot 📦🚀

@chhoumann
chhoumann deleted the chhoumann/168-diarization branch June 29, 2026 05:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Feature request: Diarization per speaker

1 participant