Skip to content

fix(transcription): resolve three deepsec transcription-pipeline bugs - #224

Merged
chhoumann merged 3 commits into
masterfrom
chhoumann/deepsec-transcription-bugs
Jun 29, 2026
Merged

chhoumann merged 3 commits into
masterfrom
chhoumann/deepsec-transcription-bugs

Conversation

@chhoumann

Copy link
Copy Markdown
Owner

Summary

Fixes three deepsec-confirmed correctness bugs in the transcription pipeline. Each is an independent commit.

1. Failed chunks saved as a "completed" transcript, blocking retry — other-silent-failure

src/services/TranscriptionService.ts

When a chunk exhausted its retries the worker wrote an [Error transcribing chunk N] placeholder and counted it as completed, and buildTranscriptBody() only rejected a fully empty body. A transcript made of error placeholders therefore passed the guard and was saved as "Transcription completed and saved." Because the existence check short-circuits on a later attempt ("you've already transcribed this episode"), the user was left with a corrupt/partial transcript - in the worst case a file containing only [Error transcribing chunk 0] - and no straightforward way to re-run.

Fix:

  • transcribeChunks() now reports how many chunks failed.
  • buildTranscriptBody() strips the error placeholders to decide whether any real speech was transcribed. If nothing real remains (every chunk failed, or the only successes were empty) it throws, so no file is written and the episode stays retryable - mirroring the diarization provider's existing all-chunks-failed contract. This also closes a mixed-failure hole (some chunks fail, surviving chunks return empty text) that a simple "all chunks failed" count would miss.
  • When some chunks fail but real speech remains, the otherwise-good transcript is kept and the final notice reports the partial failure (count + how to retry) instead of a clean success.

2. Byte-splitting container/lossless audio produced undecodable chunks — other-logic-bug (audioChunker)

src/services/audioChunker.ts

createChunkFiles() only routed m4a over the decode+WAV re-encode path; every other format over the 20 MB request limit fell through to raw byte-splitting. MP3 frames are self-synchronizing so that is fine, but container/header-dependent formats (ogg, webm, flac, wav, m4a) carry codec setup only at the stream start, so non-first byte slices are undecodable and the API rejects or garbles them (which then fed the placeholder path in bug 1).

Fix: restrict raw byte-splitting to MP3; route every other format through the existing decode+WAV path. When WAV conversion yields nothing (no Web Audio decoder, or an undecodable codec) throw a clear, retryable error instead of byte-splitting a container into corrupt chunks. This also replaces the previous m4a fallback that silently byte-split into corrupt chunks when conversion failed.

3. Speaker label mangled by String.replace $-patterns — other-logic-bug (diarization segments)

src/services/diarization/segments.ts

formatSpeakerLabel() passed the provider-supplied speaker label as the replacement string to String.replace, which interprets $$, $&, $`, $', $n. A label containing one of those would be rewritten with matched/surrounding text. Fixed by using a replacer function so the label is inserted verbatim. (Current providers emit A/B/numeric labels, so this is a latent defect, fixed for completeness.)

Tests

  • New unit tests for each fix, including ones proving the bad input is now blocked: all-chunks-failed and mixed-fail+empty-success throw (no file written); ogg routes to WAV / corrupt byte-split is no longer produced / undecodable formats throw; a $-laden speaker label is inserted verbatim.
  • npm run lint, typecheck, build all pass; npm run test passes for all changed areas (audioChunker 13, TranscriptionService 14, segments 15). The full suite is green when the machine isn't saturated; under heavy parallel load one pre-existing timing-sensitive test (PodcastView.integration.test.ts, feed-cache, unrelated to these files) can hit its 5000ms timeout - it passes in isolation.

Runtime verification

Confirmed in an isolated Obsidian (Electron) e2e vault that AudioContext.decodeAudioData is available and decodes audio to a valid AudioBuffer - so the non-mp3 → WAV path this PR relies on is reachable at runtime (jsdom has no Web Audio, which is why the unit tests stub it).

Notes / out of scope

  • The download layer defaults an undetected extension to mp3, so a truly unidentifiable non-mp3 file >20 MB could still be byte-split - pre-existing behavior, now caught by bug 1's retryable failure. A WebM/EBML magic-number entry in detectAudioFileExtension would close that gap (follow-up).

Does not merge automatically - opening for review.

A chunk that exhausted its retries wrote an "[Error transcribing chunk N]"
placeholder and was counted as completed, and buildTranscriptBody only
rejected a fully-empty body. A transcript made of error placeholders was
therefore saved as "completed", and the existence check then blocked any
retry, leaving the user with a corrupt transcript and no easy way to re-run.

transcribeChunks now reports how many chunks failed. buildTranscriptBody
strips the error placeholders to decide whether any real speech was
transcribed: if nothing real remains (every chunk failed, or the only
successes were empty) it throws, so no file is written and the episode stays
retryable - mirroring the diarization provider's all-chunks-failed contract.
When some chunks fail but real speech remains, the transcript is kept and a
warning notice reports the partial failure instead of a clean success.

Fixes deepsec finding other-silent-failure.
…itting

createChunkFiles only routed m4a over the decode+WAV path; every other
format over the 20 MB request limit fell through to raw byte-splitting. MP3
frames are self-synchronizing so that is fine, but container/header-dependent
formats (ogg, webm, flac, wav, m4a) carry codec setup only at the stream
start, so non-first byte slices are undecodable and the API rejects or
garbles them.

Restrict raw byte-splitting to MP3; route every other format through the
existing decode+WAV re-encode path. When WAV conversion yields nothing (no
Web Audio decoder, or an undecodable codec) throw a clear error rather than
byte-splitting a container into corrupt chunks, keeping the run retryable.

Fixes deepsec finding other-logic-bug (audioChunker).
formatSpeakerLabel passed the speaker label as the replacement string to
String.replace, which interprets special patterns ($$, $&, $`, $', $n). A
label containing one of those would be rewritten with matched/surrounding
text instead of inserted literally. Use a replacer function so the label is
inserted verbatim.

Fixes deepsec finding other-logic-bug (diarization segments).
@cloudflare-workers-and-pages

Copy link
Copy Markdown

Deploying podnotes with  Cloudflare Pages  Cloudflare Pages

Latest commit: dde9356
Status: ✅  Deploy successful!
Preview URL: https://101feb8a.podnotes.pages.dev
Branch Preview URL: https://chhoumann-deepsec-transcript.podnotes.pages.dev

View logs

@chhoumann
chhoumann merged commit 83c34e7 into master Jun 29, 2026
3 checks passed
@chhoumann
chhoumann deleted the chhoumann/deepsec-transcription-bugs branch June 29, 2026 07:23
github-actions Bot pushed a commit that referenced this pull request Jul 9, 2026
## [2.17.3](2.17.2...2.17.3) (2026-07-09)

### Bug Fixes

* **feed/search:** parse feed once + content-based search cache ([#225](#225)) ([053d51f](053d51f)), closes [#149](#149)
* make episode identity key collision-resistant and prototype-safe ([#226](#226)) ([a5683db](a5683db))
* **opml:** correct import progress math and saved-count reporting ([#221](#221)) ([a79e529](a79e529))
* **security:** validate feed/URI URLs and cap download size ([#223](#223)) ([edef281](edef281))
* **template:** neutralize feed-controlled note injection ([#228](#228)) ([ef4ecbd](ef4ecbd))
* **timestamp:** escape live table-cell pipe after an escaped backslash ([#227](#227)) ([a34dfca](a34dfca))
* **transcription:** resolve three deepsec transcription-pipeline bugs ([#224](#224)) ([83c34e7](83c34e7))
@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 2.17.3 🎉

The release is available on GitHub release

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant