fix(transcription): resolve three deepsec transcription-pipeline bugs - #224
Merged
Merged
Conversation
A chunk that exhausted its retries wrote an "[Error transcribing chunk N]" placeholder and was counted as completed, and buildTranscriptBody only rejected a fully-empty body. A transcript made of error placeholders was therefore saved as "completed", and the existence check then blocked any retry, leaving the user with a corrupt transcript and no easy way to re-run. transcribeChunks now reports how many chunks failed. buildTranscriptBody strips the error placeholders to decide whether any real speech was transcribed: if nothing real remains (every chunk failed, or the only successes were empty) it throws, so no file is written and the episode stays retryable - mirroring the diarization provider's all-chunks-failed contract. When some chunks fail but real speech remains, the transcript is kept and a warning notice reports the partial failure instead of a clean success. Fixes deepsec finding other-silent-failure.
…itting createChunkFiles only routed m4a over the decode+WAV path; every other format over the 20 MB request limit fell through to raw byte-splitting. MP3 frames are self-synchronizing so that is fine, but container/header-dependent formats (ogg, webm, flac, wav, m4a) carry codec setup only at the stream start, so non-first byte slices are undecodable and the API rejects or garbles them. Restrict raw byte-splitting to MP3; route every other format through the existing decode+WAV re-encode path. When WAV conversion yields nothing (no Web Audio decoder, or an undecodable codec) throw a clear error rather than byte-splitting a container into corrupt chunks, keeping the run retryable. Fixes deepsec finding other-logic-bug (audioChunker).
formatSpeakerLabel passed the speaker label as the replacement string to String.replace, which interprets special patterns ($$, $&, $`, $', $n). A label containing one of those would be rewritten with matched/surrounding text instead of inserted literally. Use a replacer function so the label is inserted verbatim. Fixes deepsec finding other-logic-bug (diarization segments).
Deploying podnotes with
|
| Latest commit: |
dde9356
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://101feb8a.podnotes.pages.dev |
| Branch Preview URL: | https://chhoumann-deepsec-transcript.podnotes.pages.dev |
github-actions Bot
pushed a commit
that referenced
this pull request
Jul 9, 2026
## [2.17.3](2.17.2...2.17.3) (2026-07-09) ### Bug Fixes * **feed/search:** parse feed once + content-based search cache ([#225](#225)) ([053d51f](053d51f)), closes [#149](#149) * make episode identity key collision-resistant and prototype-safe ([#226](#226)) ([a5683db](a5683db)) * **opml:** correct import progress math and saved-count reporting ([#221](#221)) ([a79e529](a79e529)) * **security:** validate feed/URI URLs and cap download size ([#223](#223)) ([edef281](edef281)) * **template:** neutralize feed-controlled note injection ([#228](#228)) ([ef4ecbd](ef4ecbd)) * **timestamp:** escape live table-cell pipe after an escaped backslash ([#227](#227)) ([a34dfca](a34dfca)) * **transcription:** resolve three deepsec transcription-pipeline bugs ([#224](#224)) ([83c34e7](83c34e7))
Contributor
|
🎉 This PR is included in version 2.17.3 🎉 The release is available on GitHub release Your semantic-release bot 📦🚀 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes three deepsec-confirmed correctness bugs in the transcription pipeline. Each is an independent commit.
1. Failed chunks saved as a "completed" transcript, blocking retry —
other-silent-failuresrc/services/TranscriptionService.tsWhen a chunk exhausted its retries the worker wrote an
[Error transcribing chunk N]placeholder and counted it as completed, andbuildTranscriptBody()only rejected a fully empty body. A transcript made of error placeholders therefore passed the guard and was saved as "Transcription completed and saved." Because the existence check short-circuits on a later attempt ("you've already transcribed this episode"), the user was left with a corrupt/partial transcript - in the worst case a file containing only[Error transcribing chunk 0]- and no straightforward way to re-run.Fix:
transcribeChunks()now reports how many chunks failed.buildTranscriptBody()strips the error placeholders to decide whether any real speech was transcribed. If nothing real remains (every chunk failed, or the only successes were empty) it throws, so no file is written and the episode stays retryable - mirroring the diarization provider's existing all-chunks-failed contract. This also closes a mixed-failure hole (some chunks fail, surviving chunks return empty text) that a simple "all chunks failed" count would miss.2. Byte-splitting container/lossless audio produced undecodable chunks —
other-logic-bug(audioChunker)src/services/audioChunker.tscreateChunkFiles()only routed m4a over the decode+WAV re-encode path; every other format over the 20 MB request limit fell through to raw byte-splitting. MP3 frames are self-synchronizing so that is fine, but container/header-dependent formats (ogg, webm, flac, wav, m4a) carry codec setup only at the stream start, so non-first byte slices are undecodable and the API rejects or garbles them (which then fed the placeholder path in bug 1).Fix: restrict raw byte-splitting to MP3; route every other format through the existing decode+WAV path. When WAV conversion yields nothing (no Web Audio decoder, or an undecodable codec) throw a clear, retryable error instead of byte-splitting a container into corrupt chunks. This also replaces the previous m4a fallback that silently byte-split into corrupt chunks when conversion failed.
3. Speaker label mangled by
String.replace$-patterns —other-logic-bug(diarization segments)src/services/diarization/segments.tsformatSpeakerLabel()passed the provider-supplied speaker label as the replacement string toString.replace, which interprets$$,$&,$`,$',$n. A label containing one of those would be rewritten with matched/surrounding text. Fixed by using a replacer function so the label is inserted verbatim. (Current providers emitA/B/numeric labels, so this is a latent defect, fixed for completeness.)Tests
$-laden speaker label is inserted verbatim.npm run lint,typecheck,buildall pass;npm run testpasses for all changed areas (audioChunker 13, TranscriptionService 14, segments 15). The full suite is green when the machine isn't saturated; under heavy parallel load one pre-existing timing-sensitive test (PodcastView.integration.test.ts, feed-cache, unrelated to these files) can hit its 5000ms timeout - it passes in isolation.Runtime verification
Confirmed in an isolated Obsidian (Electron) e2e vault that
AudioContext.decodeAudioDatais available and decodes audio to a validAudioBuffer- so the non-mp3 → WAV path this PR relies on is reachable at runtime (jsdom has no Web Audio, which is why the unit tests stub it).Notes / out of scope
mp3, so a truly unidentifiable non-mp3 file >20 MB could still be byte-split - pre-existing behavior, now caught by bug 1's retryable failure. A WebM/EBML magic-number entry indetectAudioFileExtensionwould close that gap (follow-up).Does not merge automatically - opening for review.