Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 34 additions & 0 deletions docs/docs/transcripts.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,3 +24,37 @@ To create a transcript:
## Transcript Template

The transcript template works similarly to the [note template](./templates.md#note-template), but with the added `{{template}}` placeholder.

## Speaker Diarization

By default the transcription uses OpenAI's Whisper model, which produces plain text with **no speaker labels**. Speaker diarization is an opt-in setting that instead labels each segment of the transcript by speaker, e.g.:

```
**A:** Welcome to the show.

**B:** Thanks for having me.
```

### Enabling it

In the **Transcript settings** section, turn on **Speaker diarization** and choose a provider:

- **OpenAI** (`gpt-4o-transcribe-diarize`): reuses the OpenAI API key you already entered above, so there is nothing else to configure. Because OpenAI caps each request at 25 MB, a long episode is split into chunks that are diarized independently — so on long episodes the speaker labels can change across chunk boundaries (the same person may be labelled `A` in one chunk and `B` in the next). A typical-length episode fits in a single request and is fully consistent.
- **Deepgram**: sends the whole episode in one request, so speaker labels stay consistent across the entire episode. This requires a separate **Deepgram API key**, which you can create at [deepgram.com](https://deepgram.com) (new accounts include free credit). Your Deepgram key is stored separately from your OpenAI key and is only used for diarization.

Diarization is off by default, so existing transcripts and the plain-Whisper workflow are unchanged unless you enable it.

### Speaker label format

The **Speaker label format** setting controls the prefix added before each speaker's turn. Use the `{{speaker}}` placeholder for the speaker's label:

- OpenAI labels speakers `A`, `B`, `C`, …
- Deepgram labels speakers `1`, `2`, `3`, …

The default is `**{{speaker}}:** `, which renders as `**A:**`-style bold prefixes. To spell out the word "Speaker", set it to `**Speaker {{speaker}}:** ` (rendering `**Speaker A:**`); `> {{speaker}}: ` would instead put each turn in a blockquote.

The labelled transcript replaces the usual `{{transcript}}` value in your [transcript template](#transcript-template), so you don't need to change your template to use diarization.

### Cost

Diarization providers bill per minute/hour of audio (separately from any plain-Whisper usage). As of mid-2026, OpenAI's diarize model is roughly $0.006 per minute, and Deepgram's diarized pre-recorded transcription is roughly $0.0068 per minute. Check each provider's current pricing before transcribing long back-catalogues.
6 changes: 6 additions & 0 deletions scripts/provision-obsidian-e2e-vault.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -99,10 +99,16 @@ export const DEFAULT_PODNOTES_DATA = {
episodes: [],
},
openAIApiKey: "",
diarizationApiKey: "",
transcript: {
path: "transcripts/{{podcast}}/{{title}}.md",
template:
"# {{title}}\n\nPodcast: {{podcast}}\nDate: {{date}}\n\n{{transcript}}",
diarization: {
enabled: false,
provider: "openai",
speakerTemplate: "**{{speaker}}:** ",
},
},
feedCache: {
enabled: true,
Expand Down
9 changes: 9 additions & 0 deletions src/constants.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
import type { IPodNotesSettings } from "src/types/IPodNotesSettings";
import type { Playlist } from "./types/Playlist";
import { DEFAULT_SPEAKER_TEMPLATE } from "./services/diarization/segments";

export const VIEW_TYPE = "podcast_player_view";

Expand Down Expand Up @@ -154,10 +155,18 @@ export const DEFAULT_SETTINGS: IPodNotesSettings = {
episodes: [],
},
openAIApiKey: "",
diarizationApiKey: "",
transcript: {
path: "transcripts/{{podcast}}/{{title}}.md",
template:
"# {{title}}\n\nPodcast: {{podcast}}\nDate: {{date}}\n\n{{transcript}}",
// Diarization is off by default so existing behaviour (plain Whisper) is
// unchanged; enabling it routes audio to the chosen provider (#168).
diarization: {
enabled: false,
provider: "openai",
speakerTemplate: DEFAULT_SPEAKER_TEMPLATE,
},
},
feedCache: {
enabled: true,
Expand Down
11 changes: 10 additions & 1 deletion src/main.ts
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,9 @@ import { DEFAULT_SETTINGS, VIEW_TYPE } from "src/constants";
import {
migrateDownloadPath,
migrateNoteSettings,
migrateTranscriptSettings,
} from "src/settingsMigrations";
import { requiredTranscriptionKeyPresent } from "src/services/diarization";
import { PodNotesSettingsTab } from "src/ui/settings/PodNotesSettingsTab";
import { MainView } from "src/ui/PodcastView";
import { QueueReorderModal } from "src/ui/QueueReorderModal";
Expand Down Expand Up @@ -482,7 +484,8 @@ export default class PodNotes extends Plugin implements IPodNotes {
name: "Transcribe current episode",
checkCallback: (checking) => {
const canTranscribe =
!!this.api.podcast && !!this.settings.openAIApiKey?.trim();
!!this.api.podcast &&
requiredTranscriptionKeyPresent(this.settings);

if (checking) {
return canTranscribe;
Expand Down Expand Up @@ -721,6 +724,12 @@ export default class PodNotes extends Plugin implements IPodNotes {
// default, preserving any path/template the user configured (#160). Returns
// a fresh object, so DEFAULT_SETTINGS.note is never mutated.
this.settings.note = migrateNoteSettings(loadedData?.note);
// Backfill the diarization defaults onto the stored transcript object so an
// existing user (who has only { path, template } persisted) gets a valid
// transcript.diarization instead of undefined (#168).
this.settings.transcript = migrateTranscriptSettings(
loadedData?.transcript,
);
}

async saveSettings() {
Expand Down
45 changes: 45 additions & 0 deletions src/services/TranscriptionService.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -300,4 +300,49 @@ describe("TranscriptionService", () => {
await service.transcribeCurrentEpisode();
});
});

describe("createChunkFiles small-file fast path (#168 / PR #204 review)", () => {
type ChunkFn = (args: {
buffer: ArrayBuffer;
basename: string;
extension: string;
mimeType: string;
}) => Promise<File[]>;

function createChunkFiles(): ChunkFn {
const service = new TranscriptionService(createMockPlugin());
return (service as unknown as { createChunkFiles: ChunkFn })
.createChunkFiles.bind(service);
}

test("sends a small m4a as a single original file instead of WAV-splitting it", async () => {
// A 1 KB m4a is well under the 20 MB chunk size, so it must not be
// converted to WAV (which would balloon it into many chunks and reset
// diarization speaker labels). It should be one intact .m4a file.
const files = await createChunkFiles()({
buffer: new ArrayBuffer(1024),
basename: "episode",
extension: "m4a",
mimeType: "audio/mp4",
});

expect(files).toHaveLength(1);
expect(files[0].name).toBe("episode.m4a");
expect(files[0].type).toBe("audio/mp4");
expect(files[0].name).not.toContain(".wav");
});

test("still sends a small mp3 as a single original file (unchanged)", async () => {
const files = await createChunkFiles()({
buffer: new ArrayBuffer(2048),
basename: "episode",
extension: "mp3",
mimeType: "audio/mpeg",
});

expect(files).toHaveLength(1);
expect(files[0].name).toBe("episode.mp3");
expect(files[0].size).toBe(2048);
});
});
});
125 changes: 109 additions & 16 deletions src/services/TranscriptionService.ts
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,15 @@ import {
} from "../utility/enforceMaxPathLength";
import { ensureFolderExists } from "../utility/ensureFolderExists";
import type { Episode } from "src/types/Episode";
import {
type DiarizationAudio,
type DiarizationProviderId,
type DiarizedSegment,
diarizeWithDeepgram,
diarizeWithOpenAI,
renderDiarizedTranscript,
requiredTranscriptionKeyPresent,
} from "./diarization";

function TimerNotice(heading: string, initialMessage: string) {
let currentMessage = initialMessage;
Expand Down Expand Up @@ -71,9 +80,14 @@ export class TranscriptionService {
}

async transcribeCurrentEpisode(): Promise<void> {
if (!this.plugin.settings.openAIApiKey?.trim()) {
if (!requiredTranscriptionKeyPresent(this.plugin.settings)) {
const diarization = this.plugin.settings.transcript.diarization;
const needsDeepgram =
diarization?.enabled && diarization.provider === "deepgram";
new Notice(
"Please add your OpenAI API key in the transcript settings first.",
needsDeepgram
? "Please add your Deepgram API key in the transcript settings to use Deepgram diarization."
: "Please add your OpenAI API key in the transcript settings first.",
);
return;
}
Expand Down Expand Up @@ -162,19 +176,18 @@ export class TranscriptionService {
} = await getEpisodeAudioBuffer(episode);
const mimeType = this.getMimeType(fileExtension);

notice.update("Creating audio chunks...");
const files = await this.createChunkFiles({
buffer: fileBuffer,
basename,
extension: fileExtension,
mimeType,
});

notice.update("Starting transcription...");
const transcription = await this.transcribeChunks(files, notice.update);
const transcriptBody = await this.buildTranscriptBody(
{
buffer: fileBuffer,
mimeType,
extension: fileExtension,
basename,
},
notice.update,
);

notice.update("Saving transcription...");
await this.saveTranscription(episode, transcription);
await this.saveTranscription(episode, transcriptBody);

notice.stop();
notice.update("Transcription completed and saved.");
Expand All @@ -189,6 +202,71 @@ export class TranscriptionService {
}
}

/**
* Produce the transcript body that fills the template's `{{transcript}}` tag.
* Plain Whisper returns one run-on block, so it is reflowed after sentence
* periods for readability; diarization returns speaker-labeled turns already
* separated into paragraphs, so it is rendered as-is. Diarization that yields
* no speech is treated as a failure rather than writing an empty transcript.
*/
private async buildTranscriptBody(
audio: DiarizationAudio,
updateNotice: (message: string) => void,
): Promise<string> {
const diarization = this.plugin.settings.transcript.diarization;

if (diarization?.enabled) {
const segments = await this.diarize(
audio,
diarization.provider,
updateNotice,
);
if (segments.length === 0) {
throw new Error("Diarization returned no speech segments.");
}
return renderDiarizedTranscript(segments, diarization.speakerTemplate);
}

updateNotice("Creating audio chunks...");
const files = await this.createChunkFiles(audio);
updateNotice("Starting transcription...");
const transcription = await this.transcribeChunks(files, updateNotice);
return transcription.replace(/\.\s+/g, ".\n\n");
}

/** Route the episode audio to the configured diarization provider (#168). */
private async diarize(
audio: DiarizationAudio,
provider: DiarizationProviderId,
updateNotice: (message: string) => void,
): Promise<DiarizedSegment[]> {
if (provider === "deepgram") {
const apiKey = this.plugin.settings.diarizationApiKey?.trim();
if (!apiKey) {
throw new Error("Missing Deepgram API key for diarization.");
}
// Deepgram ingests the whole file in one request, so it needs no
// chunking — which is exactly why its speaker labels stay consistent
// across the entire episode.
return diarizeWithDeepgram({
audio,
apiKey,
onProgress: updateNotice,
});
}

// OpenAI diarization shares Whisper's 25 MB/request cap, so reuse the same
// chunking. Speaker labels can differ across chunks on a long episode.
updateNotice("Creating audio chunks...");
const chunkFiles = await this.createChunkFiles(audio);
Comment thread
chhoumann marked this conversation as resolved.
return diarizeWithOpenAI({
getClient: () => this.getClient(),
chunkFiles,
maxRetries: this.MAX_RETRIES,
onProgress: updateNotice,
});
}

private async createChunkFiles({
buffer,
basename,
Expand All @@ -200,6 +278,19 @@ export class TranscriptionService {
extension: string;
mimeType: string;
}): Promise<File[]> {
// A file that already fits in a single request needs neither conversion nor
// splitting: send the original (compressed) bytes. OpenAI accepts m4a/mp3/etc.
// directly under the upload limit, so this skips the m4a->WAV path below for
// small m4a episodes — that path only exists to SAFELY SPLIT an m4a (which
// can't be byte-split) and would otherwise balloon a small m4a into many
// uncompressed WAV chunks, multiplying requests and, for diarization,
// resetting speaker labels at artificial chunk boundaries (#168 / PR #204 review).
if (buffer.byteLength <= this.CHUNK_SIZE_BYTES) {
return [
new File([buffer], `${basename}.${extension}`, { type: mimeType }),
];
}

if (this.shouldConvertToWav(extension, mimeType)) {
const wavChunks = await this.convertToWavChunks(buffer, basename);
if (wavChunks.length > 0) {
Expand Down Expand Up @@ -486,14 +577,16 @@ export class TranscriptionService {

private async saveTranscription(
episode: Episode,
transcription: string,
transcriptBody: string,
): Promise<void> {
const transcriptPath = this.getTranscriptPath(episode);
const formattedTranscription = transcription.replace(/\.\s+/g, ".\n\n");
// transcriptBody is already formatted by buildTranscriptBody (sentence
// reflow for Whisper, speaker turns for diarization), so it is templated
// verbatim here.
const transcriptContent = TranscriptTemplateEngine(
this.plugin.settings.transcript.template,
episode,
formattedTranscription,
transcriptBody,
);

const vault = this.plugin.app.vault;
Expand Down
Loading
Loading