Skip to content

Repository files navigation

Call Speaker Diarization Kit

A reference pipeline for a common small-business phone problem: several people share one phone line or one main business number, and you need to know who was actually talking on a given call: for QA, coaching, CRM logging, dispute resolution, or just knowing which rep to follow up with.

Generalized from a client engagement. Case study: https://patrickgibbs.dev/work/case/speaker-id/

This repo demonstrates the end-to-end pattern:

  1. Enroll each known team member's voice once, from a short reference clip, into a "voiceprint" (a speaker embedding vector).
  2. Identify the speaker(s) on a new call recording by comparing its embedding against every enrolled voiceprint with cosine similarity, and refusing to guess when nothing matches confidently.
  3. Transcribe and tag the call, merging speech-to-text output with speaker identity into one structured, queryable JSON record per call.

It is written to be read end-to-end as engineering documentation of a working pattern, not just API glue. The similarity math, thresholding logic, and edge-case handling (ambiguous match, no match, low confidence) are real and runnable. The actual neural network calls (the speaker-embedding model and the transcription model) are mocked/stubbed with clear markers, so the repo runs with zero downloads and zero GPU. Swap in real model calls at the two marked integration points and the rest of the pipeline is unchanged.

Why this pattern, not full diarization-from-scratch

Generic diarization ("cluster the audio into N anonymous speakers") is a different, harder problem than what most small businesses actually need. Most of the time the set of possible speakers is small and known in advance (the 3-8 people who work the phones). That turns an open-ended clustering problem into a much more tractable 1:N identification-against-enrolled-set problem: closer to speaker verification than to blind diarization. This kit targets that common, cheaper, more reliable case.

Pipeline overview

 reference clips                    new call recording
 (one per known speaker)                    │
        │                                   │
        ▼                                   ▼
 ┌──────────────────┐              ┌──────────────────────┐
 │  enrollment/      │              │  identify/            │
 │  enroll_speaker.py│              │  identify_speaker.py  │
 │                   │              │                       │
 │  audio → embedding│              │  audio → embedding    │
 │  → voiceprint.json│              │  → cosine similarity  │
 └────────┬──────────┘              │  vs every voiceprint  │
          │                         │  → best match + score │
          ▼                         └──────────┬────────────┘
   voiceprints/*.json                           │
          ▲                                     │
          └─────────────loaded by───────────────┘
                                                 ▼
                                   ┌───────────────────────────┐
                                   │  transcribe-and-tag/       │
                                   │  transcribe_and_tag.py     │
                                   │                             │
                                   │  faster-whisper transcript  │
                                   │  segments + per-segment     │
                                   │  speaker identification     │
                                   │  → structured call JSON     │
                                   └───────────────────────────┘

Repo layout

call-speaker-diarization-kit/
├── README.md                  ← you are here
├── ACCURACY-NOTES.md           ← threshold tuning, real-world failure modes
├── LICENSE                     ← MIT
├── requirements.txt
├── .gitignore
├── enrollment/
│   └── enroll_speaker.py       ← reference clip -> voiceprint
├── identify/
│   ├── identify_speaker.py     ← call audio -> best-match speaker + confidence
│   └── demo_identify.py        ← runnable mocked-embedding demo (no audio needed)
├── transcribe-and-tag/
│   └── transcribe_and_tag.py   ← faster-whisper + identify -> structured JSON
└── voiceprints/                ← enrolled speaker voiceprints land here (gitignored data, .gitkeep only)

Quickstart

pip install -r requirements.txt

# 1. Enroll each known speaker from a short (10-30s), clean reference clip
python enrollment/enroll_speaker.py --speaker-id alex   --audio samples/alex_reference.wav
python enrollment/enroll_speaker.py --speaker-id jordan --audio samples/jordan_reference.wav

# 2. Run the no-audio-required mocked demo to see the matching + threshold logic
python identify/demo_identify.py

# 3. Identify the speaker on a real call recording
python identify/identify_speaker.py --audio samples/incoming_call.wav

# 4. Transcribe a call and tag each segment with a speaker identity
python transcribe-and-tag/transcribe_and_tag.py --audio samples/incoming_call.wav

What's real vs. what's mocked

Component Status in this repo
Cosine similarity matching Real, runnable, unit-testable
Confidence thresholding + "no confident match" handling Real, runnable
Voiceprint storage format (JSON) Real
Speaker embedding model call (Resemblyzer/SpeechBrain-style) Mocked: clearly marked # MOCK: block; swap in real model inference here
faster-whisper transcription call Mocked: clearly marked # MOCK: block; swap in real WhisperModel(...).transcribe(...) here
Structured per-call JSON output shape Real, matches what a production version would emit

Swapping in real models

  • Speaker embeddings: replace the embed_audio() stub in enrollment/enroll_speaker.py and identify/identify_speaker.py with a call to a pretrained speaker-embedding model, e.g. Resemblyzer's VoiceEncoder().embed_utterance(wav), or a SpeechBrain ECAPA-TDNN speaker recognition model. Both return a fixed-length float vector per utterance; everything downstream (storage, cosine similarity, thresholding) is embedding-model-agnostic as long as you enroll and identify with the same model and vector dimension.
  • Transcription: replace the transcribe_audio() stub in transcribe-and-tag/transcribe_and_tag.py with faster-whisper's WhisperModel("base").transcribe(audio_path, word_timestamps=True), which already returns segment-level timestamps in the same shape this repo mocks.

Tests

There is no test framework in this repo. What exists is a runnable demo that asserts its own three scenarios:

python identify/demo_identify.py

It runs with a plain Python 3.10+ install and no third party packages (requirements.txt ships every dependency commented out on purpose), and it exits non-zero if the threshold logic stops behaving as documented. Verified passing: all three scenarios (confident match, ambiguous match, no confident match) behaved as expected.

Known limits

  • DEFAULT_CONFIDENCE_THRESHOLD = 0.75 is a demonstration value, not a validated production threshold. See ACCURACY-NOTES.md for what moves it: background noise, phone codec compression, and whether enrollment clips and call audio travel the same path.
  • The speaker embedding and transcription model calls are mocked, so accuracy numbers are not available from this repo. The matching, thresholding, and output-shape logic around them is real.
  • This is 1:N identification against a known enrolled set, not blind diarization. It answers "which of our people is this", not "how many unknown speakers were on this call".
  • Enrollment and identification must use the same embedding model and vector dimension. Switching models invalidates existing voiceprints.

License

MIT, see LICENSE.

About

Speaker identification on a shared phone line, with confidence thresholds. Generalized from a client engagement.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages