Reusable Hebrew text normalization, aligned-audio segmentation, quality filtering, and
dataset auditing tools. The Pocket TTS profiles are designed for CrowdRecital recordings
containing audio.wav, transcript.aligned.json, and metadata.json per recording.
The repository contains code and configuration only. Raw audio, prepared datasets, cached parts, and model artifacts are intentionally ignored and must be transferred separately.
- Python 3.11 or 3.12
- ffmpeg available on
PATH - enough disk for the prepared audio (the 38.9-hour CrowdRecital build measured about 6.7 GB)
Linux/server setup:
git clone https://github.com/asaelbarilan/hebrew-tts-data-tools.git
cd hebrew-tts-data-tools
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
ffmpeg -versionWindows setup:
py -3.12 -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
ffmpeg -versionFor development and tests, install requirements-dev.txt instead.
Two checked-in profiles are provided:
configs/prepare_pocket_tts_hebrew.yaml: proven student profile, 4-16 seconds with an approximately 8-second target.configs/prepare_pocket_tts_hebrew_teacher.yaml: experimental teacher profile, 4-40 seconds with an approximately 23-second measured mean. Training must use--head-samples 500for full coverage of its longest clips.
Set portable paths and run a representative pilot before the full build:
export CROWD_RECITAL_ROOT=/data/crowd-recital
export POCKET_TTS_PREPARED_DIR=/data/prepared/CrowdRecital_pockettts_8s
python data_prep/prepare_ivritai.py \
--config configs/prepare_pocket_tts_hebrew.yaml \
--max_source_entries 50 \
--source_selection spreadInspect the pilot, then remove --max_source_entries for the full build:
python data_prep/prepare_ivritai.py \
--config configs/prepare_pocket_tts_hebrew.yamlThe preparation job saves each completed chunk beside the output in a _parts directory.
Rerunning the identical command reuses completed chunks after an interruption. Do not reuse
the parts directory after changing the source corpus or segmentation profile; choose a new
output path instead.
For the teacher profile:
export POCKET_TTS_PREPARED_DIR=/data/prepared/CrowdRecital_pockettts_teacher_23s
python data_prep/prepare_ivritai.py \
--config configs/prepare_pocket_tts_hebrew_teacher.yaml \
--max_source_entries 50 \
--source_selection spread- aligned slicing on transcript segment boundaries;
- minimum session/sample and segment quality of 0.6;
- Hebrew number/text normalization after segmentation;
- removal of wiki
==markers while retaining spoken heading words; - removal of normalized rows above a 0.15 Latin-letter ratio;
- propagation of
user_idas the speaker identity andbiological_sexfor auditing; - no internal random validation split. The Pocket TTS repository creates a deterministic, speaker-disjoint split and different-clip/same-speaker prompt pairs afterward.
source_id is constant (recital) in this corpus. Never use it as the train/validation
grouping key; use user_id.
Run regression tests:
python -m pytest tests -qListen to prepared samples and compare normalized/raw text:
python ui/output_samples_app.py --data_path "$POCKET_TTS_PREPARED_DIR"The UI prints its localhost URL. On a remote server, forward that port over SSH or audit a local copy of the dataset.
Quick normalizer check:
python scripts/smoke_test_normalizer.pyPocket TTS needs to know when each word was said, not only what was said. From its
training/README.md:
it's not enough to have pairs of
(speech audio, transcript), we also need to know when exactly each word was said. This is because for training, the audio is cut into smaller parts, and we need to know exactly what text corresponds to what part of the audio.
The trainer does run without word timings -- it falls back to a random window for the voice prompt -- but it then loses the trailing-silence trim, which Kyutai identify as the reason a model "emits silence instead of EOS, so generations never terminate".
CrowdRecital ships timings in transcript.aligned.json. Most other corpora do not, so align
them yourself:
# Confirm a candidate model is a Hebrew char-level CTC head before spending GPU time
PYTHONPATH=/path/to/pocket-tts python -m data_prep.align_hebrew in.jsonl out.jsonl --check-model
# Align (input/output are Pocket TTS manifests)
PYTHONPATH=/path/to/pocket-tts python -m data_prep.align_hebrew manifest.jsonl manifest_aligned.jsonl --model imvladikon/wav2vec2-xls-r-300m-hebrew --device cuda --batch-size 8--resume appends and skips utterances already aligned, so preemptible jobs continue where
they stopped.
Wav2Vec2 is multilingual (XLS-R); only the CTC head and its 27-character vocabulary are Hebrew-specific. Measured against CrowdRecital's own timings, 40 utterances / 316 words:
| boundary | median | p90 |
|---|---|---|
| word end | 19 ms | 71 ms |
| word start | 74 ms | 158 ms |
One frame at Pocket TTS's 12.5 Hz is 80 ms, so word ends land within a quarter of a frame.
Ends are the accurate side, which is fortunate: last_word_end is what drives the silence
trim. 5.7% of words returned no timing.
Caveat: CrowdRecital's timings are themselves automatic, so this measures agreement between two aligners, not accuracy against hand labels.
This tool produces the segmented Hugging Face dataset. Continue with the training repository:
git clone https://github.com/asaelbarilan/pocket-tts.gitFollow HEBREW_SERVER_RUNBOOK.md there to create speaker-disjoint manifests, train the Hebrew
SentencePiece model, precompute model-specific Mimi latents, and launch either the 6-layer
student adaptation or an experimental 24-layer teacher adaptation.