(Work in Progress)
Here is a collection of my own reading in speech recognition and related topics.
I often track -
The five most important ones which everyone should read -
-
Unsupervised Cross-lingual Representation Learning for Speech Recognition (Conneau et al., 2020).
-
Self-training and Pre-training are Complementary for Speech Recognition (Xu et al., 2020).
-
Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training (Hsu, et al., 2021).
-
Simple and Effective Zero-shot Cross-lingual Phoneme Recognition (Xu et al., 2021).
The models
Later development
-
XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale (The repo)
-
Seamless M4T: Massively Multilingual & Multimodal Machine Translation
Language model integration
-
huggingsound One of the great works from Prof. Jonatas Grosman.
-
Directly using the Wav2Vec2ProcessorWithLM class from Patrick Von Platen.
-
Training KenLM You should also read Kenneth Heafield's website, it has a wealth of information.
SFT
Work in Progress.
The original paper
Open Whisper-style Speech Model (OWSM)
- Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data (Peng, 2023) - An impressive effort from CMU WavLab to replicate Whisper enc-dec style training.
- OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer (Peng, 2024)
Other Whisper-style models:
- SenseVoice Page supports Mandarin, Cantonese, Japanese and Korean.
Interesting Variants:
- Grave's CTC paper,his thesis with better derivation. You may also want to look at this explainer
- Grave's RNNT paper - it's tough to explain the idea well though.
- AED - perhaps LAS but modern systems often jointly train CTC and AED.
- LALM - Many come to mind. Perhaps AudioPALM comes to mind first.
- A great survey paper would be End-to-End Speech Recognition: A Survey
- Open SLR Collections of free speech dataset.
- LDC OG of language resources. Fees for non-members can be hefty.
Multilingual
- Babel (Where is it now?)
- Voxpopuli Collected from 2009-2020 European Parliament event recordings. >400k hours of data. (Paper)
- Multilingual LibriSpeech (MLS) A multilingual version of libri-light. It's still heavily tilted towards English, but it also contains significant amount of German, Spanish and 6 other languages. (Paper)
- Common Voice A multilingual dataset. When you test on CV, remember that there are multiple versions of the dataset. On HuggingFace, also know that some of these databases are gated. (i.e., requires login)
- FLEURS Standard multilingual dataset for ASR and LID purpose.
English-only
- Librispeech One of the golden benchmarks in ASR. (Paper)
- Libri-Light In a sense, it is the extension of Librispeech but with 60k hour of unlabelled data. wav2vec2's models prefixed with -Lv60 are speech representation, for example, are all pre-trained with this dataset. (Paper)
- LibriSpeech-PC LibriSpeech with punctuation and capitalization (PC) (Paper)
- Libri-Heavy Labeled version of libri-light also annotated with punctuation and context(Paper) All segments are short (<20s). The group also release a version with long duration called libriheavy-long.
Portuguese
- NURC-SP 239.30 hours of audio from the NURC project (Paper) You may also find the information from TaRSilla Project and the original NURC project useful
- CORAA 290.77 hours of spontaneous speech. (Paper)
- The BP Dataset 400 hours. (Paper)
Cantonese
- CommonVoice and MLS both have subsets on Cantonese
- MDCC Dataset (The Paper is also a survey on different Cantonese dataset.)
Vietnamese
- CommonVoice dataset contains a Vietnamese portion
- VIVOS
- FOSD - Vietnamese
- VLSP
Other awesome lists:
Hugging Face: (Note: You often need to hack the code to get it working.)
- ISCA Archive If you want to search for all Interspeech conference papers. (Or Eurospeech/ICSLP if you still remember them...)
- ICASSP Another OG yearly conference on speech. Sad. Not all the papers are archived. So you may need to try your luck to see if the authors put them on archive.
The two are specific to speech. These days peeps love to publish on AAAI and NeurIPS. You know where to find them already.
- IEEE/ACM Transactions on Audio, Speech, and Language Processing It used to be called IEEE Transactions on Audio, Speech, Processing, or "ASP" in short. Now it's called "ASLP" I guess.
Cool TTS links
Important techniques (unsorted)
Great Explainers
- MCMC without all the BS
- A (Long) Peek into Reinforcement Learning by Lilian Weng
- Policy Gradient Algorithms
- Flow Matching
- Gumbel Softmax
- Optimal Transport by Alex Williams
- CTC by Awni Hannun
Matrix Calculus If you are never confused about the gradient derivation in our field, you probably haven't looked deep enough into the math...
- Symbolic Matrix Differentiation Solver
- Matrix Cookbook
- Matrix-based Approaches which Prof. Magnus gave a great expose. You might also need the reverse of vectorization to get the Math right.
- Robot Chinwag's Guide on Tensor Calculus If you feel the matrix-based approach is deeply dissatisfying...
- Chinese - 中央研究院現代漢語標記語料庫
- Cantonese
See LLM.md.