Motivation
I've received a few questions about how to adjust VAD (Voice Activity Detection) sensitivity for different environments and use-cases. Right now, there's no documentation or guide for tuning these parameters, which makes it hard for users to get reliable results, especially with background noise or low-volume speakers. A clear guide would help users adapt VAD behavior to their needs.
Proposed approach
- Document the main VAD sensitivity parameters, with descriptions and expected value ranges.
- Provide sample configs for common scenarios (quiet room, noisy office, etc).
- Include troubleshooting tips for common issues like false triggers or missed speech.
Motivation
I've received a few questions about how to adjust VAD (Voice Activity Detection) sensitivity for different environments and use-cases. Right now, there's no documentation or guide for tuning these parameters, which makes it hard for users to get reliable results, especially with background noise or low-volume speakers. A clear guide would help users adapt VAD behavior to their needs.
Proposed approach