A self-hosted web app that turns Finnish text into speech. The default engine is Piper (fast, CPU-only, runs natively on Apple Silicon), with Chatterbox-Finnish available as an optional higher-quality (but slower, heavier) engine.
Plenty of "free TTS" websites exist (mostly thin wrappers around commercial APIs, with character limits, ads, or a push toward a paid tier), and even solid options like ElevenLabs cap Finnish at a monthly free quota. This project isn't trying to be "yet another TTS site" - it's meant to be:
- Private: with Piper, your text never leaves this machine. Even with edge-tts, text goes straight to Microsoft, not to some unknown third-party wrapper site.
- No character limits, ads, or accounts - because you're hosting it yourself.
- Embeddable directly into your own site/workflow, not a separate tool you have to go visit.
- Three quality/speed tiers (fast Piper, high-quality Chatterbox, free-cloud Noora) with no vendor lock-in.
./scripts/setup.sh # once - creates a venv, installs dependencies
./scripts/run.sh # starts the serverThen open http://localhost:8000. The default voice (Harri, medium quality) ships with the project, so it works immediately - nothing to download first.
To let others on the same home/office network use it, share the IP address run.sh prints (something like http://192.168.1.23:8000). To reach it from outside your local network, Tailscale or Cloudflare Tunnel are simple options - the machine needs to stay on and awake since it's acting as the server, not an always-on host.
backend/
main.py FastAPI: /api/voices, /api/synthesize
tts_engine.py text-to-sentence chunking + Piper + Chatterbox (optional)
requirements.txt
frontend/
index.html UI (no build step, served directly)
voices/fi_FI/ Piper voice (harri-medium ships with the project)
engines/ clone Chatterbox-Finnish here (optional)
scripts/
setup.sh, run.sh, download_voices.sh
Text you submit gets split into sentences server-side (carefully, so Finnish abbreviations like "esim.", "mm.", and dates like 13.8.2026 aren't mistaken for sentence endings), each piece is synthesized separately, then stitched into one WAV file with a short silence between chunks. This chunking is what lets long text (an article, a book chapter) work without issues.
| Piper (default) | Chatterbox-Finnish (optional) | |
|---|---|---|
| Quality | Good, consistent | Higher, more natural |
| Speed | Very fast (RTF ≈ 0.25 on a typical CPU) | Slower, especially without CUDA |
| Hardware | CPU only | GPU preferred (CUDA or MPS) |
| Setup | Ships ready to go | Manual, a few steps (see below) |
| Voice license | CC0 (fully free) | MIT |
For "a shared tool that also handles long text well," Piper is the more sensible default: multiple people can request audio at once without anyone waiting on a GPU. Add Chatterbox as a "high quality" mode for when you personally want the best possible rendering of something that matters (e.g. a book chapter).
synthesize_chatterbox is written to match the exact loading pattern published in the official model card, but expect to test/debug it on your first real run.
# 1. Clone the repo (large, uses Git LFS)
git lfs install
git clone https://huggingface.co/Finnish-NLP/Chatterbox-Finnish engines/chatterbox-finnish
# 2. Install the extra dependencies (torch is heavy)
source .venv/bin/activate
pip install torch torchaudio safetensors
# get the exact torch install command (CPU/MPS) from pytorch.org
# 3. Per that repo's own README, make sure engines/chatterbox-finnish/pretrained_models
# and engines/chatterbox-finnish/models/*.safetensors exist (it may have its own
# setup script - check its README)After this, /api/voices should report chatterbox-finnish as available: true and it'll be selectable in the UI.
Important notes:
- This model needs a short reference audio file (
audio_prompt_path) for voice cloning; the current code looks for a.wavfile underengines/chatterbox-finnish/samples/. If it can't find one, add your own sample or pass a path viareference_audioin the API request. - macOS/Apple Silicon: the Chatterbox package officially supports the
mpsdevice, but user reports show some operators (e.g.aten::_fft_r2c) aren't implemented on MPS yet. The current code automatically retries once on CPU if MPS generation errors out (slower, but works). - If you still hit import errors, the exact path in
src/chatterbox_/tts.pymay differ slightly between repo versions - compare against the actual files you cloned. - Memory: this model is large and, on Apple Silicon, shares system RAM with everything else (no dedicated VRAM). On a 16GB Mac, even a single short sentence can push into swap - treat this engine as an occasional, short-passage tool, not a daily driver. The app enforces a low character cap on this engine for exactly this reason.
- This project's code: MIT - see LICENSE.
- Piper: MIT. Harri voice: CC0 (fully free, commercial use included).
- Chatterbox (base): MIT. Chatterbox-Finnish: MIT.
- edge-tts: MIT (it calls a free, unofficial Microsoft service - see below).
- If you ever add the
asmovoice (F5-TTS based): that one is CC BY-NC 4.0, non-commercial only.
This covers running the app behind an nginx location block on a domain/subdomain you already manage, without creating a new server block or SSL certificate. The values below are illustrative - substitute your own domain, port, and path throughout; deploy/nginx-finnish-tts.conf and deploy/finnish-tts.service use explicit <PLACEHOLDER> markers for exactly this reason.
No GPU needed - Piper is CPU-only by design. On a GPU-less VPS, skip Chatterbox-Finnish and stick with Piper (and edge-tts's Noora, which is also lightweight).
# On the VPS (Ubuntu/Debian assumed):
sudo apt update && sudo apt install -y python3-venv python3-pip
# Transfer the project via scp or git, then:
cd finnish-tts-app
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install -r backend/requirements.txtPick a port. Check what's already listening before choosing one - especially on a server that already runs other services:
sudo ss -tlnp | grep ":8000" # replace 8000 with whatever port you're consideringUse whatever comes back free, and set it consistently in the systemd service and the nginx block below.
Make the service persistent with systemd (survives SSH disconnects and reboots):
# Edit <USER>, <PATH_TO_PROJECT>, and <PORT> in deploy/finnish-tts.service first
sudo cp deploy/finnish-tts.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now finnish-tts
journalctl -u finnish-tts -f # watch the logsThe app listens on 127.0.0.1:<your port> only, local, not reachable from outside. To embed it in a div on an existing subdomain without touching that subdomain's other paths, see deploy/nginx-finnish-tts.conf - it's two steps: one limit_req_zone line in http{} (rate limiting, once for all of nginx), and a location block inside that subdomain's existing server file, pointing at your chosen port. Before and after applying:
sudo nginx -t # check syntax before reloading
sudo systemctl reload nginx
# then re-check that the domain's other existing paths still load exactly as beforeAnd on your own page:
<iframe src="https://your-domain.example/your-path/"
style="width:100%; height:640px; border:0;" loading="lazy"></iframe>Whatever port you pick doesn't need to be opened in any firewall - only nginx (locally) talks to it, never the outside world.
Security: rate limiting (6 requests/minute per IP, adjust to taste) is set directly in the nginx config, no extra Python dependency needed.
Restricting which voices are exposed: set the ENABLED_VOICES environment variable in the systemd service (comma-separated voice ids, e.g. edge-noora) if a given deployment should only offer specific voices - useful for a public-facing instance where you'd rather not expose every option. Leave it unset to expose everything (the default).
- "Voice model missing": make sure
voices/fi_FI/fi_FI-harri-medium.onnxexists; if not, run./scripts/download_voices.sh fi_FI harri medium. - Port already in use: pick a different one, e.g.
PORT=8001 ./scripts/run.sh. - Voice sounds robotic/off: try the
harri-lowvoice for comparison; if Chatterbox is set up, compare against that too.