Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Suomeksi ääneen — Finnish Text-to-Speech Reader

فارسی →

A self-hosted web app that turns Finnish text into speech. The default engine is Piper (fast, CPU-only, runs natively on Apple Silicon), with Chatterbox-Finnish available as an optional higher-quality (but slower, heavier) engine.

Why this, not just another TTS website?

Plenty of "free TTS" websites exist (mostly thin wrappers around commercial APIs, with character limits, ads, or a push toward a paid tier), and even solid options like ElevenLabs cap Finnish at a monthly free quota. This project isn't trying to be "yet another TTS site" - it's meant to be:

  • Private: with Piper, your text never leaves this machine. Even with edge-tts, text goes straight to Microsoft, not to some unknown third-party wrapper site.
  • No character limits, ads, or accounts - because you're hosting it yourself.
  • Embeddable directly into your own site/workflow, not a separate tool you have to go visit.
  • Three quality/speed tiers (fast Piper, high-quality Chatterbox, free-cloud Noora) with no vendor lock-in.

Quick start

./scripts/setup.sh   # once - creates a venv, installs dependencies
./scripts/run.sh      # starts the server

Then open http://localhost:8000. The default voice (Harri, medium quality) ships with the project, so it works immediately - nothing to download first.

To let others on the same home/office network use it, share the IP address run.sh prints (something like http://192.168.1.23:8000). To reach it from outside your local network, Tailscale or Cloudflare Tunnel are simple options - the machine needs to stay on and awake since it's acting as the server, not an always-on host.

Project structure

backend/
  main.py          FastAPI: /api/voices, /api/synthesize
  tts_engine.py     text-to-sentence chunking + Piper + Chatterbox (optional)
  requirements.txt
frontend/
  index.html        UI (no build step, served directly)
voices/fi_FI/        Piper voice (harri-medium ships with the project)
engines/              clone Chatterbox-Finnish here (optional)
scripts/
  setup.sh, run.sh, download_voices.sh

How it works

Text you submit gets split into sentences server-side (carefully, so Finnish abbreviations like "esim.", "mm.", and dates like 13.8.2026 aren't mistaken for sentence endings), each piece is synthesized separately, then stitched into one WAV file with a short silence between chunks. This chunking is what lets long text (an article, a book chapter) work without issues.

Two engines

Piper (default) Chatterbox-Finnish (optional)
Quality Good, consistent Higher, more natural
Speed Very fast (RTF ≈ 0.25 on a typical CPU) Slower, especially without CUDA
Hardware CPU only GPU preferred (CUDA or MPS)
Setup Ships ready to go Manual, a few steps (see below)
Voice license CC0 (fully free) MIT

For "a shared tool that also handles long text well," Piper is the more sensible default: multiple people can request audio at once without anyone waiting on a GPU. Add Chatterbox as a "high quality" mode for when you personally want the best possible rendering of something that matters (e.g. a book chapter).

Optional: enabling Chatterbox-Finnish

⚠️ I couldn't test this part in my own sandbox (no GPU, no Hugging Face access there), so synthesize_chatterbox is written to match the exact loading pattern published in the official model card, but expect to test/debug it on your first real run.

# 1. Clone the repo (large, uses Git LFS)
git lfs install
git clone https://huggingface.co/Finnish-NLP/Chatterbox-Finnish engines/chatterbox-finnish

# 2. Install the extra dependencies (torch is heavy)
source .venv/bin/activate
pip install torch torchaudio safetensors
# get the exact torch install command (CPU/MPS) from pytorch.org

# 3. Per that repo's own README, make sure engines/chatterbox-finnish/pretrained_models
#    and engines/chatterbox-finnish/models/*.safetensors exist (it may have its own
#    setup script - check its README)

After this, /api/voices should report chatterbox-finnish as available: true and it'll be selectable in the UI.

Important notes:

  • This model needs a short reference audio file (audio_prompt_path) for voice cloning; the current code looks for a .wav file under engines/chatterbox-finnish/samples/. If it can't find one, add your own sample or pass a path via reference_audio in the API request.
  • macOS/Apple Silicon: the Chatterbox package officially supports the mps device, but user reports show some operators (e.g. aten::_fft_r2c) aren't implemented on MPS yet. The current code automatically retries once on CPU if MPS generation errors out (slower, but works).
  • If you still hit import errors, the exact path in src/chatterbox_/tts.py may differ slightly between repo versions - compare against the actual files you cloned.
  • Memory: this model is large and, on Apple Silicon, shares system RAM with everything else (no dedicated VRAM). On a 16GB Mac, even a single short sentence can push into swap - treat this engine as an occasional, short-passage tool, not a daily driver. The app enforces a low character cap on this engine for exactly this reason.

Licenses

  • This project's code: MIT - see LICENSE.
  • Piper: MIT. Harri voice: CC0 (fully free, commercial use included).
  • Chatterbox (base): MIT. Chatterbox-Finnish: MIT.
  • edge-tts: MIT (it calls a free, unofficial Microsoft service - see below).
  • If you ever add the asmo voice (F5-TTS based): that one is CC BY-NC 4.0, non-commercial only.

Deploying to a VPS (inside an existing subdomain, with your own nginx)

This covers running the app behind an nginx location block on a domain/subdomain you already manage, without creating a new server block or SSL certificate. The values below are illustrative - substitute your own domain, port, and path throughout; deploy/nginx-finnish-tts.conf and deploy/finnish-tts.service use explicit <PLACEHOLDER> markers for exactly this reason.

No GPU needed - Piper is CPU-only by design. On a GPU-less VPS, skip Chatterbox-Finnish and stick with Piper (and edge-tts's Noora, which is also lightweight).

# On the VPS (Ubuntu/Debian assumed):
sudo apt update && sudo apt install -y python3-venv python3-pip

# Transfer the project via scp or git, then:
cd finnish-tts-app
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install -r backend/requirements.txt

Pick a port. Check what's already listening before choosing one - especially on a server that already runs other services:

sudo ss -tlnp | grep ":8000"   # replace 8000 with whatever port you're considering

Use whatever comes back free, and set it consistently in the systemd service and the nginx block below.

Make the service persistent with systemd (survives SSH disconnects and reboots):

# Edit <USER>, <PATH_TO_PROJECT>, and <PORT> in deploy/finnish-tts.service first
sudo cp deploy/finnish-tts.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now finnish-tts
journalctl -u finnish-tts -f    # watch the logs

The app listens on 127.0.0.1:<your port> only, local, not reachable from outside. To embed it in a div on an existing subdomain without touching that subdomain's other paths, see deploy/nginx-finnish-tts.conf - it's two steps: one limit_req_zone line in http{} (rate limiting, once for all of nginx), and a location block inside that subdomain's existing server file, pointing at your chosen port. Before and after applying:

sudo nginx -t              # check syntax before reloading
sudo systemctl reload nginx
# then re-check that the domain's other existing paths still load exactly as before

And on your own page:

<iframe src="https://your-domain.example/your-path/"
        style="width:100%; height:640px; border:0;" loading="lazy"></iframe>

Whatever port you pick doesn't need to be opened in any firewall - only nginx (locally) talks to it, never the outside world.

Security: rate limiting (6 requests/minute per IP, adjust to taste) is set directly in the nginx config, no extra Python dependency needed.

Restricting which voices are exposed: set the ENABLED_VOICES environment variable in the systemd service (comma-separated voice ids, e.g. edge-noora) if a given deployment should only offer specific voices - useful for a public-facing instance where you'd rather not expose every option. Leave it unset to expose everything (the default).

Quick troubleshooting

  • "Voice model missing": make sure voices/fi_FI/fi_FI-harri-medium.onnx exists; if not, run ./scripts/download_voices.sh fi_FI harri medium.
  • Port already in use: pick a different one, e.g. PORT=8001 ./scripts/run.sh.
  • Voice sounds robotic/off: try the harri-low voice for comparison; if Chatterbox is set up, compare against that too.

About

Self-hosted Finnish text-to-speech web app — Piper, Chatterbox-Finnish, and free Azure voices, embeddable anywhere

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages