Questions & answers about Fish Audio

Frequently asked questions

Is Fish Audio free to use?
Fish Audio offers a free tier for testing text-to-speech and voice cloning in the browser, plus the fully open-source Fish Speech model you can self-host at no cost on GitHub. Paid API plans add higher character limits, commercial usage rights and priority generation for production workloads at per-character pricing.
How much reference audio does Fish Audio need to clone a voice?
Fish Audio clones a voice from a short reference sample of roughly 10 to 30 seconds. The Fish Speech S2 model captures timbre, speaking style and emotional tendencies from that clip, producing a consistent cloned voice without additional fine-tuning or hours of training data.
How does Fish Audio compare to ElevenLabs?
Fish Audio is built on the open-source Fish Speech S2 model, so you can self-host it for free, while ElevenLabs is closed and subscription-only. Fish Audio supports 80+ languages with sub-150ms latency and emotion control; teams choosing it cite open weights, lower cost and on-premise deployment as the main reasons.
What languages does Fish Audio support?
Fish Audio supports more than 80 languages, including English, Chinese, Japanese, Korean, French, German, Arabic and Spanish, with new languages added over time. The S2 Pro model was trained on over 10 million hours of multilingual audio, enabling cross-lingual voice cloning where one cloned voice speaks several languages.
Does Fish Audio have an API for developers?
Yes. Fish Audio provides a text-to-speech and voice-cloning API with SDKs (including JavaScript and Python) that accept reference audio and text and return generated speech. Developers integrate it for narration, podcasts, courses and apps, with per-character pricing and low-latency streaming for real-time use cases.