← MAIN AI DASH
🐟
Fish Audio
Open-source TTS and voice cloning with leading multilingual benchmarks.
Visit Fish Audio →
9.8
Rating / 10
50K daily visits
Overview
Fish Audio is an open-source AI voice synthesis platform offering text-to-speech, voice cloning, and speech-to-text capabilities through its S2 model family. As of mid-2026, the platform is widely cited in independent benchmarks for achieving top-tier word-error rates — the S2 model scores 0.54% WER on Chinese voice-cloning tests and 1.07% on English, outperforming Seed-TTS, CosyVoice 3, and Qwen3-TTS in head-to-head evaluations. The platform supports 100+ languages with automatic detection, offers a web interface for non-technical users, and provides a developer API priced at $15 per million UTF-8 bytes (roughly 180,000 English words or ~12 hours of speech). A speech-to-text API runs at $0.36 per audio hour. Pricing for the web platform is credit-based: Free ($0/month, ~8,000 monthly credits, ~7 minutes of generation, 500-character limit per generation, 3 public voice slots, no commercial rights), Plus ($11/month billed annually, 250,000 credits, ~200 minutes, 15,000-character limit, private voice slots), Pro ($75/month annually, 2M credits, ~1,620 minutes, 3 team seats), Max ($749/month annually, 25M credits, ~6,250 minutes, 10 team seats), and Enterprise (custom, zero data retention, on-premise deployment). Commercial usage rights require careful verification of current terms, as the FAQ wording around monetized content has been reported as ambiguous. User feedback is broadly positive on voice naturalness and cloning fidelity — the S2 model's LLM-as-a-Judge evaluation scores highly on speech naturalness and instruction following. Developers praise the simple API structure and competitive per-byte pricing for high-volume applications. However, common complaints include the restrictive 500-character limit on the free tier, ambiguous commercial licensing terms that create uncertainty for monetized YouTube or game content, and the fact that the highest-quality voice cloning still requires clean source audio — noisy or low-quality samples produce noticeably synthetic output. For developers building voice agents or audiobook pipelines, Fish Audio offers strong value; for casual users, the free tier is more of a brief demo than a working tool.
✅ Benefits
  • Leading benchmark performance — S2 achieves 0.54% WER on Chinese and 1.07% on English voice-cloning tests, outperforming Seed-TTS and Qwen3-TTS in independent evaluations
  • Developer-friendly API at $15 per million UTF-8 bytes (~12 hours of speech), making it one of the most cost-effective options for high-volume TTS and voice-agent applications
⚠️ Drawbacks
  • Free tier is heavily restricted to ~7 minutes of generation with a 500-character cap per request, making it a brief demo rather than a practical ongoing tool
  • Commercial usage terms are reported as ambiguous in the FAQ, creating legal uncertainty for monetized content on YouTube, games, and branded material