AI text-to-speech and voice-cloning studio with emotion controls, 2M+ voices and speech-to-text via web or API.
What it does
Fish Audio is an AI voice platform offering expressive text-to-speech, voice cloning and speech-to-text. Its S2 model supports emotion and effect tags, real-time generation and a library of over two million voices, aimed at creators, developers and teams. It powers video voiceovers, audiobooks, character voices and conversational agents through web tools and an API.
How to use: Users can discover and use pre-built voice models or build their own. The platform offers a text-to-speech toolkit where users can input text and select a voice model to generate speech.
Core features
Best for
Reviews
Big-picture takes: what it's for and whether it delivers. High-engagement YouTube videos — not sponsored.
Tutorials
Step-by-step: exactly how to get things done with it.