Key facts
- Fish Audio has raised $50 million in seed funding.
- The funding round was led by Coreline Ventures and Capital Today.
- The company has over 8 million users and $21 million in annual recurring revenue.
- Fish Audio develops AI voice models for creators and enterprises.
- The startup plans to release new audio understanding and speech-to-speech models.
Palo Alto-based Fish Audio has secured $50 million in seed funding to advance its development of AI voice models for both creative and enterprise applications. The company, which began as a personal project by former NVIDIA researcher Shijia Liao, has rapidly grown to serve over 8 million users and generate $21 million in annual recurring revenue.
The funding round was led by Coreline Ventures and Capital Today, with participation from several other venture capital firms. Fish Audio's technology aims to provide more expressive and steerable synthetic voices, catering to diverse needs ranging from AI avatars and video game characters to automated customer support and sales operations.
Since its launch, Fish Audio has released five models, including speech generation and speech-to-text capabilities. While some models are open-source, its premium S2.1 Pro model is available via a paid API, alongside monthly plans for creators and enterprise solutions. The company has faced scrutiny regarding voice data consent, but has since implemented an automated take-down process for alleged unauthorized voice uploads.
Looking ahead, Fish Audio plans to introduce an audio understanding model and a speech-to-speech model within the year. The company competes in a crowded market with established players like ElevenLabs and WellSaid, but aims to differentiate itself through fine-grained controls for developers and cost-efficient model training.
