Monday, August 3, 2026

Startups & Funding

Fish Audio raises $52M seed for AI voice models

Palo Alto startup Fish Audio raised $52 million in a seed round to expand its AI voice models for creators and enterprises.

Fish Audio co-founders Shijia Liao and Rissa Cao sitting in chairs in front of a window.
Photo: Fish Audio

Fish Audio, a Palo Alto-based startup building AI voice models, said on Tuesday it has raised $52 million in a seed round led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. The company said more than 8 million people now use the open-source or hosted versions of its models, generating annual recurring revenue of $21 million.

Fish Audio was started by former Nvidia researcher Shijia Liao, who, frustrated by non-expressive synthetic voices on the market, trained a voice-generation model on a single GPU and open-sourced it. The resulting Fish Speech repository on GitHub now has more than 31,000 stars and is used by indie developers, video game designers, and creators. The startup has released five models in the past year — four for speech generation and one for speech-to-text — open-sourcing three of the speech-generation models while keeping its newest, S2.1 Pro, available only through a paid API. Its library now spans more than 15,000 natural language controls. Enterprise customers including HeyGen and Sanas already use its APIs and platform, alongside paid monthly plans for creators and teams that include voice-cloning features.

“Every enterprise has different use cases and different preferences,” said Fish Audio CEO and co-founder Rissa Cao, citing HeyGen’s need for realism in AI avatars, gaming studios’ demand for expressive character voices, and voice-agent companies’ need for natural, low-latency voices for calls.

Fish Audio has partly built its voice library by inviting users to submit their own voices for training, compensating them when used. A few months ago, some creators alleged their voices were uploaded without consent; the company had a DMCA takedown process in place, but takedowns took a long time. Cao told TechCrunch the company has since automated that process — creators who submit a voice sample or a contract proving ownership can now get an uploaded voice removed in less than three minutes. That doesn’t stop someone from uploading a voice without the owner’s knowledge in the first place; the voice stays live until the owner finds out and files for removal.

Cao said the seed funds will support a planned audio-understanding model and a speech-to-speech model, both due this year, as Fish Audio competes with ElevenLabs, WellSaid, Cartesia, Speechify, Async, and Krisp for creator and enterprise budgets.

Why it matters

Fish Audio’s bet is that automated consent tooling, not just model quality, will decide which voice-AI platform creators and enterprises trust with their vocal likeness.