Business

Fish Audio raises $50M seed to expand expressive, controllable AI voice models

Fish Audio, the Palo Alto startup behind open-source voice-generation tools, secured $50 million in a seed round to scale more expressive and steerable speech models for creators and enterprises.

Fish Audio raises $50M seed to expand expressive, controllable AI voice models

Palo Alto–based Fish Audio announced it has raised $50 million in a seed round led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners and HF0.

Since launching last year, the startup says more than 8 million people have used either the open-source or hosted versions of its models, and it now reports $21 million in annual recurring revenue (ARR). Fish Audio has built a library of over 15,000 natural language controls intended to serve both creative use cases that demand expressive voices and enterprise customers that need more steerable voice agents for support and sales operations.

Origins and product lineup

Fish Audio began as a small project by former NVIDIA researcher Shijia Liao, who, unsatisfied with non-expressive synthetic voices on the market, trained a voice-generation model on a single GPU and open-sourced it. The Fish Speech repository on GitHub has accumulated more than 31,000 stars and is used by indie developers, video game designers and creators.

Over the past year the company released five models: four speech-generation models and one speech-to-text model. Three of the speech-generation models are open-source; the latest S2.1 Pro model is available only via the company’s paid API.

Fish Audio offers paid monthly plans that unlock a set number of generation minutes and voice-cloning features for creators and teams, and it provides an enterprise version of its APIs and platform. The company says organizations such as HeyGen, Sanas and Plaud are already customers.

Community sourcing, consent issues and takedowns

One way Fish Audio expanded its voice library was by asking users to submit their own voices for model training and compensating them if their voices were used. That approach led to controversy months ago when some creators alleged their voices had been uploaded without consent.

Fish Audio had a DMCA content take-down process in place, but takedowns reportedly took a long time. CEO and co-founder Rissa Cao told TechCrunch the company has automated the takedown workflow: creators can submit a short voice sample or a contract to prove ownership, and the company will remove a contested voice from its platform in under three minutes.

However, automated takedown capability does not prevent someone from uploading an artist’s voice without their knowledge; until the artist discovers the upload and files a removal request, the voice can remain on and be used via the platform.

Oskue Honda, a partner at Coreline Ventures, told TechCrunch that a community-driven model only succeeds if creators trust the platform. He said consent, transparency and attribution must be built into the product rather than treated as afterthoughts, and advocated for verified voice ownership, clear licensing terms, straightforward reporting and takedown processes, and eventual revenue-sharing models so creators financially benefit when their voices are licensed or used commercially.

Why they raised capital and what’s next

Cao said the project ran efficiently as an open-source offering and didn’t require outside capital at first, but the company sought investment to develop more advanced models and to accommodate growing enterprise interest.

Fish Audio plans to release an audio-understanding model this year and is also developing a speech-to-speech model.

The speech-generation market is crowded, with competitors such as ElevenLabs, WellSaid, Cartesia, Speechify, Async (formerly Podcastle) and Krisp. Rico Mallozzi, a partner at 359 Capital, told TechCrunch that fine-grained developer controls and cost-efficient model training will help Fish Audio better compete with larger, well-funded AI labs, and praised the team’s technical work narrowing the gap between artificial-sounding and human-like voices.