FUNDING

Fish Audio Raises $52 Million to Scale its AI Speech Technology

The company has grown to more than 8 million users and $21 million in annual recurring revenue as it develops AI voice technology for creators, developers, and enterprises.

By Donna Joseph
July 29, 2026 12:18 AM • Updated July 29, 2026
Fish Audio Raises $52 Million to Scale its AI Speech Technology Photo by SBR

Summary
  • Fish Audio has raised $52 million in seed funding as its AI speech technology reaches more than 8 million users and generates $21 million in annual recurring revenue.
  • It serves creators and enterprises with AI voice technology, offering more than 15,000 natural-language controls, voice cloning, speech generation, and speech-to-text capabilities.
  • The company is expanding its technology while addressing voice ownership concerns, with an audio-understanding system and speech-to-speech system in development and an automated process for removing voices uploaded without permission.

PALO ALTO, Calif., July 28, 2026 — AI-generated voices are becoming part of creative work and business operations, but different users expect different things from synthetic speech. Creators want voices that can express emotion and personality, while businesses using AI for customer support and sales operations need greater control over how those voices sound and respond. Fish Audio is building technology for both groups. The Palo Alto, California-based startup has developed a library of more than 15,000 natural-language controls that can be used to control AI-generated speech. Since launching last year, Fish Audio says more than 8 million people have used its open-source or hosted products, and the company now generates $21 million in annual recurring revenue. On Tuesday, the startup announced a $52 million seed round led by Coreline Ventures and Capital Today. 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0 also participated.

Fish Audio began with a much smaller project. Former Nvidia researcher Shijia Liao became frustrated with the lack of expression in synthetic voices available at the time, so he trained a voice-generation system on a single GPU and released it as open source. The project grew into Fish Audio, and the Fish Speech repository on GitHub now has more than 31,000 stars. Independent developers, video game designers, and creators use the technology. During the past year, Fish Audio has released five systems: four for speech generation and one for speech-to-text. Three speech-generation systems are open source, while the newest S2.1 Pro is available through the company’s paid API. Fish Audio also offers monthly plans for creators and organizations that include a set number of generation minutes and voice-cloning features. Businesses can access an enterprise version of its APIs and platform.

One Technology, Many Voice Needs

The range of Fish Audio’s customers shows why control matters in AI-generated speech. HeyGen uses the company’s voices for AI avatars and wants realistic speech, according to Fish Audio CEO and co-founder Rissa Cao. A gaming studio may instead need expressive voices for characters, while voice-agent companies such as LiveKit require natural-sounding speech with low latency during calls. Fish Audio has built its library of more than 15,000 natural-language controls to serve these different requirements. The company has also expanded its voice library by allowing users to submit their own voices for training and compensating them when those voices are used. HeyGen and Sanas are among the organizations that Fish Audio says already use its technology, giving the startup customers across AI avatars, voice agents and other applications.

Fish Audio’s business now spans open-source software, paid subscriptions and enterprise services. Creators and organizations can purchase monthly plans that provide a set number of voice-generation minutes and voice-cloning features, while businesses can use the company’s APIs and platform. That commercial expansion is one reason Fish Audio decided to raise outside capital. Cao said the company was able to operate without external funding when it was focused on an open-source project and creator plans. The company wanted to develop more advanced technology and serve enterprise customers, however, and investor interest was growing. The $52 million seed round gives Fish Audio capital to support that work as it expands its products beyond its original open-source foundation.

Podcast Thumbnail

Who Owns an AI-Generated Voice?

The growth of AI voice technology has also created questions about consent and ownership. Fish Audio has invited users to submit their voices for training and compensates them when their voices are used, but the company faced complaints from creators who alleged that their voices had been uploaded without permission. Fish Audio had a DMCA takedown process for handling those complaints, but some creators said removals took too long. Cao said the company has now automated the process. A creator can submit a short voice sample or a contract showing that an uploaded voice belongs to them, and Fish Audio says the voice can be removed from the platform in less than three minutes. The faster system addresses the time required for a takedown, but it does not prevent someone from uploading another person’s voice without permission. Until the creator discovers the unauthorized use and requests removal, the voice can remain on the platform.

The issue has placed creator trust alongside technical performance as an important consideration for Fish Audio’s community-driven system. Osuke Honda, a partner at Coreline Ventures, said such a system can only work when creators trust the platform. He called for verified voice ownership, clear licensing terms, simple reporting and takedown procedures, and eventually revenue-sharing arrangements that allow creators to benefit financially when their voices are licensed or used commercially. For Fish Audio, these questions are closely connected to the way its voice library is built. The company is not only developing speech-generation technology; it is also operating a platform where users can contribute voices that may become part of that technology. As AI-generated voices become more capable, questions about permission, ownership, and compensation will remain relevant to creators whose voices are used in synthetic speech.

Image credit: Fish Audio

New Systems Expand Fish Audio’s Reach

Fish Audio is now preparing to move into additional areas of audio technology. The company plans to release an audio-understanding system this year and is also developing a speech-to-speech system. These products will add to a portfolio that already includes speech-generation and speech-to-text technology. The company is competing with established names in the sector, including ElevenLabs, WellSaid, Cartesia, Speechify, Async, formerly known as Podcastle, and Krisp. Creators and enterprises are the primary customers for these products, while large AI companies are also developing their own speech technology. Fish Audio believes its detailed controls for developers and cost-efficient training can help it compete with larger AI companies. Rico Mallozzi, a partner at 359 Capital, said Fish Audio has developed advanced speech technology with a relatively small group of researchers and engineers and has made progress in closing the gap between artificial-sounding and human-like voices.

Fish Audio now has a different set of resources from the ones available when Liao trained the first voice-generation system on a single GPU. More than 8 million people have used its open-source or hosted products, the company reports $21 million in annual recurring revenue, and the new $52 million funding round gives it capital to develop new technology and serve enterprise customers. The company’s next products will expand its work in audio, while its existing services give creators, developers and businesses access to voice generation and voice cloning technology. At the same time, Fish Audio must address concerns around unauthorized voice uploads and provide creators with reliable ways to control how their voices are used. The company’s growth began with a simple idea: make synthetic voices more expressive. Today, Fish Audio is building a larger business around that idea, with its technology reaching creative projects, AI avatars, voice agents, and enterprise applications.

Fish Audio’s business now spans open-source software, paid subscriptions, and enterprise services. Creators and organizations can purchase monthly plans that provide a set number of voice-generation minutes and voice-cloning features, while businesses can use the company’s APIs and platform.


What To Read Next

Encore Raises $30 Million to Expand Adoption of AI Agents That Learn From Customer Conversations

Encore Raises $30 Million to Expand Adoption of AI Agents That Learn From Customer Conversations

The startup enables companies to deploy AI agents that communicate with customers through voice or text. These agents can manage customer interactions directly or assist employees by recommending responses and engagement strategies during conversations.
Enigma Raises $71 Million to Explore a New Kind of Robotic Intelligence
The Israeli startup is putting more than 100 AI-powered robots online to study how people naturally communicate with machines.
Corgi’s Valuation Set to Double as AI-Powered Insurance Startup Secures Another Round
For some of its insurance operations, Corgi uses a structure known as a Risk Retention Group, or RRG. An RRG allows businesses or individuals with similar liabilities to pool their resources and self-insure collectively.

Business