Fish Audio Raises $52 Million Seed Round To Expand Expressive Voice AI Platform

Fish Audio has raised $52 million in seed funding to expand its real-time text-to-speech, voice cloning and voice-agent platform. The round was led by Coreline Ventures and Capital Today, with participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, Alphalist Partners and angel investors.

The financing follows Fish Audio’s first year of commercial operations, during which it grew to approximately $21 million in annual recurring revenue and more than eight million users. The platform serves creators, developers and enterprises seeking realistic synthetic voices, multilingual support and greater control over emotional expression.

Fish Audio grew out of Fish Speech, an open-source project created by co-founder and Chief Scientist Shijia Liao. The project has accumulated more than 31,000 GitHub stars and developed a following among game developers, independent software builders and digital content creators.

The company’s technology can clone a voice from a five-second sample in approximately 15 seconds, supports more than 83 languages and provides more than 15,000 natural-language controls for word-level emotion. Its S2.1 Pro model was preferred by approximately 67% of listeners over competing products in company-reported blind tests.

Fish Audio also offers on-premises deployment, zero-data-retention options and HIPAA-compliant configurations for enterprises with security, privacy or regulatory requirements. The company plans to use the financing to expand beyond text-to-speech into voice-native language models, speech-to-speech systems and other audio-focused AI capabilities.

Additional capital will also support enterprise sales, developer tools and integrations with platforms such as LiveKit and Retell. Fish Audio is headquartered in Palo Alto and serves customers including HeyGen, Telnyx, OpenArt, Sanas and other voice and media technology companies.

KEY QUOTES:

“We built Fish Audio because we wanted voice AI that sounded human, not chunky or robotic, and we wanted that quality to be accessible at any scale.”

“We make high-quality, human-sounding voices available to every user, from beginner creatives to million-dollar enterprises, so communication is not only more efficient, but more trustworthy. We’ve always believed that if we kept making the models better, people would notice. Eight million of them did. There’s a lot of work left to do, and now we have the resources to do it.”

Rissa Cao, Co-Founder and CEO of Fish Audio

“Voice is becoming the default interface for AI, and Fish Audio is unlocking this opportunity to a new generation of creators, developers, and enterprises.”

“In its short history, Fish Audio has built an unbeatable track record of pushing the envelope on performance, multilingual support, emotional expression, and cost. All factors that have quickly made Fish Audio the default choice for creators, developers, and now enterprises globally, and we expect them to continue to lead the way.”

Osuke Honda, Managing Partner at Coreline Ventures