Shijia Liao Fish Audio is a founder story that started with one person, one gaming GPU, and a personal frustration with how robotic AI voices sounded. Liao is the co-founder and chief scientist of Fish Audio, a Palo Alto voice AI company that just closed a $52 million seed round led by Coreline Ventures and Capital Today, reaching $21 million in annual recurring revenue and more than 8 million users.
Before any institutional funding existed, Liao was a former NVIDIA video researcher training voice models alone in his bedroom, driven by frustration as a lifelong anime and VTuber fan with how flat synthetic voices sounded. Here is how that open-source side project became one of the fastest-growing companies in AI voice technology.
Who Is Shijia Liao?
Shijia Liao is the co-founder and chief scientist of Fish Audio, whose legal entity is registered as Hanabi AI Inc., according to SiliconANGLE’s coverage of the company’s funding round. Professional profile data lists his education at the University of Maryland. Before Fish Audio, Liao worked as a video researcher at NVIDIA, and describes himself, across multiple interviews given around the funding round, as a lifelong fan of VTubers and Japanese anime.
That personal interest wasn’t incidental to Fish Audio’s origin. Liao has said he grew tired of listening to the flat, monotonous synthetic voices produced by early AI voice models, a frustration rooted specifically in his own experience as a consumer of anime and VTuber content, where expressive, characterful voice performance is central to the medium.
From a Single GPU to 31,000 GitHub Stars
Liao began training his own voice-generation models using nothing more than a single gaming GPU in his bedroom. Rather than keeping the resulting work private, he open-sourced it as Fish Speech, a project that grew quickly on GitHub, accumulating more than 31,000 stars and attracting a following among indie developers, video game designers, and content creators looking for more expressive synthetic voices than existing commercial tools offered.
That open-source traction, built entirely before Fish Audio existed as a formal company, gave Liao and eventual co-founder Rissa Cao, who serves as CEO, a working technical foundation and an engaged user base before they ever raised outside capital, according to TechCrunch’s reporting on the round. Liao continues to publish original research under the Fish Audio name, according to the company’s own blog.
The Problem Fish Audio Set Out to Solve
Synthetic voice technology has historically struggled with expressiveness: AI-generated voices that could speak clearly but sounded flat, monotonous, and clearly artificial, particularly compared to the nuanced vocal performances found in anime, gaming, and other character-driven media. That gap mattered increasingly as more creators, from indie game developers to individual content creators, needed access to expressive voice technology without the budget of a major studio.
Cao has framed Fish Audio’s mission around accessibility as much as quality. “We make high-quality, human-sounding voices available to every user,” Cao said in comments reported by SiliconANGLE, positioning the company’s ambition as spanning both individual creators and larger enterprise customers rather than serving either group exclusively.
Building Fish Audio: Voice Cloning, Emotion Control, and Open Source
Fish Audio’s platform can clone a voice from a five-second audio clip in roughly 15 seconds, supports more than 83 languages, and offers word-level emotion control through more than 15,000 natural language controls, according to the company’s own funding announcement. The company has launched five models over the past year, four for speech generation and one for speech-to-text, open-sourcing three of the speech-generation models while keeping its newest flagship model, S2.1 Pro, available only through a paid API.
The company says S2.1 Pro was preferred by 67% of listeners over competitor models in blind listening tests. That figure comes from Fish Audio’s own internal testing rather than an independently conducted or third-party audited benchmark, worth noting given how central voice quality comparisons are to the company’s positioning against established competitors.
Inside Fish Audio’s $52 Million Seed Round
Shijia Liao Shijia Liao Fish Audio closed a $52 million seed round, announced July 28, 2026, led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. The round is unusually large for a seed stage, reflecting both the company’s already-substantial revenue and its rapid user growth heading into the raise.
Cao has framed the company’s growth as validation of a straightforward strategy: continuing to improve the underlying models and trusting that users would notice, a philosophy that appears to have paid off given the company’s current scale. The funding will support expanding Fish Audio’s model lineup beyond its current text-to-speech capabilities, according to SiliconANGLE’s reporting on the round.
Growth and Traction: 8 Million Users and $21 Million ARR
Fish Audio says it has reached more than 8 million users and $21 million in annual recurring revenue, figures reported consistently across multiple outlets covering the funding round, including TechCrunch, MLQ News, and SiliconANGLE. That scale is notable for a company that began as one person’s open-source side project rather than a venture-backed effort from its earliest days.
The company serves both individual creators and enterprise customers, offering on-premises deployment and HIPAA-compliant configurations for the latter, a combination that lets Fish Audio monetize its large open-source-derived user base while also pursuing larger enterprise contracts that typically carry stricter compliance requirements.
The Investors Backing Shijia Liao
Coreline Ventures and Capital Today co-leading Shijia Liao Fish Audio’s $52 million seed round, with five additional participating firms, reflects substantial investor conviction in the company’s growth trajectory heading into the raise. The size and structure of the round, a large seed rather than a traditional Series A, suggests investors were comfortable underwriting the company’s already-demonstrated revenue and user traction rather than treating it as a pre-revenue bet.
That combination of firms, spanning both traditional venture capital and firms like HF0, a residency-style program for AI founders, suggests a syndicate assembled partly around Fish Audio’s specific position at the intersection of open-source developer traction and commercial enterprise potential.
Fish Audio’s Position in the AI Voice Race
Fish Audio enters a competitive AI voice generation market that includes well-funded incumbents such as ElevenLabs, and its own materials directly position its S2.1 Pro model as outperforming unnamed competitors in blind listening tests. The company’s open-source origins and continued partial open-sourcing of its models differentiate it from more closed, enterprise-first competitors, giving it a distinct developer and indie-creator community that predates its venture funding entirely.
That community-first origin, an engaged open-source user base built before any company existed, is arguably Fish Audio’s clearest structural advantage against competitors that built their user base primarily through paid marketing and enterprise sales from the start.
What’s Next for Fish Audio
With fresh seed capital, Fish Audio’s near-term priorities center on expanding its model lineup beyond its current text-to-speech focus, according to reporting on the round. The company’s dual customer base, creators and enterprises, suggests continued investment in both open-source community tools and higher-compliance enterprise offerings simultaneously.
Liao’s continued involvement in publishing original research, rather than moving entirely into an executive role, suggests Fish Audio intends to maintain its technical credibility within the open-source and developer community that gave the company its earliest traction, even as it scales its commercial operations.
Lessons for Entrepreneurs From Shijia Liao’s Journey
The Shijia Liao Fish Audio story offers a clear lesson for founders building AI products in categories closely tied to their own personal interests: his frustration as an anime and VTuber fan with flat synthetic voices gave him both the motivation to start and a built-in understanding of what expressive voice technology actually needed to deliver for that specific audience. Second, open-sourcing Fish Speech before any company existed let Liao build a genuine, engaged developer community and battle-test the technology against real-world use cases well before venture capital ever entered the picture.
Third, Fish Audio’s decision to keep its most advanced model, S2.1 Pro, behind a paid API while continuing to open-source earlier models shows a workable path to monetizing open-source traction without abandoning the community that built it.
Frequently Asked Questions
Who is Shijia Liao? Shijia Liao is the co-founder and chief scientist of Fish Audio. A former NVIDIA video researcher, he began training voice-generation models on a single gaming GPU as a personal project before open-sourcing the resulting work as Fish Speech.
What does Fish Audio do? Fish Audio builds AI voice generation models that can clone a voice from a five-second clip, support more than 83 languages, and offer word-level emotion control, serving both individual creators and enterprise customers.
How much funding has Fish Audio raised? Fish Audio raised $52 million in seed funding, announced July 28, 2026, led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
Who co-founded Fish Audio with Shijia Liao? Fish Audio was co-founded by Shijia Liao, who serves as chief scientist, and Rissa Cao, who serves as CEO.
How many users does Fish Audio have? Fish Audio says it has more than 8 million users and $21 million in annual recurring revenue as of its July 2026 seed funding announcement.
What is Fish Speech? Fish Speech is the open-source voice generation project Liao created before Fish Audio existed as a company, which grew to more than 31,000 stars on GitHub and attracted a following among indie developers and content creators.
Conclusion
Shijia Liao Fish Audio is a reminder that some of AI’s most commercially significant companies begin not with a business plan, but with one person’s frustration and a single GPU. With $52 million in fresh seed funding, more than 8 million users, and $21 million in annual recurring revenue built substantially on top of an open-source community Liao created before the company existed, Fish Audio is testing whether that same community-first approach can scale into sustained competition with much larger, earlier-funded voice AI incumbents.
For more founder stories in AI infrastructure and open-source-to-commercial paths, see Mahesh Sathiamoorthy’s Bespoke Labs and Nathan Arnold’s Photon Queue on Denote Press.
