Fish Audio, a Palo Alto-based startup focused on AI voice generation, has raised $50 million in a seed round led by Coreline Ventures and Capital Today, with support from several other investors. The company says the funding will accelerate its work on voice models for both creative users and enterprise teams.
Founded by former NVIDIA researcher Shijia Liao, Fish Audio began as an open-source project built to make synthetic speech sound more natural and expressive. Since then, the company has grown quickly, reporting more than 8 million users across its open-source and hosted products, along with $21 million in annual recurring revenue.
Its platform now includes a library of more than 15,000 natural language controls, designed to give users finer command over tone, style, and delivery. Fish Audio has released five models over the past year, including speech generation systems and a speech-to-text model. Three of its speech generation models remain open source, while its latest S2.1 Pro model is available through a paid API.
The company also serves creators and teams through subscription plans that include generation minutes and voice cloning tools, while its enterprise offering is already used by organizations such as HeyGen, Sanas, and Plaud. Fish Audio says it is also working on an audio understanding model and a speech-to-speech system, both planned for release this year.
As the AI voice market expands, Fish Audio is betting that control, efficiency, and creator-friendly tools will help it stand out. Its next phase could shape how natural voice technology is used across entertainment, productivity, and digital communication in the years ahead.