Fish Audio, a Palo Alto-based startup specialising in AI-generated voice models, announced on Tuesday that it has raised $50 million in a seed funding round. The investment was led by Coreline Ventures and Capital Today, with additional participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
Since its launch last year, Fish Audio has attracted over 8 million users for its open-source or hosted models and currently reports an annual recurring revenue of $21 million. The company aims to cater to various use cases, from expressive models for creative applications to steerable voices for enterprise customer support and sales operations, utilising a library of over 15,000 natural language controls.
The startup was founded by former NVIDIA researcher Shijia Liao, who initially open-sourced a voice generation model. Fish Audio has released five models in the past year, including four speech generation models and one speech-to-text model. Its latest S2.1 Pro model is exclusively available through a paid API.
Fish Audio plans to release an audio understanding model this year and is also developing a speech-to-speech model. The company has automated its content take-down process, allowing creators to prove ownership of uploaded voices and have them removed from the platform in under three minutes.