Coqui and Hugging Face Launch Open-Access XTTS Voice AI Model
AI startup Coqui partners with Hugging Face to release XTTS, a generative voice AI model that clones voices from just three seconds of audio. The open-access model supports 13 languages and offers cross-language cloning capabilities.
AI startup Coqui partners with Hugging Face to release XTTS, an open-access foundation model for generative voice AI. The model builds on the success of Coqui's previous releases and currently ranks as the top trending repository on GitHub and the top trending space on Hugging Face. This collaboration highlights a shared mission to democratize artificial intelligence technology.
XTTS offers groundbreaking features, including the ability to clone a voice from a mere three-second audio clip while transferring emotion and style. The model supports 13 languages, including English, Spanish, French, Arabic, and Mandarin Chinese, and enables cross-language voice cloning. Additionally, it delivers high-quality audio with a superior 24khz sampling rate for refined speech generation.
Coqui, founded by former Mozilla machine learning experts, recently closes a $3.3 million seed funding round to support its open-science vision. Company leadership emphasizes a commitment to building ubiquitous foundation models for voice AI, and both companies express enthusiasm for future model launches on the Hugging Face platform.