OpenAI Introduces Voice Cloning Model With Strict Limited Access

OpenAI reveals Voice Engine, an AI model that clones a human voice from just a 15-second audio sample. The company restricts access to a small group of trusted partners due to the significant risks involved.

OpenAI unveils Voice Engine, a new artificial intelligence model that requires only a 15-second audio sample to clone a human voice. The technology currently powers the text-to-speech API and the voice features found in ChatGPT, but its ability to replicate any speaker's unique vocal characteristics presents massive implications for the spoken audio market.

The voice cloning capability threatens to disrupt numerous industries, putting direct pressure on competing AI audio startups like ElevenLabs and Meta. Beyond commercial applications for podcasters, voice actors, and customer service agents, OpenAI highlights the model's potential to provide non-robotic, personalized voices for non-verbal individuals and to support educational programs for those with speech impairments.

Despite its broad capabilities, OpenAI restricts Voice Engine to a small group of trusted partners, including educational technology company Age of Learning and visual storytelling platform HeyGen. This cautious rollout reflects the company's awareness of the serious risks associated with voice replication, as they actively seek to prevent misuse while exploring the technology's beneficial applications.

Read More at the original source →