OpenAI Unveils Voice Engine for 15-Second Voice Cloning

OpenAI introduces Voice Engine, a tool that creates synthetic voice copies from just 15 seconds of audio. The company delays public release to focus on safety and abuse mitigations.

OpenAI introduces Voice Engine, a new text-to-speech expansion that allows users to clone any voice from a mere 15-second audio sample. The same underlying model already powers the read-aloud feature in ChatGPT and Spotify's multilingual podcast dubbing for hosts like Lex Fridman. Despite these existing applications, OpenAI keeps the broader Voice Engine tool in a limited preview without a set public release date.

The company states that this cautious approach gives it time to study how the technology is used and abused in the real world. OpenAI product staff member Jeff Harris emphasizes that the goal is to ensure everyone feels good about the deployment and to implement proper safety mitigations against the dangers of deepfakes. By holding back a full launch, OpenAI aims to navigate the risky landscape of synthetic media responsibly.

Details about the model's training data remain sparse, with Harris only confirming a mix of licensed and publicly available information. This secrecy comes as OpenAI faces ongoing IP lawsuits regarding its use of copyrighted material to train other generative AI systems. While the company offers opt-out mechanisms for its image generators and licensing agreements with some publishers, it currently provides no such opt-out option for voice data used in Voice Engine.

Read More at the original source →