Amazon Unveils Alexa Feature to Clone Voices from Short Audio Clips

Amazon announces a new AI capability that synthesizes a short audio clip into longer speech, allowing Alexa to mimic the voice of a deceased loved one to read stories. The technology requires only one minute of recording to produce a high-quality voice.

At the re:Mars conference in Las Vegas, Amazon reveals a groundbreaking new feature for Alexa that synthesizes short audio clips into longer speech. In a striking demonstration, the smart assistant uses the voice of a deceased grandmother to read a bedtime story to her grandson. This technology achieves highly impressive audio output using just one minute of reference audio.

Rohit Prasad, Amazon's Head Scientist for Alexa, explains that the system frames the process as a voice conversion task rather than traditional speech generation. This innovative approach allows the AI to produce a high-quality voice with less than a minute of recording, bypassing the traditional requirement of hours spent in a recording studio. Prasad enthusiastically states that such advancements prove society is living in the golden era of AI.

While the scenario presents a heartwarming use case, Amazon provides no timeline or specific details regarding an official release. The lack of clarity naturally invites intense scrutiny regarding potential applications of this voice cloning technology beyond reading bedtime stories. As the capabilities of artificial intelligence expand, questions about ethical boundaries and misuse inevitably follow.

Read More at the original source →