Amazon Showcases Alexa Feature That Clones Voices From One Minute of Audio
Amazon reveals a new Alexa capability that clones a deceased relative's voice to read stories using just one minute of audio. The impressive technology frames voice replication as a conversion task rather than a traditional generation path.
Amazon head scientist Rohit Prasad unveils a new Alexa feature at the re:Mars conference that uses synthetic speech to replicate a deceased grandmother's voice for reading a bedtime story. The demonstration shows a child asking the voice assistant to continue reading The Wizard of Oz using the cloned voice of his passed relative. This emotional display highlights how artificial intelligence simulates human connections through highly personalized audio experiences.
The technology stands out because it generates a high-quality voice clone from less than a minute of sample audio instead of the industry standard of hours in a recording studio. Prasad explains that Amazon achieves this breakthrough by framing the challenge as a voice conversion task rather than a standard speech generation path. This approach vastly reduces the resources needed and opens the door for everyday users to create custom personal voices for their smart speakers.
While the demonstration draws comparisons to brand-specific Alexa voices and similar smart speakers like the Takara Tomy Coemo, it raises unique questions about using AI to interact with the voices of people who have passed away. Despite the impressive technological leap and the metaphorical fireworks at the end of the presentation, the feature currently lacks an official name or a set timeline for a public rollout. Amazon positions this experiment as a prime example of the current golden era of AI turning science fiction into reality.