Microsoft Releases Custom Neural Voice for Branded AI Speech

Microsoft makes its Custom Neural Voice text-to-speech service generally available, allowing brands to create unique synthetic voices. Strict application controls remain in place to prevent misuse of the AI technology.

Microsoft announces limited access to its Custom Neural Voice, a text-to-speech feature within Azure Cognitive Services that allows developers to create unique, branded synthetic voices. Companies like AT&T, Duolingo, and Progressive already use this technology to build customized speech solutions for their customers. Although the service is now generally available, Microsoft requires customers to apply for access and implements technical controls to prevent misuse.

The underlying technology relies on three major components to generate highly realistic audio. First, a Text Analyzer converts written input into a sequence of phonemes. Next, a Neural Acoustic Model processes these phonemes to predict acoustic features like timbre, speed, and intonations. Finally, a Neural Vocoder translates those predicted features into audible sound waves.

Because the system trains on deep neural networks using real voice recordings, users can finely tune the engine to match specific brand identities. To utilize this feature, customers need an active Azure account and subscription. Once Microsoft approves their application, developers can start a custom voice project, upload their audio data, train and test the model, and deploy the final synthetic voice.

Read More at the original source →