Microsoft Launches Custom Neural Voice for Branded AI Speech

Microsoft makes its Custom Neural Voice text-to-speech service generally available, allowing developers to create unique synthetic voices for their brands. The company requires customers to apply for access to prevent misuse of the technology.

Microsoft announces limited access to its Custom Neural Voice, a text-to-speech feature within Azure Cognitive Services that enables developers to build unique, branded synthetic voices. The service is now generally available, yet Microsoft requires customers to apply and receive approval before using it to prevent potential misuse. Since its initial preview, companies like AT&T, Duolingo, and Progressive already use the technology to create custom speech solutions for their users.

The underlying Neural TTS technology relies on three primary components to generate highly realistic audio. First, the Text Analyzer processes input text and converts it into a sequence of phonemes. Next, the Neural Acoustic Model takes these phonemes and predicts acoustic features like timbre, speed, intonation, and stress patterns. Finally, the Neural Vocoder transforms those predicted acoustic features into audible sound waves.

Microsoft trains these neural voice models using deep neural networks and real voice recording samples, giving customers the ability to adapt the engine to their specific scenarios. To utilize the service, organizations need an active Azure account and subscription. Once Microsoft approves their application, developers can start a custom voice project, upload their audio data, train and test their model, and deploy the final synthetic voice.

Read More at the original source →