Microsoft AI Lab Launches In-House Voice and Language Models

Microsoft AI Lab introduces MAI-Voice-1 for rapid speech generation and MAI-1-preview, its first fully in-house foundation language model. Both models aim to enhance consumer applications like Copilot without relying on third-party technology.

Microsoft AI lab officially unveils MAI-Voice-1 and MAI-1-preview, marking a major step in the company's independent artificial intelligence development. These two models serve distinct but complementary purposes, with MAI-Voice-1 handling speech generation and MAI-1-preview acting as a general-purpose language model. By building these systems entirely in-house, Microsoft eliminates third-party dependencies and asserts greater control over its AI ecosystem.

MAI-Voice-1 stands out as a highly efficient speech synthesis model that generates one minute of natural-sounding audio in under a second using just a single GPU. Built on a transformer-based architecture, it supports both single-speaker and multi-speaker scenarios across multiple languages. The technology is already powering voice updates in Microsoft Copilot Daily and is available for public testing in Copilot Labs for tasks like creating audio stories.

Meanwhile, MAI-1-preview represents Microsoft's first end-to-end, self-built foundation language model, trained on a massive cluster of approximately 15,000 NVIDIA H100 GPUs. It utilizes a mixture-of-experts architecture and focuses heavily on instruction-following and conversational tasks for everyday consumer use. Microsoft is currently rolling out access to this text-based model in select Copilot scenarios and plans a gradual expansion as it gathers user feedback.

Read More at the original source →