Mistral Launches Open Source Voxtral TTS to Challenge Voice AI Rivals

French AI startup Mistral introduces Voxtral TTS, a highly efficient open source text-to-speech model that supports nine languages and requires less than five seconds of audio for voice cloning.

French AI company Mistral releases Voxtral TTS, a new open source text-to-speech model designed for voice AI assistants and enterprise customer support. The compact model runs efficiently on edge devices like smartphones and smartwatches, offering state-of-the-art performance at a fraction of the cost of competing solutions. With this release, Mistral directly challenges established voice AI providers such as ElevenLabs, Deepgram, and OpenAI.

Voxtral TTS supports nine languages, including English, French, Spanish, and Hindi, and allows users to clone a custom voice with less than five seconds of audio. The system captures subtle accents, inflections, and intonations while maintaining the voice's unique characteristics across different languages, making it highly useful for dubbing and real-time translation. Built for real-time performance, the model boasts a time-to-first-audio of 90 milliseconds and renders a 10-second audio clip in just 1.6 seconds.

This new release expands Mistral's growing portfolio of voice products following its earlier launch of transcription models earlier in the year. The company plans to build a comprehensive end-to-end platform that processes multimodal inputs like audio, text, and images to power advanced agentic systems. By keeping the model open source, Mistral aims to attract enterprise customers who want to deeply customize their voice solutions without being locked into proprietary ecosystems.

Read More at the original source →