Alibaba Releases Qwen3-Omni as Free Open Source Multimodal AI
Alibaba's Qwen team launches Qwen3-Omni, a fully open source multimodal AI model that processes text, image, audio, and video under an Apache 2.0 license. The model outperforms proprietary rivals like GPT-4o in several benchmarks while offering extensive language support and low-latency inference.
Alibaba's Qwen team releases Qwen3-Omni as a fully open source, natively end-to-end multimodal AI model. Available under an Apache 2.0 license, the model allows developers and enterprises to freely download, modify, and deploy the system for commercial applications. It processes text, image, audio, and video inputs and generates both text and audio outputs, creating a powerful free alternative to proprietary systems from US tech giants.
The model supports 119 languages for text, 19 for speech input, and 10 for speech output, including specific dialects like Cantonese. It utilizes a unique Thinker–Talker architecture where the Thinker component handles reasoning and multimodal understanding while the Talker generates natural speech. A Mixture-of-Experts design ensures fast inference, achieving streaming latency as low as 234 milliseconds for audio and 547 milliseconds for video.
Benchmark results reveal that Qwen3-Omni surpasses GPT-4o and Gemini 2.0 Flash across multiple text, reasoning, speech, and vision metrics. The system offers three distinct variants: an Instruct Model for full capabilities, a Thinking Model for text-only reasoning, and a Captioner Model for low-hallucination audio captioning. Practical applications span multilingual transcription, translation, OCR, video understanding, and real-time AI assistants, marking a major shift in the open source AI landscape.