Mistral Unveils Open-Source Mistral 3 Model Family With Frontier Large MoE

Mistral launches the Mistral 3 family, featuring the powerful Mistral Large 3 mixture-of-experts model and three highly efficient small dense models. All new models release under the Apache 2.0 license to maximize accessibility for developers.

Mistral officially introduces the Mistral 3 model family, headlined by Mistral Large 3, a sparse mixture-of-experts architecture containing 675 billion total parameters and 41 billion active parameters. The company also releases three compact dense models—14B, 8B, and 3B—known as the Ministral series. Every model in this new generation operates under the permissive Apache 2.0 license, enabling widespread community use and enterprise customization.

Mistral Large 3 ranks as one of the top open-weight models globally, achieving parity with leading instruction-tuned competitors on general prompts while adding image understanding capabilities and exceptional multilingual conversation skills. The model trains from scratch on 3,000 NVIDIA H200 GPUs and claims the second spot in the open-source non-reasoning category on the LMArena leaderboard. Mistral promises that a dedicated reasoning version of this flagship model is arriving soon.

To ensure broad accessibility, Mistral collaborates with NVIDIA, vLLM, and Red Hat to provide highly optimized deployment options. Developers run the massive Mistral Large 3 efficiently on a single 8x A100 or H100 node, or on Blackwell NVL72 systems, thanks to a compressed NVFP4 format checkpoint. Across the entire Mistral 3 lineup, the training process leverages NVIDIA Hopper GPUs and HBM3e memory to handle these demanding frontier-scale workloads.

Read More at the original source →