Meta Unveils Llama 4 Herd With Native Multimodal Capabilities
Meta launches Llama 4 Scout and Maverick as the first open-weight natively multimodal models using a mixture-of-experts architecture. The new models deliver industry-leading context windows and outperform major competitors across various benchmarks.
Meta introduces the Llama 4 herd, marking a significant leap in natively multimodal artificial intelligence. This new suite features Llama 4 Scout and Llama 4 Maverick as the first open-weight models built on a mixture-of-experts (MoE) architecture. Both models boast 17 billion active parameters and natively process text, images, and video to help developers build highly personalized experiences.
Llama 4 Scout utilizes 16 experts and delivers impressive performance while fitting entirely on a single NVIDIA H100 GPU. It features an industry-leading context window of 10 million tokens and outperforms competitors like Gemma 3 and Gemini 2.0 Flash-Lite across multiple benchmarks. Meanwhile, Llama 4 Maverick employs 128 experts to beat GPT-4o and Gemini 2.0 Flash, achieving comparable reasoning and coding results to DeepSeek v3 at less than half the active parameters.
These efficient models derive their intelligence from Llama 4 Behemoth, a massive 288 billion active parameter model that currently outsmarts GPT-4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on several STEM benchmarks. Although Behemoth is still in training, Meta makes Scout and Maverick available for download today on llama.com and Hugging Face, while also integrating the technology into Meta AI across WhatsApp, Messenger, and Instagram.