Google Unveils Gemini 3.1 Flash-Lite for High-Volume AI Tasks

Google launches Gemini 3.1 Flash-Lite, a fast and highly cost-efficient AI model designed for large-scale developer workloads. The new model delivers significantly faster speeds than its predecessor while maintaining strong performance quality.

Google introduces Gemini 3.1 Flash-Lite as its fastest and most cost-efficient model in the Gemini 3 series. The company designs this lightweight model specifically to handle high-volume developer workloads at scale without sacrificing output quality. Developers access the new model in preview through the Gemini API in Google AI Studio, while enterprise customers connect via Vertex AI.

The pricing structure makes this model highly accessible, costing just $0.25 per one million input tokens and $1.50 per one million output tokens. According to the Artificial Analysis benchmark, 3.1 Flash-Lite outperforms the previous 2.5 Flash model with a 2.5 times faster Time to First Answer Token and a 45 percent increase in overall output speed. This low latency allows developers to build highly responsive, real-time user experiences.

Despite its low cost and fast speeds, the model achieves impressive quality metrics, including an Elo score of 1432 on the Arena.ai Leaderboard. Google recommends 3.1 Flash-Lite for a variety of practical applications, such as language translation, content moderation, user interface generation, and creating simulations. The model consistently outperforms similar models in its tier, proving that cost-efficiency does not require a compromise on intelligence.

Read More at the original source →