Google Makes Gemini 2.5 Thinking Models Generally Available
Google announces the general availability of Gemini 2.5 Pro and Flash, while introducing a new cost-effective Flash-Lite preview model. The entire 2.5 family features "thinking" capabilities that allow developers to control how much the AI reasons before responding.
Google officially announces the general availability of its Gemini 2.5 Pro and Gemini 2.5 Flash models, bringing stable releases to developers. Both models operate as "thinking" models, meaning they reason through their internal processes before generating a final response to enhance overall accuracy. Developers have direct control over the thinking budget through an API parameter, allowing them to dictate exactly how much processing the model performs before delivering an answer.
In addition to the stable releases, Google introduces Gemini 2.5 Flash-Lite in preview, offering the lowest latency and cost within the 2.5 family. This model serves as a highly cost-effective upgrade from previous 1.5 and 2.0 Flash iterations, delivering faster time to first token and higher tokens per second. Flash-Lite is specifically optimized for high-throughput tasks like large-scale classification and summarization, and it fully supports native tools such as Grounding with Google Search, Code Execution, and function calling.
Because Flash-Lite prioritizes speed and efficiency, its thinking capability remains off by default, unlike its more robust siblings. Furthermore, Google updates the pricing structure for the stable version of Gemini 2.5 Flash to eliminate the previous confusion caused by separate "thinking" and "non-thinking" price tiers. These collective updates provide developers with a flexible, scalable, and clearly priced suite of reasoning models tailored for various application requirements.