Google Updates Gemini 1.5 Models With Lower Costs and Faster Speeds
Google releases updated production-ready Gemini 1.5 Pro and Flash models that feature significantly reduced pricing, higher rate limits, and improved benchmark scores in math and vision tasks.
Google releases two updated production-ready Gemini models, Gemini-1.5-Pro-002 and Gemini-1.5-Flash-002, bringing significant improvements for developers. The company cuts the price of 1.5 Pro by over 50 percent for prompts under 128K tokens while doubling the rate limits for 1.5 Flash and tripling them for 1.5 Pro. Additionally, both models deliver twice as fast output and three times lower latency compared to their previous versions.
The updated models show notable performance gains across several key benchmarks. They achieve a seven percent increase on the challenging MMLU-Pro benchmark and a roughly 20 percent improvement on MATH and HiddenMath evaluations. The models also perform two to seven percent better on tasks involving visual understanding and Python code generation, making them more capable of handling long documents, large codebases, and video content.
Google also refines the overall user experience by adjusting default filter settings and improving response helpfulness. The models now produce fewer refusals and adopt a more concise response style based on developer feedback, which makes them easier to use and helps reduce operational costs. Developers can access these models for free through Google AI Studio and the Gemini API, or utilize them on Vertex AI for larger enterprise deployments.