Google Expands Gemini 1.5 Pro to Over 180 Countries with Audio Features
Google makes its Gemini 1.5 Pro model available in over 180 countries, introducing native audio understanding, system instructions, and JSON mode to give developers greater control.
Google expands access to its Gemini 1.5 Pro model by launching a public preview in over 180 countries through the Gemini API. Building on the success of its one-million-token context window, the updated model brings powerful new capabilities to developers worldwide. Users easily access these features by grabbing an API key in Google AI Studio.
The update introduces native audio understanding, allowing the model to process speech directly alongside text and images. Developers can upload lengthy audio recordings, such as lectures, and use the model to generate quizzes or extract key insights. Additionally, Gemini 1.5 Pro currently reasons across both video frames and audio tracks within Google AI Studio, with API support for this multimodal video processing arriving soon.
To give developers more control over model outputs, Google adds highly requested features like system instructions and JSON mode. System instructions let users define specific roles, formats, and rules to guide the AI's behavior, while JSON mode forces the model to return strictly structured data for easier application integration. Google also releases a next-generation text embedding model that outperforms comparable options on the market.