Google Expands Gemini 1.5 Pro Access to Over 180 Nations
Google brings Gemini 1.5 Pro to a global audience with new native audio understanding and developer tools like system instructions and JSON mode. The update also introduces a next-generation text embedding model.
Google makes its Gemini 1.5 Pro model available in over 180 countries through a public preview in the Gemini API. This massive expansion allows developers worldwide to access the model's impressive one million token context window. Alongside the broader rollout, Google introduces a new File API to simplify file handling for building applications.
The update brings native audio, or speech, understanding directly to the model for the first time. Developers can now upload audio recordings, such as long lectures, and the model processes the content to perform tasks like generating quizzes with answer keys. Additionally, the model gains the ability to reason across both image frames and audio tracks within uploaded videos.
To give developers greater control, Google adds highly requested features including system instructions and JSON mode. System instructions allow users to define specific roles, formats, and rules to guide the model's behavior, while JSON mode forces the model to output strictly structured JSON objects for easier data extraction. Google also releases a new next-generation text embedding model that outperforms similar models on the market.