Google Launches Gemini 1.5 Pro With Unprecedented One Million Token Context Window
Google introduces the Gemini 1.5 Pro AI model featuring a massive one million token context window capable of processing entire books or hours of video. The experimental model currently faces latency issues but offers a major leap in long-context understanding.
Google introduces the Gemini 1.5 Pro, a new generative AI model that succeeds Gemini 1.0 Pro with a massively upgraded context window. The new model handles up to one million tokens, which equates to roughly 700,000 words, 30,000 lines of code, 11 hours of audio, or an hour of video. This represents a 35-fold increase over the previous version and gives the model the longest context window of any large-scale foundation model currently available.
The multimodal capabilities of Gemini 1.5 Pro allow it to perform complex search and analysis tasks across extremely long texts and large video files. CEO Sundar Pichai notes that the 1.5 Pro achieves quality comparable to the larger 1.0 Ultra model while using significantly less computing power. However, processing these massive inputs takes between 20 seconds and a minute, presenting a notable latency issue that Google is actively working to resolve.
The full-scale version with the one million token capacity remains in a private preview phase for select enterprise customers using Vertex AI and AI Studio. A more accessible version with a 128,000 token context window is currently available, matching the capacity of competing high-end models like GPT-4-Turbo. While the experimental tier is free during this preview period, Google plans to introduce paid pricing tiers as the technology matures.