Google Unveils Gemini Live for Real-Time Voice and Vision AI
Google introduces Gemini Live, a new conversational AI experience that combines real-time voice chats with visual understanding through smartphone cameras. Powered by DeepMind's Project Astra, the feature adapts to users' speech patterns and analyzes surrounding environments.
Google introduces a new experience for its Gemini chatbot called Gemini Live, allowing users to have fluid, in-depth voice conversations on their smartphones. People can interrupt the AI to ask clarifying questions, and the system adapts to their specific speech patterns in real time to create a highly intuitive dialogue. This upgraded interface essentially evolves the capabilities of Google Lens and Google Assistant into a single, more responsive tool.
Gemini Live also sees and responds to the user's physical surroundings using the smartphone's camera. By leveraging newer generative AI techniques, Google achieves superior and less error-prone image analysis compared to previous technologies. An enhanced speech engine gives the chatbot a more emotionally expressive and realistic voice, making the back-and-forth interaction feel much more natural.
These technical innovations stem from Project Astra, a new DeepMind initiative focused on creating AI agents capable of real-time, multimodal understanding. DeepMind CEO Demis Hassabis explains the goal is to build a universal agent that sees and hears what users do to better understand their context. Although Gemini Live does not launch until later this year, it promises to answer questions about anything currently visible through a user's smartphone camera.