Google Lens Gains AI-Powered Video Question Answering
Google Lens now allows users to ask questions about moving video recordings, powered by a customized Gemini model. The feature beats similar upcoming tools from Meta and OpenAI to market.
Google updates its visual search tool, Lens, with a new feature that answers questions about video recordings in near-real-time. English-speaking users on Android and iOS devices now capture video through the Google app and ask questions about objects in their surroundings. Director of product management Lou Wang explains that a customized Gemini AI model powers this experience by making sense of both the video content and the user's spoken questions.
To use this new video analysis capability, users must join Google's Search Labs program and opt into the "AI Overviews and more" experimental features. By holding down the smartphone shutter button, users activate the video-capturing mode and speak their questions aloud. The AI identifies the most relevant and interesting frames in the video to ground its answers, which come from Google Search's AI Overviews feature.
This launch positions Google ahead of competitors like Meta and OpenAI, who recently previewed similar real-time video understanding tools for their own hardware and software. However, Google's implementation remains asynchronous rather than allowing for a continuous, real-time conversation. Wang notes that this update stems directly from observing how people naturally attempt to use visual search technology to satisfy their curiosity about the world around them.