OpenAI Brings Real-Time Video Understanding to ChatGPT After Long Delay
OpenAI finally launches real-time video understanding in ChatGPT's Advanced Voice Mode, allowing paid users to point their cameras at objects for instant analysis. The rollout excludes Enterprise and Edu subscribers until January and has no timeline for European users.
OpenAI finally releases real-time video understanding for ChatGPT, bringing a highly anticipated feature to its Advanced Voice Mode nearly seven months after the initial demonstration. Subscribers to ChatGPT Plus, Team, or Pro point their phone cameras at physical objects or share their device screens to receive near real-time spoken responses. The AI acts as a visual assistant that explains settings menus, analyzes drawings, and helps solve math problems on the fly.
Accessing the new visual capabilities requires users to tap the voice icon next to the chat bar and then select the newly added video icon on the bottom left of the screen. For screen sharing, users navigate to the three-dot menu and choose the "Share Screen" option. The rollout begins on Thursday and continues over the next week, though the release shows a staggered approach across different user tiers.
Not all subscribers receive immediate access to the feature. OpenAI delays the launch for ChatGPT Enterprise and Edu subscribers until January and currently offers no timeline for users in the EU, Switzerland, Iceland, Norway, or Liechtenstein. Despite the impressive demonstrations, the technology still shows flaws, as a recent broadcast reveals the AI making mistakes on geometry problems and experiencing occasional hallucinations during visual analysis tasks.