OpenAI Launches GPT-4o Model With Real-Time Voice and Vision

OpenAI introduces GPT-4o, an omni-modal AI model that processes text, speech, and video simultaneously. The upgrade brings real-time voice conversations and enhanced vision capabilities to ChatGPT.

OpenAI unveils GPT-4o, a new flagship AI model where the "o" stands for "omni" because it handles text, speech, and video natively. The company plans to roll out this GPT-4-level intelligence iteratively across its developer and consumer products over the coming weeks. OpenAI CTO Mira Murati states that this multimodal reasoning represents the future of human-machine interaction.

The update significantly improves the ChatGPT experience by replacing the old transcription-based voice mode with real-time, native speech. Users can now converse with the assistant naturally, interrupt it while speaking, and hear responses delivered in various emotive styles. Additionally, GPT-4o upgrades the platform's vision capabilities so it can quickly analyze photos or desktop screens to answer complex questions about software code or identify objects.

Murati notes that these features will continue to evolve as the model learns to process live video feeds, such as watching and explaining a sports game in real time. The ultimate goal is to make the interaction feel completely natural so users focus entirely on collaborating with the AI rather than navigating a user interface.

Read More at the original source →