OpenAI Launches GPT-4o With Natural Voice And Vision Features
OpenAI introduces GPT-4o, a fully multimodal model that brings advanced voice assistant and vision capabilities to both free and paid ChatGPT users. The new model processes speech, images, and video natively for more natural conversations.
OpenAI unveils GPT-4o during its Spring Update event, bringing a groundbreaking natural voice assistant and advanced vision capabilities to ChatGPT. This new model is fully multimodal by design, meaning it natively understands and processes speech, images, and video content without needing to convert them to text first. The upgraded voice feature sounds highly emotional and conversational, marking a significant leap forward in real-time AI interaction.
The update brings major benefits to free ChatGPT users, as OpenAI opens up features previously reserved for paying subscribers. Free users now gain access to image and document analysis, data analytics, and custom GPT chatbots powered by the new GPT-4o model. While Plus subscribers receive immediate access, the company gradually rolls out these features to all users across mobile, desktop, and web platforms over the coming weeks.
Despite the impressive new capabilities, GPT-4o does not outperform the standard GPT-4 on traditional text-based tasks. Its true advantage lies in live speech, video analysis, and real-time translation across multiple languages. OpenAI keeps future plans under wraps during this event, offering no new release details for the upcoming GPT-5 model or the highly anticipated Sora AI video generator.