OpenAI Launches GPT-4o with Real-Time Voice and Vision Features

OpenAI introduces GPT-4o, a native multimodal AI model that processes text, audio, and images simultaneously with human-like response speeds. The new flagship model rolls out to free and paid ChatGPT users alongside a new desktop application.

OpenAI unveils GPT-4o, a new flagship artificial intelligence model that processes text, audio, and image inputs simultaneously in real time. The "o" in the model's name stands for "omni," reflecting its ability to handle multiple types of data natively rather than relying on separate transcription and text-to-speech systems. This native processing allows the AI to respond to voice inputs with an average latency of 320 milliseconds, which closely matches typical human response times during conversations.

The new model matches GPT-4 Turbo's performance on English text while showing significant improvements in non-English languages. OpenAI Chief Technology Officer Mira Murati explains that previous voice modes required three separate models working together, which added latency and broke the sense of immersion. GPT-4o eliminates this issue by handling everything natively, creating a much more natural human-computer interaction that feels like speaking to another person.

OpenAI makes GPT-4o available to ChatGPT users for free and introduces a new desktop application for MacOS. Live demonstrations show the model displaying a wide range of emotional responses in its voice, including chuckling, soft sighs, and expressive tone changes, allowing users to direct the AI to tell stories with specific moods or vocal styles.

Read More at the original source →