OpenAI Launches Free GPT-4o Omnimodel for Real-Time Voice and Video

OpenAI introduces GPT-4o, a unified AI model that enables real-time voice, video, and text interactions for free. The new "omnimodel" promises faster responses and more natural conversations than previous versions.

OpenAI unveils GPT-4o, a new flagship AI model that allows users to interact through real-time voice conversations, live video streams, and text within a single system. Unlike previous iterations that relied on separate models for different inputs, this "omnimodel" combines all capabilities to deliver faster response times and smoother task transitions. The company makes this powerful tool available for free to all users via its app and web interface.

During a live demonstration led by CTO Mira Murati, the model shows impressive conversational abilities that easily surpass traditional virtual assistants like Siri or Alexa. Users seamlessly interrupt the AI mid-sentence, prompting it to stop, listen, and adjust its response accordingly. The system also demonstrates remarkable versatility by instantly changing its tone and vocal style on command, shifting effortlessly from a dramatic storyteller to a convincing robot voice.

This launch strategically precedes Google's major I/O conference, highlighting OpenAI's push to dominate the consumer AI assistant market. By removing the silos between text, audio, and visual processing, GPT-4o creates a much more natural and collaborative interaction between humans and machines. Paid subscribers still receive the benefit of higher usage limits, but the baseline free access represents a significant shift in making advanced multimodal AI available to the general public.

Read More at the original source →