OpenAI Launches Hyper-Realistic ChatGPT Voice Feature for Plus Users
OpenAI is rolling out an alpha version of its highly anticipated Advanced Voice Mode to a select group of ChatGPT Plus users. The new feature delivers ultra-realistic, low-latency audio conversations using the native multimodal capabilities of GPT-4o.
OpenAI officially begins rolling out ChatGPT's Advanced Voice Mode to a small group of ChatGPT Plus users, marking the first public access to GPT-4o's hyperrealistic audio responses. The company plans a gradual expansion to all Plus subscribers by fall 2024, following a June delay implemented to improve safety measures. This initial release does not include the video and screensharing capabilities that OpenAI showcased during its spring update.
Unlike the existing Voice Mode that relies on three separate models to transcribe, process, and generate speech, the new Advanced Voice Mode uses GPT-4o's native multimodal architecture. This unified approach creates significantly lower latency, allowing for natural, real-time conversations. Additionally, OpenAI claims the system detects and responds to emotional intonations in the user's voice, such as excitement, sadness, or singing.
The launch follows a controversy that erupted in May when the demo voice "Sky" drew widespread comparisons to actress Scarlett Johansson's performance in the movie "Her." Johansson stated she rejected multiple licensing requests from CEO Sam Altman and hired legal counsel, prompting OpenAI to remove the voice while denying any intentional resemblance. For now, OpenAI monitors the alpha release closely, sending in-app alerts and email instructions to the selected users who gain early access.