OpenAI Rolls Out Hyperrealistic Advanced Voice Mode to Select ChatGPT Plus Users

OpenAI begins a limited alpha release of ChatGPT's highly anticipated Advanced Voice Mode, which promises human-like conversations with ultra-low latency. The initial rollout excludes video and screen-sharing features as the company closely monitors user interactions.

OpenAI launches a limited alpha release of ChatGPT’s Advanced Voice Mode, giving a small group of ChatGPT Plus users access to GPT-4o’s hyperrealistic audio capabilities. The company plans a gradual rollout to all Plus subscribers later in the fall of 2024, following a June delay implemented to improve safety measures. This initial release focuses solely on audio, as the impressive video and screen-sharing features showcased in the spring remain under development for a later date.

Unlike the previous Voice Mode that relied on three separate models to translate, process, and generate speech, GPT-4o handles everything natively as a multimodal system. This unified approach creates significantly lower latency, allowing for smooth, natural conversations that closely mimic human interaction. Additionally, OpenAI claims the advanced system detects and responds to emotional intonations in the user's voice, such as sadness, excitement, or even singing.

The new release follows a controversy in which the original demo voice "Sky" drew widespread comparisons to actress Scarlett Johansson, leading to legal action and the eventual removal of that specific voice. OpenAI is taking a cautious approach with this rollout, notifying selected alpha testers through in-app alerts and emails while closely monitoring how they use the new feature. The company notes that it conducted extensive external testing with over 100 experts prior to this pilot launch.

Read More at the original source →