OpenAI Expands ChatGPT with Voice Conversations and Image Understanding

OpenAI enhances ChatGPT with voice interaction capabilities and multimodal image processing powered by GPT-4V and DALL-E 3. The company rolls out these features gradually while addressing new safety concerns.

OpenAI introduces new voice and image capabilities to ChatGPT, allowing users to hold spoken conversations on the mobile app and upload images for analysis. The system relies on GPT-4 Vision (GPT-4V) to process visual inputs and integrates the latest DALL-E 3 model to generate images directly within the chat interface. Furthermore, the voice feature utilizes OpenAI's Whisper model for speech recognition and a new text-to-speech engine that offers five distinct voice options.

The addition of image understanding transforms ChatGPT into a multimodal tool, enabling it to describe photos, analyze visual data, and assist vision-impaired users through applications like Be My AI. By integrating DALL-E 3, the chatbot helps users craft better prompts to create highly specific images. These combined features bridge the gap between text, audio, and visual modalities in a single AI assistant.

OpenAI deploys these advancements gradually due to the expanded risk surface that multimodal models introduce compared to text-only systems. The company conducts extensive red teaming and beta testing to mitigate potential dangers, noting that combining vision and text creates novel capabilities and unique challenges. Over several months, thousands of beta testers and developers provide feedback to help OpenAI evaluate and refine the model's behavior and safety guardrails.

Read More at the original source →