OpenAI Expands ChatGPT Beyond Text With Voice and Image Features
OpenAI brings voice conversations and image searches to ChatGPT, transforming the text-based chatbot into a more interactive assistant. The update arrives as major tech giants intensify their competition in the generative AI space.
OpenAI expands ChatGPT beyond its original text-based format by introducing new voice and image capabilities to the generative AI assistant. Users now have the ability to engage in spoken conversations with the chatbot, requesting things like bedtime stories or verbal answers to their questions. Additionally, the platform supports image-based searches, allowing individuals to upload pictures to receive explanations or step-by-step instructions.
This major update combines the familiar concept of voice assistants with OpenAI's powerful large language models. The voice feature relies on a new text-to-speech model that produces human-like audio from just a few seconds of sampled speech, utilizing five distinct voices created in collaboration with professional voice actors. The system uses OpenAI's open source Whisper technology to accurately transcribe the user's spoken words into text.
The announcement coincides with heightened competition in the generative AI sector, notably on the same day Amazon commits up to $4 billion to OpenAI rival Anthropic. As part of this rollout, Spotify serves as a launch partner and uses the voice technology to help podcasters translate their English shows into Spanish, French, or German while preserving their original vocal identity. OpenAI limits the initial release to specific partners, including notable podcasters, to carefully manage the deployment of this advanced technology.