OpenAI Grants ChatGPT Voice and Vision Capabilities in Major Update

OpenAI equips ChatGPT with the ability to process images, understand spoken words, and respond with synthetic voices. The rollout targets paying users as tech giants fiercely compete to integrate generative AI into everyday consumer life.

OpenAI announces a major update that allows ChatGPT to see, hear, and speak, marking its most significant enhancement since the launch of GPT-4. Users access voice conversations and choose from five different synthetic voices directly within the mobile app. Additionally, users share images with the chatbot and highlight specific areas to receive detailed analysis.

The company rolls out these new features to paying subscribers over the next two weeks, with voice functionality restricted to iOS and Android applications while image processing works across all platforms. This aggressive feature push reflects the high stakes of the ongoing artificial intelligence arms race among industry leaders like OpenAI, Microsoft, Google, and Anthropic.

Despite the technological leap, experts raise valid concerns regarding the potential misuse of AI-generated synthetic voices for creating convincing deepfakes and bypassing cybersecurity systems. OpenAI addresses these worries by confirming that all synthetic voices originate from direct collaborations with professional voice actors rather than scraped audio from strangers.

Read More at the original source →