OpenAI Releases DALL-E 3 and Text-to-Speech APIs for Developers
OpenAI introduces a new DALL-E 3 API for image generation and an Audio API for text-to-speech capabilities during its first developer conference. Both tools come with specific limitations and safety features.
OpenAI releases the DALL-E 3 API, allowing developers to integrate its latest text-to-image model into their own applications. The API generates images in various resolutions up to 1792×1024 and starts at a price of $0.04 per image. However, this new version currently lacks the editing and variation capabilities that are available in the older DALL-E 2 API.
Alongside the image generator, OpenAI launches a new text-to-speech API called the Audio API. This tool provides six distinct preset voices and costs $0.015 per 1,000 characters of input text. The company claims the synthesized speech sounds highly natural, which enables new use cases like voice assistants and language learning applications.
Both new APIs include specific guardrails and limitations for developers. The DALL-E 3 API automatically rewrites user prompts for safety and added detail, which potentially reduces precision, while the Audio API does not allow direct control over the emotional tone of the generated voices. Additionally, OpenAI updates its open-source speech recognition model, Whisper large-v3, with improved multilingual performance.