Google Unveils Imagen: A Photorealistic Text-to-Image AI Model

Google Research introduces Imagen, a new text-to-image diffusion model that combines large language models with high-fidelity image generation to achieve unprecedented photorealism. The model outperforms competitors in human evaluations and sets a new state-of-the-art benchmark.

Google Research introduces Imagen, a state-of-the-art text-to-image diffusion model that achieves an unprecedented level of photorealism and language understanding. The system builds on the strengths of large transformer language models for text comprehension and diffusion models for high-fidelity image generation.

The researchers discover that scaling up the language model significantly boosts image quality and text alignment, proving even more effective than simply increasing the size of the image diffusion model. Imagen achieves a remarkable FID score of 7.27 on the COCO dataset without ever training on it, and human evaluators rate the generated images as on par with actual COCO data.

To rigorously test this technology, Google introduces DrawBench, a comprehensive benchmark for comparing text-to-image models. In side-by-side comparisons against systems like DALL-E 2 and Latent Diffusion Models, human raters consistently prefer Imagen for both sample quality and accurate image-text alignment.

Read More at the original source →