Google Unveils Imagen AI Text-to-Image Generator With Impressive Realism
Google introduces Imagen, a new artificial intelligence system that creates highly photorealistic images from written text descriptions. The company keeps the tool private due to concerns about social biases and harmful stereotypes present in its training data.
Google Research unveils Imagen, a new artificial intelligence system that creates highly detailed images from written text prompts. The company claims this diffusion model offers an unprecedented level of photorealism and a deep understanding of language compared to similar tools like OpenAI's DALL-E 2. To prove its capabilities, Google creates a benchmark called DrawBench, where human raters consistently prefer Imagen over competing models in terms of both image quality and accuracy to the text prompts.
Despite the impressive examples on the Imagen website, viewers should note that these visuals represent a curated selection of the system's best work. Like other text-to-image generators, Imagen relies on massive datasets scraped directly from the internet to learn how to create its pictures. This uncurated approach to training data inherently introduces a variety of complex problems for the resulting AI.
Because the underlying datasets contain toxic language, racist slurs, and harmful social stereotypes, Google decides Imagen is not safe for public release. The researchers openly admit that their filtering efforts do not remove all the undesirable content, meaning the AI inherits the oppressive viewpoints and biases of the internet. Until these significant ethical and safety challenges are resolved, the system remains strictly in the hands of Google's researchers.