Researchers Test DALL-E 2 Reasoning Skills With Challenging Prompts

A new study by Gary Marcus, Ernest Davis, and Scott Aaronson puts DALL-E 2 to the test with difficult prompts to evaluate its common sense and reasoning abilities. The results show that while the AI sometimes succeeds, it struggles to consistently produce accurate images for complex text inputs.

Researchers Gary Marcus, Ernest Davis, and Scott Aaronson conduct a preliminary analysis of the DALL-E 2 text-to-image system to evaluate its common sense, reasoning, and comprehension of complex texts. The team specifically designs fourteen difficult prompts that are intentionally much harder than the typical examples recently showcased online.

The test results reveal a mixed performance from the AI image generator. For five out of the fourteen challenging prompts, at least one out of the ten generated images fully satisfies the researchers' specific requests. However, the system never manages to get a perfect score, as there is not a single prompt where all ten generated images successfully meet the criteria.

This analysis highlights the current limitations of advanced artificial intelligence in handling complex, nuanced instructions. While DALL-E 2 demonstrates impressive capabilities in basic image generation, it still lacks the robust common sense and consistent reasoning required to reliably interpret and execute highly sophisticated text prompts.

Read More at the original source →