AI Benchmark Nears Solution, Raising Doubts About AGI Testing
Top submissions to the ARC-AGI benchmark achieve a new high score of 55.5%, but the competition's organizers warn this success stems from brute-force workarounds rather than true reasoning.
The ARC-AGI benchmark, a prominent test designed to measure progress toward artificial general intelligence, sees its highest scores yet following a million-dollar competition. Created in 2019 by AI expert Francois Chollet, the test evaluates whether an AI system efficiently acquires new skills outside its training data rather than relying on memorized patterns. The top submission out of nearly 18,000 entries achieves a 55.5% score, marking a significant jump from last year's 33% but still falling short of the 85% threshold required to win the prize.
Despite the impressive score increase, Chollet and competition co-founder Mike Knoop warn that this milestone exposes flaws in the benchmark's design rather than signaling a genuine breakthrough in general intelligence. They point out that many successful submissions simply brute-force their way to solutions instead of demonstrating true novel reasoning. Chollet has long been critical of large language models, arguing that their reliance on pattern memorization prevents them from generating new reasoning when faced with unfamiliar situations.
The organizers view these results as proof that the current benchmark requires an update to prevent workarounds and better distinguish between actual intelligence and statistical guessing. Moving forward, the ARC Prize team plans to refine the test to ensure it forces AI systems to develop authentic generalization skills. This ongoing evolution highlights the immense difficulty of creating a standardized, foolproof method to measure true machine intelligence.