Anthropic Launches Claude 3 Models to Challenge GPT-4 Dominance
Anthropic introduces the Claude 3 model family, featuring three tiers of AI models that claim to outperform the original GPT-4 on various cognitive benchmarks. However, important caveats exist regarding these performance comparisons against newer GPT-4 Turbo versions.
Anthropic releases the Claude 3 model family, introducing three state-of-the-art language models named Claude 3 Haiku, Claude 3 Sonnet, and Claude 3 Opus. These models offer escalating levels of capability, allowing users to balance intelligence, speed, and cost based on their specific application needs. Anthropic claims that Claude 3 Opus, the most powerful model in the lineup, achieves near-human levels of comprehension and fluency on complex tasks.
The company asserts that Claude 3 Opus outperforms GPT-4 on several common evaluation benchmarks, including undergraduate-level expert knowledge, graduate-level expert reasoning, and basic mathematics. However, a significant caveat accompanies these claims, as the comparisons primarily target the original GPT-4-0314 model rather than the more recent and capable GPT-4-Turbo variants. Where data exists for the newer GPT-4-1106-preview model, it consistently outperforms Claude 3 Opus.
Despite the benchmark nuances, Claude 3 Opus likely surpasses all current GPT-4 models specifically on the GPQA graduate-level reasoning benchmark. Additionally, prediction markets like Manifold show strong confidence that Claude 3 will ultimately outrank GPT-4 on the popular LMSYS Chatbot Arena Leaderboard. Anthropic also reveals that the Claude 3 family trains on synthetic data while strictly avoiding the use of any customer-generated data from previous models.