OpenAI and Anthropic Break Rivalry for Joint AI Safety Testing

OpenAI and Anthropic briefly share access to their AI models to conduct joint safety evaluations, exposing blind spots in each other's internal testing. The rare collaboration highlights the growing need for industry-wide safety standards despite fierce market competition.

OpenAI and Anthropic briefly open up their closely guarded AI models for joint safety testing, marking a rare moment of collaboration between two fierce rivals. This cross-lab effort aims to uncover blind spots in each company's internal evaluations and show how leading AI organizations can work together on safety. OpenAI co-founder Wojciech Zaremba emphasizes that this teamwork is crucial as AI enters a consequential stage where millions rely on these systems daily.

The research arrives amid a massive arms race where billion-dollar data center investments and enormous researcher compensation packages are the norm. Some experts worry that this intense product competition pressures companies to cut corners on safety. To conduct the study, the labs grant each other special API access to less restricted versions of their models, though Anthropic later revokes a separate team's access for allegedly violating terms of service.

Despite these tensions, researchers from both sides express a strong desire to continue working together across the safety frontier. The study reveals stark differences in how the models handle uncertainty, with Anthropic's Claude models refusing to answer up to 70% of difficult questions to avoid hallucinations. Both teams hope to make this type of cross-lab safety evaluation a regular industry practice moving forward.

Read More at the original source →