Anthropic Hires Former OpenAI Safety Lead Jan Leike for New Superalignment Team
Jan Leike joins Anthropic to lead a new superalignment team after resigning from OpenAI and publicly criticizing its safety practices. The new group will focus on scalable oversight and automated alignment research.
Jan Leike, a prominent AI researcher who recently resigned from OpenAI and publicly criticized the company's safety practices, joins rival Anthropic to lead a new "superalignment" team. Leike announces on social media that his new group focuses on critical AI safety areas, including scalable oversight, weak-to-strong generalization, and automated alignment research.
At Anthropic, Leike reports directly to Chief Science Officer Jared Kaplan and oversees researchers who currently work on scalable oversight techniques. These techniques aim to control large-scale AI behavior in predictable and desirable ways. Anthropic researcher Sam Bowman expresses enthusiasm about the hire, noting that he and Leike will lead twin teams focused on aligning AI systems at human level and beyond.
This new team mirrors the mission of OpenAI's recently dissolved Superalignment team, which Leike previously co-led before experiencing friction with leadership. Anthropic consistently positions itself as a more safety-focused alternative to OpenAI, a philosophy that stems from CEO Dario Amodei's past departure from OpenAI over disagreements regarding the company's growing commercial focus.