OpenAI Creates Dedicated Team to Tame Superintelligent AI

OpenAI launches a new Superalignment team, dedicating 20% of its computing power to solve the technical challenges of controlling highly advanced AI systems that could arrive within the decade.

OpenAI forms a new Superalignment team led by chief scientist and co-founder Ilya Sutskever to tackle the challenge of steering and controlling superintelligent AI. Sutskever and alignment lead Jan Leike predict that AI systems exceeding human intelligence could arrive within the next decade. Because humans cannot reliably supervise AI that is much smarter than them, the company admits that current alignment methods like reinforcement learning from human feedback will not be enough to prevent these systems from going rogue.

To solve this massive problem, OpenAI dedicates 20% of its total secured computing power to the new team over the next four years. The group consists of top scientists and engineers from the company's previous alignment division alongside researchers from other internal organizations. Their primary objective is to build a human-level automated alignment researcher to automate the process of keeping future AI models safe and under control.

The team plans to train AI systems using human feedback and teach them to evaluate other AI systems, ultimately creating an AI that performs alignment research autonomously. OpenAI hypothesizes that AI can make faster and better progress on alignment research than humans can. As these automated systems take over more of the heavy lifting, human researchers will shift their focus to reviewing the AI-generated research rather than generating it themselves.

Read More at the original source →