OpenAI Creates Dedicated Team to Tame Superintelligent AI

OpenAI launches a new Superalignment team, dedicating 20% of its computing power to solve the challenge of controlling AI systems that surpass human intelligence.

OpenAI forms a new Superalignment team led by chief scientist and co-founder Ilya Sutskever to develop methods for steering and controlling superintelligent AI. Sutskever and alignment lead Jan Leike predict that AI exceeding human intelligence could arrive within the next decade, posing significant risks if left unchecked. They emphasize that current alignment techniques, such as reinforcement learning from human feedback, fail because humans cannot reliably supervise systems much smarter than themselves.

To tackle this massive challenge, the new team receives 20% of the company's total secured compute power and aims to solve the core technical problems of superintelligence alignment within four years. The researchers plan to achieve this by building an automated, human-level alignment researcher that uses human feedback to learn how to evaluate other AI systems. This approach relies on the hypothesis that AI can make faster and better alignment research progress than humans can alone.

Ultimately, OpenAI envisions a future where AI systems take over more of the alignment workload, conceiving and developing better techniques to ensure their own successors remain safe and aligned with human values. Human researchers will gradually shift their focus from generating this research to reviewing the alignment work done by the AI systems. Despite this ambitious roadmap, the team acknowledges that no method is foolproof when it comes to controlling potentially rogue superintelligence.

Read More at the original source →