OpenAI Dedicates Major Computing Power to Prevent Rogue Superintelligent AI
OpenAI forms a new Superalignment team, dedicating a fifth of its computing power to ensure future superintelligent AI systems remain safe and aligned with human values.
OpenAI creates a new Superalignment team led by chief scientist and co-founder Ilya Sutskever to tackle the potential dangers of superintelligent AI. The company dedicates a fifth of its total computing power to this group, which focuses on building technical guardrails to keep advanced AI systems aligned with human needs. OpenAI predicts that AI could surpass human intelligence within the next decade, making this proactive safety research essential.
The new team aims to solve the major flaw in current AI safety methods, which rely on human supervision. Because humans will not be able to reliably supervise AI systems that are much smarter than them, current techniques like reinforcement learning from human feedback will not scale to superintelligence. To overcome this, the Superalignment team plans to develop an automated AI evaluator that judges other AI models based on learned human preferences.
OpenAI sets an ambitious goal to create this technical solution within four years. The company promises to share the findings of this research broadly to improve the safety of non-OpenAI models as well. By combining this technical work with input from interdisciplinary experts, OpenAI hopes to address both the machine learning and broader societal challenges of superintelligence.