OpenAI Outlines Three-Part Strategy to Align Future AI Systems

OpenAI adopts an iterative, empirical approach to AI alignment, focusing on scalable training signals based on human intent. The company relies on three main pillars, including using human feedback to train models like InstructGPT.

OpenAI takes an iterative, empirical approach to artificial general intelligence (AGI) alignment by actively testing techniques on highly capable AI systems. Researchers study how these alignment methods scale and exactly where they break down in order to refine their safety protocols. The organization believes that even without fundamentally new theories, they can build sufficiently aligned systems to significantly advance alignment research itself.

The core strategy revolves around engineering a scalable training signal that keeps advanced AI aligned with human intent. This high-level approach rests on three main pillars: training AI systems using human feedback, training AI systems to assist human evaluation, and training AI systems to conduct alignment research. Reinforcement learning from human feedback currently serves as the primary technique for aligning deployed language models like InstructGPT.

Because unaligned AGI poses substantial risks to humanity, OpenAI commits to openly sharing its alignment research whenever it is safe to do so. The company wants to remain transparent about the practical effectiveness of its alignment techniques and hopes to ensure that all AGI developers eventually adopt the world's best alignment methods. However, OpenAI deliberately avoids the broader sociotechnical challenges of alignment, such as deciding exactly whose values the AI should follow.

Read More at the original source →