OpenAI Unveils Framework to Thwart Catastrophic AI Risks
OpenAI introduces a new preparedness framework to track and mitigate extreme dangers posed by advanced AI models. The document grants the board of directors final authority to block the release of risky technology.
OpenAI releases a 27-page Preparedness Framework that details how the company tracks, evaluates, and protects against catastrophic risks from advanced AI models. The document addresses worst-case scenarios such as mass cybersecurity disruptions and the creation of biological, chemical, or nuclear weapons. This framework highlights the company's proactive stance on AI safety as it continues to develop powerful new technologies.
The new system establishes a clear chain of command regarding the release of AI models. While company leadership makes the initial deployment decisions, the board of directors retains the final say and the right to reverse any executive decisions. Before reaching a board veto, models must pass multiple safety checks administered by a dedicated preparedness team led by MIT professor Aleksander Madry.
Madry and his team of researchers evaluate potential dangers and synthesize their findings into detailed risk scorecards. These scorecards categorize threats as low, medium, high, or critical, establishing strict thresholds for development and deployment. OpenAI states that only models with a post-mitigation score of medium or below can be deployed, and only models scoring high or below can be developed further. The company notes that this beta document will be updated regularly based on ongoing feedback.