OpenAI Safety has updated its Preparedness Framework, a system designed to measure and protect against severe harm from frontier AI capabilities. This update refines the approach to identifying and minimizing risks, providing clearer operational guidance for evaluation, governance, and disclosure of safeguards. The framework also introduces new research categories to anticipate emerging capabilities.
The update incorporates insights from internal testing, external experts, and field observations. OpenAI plans to continue publishing Preparedness findings with each frontier model release, as it has done for GPT-4o, OpenAI o1, Operator, o3-mini, deep research, and GPT-4.
Key Points
- The updated framework introduces a sharper focus on specific risks and stronger requirements for minimizing them.
- It provides clearer operational guidance on how OpenAI evaluates, governs, and discloses its safeguards.
- Future-facing research categories are included to understand emerging capabilities.
- A structured risk assessment process prioritizes capabilities that are plausible, measurable, severe, net new, and instantaneous or irremediable.
- Tracked Categories include Biological and Chemical capabilities, Cybersecurity capabilities, and AI Self-improvement capabilities.
- New Research Categories cover areas like Long-range Autonomy, Sandbagging, Autonomous Replication and Adaptation, Undermining Safeguards, and Nuclear and Radiological risks.
- Persuasion risks are handled outside the framework, through mechanisms like the Model Spec and restrictions on political campaigning.
- Capability levels are streamlined to High (amplifying existing harm pathways) and Critical (introducing unprecedented harm pathways).
- The Safety Advisory Group (SAG) reviews safeguards and makes recommendations to OpenAI Leadership.
- Scalable, automated evaluations are being developed to keep pace with more frequent model improvements.
Context
According to OpenAI, the updated framework refines the criteria for prioritizing high-risk capabilities. It uses a structured risk assessment to determine if a frontier capability could lead to severe harm, categorizing risks based on whether they are plausible, measurable, severe, net new, and instantaneous or irremediable. The framework distinguishes between established "Tracked Categories" with mature evaluations and ongoing safeguards, and new "Research Categories" for potential risks that require further threat modeling and evaluation development.
Why It Matters
This update signals OpenAI's evolving strategy for managing the safety of increasingly capable AI systems. For builders and researchers, understanding this framework provides insight into the types of risks OpenAI prioritizes and the mechanisms it employs for evaluation and mitigation, which can inform their own safety considerations when developing or deploying advanced AI.
What To Do
- Review the updated framework to understand OpenAI's current approach to severe harm risks.
- Note the distinction between "Tracked Categories" and "Research Categories" to identify areas of current and future safety focus.
- Consider how the defined capability levels (High and Critical) might influence deployment decisions for frontier models.
- Watch for future Preparedness findings published with new model releases to observe the framework in practice.
