OpenAI Safety published an update detailing the company's safety practices and its participation in the AI Seoul Summit. The company outlined its approach to integrating safety measures throughout the development process, from initial design to deployment. This includes a commitment to evolving practices for future, more capable models.
Key Points
- OpenAI shared ten active safety practices that are continuously improved.
- The company joined industry leaders, government officials, and civil society members at the AI Seoul Summit to discuss AI safety.
- OpenAI agreed to additional Frontier AI Safety Commitments, which include publishing safety frameworks like the Preparedness Framework.
- Before release, models undergo empirical red-teaming and testing; a new model will not be released if it crosses a "Medium" risk threshold from the Preparedness Framework without sufficient safety interventions.
- Over 70 external experts assessed risks associated with GPT-4o through external red teaming.
- OpenAI uses GPT-4 for content policy development and content moderation decisions.
- The company partners with organizations like Thorn's Safer to detect and report Child Sexual Abuse Material in image tools.
Context
According to OpenAI Safety, the company views safety as a continuous investment across multiple time horizons, from current models to future, more capable systems. This work is integrated across OpenAI, with increasing investment over time. The company emphasizes a balanced, scientific approach where safety measures are built into the development process from the outset.
Why It Matters
These disclosures provide insight into the specific mechanisms and partnerships OpenAI employs to manage risks associated with its AI models. Builders and practitioners can observe the types of commitments and internal frameworks that guide the development and deployment of frontier AI systems, influencing how future models may be evaluated and released.
What To Do
- Note the Frontier AI Safety Commitments and watch for further details on their implementation.
- Review the mentioned Preparedness Framework to understand OpenAI's internal risk assessment thresholds.
- Observe how external red-teaming efforts, such as those for GPT-4o, inform model evaluations.
- Consider the role of GPT-4 in content policy and moderation, and how this might influence feedback loops for policy refinement.
