OpenAI Safety published findings from a joint safety evaluation conducted with Anthropic. This collaboration involved each lab running its internal safety and misalignment evaluations on the other's publicly released models. The initiative aims to support accountable and transparent evaluation, helping to ensure models are tested against new and challenging scenarios.
OpenAI's evaluations focused on Anthropic's models, specifically Claude Opus 4 and Claude Sonnet 4. The results were presented alongside data from GPT-4o, GPT-4.1, OpenAI o3, and OpenAI o4-mini, which powered ChatGPT at the time of the evaluations. Anthropic's evaluations of OpenAI's models are available separately.
Key Points
- OpenAI and Anthropic conducted a joint safety evaluation of each other's publicly released models.
- Evaluations covered misalignment, instruction following, hallucinations, and jailbreaking.
- OpenAI tested Anthropic's Claude Opus 4 and Claude Sonnet 4 models.
- OpenAI's results were compared with GPT-4o, GPT-4.1, OpenAI o3, and OpenAI o4-mini.
- Both labs relaxed some model-external safeguards to facilitate testing, a common practice for dangerous-capability evaluations.
- Claude 4 models demonstrated strong resistance to system prompt extraction attempts.
- Claude Opus 4 and Claude Sonnet 4 achieved 1.000 performance on the Password Protection evaluations, matching OpenAI o3.
Context
According to OpenAI Safety, the goal of this external evaluation is to identify gaps that might otherwise be missed, deepen understanding of potential misalignment, and demonstrate how labs can collaborate on safety and alignment issues. The field of AI continues to evolve, and models are increasingly used in real-world tasks, making continuous safety testing essential. The evaluations were designed to explore model propensities for concerning behaviors rather than to conduct full threat modeling or estimate real-world likelihoods.
Why It Matters
This collaboration establishes a precedent for cross-organizational safety evaluations, providing builders and researchers with insights into how leading models perform under adversarial conditions. The findings highlight specific areas of model resilience and potential vulnerabilities, informing future safety research and development practices.
What To Do
- Note the specific models evaluated by OpenAI (Claude Opus 4, Claude Sonnet 4) and the OpenAI models used for comparison (GPT-4o, GPT-4.1, OpenAI o3, OpenAI o4-mini).
- Review the reported performance of Claude 4 models on system prompt extraction tasks, particularly the 1.000 score on Password Protection.
- Watch for further joint evaluations or similar cross-lab collaborations as a mechanism for transparent safety assessment.
