OpenAI Safety has published the proceedings from a workshop titled "Confidence-Building Measures for Artificial Intelligence." This event, a collaboration between OpenAI's Geopolitics Team and the Berkeley Risk and Security Lab at the University of California, convened a multistakeholder group. The workshop aimed to identify tools and strategies to mitigate potential risks that foundation models could introduce to international security.
The workshop's focus was on applying confidence-building measures (CBMs) to the context of artificial intelligence. CBMs, which originated during the Cold War, are actions designed to reduce hostility, prevent conflict escalation, and improve trust between parties. Participants noted the flexibility of CBMs as a key instrument for navigating the rapid changes in the foundation model landscape.
Key Points
- The workshop was hosted by the Geopolitics Team at OpenAI and the Berkeley Risk and Security Lab at the University of California.
- Foundation models could introduce pathways for undermining state security, including accidents, inadvertent escalation, unintentional conflict, weapons proliferation, and interference with human diplomacy.
- Participants identified six specific CBMs applicable to foundation models: crisis hotlines, incident sharing, model/transparency/system cards, content provenance and watermarks, collaborative red teaming and table-top exercises, and dataset and evaluation sharing.
- Many CBMs will need to involve a wider stakeholder community because most foundation model developers are non-government entities.
- These measures can be implemented by either AI labs or relevant government actors.
Context
According to the workshop proceedings, foundation models present several potential risks to international security. These include the possibility of accidents, inadvertent escalation, unintentional conflict, the proliferation of weapons, and interference with human diplomacy. The workshop sought to adapt the concept of confidence-building measures, historically used to reduce hostility and prevent conflict, to address these emerging challenges in the AI domain.
Why It Matters
This report indicates that both AI labs and government actors may need to consider implementing specific measures to manage the international security implications of foundation models. The identified CBMs suggest areas where developers and policymakers might focus their efforts to build trust and reduce risks.
What To Do
- Review the full workshop proceedings to understand the detailed explanations of each identified CBM.
- Note the specific CBMs that can be implemented by AI labs versus government actors.
- Watch for further guidance or initiatives from OpenAI or the Berkeley Risk and Security Lab regarding these measures.
- Consider how the proposed CBMs might apply to the development and deployment practices of foundation models.
