Anthropic, an AI safety and research company, recently submitted its response to the NTIA’s Request for Comment on AI Accountability. The submission details Anthropic's perspective on the necessary processes and infrastructure for ensuring AI accountability, particularly for highly capable and general-purpose AI models.
Key Points
- Anthropic recommends increased funding for AI model evaluation research to develop rigorous, standardized evaluations.
- Companies deploying AI systems should be required to disclose evaluation methods and results, with provisions for protecting intellectual property.
- Government agencies like NIST should work to establish industry evaluation standards and best practices for AI models.
- Standard capabilities evaluations for AI systems should be developed, focusing on critical risks such as deception and autonomy.
- A process for AI developers to pre-register large training runs with their national government should be established, including model specifications and safety plans.
- Third-party auditors should be technically literate, security-conscious, and flexible to conduct robust yet lightweight assessments.
- External red teaming should be mandated as a precondition for developers releasing advanced AI systems.
- Increased funding for interpretability research is recommended, recognizing that regulations demanding interpretable models are currently infeasible but may be possible in the future.
Context
- According to Anthropic, there is currently no robust and comprehensive process for evaluating today’s advanced artificial intelligence systems. Their recommendations consider the *NTIA’s
- potential role as a coordinating body that sets standards in collaboration with other government agencies like the National Institute of Standards and Technology (NIST).
Why It Matters
Establishing effective frameworks for AI accountability requires collaboration across researchers, AI labs, regulators, auditors, and other stakeholders. The proposals aim to mitigate AI risks while realizing its benefits, emphasizing that robust accountability and auditing mechanisms are vital for ensuring AI's transformative effects are positive.