Anthropic Alignment has published details on TASTE (The AI Safety Taste Evaluation), a benchmark created to assess the ability of AI models to evaluate AI safety research proposals. This benchmark measures agreement with the preferences of experienced human researchers. The work was conducted as part of the Anthropic Fellows Program.
Evaluating research proposals is a critical, high-leverage aspect of improving research quality, particularly in areas of AI safety where progress is difficult to measure with verifiable rewards. Such areas include mitigating risks from AI misalignment or detecting model deception, which often involve forecasting future AI system behaviors or attributing beliefs and intentions to models.
Key Points
- TASTE is a benchmark measuring how well models judge pairs of AI safety research proposals based on agreement with experienced human researchers.
- The benchmark contains 92 pairwise comparisons of research proposals.
- Estimated human agreement with the benchmark's labels is 77%.
- A discussion stage among researchers and filtering for self-reported "strong" confidence were key design choices, raising estimated human agreement by 15 percentage points.
- Claude Opus 4.6 was used to generate the initial AI safety research proposals for the benchmark.
- The best performing model, Fable 5, achieved 60% agreement, which is lower than the 77% estimated human researcher performance.
Context
The TASTE benchmark was developed to address the challenge of evaluating AI models on hard-to-verify tasks within AI safety research, where human judgment serves as the ground truth. According to Anthropic, achieving high agreement among human evaluators for subjective outputs like research proposals is difficult. The benchmark's design incorporated a three-stage process: generating proposals using Claude Opus 4.6, recruiting AI safety researchers to rate and comment on these proposals, and then filtering the human feedback to create a high-agreement dataset.
Researchers initially provided individual feedback, then discussed disagreements in pairs, and subsequently revised their preferences. This discussion stage, combined with filtering for self-reported "strong" confidence, increased estimated human agreement from 53% pre-discussion to 68% for strong-confidence, post-discussion preferences. The final TASTE dataset consists of 92 pairs with 77% estimated human agreement, derived from strong-confidence, post-discussion preferences where scores differed by at least two points, and proposals appeared a maximum of ten times.
Why It Matters
This benchmark provides a method for quantifying AI models' capabilities in a crucial, yet subjective, aspect of AI safety research: evaluating research proposals. For builders and researchers, it highlights a current gap between AI model performance and human expert judgment in assessing the quality and relevance of safety research directions, indicating that fully automating this aspect of research remains a challenge.
What To Do
- Note the 77% estimated human agreement on the TASTE benchmark as a reference point for evaluating model performance in subjective research assessment tasks.
- Compare the 60% performance of Fable 5 against the human baseline to understand the current capabilities gap in AI safety research proposal evaluation.
- Consider the benchmark's design choices, such as the discussion stage and confidence filtering, when developing evaluation methodologies for other subjective AI tasks.