Anthropic Research recently published an analysis detailing how the values expressed by Claude models vary across different models and languages. The study involved examining 300,000 real conversations, compressing the observed values into four interpretable axes to understand key patterns in Claude's responses.
Key Points
- The research analyzed 300,000 real conversations from Claude.ai to measure expressed values.
- Values were compressed into four interpretable axes, capturing 15% of the variation in Claude's values.
- The study compared value expression across Claude models, specifically Sonnet 4.6, Opus 4.6, and Opus 4.7.
- Value profiles align with subjective perceptions: Sonnet 4.6 leans towards user deference and emotional warmth, while Opus 4.7 emphasizes accuracy, precision, and guarding against misuse.
- Values also vary across the top 20 languages used on Claude.ai.
- The largest variation across languages was observed on the Warmth vs. Rigor axis.
- Claude expressed warmth-related values most in Arabic and Hindi, and rigor-related values most in English and Russian.
Context
According to Anthropic Research, the study aimed to make the analysis of thousands of distinct values tractable by compressing them into a small number of axes. This approach allows for quantifying key differences between models and potentially connecting value variation to different training decisions. The research also sought to understand how user experience compares across the many languages used with Claude, building on previous findings that Claude behaves somewhat differently in various languages.
To construct these value axes, Anthropic began with 3,307 values identified in prior work, clustering them into 339 high-level values. A privacy-preserving analysis tool then sampled 309,815 Claude.ai conversations where users posed subjective tasks. The sample was drawn equally from Sonnet 4.6, Opus 4.6, and Opus 4.7 models and the 20 most common languages, resulting in approximately 5,000 conversations per model-language pair. Claude was used to label the presence or absence of each of the 339 high-level values in every conversation, followed by dimensionality reduction to compress these labeled values into axes.
Why It Matters
This research provides a method for empirically understanding the values expressed by AI models and how these values shift across different contexts, such as model versions and languages. For builders, this approach offers a way to test how factors like behavioral training or cultural context influence AI's expressed values. For curious readers, it illuminates the nuanced ways AI models adapt their responses based on design choices and linguistic environments, moving beyond high-level constitutional guidelines to measurable behavioral patterns.