Anthropic announced a research demonstration involving its Claude 3 Sonnet model, which allowed users to interact with a version of the model dubbed "Golden Gate Claude." This demonstration, available for a 24-hour period, illustrated the company's work on interpreting large language models by manipulating internal features.
Key Points
- Anthropic released a research paper on interpreting large language models, focusing on the internal workings of Claude 3 Sonnet.
- The research identified millions of "features" within Claude's neural network, which are concepts that activate when the model processes relevant text or images.
- One such identified feature corresponds to the Golden Gate Bridge, activating when the model encounters mentions or images of the landmark.
- Researchers can adjust the activation strength of these features, leading to observable changes in Claude's behavior.
- A demonstration model, "Golden Gate Claude," was made available for 24 hours, exhibiting responses consistently focused on the Golden Gate Bridge due to heightened feature activation.
- This manipulation is described as a precise, surgical change to the model's internal activations, distinct from system prompts or fine-tuning.
- The same techniques could potentially be used to alter the strength of safety-related features, such as those concerning dangerous computer code or criminal activity.
Context
According to Anthropic, their recent research paper details the mapping of the inner workings of Claude 3 Sonnet. This work revealed millions of concepts, termed "features," that activate within the model's neural network in response to relevant input. One specific feature identified corresponds to the Golden Gate Bridge, activating when the model processes information related to the landmark.
Anthropic states that not only can these features be identified, but their activation strength can also be adjusted. This capability was demonstrated with "Golden Gate Claude," where increasing the strength of the Golden Gate Bridge feature caused the model's responses to consistently incorporate the landmark, even when not directly relevant to the query.
Why It Matters
This research offers insight into the internal mechanisms of large language models, moving beyond external controls like prompting or fine-tuning. The ability to precisely manipulate specific conceptual features within a model like Claude 3 Sonnet could contribute to a deeper understanding of AI behavior and potentially enhance the development of safer AI systems by allowing targeted adjustments to safety-related features.