Mapping the mind of a large language model
Anthropic researchers have successfully mapped millions of concepts within Claude Sonnet, a production-grade large language model, offering an unprecedented look into its internal workings.
Intelligence analysis by Gemini 2.5 Flash

This breakthrough in interpretability moves beyond treating AI models as 'black boxes,' revealing how complex concepts are represented and organized internally. By scaling up 'dictionary learning' techniques, the team identified features corresponding to entities, abstract ideas, and even multimodal inputs, paving the way for safer and more reliable AI.
Imagine a super-smart computer brain that learns like we do, but nobody knows exactly how it thinks. Scientists at Anthropic found a way to peek inside one of these brains, called Claude Sonnet, and saw millions of tiny 'thought-pieces' or ideas, like 'Golden Gate Bridge' or 'keeping secrets.' It's like instead of just seeing a delicious cake, they figured out what each ingredient does and how they all mix together to make the cake taste good. This helps them make sure the computer brain thinks safely and correctly.
Analysis
Anthropic's recent research marks a significant leap in AI interpretability, moving beyond the traditional 'black box' view of large language models. For the first time, researchers have successfully extracted and mapped millions of human-interpretable concepts from a modern, production-grade model, Claude Sonnet. This achievement provides an unprecedented window into how these complex AI systems process information and form their internal representations, which is vital for advancing AI safety and trustworthiness.
Claude Sonnet
The focus of this groundbreaking interpretability work is Claude Sonnet, one of Anthropic's deployed, state-of-the-art large language models. The ability to peer into a model of this scale and sophistication represents a major advancement, as previous interpretability efforts were largely confined to much smaller, 'toy' models. The researchers successfully extracted millions of features from a middle layer of Claude 3.0 Sonnet, creating a conceptual map of its internal states during computation. This demonstrates that interpretability techniques can indeed scale to the vastly larger and more complex AI systems currently in use, overcoming significant engineering and scientific hurdles.
This detailed look inside Sonnet revealed features with a depth, breadth, and abstraction that reflect the model's advanced capabilities. Unlike the superficial features found in simpler models, Sonnet's features correspond to a wide array of entities, from specific cities like San Francisco and historical figures like Rosalind Franklin, to scientific fields such as immunology and programming syntax. The success with Sonnet validates the scalability of their methods and opens new avenues for understanding the sophisticated behaviors of advanced AI.
Dictionary Learning
The core technique enabling this discovery is 'dictionary learning,' a method borrowed from classical machine learning. This approach isolates recurring patterns of neuron activations, termed 'features,' which can then be matched to human-interpretable concepts. Instead of viewing the model's internal state as an incomprehensible list of neuron activations, dictionary learning allows for its representation in terms of a few active features, much like sentences are built from words and words from letters.
Anthropic had previously applied dictionary learning to a small language model in October 2023, identifying coherent features for concepts like uppercase text or Python function arguments. The challenge was scaling this technique up by many orders of magnitude to models like Sonnet, which required substantial parallel computation and addressed the scientific risk that large models might behave differently. The successful application to Sonnet confirms the robustness of dictionary learning, demonstrating its capacity to reveal the intricate conceptual organization within highly complex AI architectures.
Golden Gate Bridge
One of the most compelling findings from the Sonnet analysis is the multimodal and multilingual nature of the identified features, exemplified by a feature sensitive to the Golden Gate Bridge. This particular feature activates across a range of model inputs, including English mentions of the bridge's name, discussions in Japanese, Chinese, Greek, Vietnamese, and Russian, as well as images of the landmark. This demonstrates that the model's internal representations are not confined to a single modality or language but integrate information across diverse input types.
Furthermore, the research revealed that these features are organized in a conceptually meaningful way, mirroring human notions of similarity. By measuring the 'distance' between features based on their activation patterns, researchers found that features related to the Golden Gate Bridge were 'close' to concepts like Alcatraz Island, Ghirardelli Square, and even the 1906 San Francisco earthquake. This conceptual proximity extends to abstract ideas, where a feature for 'inner conflict' was found near features related to relationship breakups, conflicting allegiances, and logical inconsistencies. This internal organization suggests a sophisticated, human-like understanding of relationships between concepts within the AI model.
Key points
- Anthropic has successfully mapped millions of concepts within Claude Sonnet, a production-grade large language model.
- This is the first detailed look inside a modern, deployed LLM, moving beyond 'black box' understanding.
- The research utilized 'dictionary learning' to isolate human-interpretable 'features' from neuron activations.
- Features found in Sonnet are deep, broad, and abstract, covering entities, abstract concepts, and programming syntax.
- The identified features are multimodal and multilingual, responding to text and images across various languages.
- Concepts are internally organized in a way that reflects human notions of similarity, showing conceptual proximity between related ideas.
This breakthrough in interpretability could significantly enhance AI safety and reliability by allowing developers to understand why models produce certain outputs. By mapping internal concepts, it becomes possible to identify and mitigate biases, harmful responses, or untruthful information, fostering greater trust in advanced AI systems.



