mechanistic interpretability
Coverage of mechanistic interpretability in the Nexus archive.
- What Anthropic’s latest AI discovery does—and doesn’t—show
Anthropic, a leading AI company, discovered a hidden internal space in large language models (LLMs) called J-space, which contains words influencing problem-solving processes without appearing in outputs. The research explores how LLMs track tasks, recognize patterns, and make decisions, revealing complex mechanisms previously unseen.
- Anthropic found a hidden space where Claude puzzles over concepts
Anthropic developed a technique called the Jacobian lens (J-lens) to uncover a hidden area named J-space within its Claude Opus 4.6 large language model. The J-space contains words related to the model's likely future responses, offering insights into its decision-making process. The company collaborated with Neuronpedia to create a public demo of the tool.