Anthropic Discovers J-space in Claude for Complex Thinking




Anthropic has released a groundbreaking study revealing J-space, an internal space for complex thinking in the Claude model. This discovery could change our understanding of how large language models operate.
Anthropic Discovers J-space in Claude for Complex Thinking
The startup Anthropic has released one of the most interesting works on interpretability in recent times. They discovered that in LLMs (large language models), there is a separate set of neural activations used as working memory for complex thinking: the so-called J-space.
We often envision LLMs as a vast jumble of numbers, where knowledge and reasoning are spread across billions of parameters. However, it seems that this is not entirely the case.
Interestingly, the human brain works in a similar way. There is a cognitive theory that claims it consists of many specialized subsystems that operate independently; but sometimes, some information enters a small shared workspace, becoming accessible to many areas of the brain and can be used for reasoning and decision-making. This is precisely the mechanism that researchers found in Claude.
Here’s how J-space works:
— The name comes from a technique developed by scientists called Jacobian Lens. This technique allows understanding, based on current activations, which tokens are more likely to appear in the future. While their appearance is not guaranteed, it indicates that the model is thinking about them and keeping them in J-space for reasoning purposes. Essentially, this is reading Claude's thoughts.
— If you ask Claude what it is thinking about, the contents of J-space align well with the model's response. Other internal representations do not possess this property. Additionally, the model can consciously change J-space. For instance, if you tell it to "think about elephants while solving a problem," that concept will appear in J-space. When solving a complex problem, intermediate steps also appear within J-space, even if they are not present in the chains of thought.
— If J-space is completely disabled, the model almost loses its abilities that require complex thinking and multi-step reasoning. For example, it can no longer compose poetry.
The most important aspect of this is that J-space can be read and modified. Anthropic provides an example: researchers deliberately implanted a hidden goal in the model during training and then observed how this goal manifested in J-space, even though it was never explicitly stated in the responses. By altering the activations in J-space, they could change the model's subsequent decisions.
Potentially, this is a pathway to creating (fully?) safe and controllable models. At the very least, today Anthropic has made the black box a bit more transparent.
Why it matters
AnalysisThis discovery could significantly alter our understanding of how large language models function and their capacity for complex thinking. Understanding J-space may aid in developing safer and more controllable AI systems.
Discuss in community
Share your questions and insights with developers