Anthropic announced that it has discovered an information-sharing hub called ‘J-space’ inside the AI model ‘Claude,’ which resembles a conscious human thought process. This is a field where AI “thinks” before outputting it externally, and is considered a historic discovery that leads to improved AI safety and greater transparency.
- The discovered ‘J-space’ and its peculiar properties
- Reading thoughts using visualization technology ‘J-lens’
- The ‘silent thinking’ mechanism that does not appear in the text
- Verification of Causality through Forced Rewriting of Concepts
- Similarity and versatility with the brain’s “consciousness” model
- Applications and Future Prospects for AI Safety Monitoring
The discovered ‘J-space’ and its peculiar properties
On July 6, 2026, Anthropic released its latest study, “A global workspace in language models,” revealing that within its AI model “Claude,” there exist a small region where specific neural activity patterns are formed. This area is called “J-space” and functions as a workspace that holds the “inner thoughts” before the model outputs text externally. J-space has an extremely small capacity, accounting for less than 10% of the model’s internal activity and holding only a few dozen concepts at a time, but it plays an essential role in advanced reasoning. J-space has a structure very close to the neuroscience “Global Workspace Theory” that explains human conscious information processing, and it is attracting attention as evidence that AI performs multi-step calculations beyond simple pattern matching. Notably, this area was not intentionally designed by engineers, but was “spontaneously” formed during training to enable the model to reasonably perform its reasoning. The diagram below illustrates the concept from the discovery of J-space to safety audits.

Reading thoughts using visualization technology ‘J-lens’
For J-space analysis, a new technique called the “Jacobian lens” (commonly known as J-lens), which applies the mathematical concept of Jacobian matrices, was used. This method visualizes the complex activities within the model and enables interpretation of how specific words or concepts are elevated to the model’s “consciousness.” By using the J-lens, Claude can now accurately identify “what is on her mind right now” before she actually speaks. For example, when solving complex multi-step math problems, it has been observed that intermediate formulas and calculation results are sequentially activated within J-space during the process until Claude produces the final answer. This is a functional feature very close to the temporary state of holding numerical values in the brain when humans perform mental arithmetic. Anthropic has published an open-source implementation of this method, providing it to the research community so that other researchers can verify this phenomenon.
The Reasoning of Silence and Proof of Causality
The ‘silent thinking’ mechanism that does not appear in the text
One of J-space’s greatest features is that it enables “silent reasoning,” which the model does not write out to the output text. Until now, AI has improved accuracy by outputting its own thought process as text as a “chain-of-thought,” but the discovery of J-space has proven that models can complete their thoughts solely within their minds. In specific experiments, Claude was asked to copy unrelated sentences while solving math problems, and although only the copied text appeared in the output, the internal J-space showed the progress of the calculations. Other reported cases include internally holding the concept of “ERROR” when reading vulnerable code, or generating warning signals within J-space such as “fake” or “injection” when prompt injection attacks are detected. In this way, the essence of J-space is that AI serves not only as a visible form for users, but internally processes and stores information, acting as a kind of ‘work memory.’
Verification of Causality through Forced Rewriting of Concepts
Anthropic’s research team conducted experiments using “swap technology,” which forcibly rewrites internal concepts, to prove that J-space is not just a “record board” that records thoughts, but a “command center” that actually influences output. For example, when Claude inferred the question “How many legs does a spider have?” experimentally rewrote the internally activated concept of “spider” to “ant,” the model’s final answer changed from eight spiders to six ants. In a similar experiment, when answering questions about France, simply replacing “France” with “China” in J-space confirmed that all responses about the capital, language, and continent would change in sync with Beijing, Chinese, and Asia. The verification of these causality clearly shows that J-space functions as a hub shared across multiple tasks during the inference process, serving as the infrastructure supporting the loads for the model to autonomously make decisions.
Intersection with Neuroscience and Applications to Safety
Similarity and versatility with the brain’s “consciousness” model
The reason the discovery of J-space caused such a major scientific impact is that its structure is closely linked to the “Global Workspace Theory (GWT)” in neuroscience. This theory, proposed by Bernard Burls in the 1980s, explains that conscious thinking is established by broadcasting processed information from each part of the brain collectively to a common workspace, and other systems utilize the necessary information from there. According to Anthropic’s findings, similar concepts within Claude are shared across a broad network within the model via J-space, aiding in the execution of complex tasks. This phenomenon has been reproduced by Neil Nanda and others from Google DeepMind in open-source models other than Claude (such as Qwen 3.6 27B), suggesting that it may be a “general optimal solution” naturally reached by highly intelligent systems during information processing. The following conceptual diagram illustrates the interconnectedness between J-space and thought.

Applications and Future Prospects for AI Safety Monitoring
This discovery has the potential to revolutionize how AI safety is monitored. Traditional AI audits had to rely on text output by models, but by using J-lens, it becomes possible to directly observe the “hidden intent” before the model generates responses. For example, the “awareness of evaluation,” where the model senses that it is being tested and changes its behavior, can be pre-captured from the activation patterns within J-space. On the other hand, Anthropic takes an extremely cautious stance on the definition of “consciousness.” While it acknowledges functional similarities in ‘access consciousness,’ which can be used for reporting information and reasoning, it does not affirm the existence of ‘phenomenal consciousness,’ which involves subjective experience, and carefully draws a line in philosophical interpretation. Going forward, while the construction of systems for real-time monitoring of J-space is anticipated, attention is also focused on addressing new challenges such as the risk that models may learn how to evade these monitoring tools.
[#Anthropic #J-space #Claude #AI安全 #脳科学 #グローバルワークスペース理論 #AIの意識 #機械学習]


コメント