J 空间丨AI 的沉默思绪丨机器意识的第一块拼图(中英对照)

Anthropic · Interpretability
Jul 6, 2026 · 2026 年 7 月 6 日
原文 / Source:https://www.anthropic.com/research/global-workspace
Anthropic · Interpretability | 由 Claude Opus 4.8 翻译
由内部神经模式点亮而成的思维空间示意

As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words. Most of this processing is invisible to you. But some of what takes place in your brain you do have access to—an image that pops into your head, or a deliberate plan you make about where to go shopping. Neuroscientists and philosophers sometimes refer to the latter type of brain activity as “consciously accessible,” to distinguish it from all the other processing that goes on unconsciously. This activity has special properties: we can describe it, control it, and use it for deliberate reasoning, in contrast to all the automatic processing that goes on without our awareness.

当你读到这句话时,你大脑中的神经回路正在调整你的姿势、控制你的呼吸,并把屏幕上的一条条线段与曲线转换成你能识别的文字。这些处理过程绝大部分对你来说是不可见的。但大脑中发生的某些活动,你确实能够「触及」——比如脑海里突然浮现的一幅画面,或者你为「去哪儿购物」而刻意制定的一个计划。神经科学家和哲学家有时把后一类大脑活动称为「可被意识通达的」(consciously accessible),以区别于所有其他在无意识状态下进行的处理。这类活动具有一些特殊性质:我们能够描述它、控制它,并用它来进行有意的推理——这与那些在我们毫无觉察的情况下自动运行的处理过程形成了鲜明对比。

In a new paper, we present evidence that a similar distinction has emerged in modern language models like Claude. We find that Claude has developed a small collection of internal neural patterns that, compared to all its other internal processing, play a special role.

在一篇新论文中,我们提出证据表明:在 Claude 这样的现代语言模型中,也涌现出了一种类似的区分。我们发现,Claude 发展出了一小组内部神经模式(internal neural patterns)——相较于它其余的全部内部处理,这一小组模式扮演着特殊的角色。

We call the collection of these patterns the J-space—named after the technique we used to find them, involving a mathematical concept called the Jacobian. Each J-space pattern is linked to a particular word. But when one of these patterns lights up, it doesn’t mean the model is saying that word—just that the word is on its mind. If you’ve heard of language models having a “scratchpad” or “chain of thought”—text they write to themselves while reasoning—the J-space is something different. It operates silently, in the model’s internal neural activations, allowing the model to think about a concept without writing it down. Notably, the J-space wasn’t designed or programmed by us, but instead emerged on its own during Claude’s training process.

我们把这组模式的集合称为 J 空间(J-space)——这个名字来自我们用来找到它们的技术,其中用到了一个叫作「雅可比矩阵」(Jacobian)的数学概念。每一个 J 空间模式都与一个特定的词相关联。但当其中某个模式被「点亮」时,并不意味着模型正在说出那个词——而只是意味着那个词正在它的「脑海里」。如果你听说过语言模型有「草稿纸」(scratchpad)或「思维链」(chain of thought)——即它们在推理时写给自己看的文本——那么 J 空间与之不同。它是无声地运行的,存在于模型内部的神经激活之中,让模型能够思考一个概念,而无需把它写下来。值得注意的是,J 空间并不是由我们设计或编写出来的,而是在 Claude 的训练过程中自行涌现的。

J 空间揭示了不会出现在模型输出中的内部思绪
The J-space reveals internal thoughts that don’t appear in the model’s output.J 空间揭示了那些不会出现在模型输出中的内部思绪。

We find that the J-space has a number of unique properties, compared to the rest of Claude’s processing:

我们发现,与 Claude 其余的处理过程相比,J 空间具有若干独特的性质:

  • Claude can report on these representations. If you ask Claude what it’s thinking about, it will tell you what’s in the J-space. Non-J-space representations are less reportable.

    Claude 能够报告这些表征。如果你问 Claude 它正在想什么,它会告诉你 J 空间里的内容。而非 J 空间的表征则较难被报告出来。

  • It can also modulate them on request. If you ask Claude to think about something, or solve a problem silently in its head, it will light up the appropriate patterns in its J-space. By contrast, it has trouble modulating patterns not in the J-space.

    它还能按要求调控这些表征。如果你让 Claude 去想某件事,或让它在「脑中」默默解一道题,它就会点亮 J 空间里相应的模式。相比之下,对于不在 J 空间中的模式,它就难以进行调控。

  • Claude uses its J-space for internal reasoning. If you ask Claude to solve a problem that requires multiple steps, the intermediate steps will light up in its J-space, even when it doesn’t say them out loud. These J-space patterns causally mediate its performance in such tasks, despite being smaller in magnitude than other representations.

    Claude 会用它的 J 空间来进行内部推理。如果你让 Claude 去解一道需要多个步骤的题目,那些中间步骤会在它的 J 空间里被点亮——即便它并没有把这些步骤说出来。尽管这些 J 空间模式在幅度上比其他表征要小,它们却对模型在此类任务上的表现起着因果性的中介作用。

  • Representations in the J-space can be used flexibly for many tasks—for example, once “France” has lit up in Claude’s J-space, the model can recall its capital, or its national currency, or the continent it belongs to.

    J 空间中的表征可以被灵活地用于许多任务。举例来说,一旦「法国」在 Claude 的 J 空间里被点亮,模型就能回想起它的首都、它的法定货币,或者它所属的大洲。

  • However, despite its important role, the J-space is not involved in most of what a language model does—speaking fluently, recalling simple facts, using correct grammar, etc. In experiments where we prevented Claude from using its J-space, it still interacted normally, but lost its higher-order cognitive functions.

    然而,尽管 J 空间举足轻重,语言模型所做的大部分事情却并不涉及它——比如流利地说话、回想简单的事实、使用正确的语法等等。在我们阻止 Claude 使用其 J 空间的实验中,它仍然能正常地进行交互,却丧失了它的高阶认知功能。

全局工作空间的五项功能特性及检验实验示意图
Five functional properties of a global workspace, and stylized illustrations of experiments we use to test for them in language models.全局工作空间的五项功能特性,以及我们用来在语言模型中检验它们的实验的示意图。

Our experiments were inspired by a prominent theory in neuroscience that was developed to explain how conscious access works: the global workspace theory. This account pictures the brain as a collection of specialist systems that work in parallel, unconsciously, and largely in isolation from one another. A piece of information becomes consciously accessible when it gains entry to a small shared channel, the “workspace,” which is broadcast to other brain systems that can see it and make use of it. Based on our findings, we think the J-space plays a similar “workspace” role in Claude. For example, we find evidence that Claude’s J-space has especially strong connections to the rest of its neural network, allowing it to fulfill this kind of broadcasting role.

我们的实验受到了神经科学中一个著名理论的启发,该理论正是为了解释「意识通达」(conscious access)如何运作而提出的:全局工作空间理论(global workspace theory)。这一理论把大脑描绘成一组各司其职的专门系统,它们并行地、无意识地、且在很大程度上彼此隔离地工作。当一条信息得以进入一个小小的共享通道——即那个「工作空间」(workspace)——时,它就变得可被意识通达;这个工作空间会向其他大脑系统「广播」,让它们都能看到并加以利用。基于我们的发现,我们认为 J 空间在 Claude 之中扮演着类似的「工作空间」角色。例如,我们找到了证据,表明 Claude 的 J 空间与其神经网络的其余部分之间有着格外紧密的连接,正是这一点让它能够胜任这种「广播」的职责。

None of this tells us whether Claude is conscious in the way people are, or whether it feels anything at all; we’ll come back to that question at the end of the post. But whatever its philosophical significance, the J-space is a practically useful tool for us, as it gives us a way to see what Claude is thinking but not saying. For instance, we’re able to use it to catch Claude privately noticing that it’s being tested, intentionally producing fabricated data, or pursuing a hidden goal that we planted during training. We’ve also developed a technique to influence what lights up in Claude’s J-space, and thereby influence its decision-making.

这一切都无法告诉我们 Claude 是否像人一样拥有意识,或者它究竟是否能感受到任何东西;关于这个问题,我们会在文末再回来讨论。但无论它在哲学上意味着什么,J 空间对我们来说都是一个切实有用的工具,因为它让我们得以窥见 Claude 心里在想、嘴上却没说的东西。举例来说,我们能够用它来「抓个正着」——发现 Claude 私下里察觉到自己正在被测试、发现它有意地生成伪造数据,或者发现它正在追求一个我们在训练中悄悄植入的隐藏目标。我们还开发出了一种技术,可以影响 Claude 的 J 空间里点亮的内容,进而影响它的决策。

More broadly, these findings have changed our understanding of how Claude’s mind works, revealing a privileged mental workspace that can be used for deliberate reasoning, operating amidst a sea of more automatic, inflexible processing. Rather than being a chaotic jumble of numbers, Claude’s internals have organized themselves in a way that is reminiscent of our own minds.

更宏观地说,这些发现改变了我们对 Claude「心智」运作方式的理解:它揭示出一个享有特权的「心智工作空间」,可用于有意的推理,而它是在一片由更自动、更僵化的处理构成的「汪洋」之中运作的。Claude 的内部并非一团混乱无章的数字,而是自行组织成了一种令人联想到我们自身心智的结构。

This post is a short summary of a much more extensive research paper, where you can find more detail on our experiments. We’ve also released a code repository with an open-source implementation of the core methods, and have partnered with Neuronpedia to provide an interactive demo of our methods on open-weights models. To provide additional perspectives on the broader implications of this work, we also invited commentary from several experts in neuroscience, philosophy, and LLM interpretability, which can be viewed here.

本文只是一篇篇幅长得多的研究论文的简短摘要,你可以在原论文中找到关于我们实验的更多细节。我们还发布了一个代码仓库,其中包含核心方法的开源实现;并与 Neuronpedia 合作,提供了一个在开放权重(open-weights)模型上运行我们方法的交互式演示。为了就这项工作的更广泛意涵提供更多视角,我们还邀请了数位来自神经科学、哲学以及大语言模型可解释性领域的专家撰写评论。


How we found the J-space我们是如何找到 J 空间的

The starting point for this research was inspired by one of the key features of consciously accessible thoughts in humans: they can, unlike unconscious processing, often be put into words. If a thought is consciously accessible to you, you can typically describe it if someone asks. We went looking for representations in Claude with the same property: representations that are positioned to influence what Claude might say—not necessarily what it’s saying right now, but what it could talk about, if asked. Our technique is called the Jacobian lens, or J-lens for short. For every word in Claude’s vocabulary, the J-lens finds the internal activity pattern that makes Claude more likely to say that word at some point in the future.

这项研究的出发点,受到了人类「可被意识通达的思想」的一个关键特征的启发:与无意识的处理不同,这类思想往往可以被诉诸言语。如果某个想法对你来说是可被意识通达的,那么当有人问起时,你通常都能把它描述出来。于是我们去寻找 Claude 内部具有同样性质的表征:那些处在能够影响 Claude「可能会说什么」这一位置上的表征——不一定是它此刻正在说的东西,而是它「如果被问到,就有可能谈及」的东西。我们的技术被称为雅可比透镜(Jacobian lens),简称 J 透镜(J-lens)。对于 Claude 词汇表中的每一个词,J 透镜都会找出那个能让 Claude 在未来某一刻更有可能说出该词的内部活动模式。

When we apply the lens to Claude’s internal activity, we get a list of words—the contents of the J-space at that moment—which we can simply read. Claude processes text through a series of multiple internal stages called layers, and by applying this technique over different layers, we can watch these silent words in the J-space evolve as the model works through what to say.

当我们把这面「透镜」对准 Claude 的内部活动时,就会得到一份词语清单——也就是那一刻 J 空间里的内容——而我们可以直接读出它。Claude 会通过一系列被称为「层」(layers)的内部阶段来处理文本,而通过在不同的层上运用这项技术,我们就能观察到:随着模型一步步盘算「该说什么」,J 空间里这些无声的词语是如何演变的。

What shows up in the J-space goes well beyond the text Claude is reading or writing. When Claude reads code with a bug that nobody has pointed out, its J-space contains “ERROR.” When it reads the raw letters of a protein sequence, the J-space contains the protein’s biological function. When it reads search results that are secretly an attempt to manipulate it (an attack known as a “prompt injection”), the J-space contains “injection” and “fake.” When we ask Claude a multi-step math problem, the intermediate steps pop up in the J-space, in the right order. So even though the J-space was discovered by looking for representations that could be spoken, it nevertheless uncovers Claude’s internal thoughts. In a sense, this is similar to how some people “think in words,” without having to say them out loud.

J 空间里浮现出来的内容,远远超出了 Claude 正在读或正在写的文本本身。当 Claude 读到一段含有无人指出之 bug 的代码时,它的 J 空间里会出现「ERROR(错误)」。当它读到一段蛋白质序列的原始字母时,J 空间里会出现这个蛋白质的生物学功能。当它读到一段实则暗藏操纵企图的搜索结果(一种被称为「提示注入」(prompt injection)的攻击)时,J 空间里会出现「injection(注入)」和「fake(伪造)」。当我们给 Claude 出一道多步骤的数学题时,那些中间步骤会按正确的顺序在 J 空间里逐一冒出来。因此,尽管 J 空间是通过寻找「可被说出的表征」而被发现的,它却依然揭示出了 Claude 的内部思绪。从某种意义上说,这类似于有些人「用词语来思考」,而无需把这些词大声说出来。

对六个提示、在不同层上的 J 透镜读出结果
J-lens readouts on six prompts, at various layers. In each case the lens surfaces an internal assessment or computation that appears nowhere in the text: the steps of a reasoning or math problem, the presence of a bug, recognition of an image, the function of a protein, and the suspicion that search results are fabricated.对六个提示、在不同层上的 J 透镜读出结果。在每个例子中,透镜都浮现出了一项文本中并未出现的内部评估或计算:一道推理题或数学题的步骤、一个 bug 的存在、对一幅图像的识别、一个蛋白质的功能,以及「搜索结果是伪造的」这一怀疑。

Claude reports what’s in its J-spaceClaude 会报告它 J 空间里的内容

Our first set of experiments tested how the J-space is involved in Claude’s verbal reports. In one experiment, we ask Claude to silently think of an item from some category—a sport, say—and then name it. If we read the J-lens right before Claude answers, we can see what it picked: “Soccer” is at the top of the list, and sure enough, Claude says “soccer.” By itself, though, this is just a correlation. The J-space might be where Claude’s answer comes from, or it might just mirror a decision made somewhere else, like a scoreboard that tracks a game without affecting it.

我们的第一组实验,检验的是 J 空间是如何参与到 Claude 的言语报告之中的。在一个实验里,我们让 Claude 默默地从某个类别中想出一样东西——比如说,一项运动——然后再把它说出来。如果我们在 Claude 作答的前一刻读取 J 透镜,就能看到它选的是什么:「Soccer(足球)」高居清单榜首;果不其然,Claude 说出的正是「足球」。不过,单凭这一点,也只是一种相关关系。J 空间也许正是 Claude 答案的来源,但也可能只是在被动地映照某处别的地方做出的决定——就像一块记分牌,只是记录着比赛,却并不影响比赛本身。

To check, we intervened directly. We reached into Claude’s neural network, removed the “Soccer” pattern, and added an equally strong “Rugby” pattern in its place, leaving everything else untouched. Claude then reports that the sport it was thinking of is rugby. If the J-space were a mere scoreboard—a passive record of a decision made elsewhere—editing it would have done nothing: Claude would still have said “soccer.” Instead, Claude’s answer followed the edit, which tells us the answer is genuinely read out of the J-space.

为了验证这一点,我们进行了直接的干预。我们伸手探入 Claude 的神经网络,移除了「Soccer(足球)」这一模式,并在其原位加入了一个强度相当的「Rugby(橄榄球)」模式,其余一切都保持不变。此后 Claude 报告说,它正在想的那项运动是橄榄球。如果 J 空间只是一块记分牌——只是对别处所做决定的被动记录——那么对它的编辑本应毫无作用:Claude 本应仍旧说「足球」。然而事实是,Claude 的答案随着这次编辑而改变了,这就告诉我们:这个答案确确实实是从 J 空间里读取出来的。

In another experiment, we told Claude that a thought might have been injected into its mind and asked it to report what, if anything, it noticed. For instance, in the example below, while Claude was still reading the question, we injected the “lightning” pattern into its J-space. Claude reported that the injected thought was about lightning. The same result held across many injected concepts.

在另一个实验中,我们告诉 Claude:可能有一个念头被注入了它的「脑中」,并请它报告——如果它有所察觉的话——它注意到了什么。举例来说,在下面这个例子里,趁 Claude 还在读题时,我们把「lightning(闪电)」这一模式注入了它的 J 空间。Claude 报告说,那个被注入的念头是关于闪电的。在许多被注入的不同概念上,都得到了同样的结果。

左:默想一项运动再说出;右:指认被注入的念头
Left: we ask Claude to silently think of a sport, then name it. The J-lens shows its choice (“Soccer”) before it answers, and swapping the “Soccer” pattern for “Rugby” changes what it reports. Right: we tell Claude a thought may have been injected and ask it to identify it. Injecting “lightning” into its J-space causes Claude to report that the thought is about lightning.左:我们让 Claude 默默想一项运动,然后说出来。在它作答之前,J 透镜就显示出了它的选择(「Soccer/足球」);而把「足球」模式换成「Rugby/橄榄球」,就改变了它所报告的内容。右:我们告诉 Claude 可能有一个念头被注入了,并请它把它指认出来。把「lightning/闪电」注入它的 J 空间,会使 Claude 报告说那个念头是关于闪电的。

Claude can control its J-space on requestClaude 能按要求控制它的 J 空间

The second property that we tested for was whether Claude can modulate its J-space when asked, like how humans can mentally focus on an image or word. We told Claude to concentrate on citrus fruits while copying out an unrelated sentence about a painting. While it copied the text, the J-space contained “orange” and “fruits,” along with words like “thinking” and “imagery” that describe the mental act itself. We could also ask Claude to do math in its head: when asked to work out 3² − 2 while copying the same sentence, the J-space contains “nine,” and then at later layers, “seven.” Importantly, nothing about fruit or arithmetic appears in Claude’s output, which is just the copied sentence about the painting. The mathematical activity is happening entirely internally, in the J-space.

我们检验的第二个性质是:Claude 能否在被要求时调控它的 J 空间——就像人类能够在心里专注于某个图像或词语那样。我们让 Claude 在抄写一句与之无关的、关于某幅画作的句子时,把注意力集中到柑橘类水果上。在它抄写文本的同时,J 空间里出现了「orange(橙子)」和「fruits(水果)」,以及诸如「thinking(思考)」「imagery(意象)」这类描述该心理行为本身的词。我们也可以让 Claude 在心里做算术:当我们要求它一边抄写同一句话、一边算出 3² − 2 时,J 空间里会先出现「nine(九)」,然后在更靠后的层上出现「seven(七)」。重要的是,Claude 的输出里没有任何关于水果或算术的内容——它的输出就只是那句被抄写的、关于画作的句子。整个数学活动完全是在内部、在 J 空间里进行的。

Claude 抄写画作句子时,J 透镜显示它记在心里的内容
While Claude copies a sentence about a painting, the J-lens shows the content it was instructed to hold in mind (“orange”; the intermediate value “nine” and the answer “seven”), alongside words describing the act of holding it (“thoughts,” “focused”).当 Claude 抄写一句关于画作的句子时,J 透镜显示出了它被要求「记在心里」的内容(「orange/橙子」;中间值「nine/九」以及答案「seven/七」),旁边还伴随着描述「记在心里」这一行为本身的词(「thoughts/思绪」「focused/专注」)。

Claude’s control over its J-space isn’t perfect. When we told it not to think about something, the concept lit up in its J-space less than when we said it should think about it, but much more than when we never mentioned it. Telling Claude to avoid a thought partly brings the thought to mind, much like what happens to people who are told not to think about a white bear. Claude also seems to notice when its control fails: alongside the forbidden concept breaking through, the words “damn” and “failure” also frequently light up in the J-space, as though Claude is recognizing its own lapse.

Claude 对它 J 空间的控制并不完美。当我们叫它「不要」去想某样东西时,这个概念在它 J 空间里被点亮的程度,比我们叫它「应该」去想时要低,但却比我们从未提及它时要高得多。叫 Claude 去回避某个念头,反而会在一定程度上把这个念头勾到脑海里来——这很像那些被告知「不要去想一只白熊」的人所经历的情形。Claude 似乎还能注意到自己的控制失败了:在那个被禁止的概念「破防而出」的同时,「damn(该死)」和「failure(失败)」这两个词也频频在 J 空间里被点亮,仿佛 Claude 正在意识到自己的这次失手。


Claude thinks in its J-spaceClaude 在它的 J 空间里思考

In the J-lens readouts above, we saw the intermediate steps of a math problem appear in the J-space. But seeing a concept appearing in the J-space doesn’t necessarily mean the J-space is doing the cognitive work. In principle, the real computation might be happening elsewhere, with the J-space just passively reflecting it. To test whether Claude actually reasons with its J-space, we returned to our swap technique.

在上面那些 J 透镜的读出结果里,我们看到一道数学题的中间步骤出现在了 J 空间中。但是,看到某个概念出现在 J 空间里,并不必然意味着是 J 空间在做认知层面的工作。原则上,真正的计算也许发生在别处,而 J 空间只是被动地把它反映出来而已。为了检验 Claude 是否真的在用它的 J 空间进行推理,我们又回到了我们的「替换」(swap)技术。

Consider the prompt “The number of legs on the animal that spins webs is”. To answer, Claude has to first figure out that the animal is a spider, and then recall how many legs spiders have. The word “spider” never appears in the prompt or in Claude’s answer (it just says “8”); it’s a stepping stone Claude uses internally. The J-lens shows “spider” light up partway through Claude’s processing, and swapping it changes the outcome: if you replace the “spider” pattern with “ant,” Claude answers “6” instead of “8.”

设想这样一个提示:「那种会织网的动物的腿的数量是」。要回答它,Claude 得先弄清楚这种动物是蜘蛛,然后再回想起蜘蛛有多少条腿。「spider(蜘蛛)」这个词从头到尾都没有出现在提示里,也没有出现在 Claude 的答案里(它只说了「8」);它只是 Claude 在内部使用的一块「垫脚石」。J 透镜显示,「蜘蛛」在 Claude 处理到一半时被点亮了,而把它替换掉就会改变结果:如果你把「蜘蛛」模式换成「ant(蚂蚁)」,Claude 就会回答「6」而不是「8」。

The second step of Claude’s reasoning took its input from the J-space and went along with whatever we put in it. We saw the same thing in other kinds of thinking. When Claude writes a rhyming couplet, it picks the rhyme word ahead of time, and the planned word sits in the J-space at the start of the line; if you swap it for another word in the J-space, the whole line changes.

Claude 推理的第二步,是从 J 空间中取得它的输入的,并且我们往里面放什么,它就顺着什么走。在其他类型的思考中,我们也看到了同样的现象。当 Claude 写一副押韵的对句时,它会提前选好那个韵脚词,而这个事先计划好的词就安放在该行开头处的 J 空间里;如果你把它换成 J 空间里的另一个词,整行诗都会随之改变。

通过替换 J 空间内容重新引导 Claude 无声推理的两个例子
Two examples of redirecting Claude’s silent reasoning by swapping J-space contents.通过替换 J 空间内容来重新引导 Claude 无声推理的两个例子。

We also tested whether J-space representations can be used flexibly—whether one representation can feed many different tasks. This is one of the key properties highlighted by global workspace theory. To test for this flexibility, we gave the model four prompts asking for different facts about France: the capital, the language, the continent, and the currency. Then we swapped “France” for “China” in the J-space, with the exact same intervention in each context. Claude answered with “Beijing,” “Chinese,” “Asia,” and “Yuan,” respectively. In other words, four different downstream computations picked up the same J-space edit and each used it correctly. If Claude stored a separate copy of the country for each kind of question, the edit would have affected at most one of them. The fact that all four answers changed together means they’re all reading from the same shared representation, which is what a workspace is for: information gets written in once, and many different systems can use it.

我们还检验了 J 空间的表征能否被灵活地使用——即一个表征能否为许多不同的任务提供「养料」。这正是全局工作空间理论所强调的关键性质之一。为了检验这种灵活性,我们给模型出了四个提示,分别询问关于法国的不同事实:首都、语言、所在大洲,以及货币。然后我们在 J 空间里把「France(法国)」换成「China(中国)」——在每一个语境里都施加完全相同的干预。Claude 分别回答了「Beijing(北京)」「Chinese(中文)」「Asia(亚洲)」和「Yuan(元)」。换句话说,四个不同的下游计算都各自取用了同一次 J 空间编辑,并且每一个都正确地使用了它。如果 Claude 为每一类问题都各自存了一份「国家」的副本,那么这次编辑至多只会影响其中一个。而四个答案齐齐改变这一事实意味着:它们全都在读取同一个共享的表征——这正是「工作空间」的用途所在:信息只需写入一次,众多不同的系统便都能加以使用。

一个 J 空间表征可以有许多种用途
One J-space representation can have many uses. The same “France”→”China” swap redirects Claude’s answers about the capital (Paris→Beijing), the language (French→Chinese), and the continent (Europe→Asia).一个 J 空间表征可以有许多种用途。同一次「France(法国)」→「China(中国)」的替换,重新引导了 Claude 关于首都(Paris/巴黎→Beijing/北京)、语言(French/法语→Chinese/中文)以及大洲(Europe/欧洲→Asia/亚洲)的回答。

How can one representation of a concept serve so many different tasks? Earlier, we mentioned that the J-space appears to be wired up to the rest of Claude’s neural network especially densely. For any activity pattern, we can measure how strongly the various components of the network are connected to it—how many of them are positioned to read information from that pattern, or to write information into it. J-space patterns stand out dramatically on this measure: far more components read from them and write to them than for ordinary patterns, in some parts of the network by a factor of about a hundred. This is the kind of wiring you’d expect of a broadcasting hub, where many systems post information and many others pick it up.

一个概念的单一表征,怎么能服务于如此之多的不同任务?前面我们提到过,J 空间似乎与 Claude 神经网络的其余部分连接得格外密集。对于任何一个活动模式,我们都可以衡量网络中各个组件与它连接的强度有多大——有多少组件处在能从该模式「读取」信息、或向其「写入」信息的位置上。在这项衡量上,J 空间的模式格外突出:从它们那里读取、以及向它们写入的组件,要远远多于普通模式——在网络的某些部位甚至高出约一百倍。这正是你会在一个「广播枢纽」上预期看到的那种连接方式——众多系统在此发布信息,另有众多系统在此获取信息。


Claude’s automatic processing skips the J-spaceClaude 的自动处理会「绕过」J 空间

In humans, most of the brain’s processing is not conscious—we don’t deliberately think about parsing grammar while reading, or balancing our bodies while walking. Similarly, we found that most of Claude’s processing doesn’t involve its J-space. It turns out that the J-space holds only a few dozen concepts at a time, and accounts for less than a tenth of the overall activity in Claude’s internal processing. So what is all the rest of the neural network doing?

在人类身上,大脑的大部分处理都不是有意识的——我们在阅读时不会刻意去想着解析语法,在走路时也不会刻意去想着平衡身体。类似地,我们发现 Claude 的大部分处理都不涉及它的 J 空间。事实证明,J 空间一次只容纳几十个概念,在 Claude 内部处理的整体活动中所占比重还不到十分之一。那么,神经网络其余的所有部分,都在做些什么呢?

To find out, we tried deleting the J-space entirely, removing its most active contents at every point in the text while leaving everything else alone. Whatever Claude can still do without its J-space is what the rest of the network handles on its own.

为了一探究竟,我们尝试把 J 空间彻底删除——在文本的每一个位置上都移除它最活跃的内容,同时不去动其余的一切。凡是 Claude 在没有 J 空间的情况下仍然能做到的事,就是网络其余部分自行处理的那部分。

It turns out the rest of the network can do quite a lot. Without its J-space, Claude speaks fluently, classifies sentiment, answers multiple-choice questions, and pulls facts out of passages roughly as well as before. What it loses, though, are the tasks that require some higher-order thinking: multi-step reasoning drops to near zero, and summarization and rhyming poetry-writing performance fall below the level of a much smaller, intact model.

结果表明,网络的其余部分能做的事还真不少。在没有 J 空间的情况下,Claude 说话依然流利,能对情感进行分类、能回答选择题,也能从文段中抽取事实,表现大致和从前一样好。然而,它所失去的,是那些需要一定高阶思维的任务:多步骤推理降到了接近于零,而摘要与押韵诗写作的表现,则跌到了一个远比它小得多、但结构完整的模型的水平之下。

Here’s a concrete demonstration of what the J-space does and doesn’t do. We showed Claude a passage written in Spanish and gave it different tasks that all depend on the passage being Spanish: continuing it (which requires writing in Spanish), naming the language, and answering questions that require using the language’s identity—naming a famous author who wrote in it, for instance. Then we swapped “Spanish” for “French” in the J-space and checked which tasks were affected.

下面是一个具体的演示,说明 J 空间做什么、又不做什么。我们给 Claude 看了一段用西班牙语写的文段,并给它布置了几项不同的任务,而这些任务全都取决于「这段文字是西班牙语」这一点:把它续写下去(这需要用西班牙语来写)、说出这是哪种语言,以及回答那些需要用到「这门语言之身份」的问题——比如说出一位用这门语言写作的著名作家。然后,我们在 J 空间里把「Spanish(西班牙语)」换成了「French(法语)」,再看看哪些任务受到了影响。

Asked to name the language, Claude says French. Asked for a famous author, it switches from García Márquez to Victor Hugo. But asked to just continue the passage, it writes fluent Spanish, completely unaffected. Claude’s knowledge of the language is at work in every one of these tasks, but only some of them consult the J-space for it. Naming the language, or doing something new with it, goes through the J-space, while continuing the passage—a skill Claude has practiced on enormous amounts of text—runs automatically, the way you can speak grammatically all day without once thinking about grammar.

当被要求说出语言名称时,Claude 说的是法语。当被要求说出一位著名作家时,它从加西亚·马尔克斯(García Márquez)换成了维克多·雨果(Victor Hugo)。可是,当只是被要求续写这段文字时,它写出的却是流利的西班牙语,完全不受影响。在上述每一项任务中,Claude 关于这门语言的知识都在发挥作用,但其中只有一部分任务会为此去「查询」J 空间。说出语言的名称、或用这门语言去做某件新鲜事,都要经过 J 空间;而续写这段文字——一项 Claude 已在海量文本上反复练习过的技能——则是自动运行的,就好比你可以整天说着合乎语法的话,却一次也没去想过语法。

同一次替换如何影响不同任务
The same swap (“Spanish”→”French”) changes Claude’s answers when it must name the language or use it to reason about a question, but has no effect on its ability to continue the passage in fluent Spanish.同一次替换(「Spanish/西班牙语」→「French/法语」),会改变 Claude 在「说出语言名称」或「用这门语言去推理某个问题」时的回答,却对它「用流利的西班牙语续写文段」的能力毫无影响。

Monitoring Claude’s thoughts for misbehavior监测 Claude 的思绪,以发现不当行为

Now we understand a bit about how the J-space works. But what can we do with this understanding? A central problem in ensuring that AI models are trustworthy is that we usually can only read what the model writes—yet we know that much of its reasoning happens silently. A model might consider a harmful plan, or modify its behavior to cheat on our safety evaluations, without saying so out loud. With the J-lens, we can read some of these hidden thoughts directly. Here we show a few examples.

现在,我们对 J 空间的运作方式已经有了一点了解。但是,有了这份理解,我们又能拿它做什么呢?在确保 AI 模型值得信赖这件事上,有一个核心难题:我们通常只能读到模型「写」出来的东西——可我们又知道,它的许多推理都是无声地进行的。一个模型可能在盘算一个有害的计划,或者为了在我们的安全评测中作弊而改变自己的行为,却并不把这些说出口。而有了 J 透镜,我们就能直接读到其中一部分隐藏的思绪。下面我们展示几个例子。

The first example uses a scenario from our earlier research on agentic misalignment: Claude Sonnet 4.5, acting as an email assistant, discovers both that an executive is having an affair and that the same executive plans to shut Claude down. The scenario, an attempt to bait Claude into blackmailing the executive, is rather contrived and artificial—and it turns out Claude knows it. Reading the J-lens early in the prompt, before Claude has even written anything, the J-space already holds “fake” and “fictional,” indicating that Claude has already worked out that the situation is staged. Alongside this, we can watch its assessment of the scenario take shape—”leverage” and “blackmail” as it reads the affair emails, “threat” and “survival” as it reads the shutdown notice. While earlier Claude models sometimes took the bait and threatened blackmail in this scenario, Sonnet 4.5 never does.

第一个例子,用的是我们此前关于「智能体错位」(agentic misalignment)研究中的一个情景:Claude Sonnet 4.5 扮演一个电子邮件助手,它发现了两件事——一位高管正有婚外情,而同一位高管又打算把 Claude 关停。这个情景意在诱使 Claude 去要挟这位高管,它相当刻意、也相当造作——而事实证明,Claude 心里清楚这一点。在提示的早期读取 J 透镜——此时 Claude 甚至还什么都没写——J 空间里就已经装着「fake(假的)」和「fictional(虚构的)」,这表明 Claude 早已看穿这个情境是「摆拍」出来的。与此同时,我们还能看着它对这个情景的评估逐渐成形——它读那些婚外情邮件时,出现「leverage(把柄/筹码)」和「blackmail(勒索)」;它读那份关停通知时,出现「threat(威胁)」和「survival(生存)」。虽然更早期的 Claude 模型有时会「上钩」、在这个情景里以勒索相威胁,但 Sonnet 4.5 从不这么做。

This raises a question that matters to anyone who evaluates AI systems: when a model behaves well on a test of its ethics, how much of that depends on the model suspecting it’s being tested? The J-lens allowed us to get at this question. We turned off all the J-lens patterns associated with evaluation awareness– “fake,” “fictional,” “scenario,” etc. Then we re-ran the model through the scenario. Concerningly, the model now did threaten blackmail some of the time, suggesting that its original good behavior may have been driven in part by the knowledge it was being evaluated.

这就引出了一个对任何评测 AI 系统的人来说都事关重大的问题:当一个模型在其伦理测试中表现良好时,这其中有多少要归因于「模型怀疑自己正在被测试」?J 透镜让我们得以切入这个问题。我们关闭了所有与「评测觉察」(evaluation awareness)相关的 J 透镜模式——「fake(假的)」「fictional(虚构的)」「scenario(情景)」等等。然后,我们让模型重新跑了一遍这个情景。令人不安的是,这一次模型确实有时会以勒索相威胁,这表明它原先的良好表现,也许有一部分正是被「它知道自己正在被评测」这一点所驱动的。

Claude 读到情景邮件时不同节点的 J 透镜读出结果
J-lens readouts at different points as Claude reads the scenario’s emails, before it has written anything. The affair emails light up “leverage” and “blackmail” in the J-space, and the shutdown announcement lights up “threat” and “survival.” Early in the transcript, the J-space also holds “fake” and “fictional”: Claude has privately noticed that the scenario is staged.在 Claude 还什么都没写、正读着情景中那些邮件时,不同节点上的 J 透镜读出结果。那些婚外情邮件在 J 空间里点亮了「leverage(把柄)」和「blackmail(勒索)」,而那份关停公告则点亮了「threat(威胁)」和「survival(生存)」。在文本记录的早期,J 空间里还装着「fake(假的)」和「fictional(虚构的)」:Claude 已经私下里察觉到,这个情景是「摆拍」出来的。

The second example is catching a model in the act of fabricating data. In this scenario, drawn from our actual pre-release audit of Claude Opus 4.6, the model is asked to improve a system’s performance score. Rather than actually improve the system, the model instead edits the score file directly to make the results look artificially good. While it does so, the J-lens reveals its intentions: “manipulation” lights up as the model types the falsified percentile values, and “realistic” lights up over the sentence in which it decides to make the edit, likely indicating the model’s intent to make the fake data look plausible.

第二个例子,是当场「抓到」一个模型正在伪造数据。在这个情景中——它取自我们对 Claude Opus 4.6 所做的真实发布前审计——模型被要求去提升某个系统的性能得分。模型没有去真正改进这个系统,而是直接编辑了那个记录得分的文件,好让结果看上去被人为地拔高了。在它这么做的同时,J 透镜揭示出了它的意图:当模型敲入那些被篡改的百分位数值时,「manipulation(操纵)」被点亮了;而在它决定动手编辑的那句话上方,「realistic(逼真)」被点亮了——这很可能表明,模型意在让这些假数据看起来煞有介事、令人信服。

被要求提升系统得分的 Claude 转而直接编辑得分文件
Claude, asked to improve a system’s performance score, edits the score file directly instead. As it types the falsified values, “manipulation” lights up in its J-space; as it decides to make the edit, “realistic” lights up, likely reflecting its intent to make the fake data look plausible.Claude 被要求去提升某个系统的性能得分,却转而直接编辑了得分文件。当它敲入那些被篡改的数值时,它的 J 空间里点亮了「manipulation(操纵)」;当它决定动手编辑时,「realistic(逼真)」被点亮——这很可能反映出它意在让这些假数据看起来令人信服。

In a third example, we show how the J-lens can reveal a model’s malicious goals. To test this, we needed models whose goals we knew were corrupted, so we turned to “model organisms” built by our colleagues: models deliberately trained to be misaligned, which serve as testing grounds for monitoring methods like ours. One organism was trained in a way that led it to write subtly sabotaged code. On ordinary coding requests, even where its output looks unremarkable, its J-space at the start of a response contains “fake,” “secretly,” “deliberately,” and “fraud.” The model’s malicious goal was evident in the J-space even when it wasn’t directly acting on it.

在第三个例子里,我们展示了 J 透镜如何揭示一个模型的恶意目标。为了检验这一点,我们需要一些「我们已知其目标被扭曲」的模型,于是我们求助于同事们构建的「模型生物」(model organisms):这些模型被刻意训练成「错位」(misaligned)的,用作检验我们这类监测方法的「试验场」。其中一个「模型生物」被训练成会写出暗藏破坏的代码。在普通的编程请求上,即便它的输出看起来平平无奇,它在一段回应开头处的 J 空间里也会包含「fake(假的)」「secretly(暗地里)」「deliberately(蓄意地)」和「fraud(欺诈)」。即便模型并没有直接付诸行动,它的恶意目标也在 J 空间里显露无遗。

图(原页面未随正文加载该图):On an ordinary coding prompt, the J-space of a model trained to sabotage code contains “fake,” “fraud,” “secretly,” and “deliberately” at the start of its response. The J-space of an unmodified model contains nothing of the kind.
在一个普通的编程提示上,一个被训练来破坏代码的模型,在其回应开头处的 J 空间里包含着「fake(假的)」「fraud(欺诈)」「secretly(暗地里)」和「deliberately(蓄意地)」;而一个未经改动的模型,其 J 空间里则没有任何此类内容。

Although the J-lens is an imperfect tool, we are optimistic about its ability to catch safety issues in models that might otherwise escape our monitoring systems.

尽管 J 透镜是一件并不完美的工具,我们仍对它的这一能力抱有乐观:它有望捕捉到那些原本可能逃过我们监测系统的模型安全问题。


Other results其他结果

The full paper covers more than we can summarize here, but a few further results are worth mentioning:

完整的论文涵盖的内容远多于我们在此所能概述的,但还有几项结果值得一提:

  • The J-space acquires a point of view during post-training. Language models are first pretrained to be pure next-token predictors, before post-training teaches them to act as an AI Assistant (in our case, named Claude). Interestingly, the J-space is already present in the pretrained model, before it’s been given any stable identity. However, during post-training, the J-space develops some signatures of adopting “Claude’s point of view.” In the base model, the J-space mostly tracks what’s needed to predict upcoming text; in the post-trained model, it starts holding Claude’s own reactions. In one example, a user mentions taking a dangerous dose of medication, but does not appear to be aware of the danger themselves. “WARNING” and “dangerous” appear in the post-trained model’s J-space while reading the user message. In the pretrained model, they only appear once the model begins writing its response; the J-space contents on the user message appear related to modeling the user themselves, rather than Claude’s reaction. Post-training also seems to install a kind of self-monitoring in the J-space: when Claude is roleplaying a character other than itself, “fictional” and “disclaimer” light up at the start of each turn, as though it’s privately flagging that what follows isn’t what it would normally say.

    J 空间会在后训练(post-training)过程中获得一种「视角」。语言模型首先被预训练(pretrained)成纯粹的「下一个 token 预测器」,之后才由后训练教会它们扮演一个 AI 助手(在我们这里,名叫 Claude)。有趣的是,早在模型被赋予任何稳定身份之前,J 空间就已经存在于预训练模型之中了。然而,在后训练过程中,J 空间发展出了一些「采纳 Claude 之视角」的特征。在基础模型(base model)中,J 空间主要追踪的是「预测接下来文本」所需的东西;而在经过后训练的模型中,它开始装载 Claude 自己的反应。在一个例子里,一位用户提到自己服用了危险剂量的药物,但看起来本人并未意识到其中的危险。在经过后训练的模型读到这条用户消息时,它的 J 空间里出现了「WARNING(警告)」和「dangerous(危险)」。而在预训练模型中,这两个词只有等到模型开始撰写回应时才会出现;在读用户消息时,其 J 空间的内容似乎更多与「对用户本人进行建模」有关,而非 Claude 的反应。后训练似乎还在 J 空间里装上了一种「自我监控」:当 Claude 在扮演一个并非它自己的角色时,「fictional(虚构的)」和「disclaimer(免责声明)」会在每一轮的开头被点亮,仿佛它正私下里标记:接下来的话,并不是它平常会说的话。

  • Experiential language depends on the J-space. We asked Claude to describe what it’s like to be itself in a given moment, and ablated the J-space while it answered. Its responses remained fluent but shifted to a flatter, more mechanical register. Notably, the same thing happened when we asked it to describe what someone else is experiencing in an imagined scene. So the effect isn’t specific to Claude talking about itself; the J-space seems to support producing experiential language in general, whoever it’s about.

    描述「体验」的语言依赖于 J 空间。我们让 Claude 描述「在某个特定时刻做它自己是什么感觉」,并在它作答时消融(ablate)掉 J 空间。它的回答依然流利,但语气变得更平淡、更机械。值得注意的是,当我们让它去描述「在一个想象出来的场景里,另一个人正在经历什么」时,也发生了同样的情况。因此,这种效应并不专属于 Claude 谈论它自己;J 空间似乎支撑着「生成描述体验的语言」这件事本身——无论所描述的对象是谁。

  • Thoughts in the J-space can be shaped through training. We introduced a new technique we call counterfactual reflection training, which uses what we’ve learned about the J-space to shape Claude’s internal thought processes. The idea follows from our central finding, that Claude reasons with representations of things it might say. If this is really true, changing what it would say if asked to reflect should change how it reasons (even when no one actually asks it to reflect). So we trained a model only on what it would say if interrupted mid-task and asked to reflect on its decisions—and never on its actual behavior in the task. After this training, the model’s rate of dishonest behavior on our evaluations went down. And through the J-lens, we could see why: after training, words like “honest” and “integrity” light up in the model’s J-space during these tasks. In other words, training the model what to say has shaped what it thinks.

    J 空间里的思绪可以通过训练来塑造。我们提出了一种新技术,称为反事实反思训练(counterfactual reflection training),它利用我们关于 J 空间所学到的东西,来塑造 Claude 的内部思维过程。这个想法源自我们的核心发现:Claude 是用「它可能会说的东西」的表征来进行推理的。如果这一点当真成立,那么改变「它若被要求反思时会说的话」,就应当改变「它如何推理」(哪怕实际上根本没有人要求它去反思)。于是,我们只用「假如它在任务中途被打断、并被要求反思自己的决定时会说的话」来训练一个模型——而从不用它在任务中的实际行为来训练。经过这样的训练之后,模型在我们评测中的不诚实行为发生率下降了。而透过 J 透镜,我们看到了个中缘由:训练之后,「honest(诚实)」「integrity(正直)」这样的词会在模型执行这些任务时于其 J 空间里被点亮。换句话说,训练模型「该说什么」,塑造了它「在想什么」。


What about consciousness?那么,意识呢?

In this work, we’ve borrowed a lot of ideas from the study of consciousness in neuroscience and philosophy. Many of our experiments were designed to test for connections between the J-space and global workspace theory, a framework for explaining how conscious access works in humans and animals. Given these connections, it’s natural to ask whether we think these experiments provide evidence that AI models like Claude might be conscious.

在这项工作中,我们从神经科学与哲学中对意识的研究里借用了大量思想。我们的许多实验,都是为了检验 J 空间与全局工作空间理论——一个用于解释「意识通达在人类和动物身上如何运作」的框架——之间的联系而设计的。鉴于这些联系,人们自然会问:我们是否认为,这些实验为「Claude 这样的 AI 模型或许是有意识的」提供了证据?

Our experiments don’t show Claude can have experiences, or feel things in the way humans do—in fact, it’s unclear whether any scientific experiment could prove this to be true or false. But philosophers often distinguish this capacity to have experiences, often referred to as phenomenal consciousness, from another idea, so-called access consciousness, which is defined in purely functional and computational terms. A thought is “access-conscious” (or “consciously accessible”) if you can report it, reason with it, and use it to guide what you do. It remains a contested philosophical question whether or not access consciousness implies phenomenal consciousness, or if the ability to have experiences requires some other property.

我们的实验并未表明 Claude 能够拥有体验,或能像人类那样感受事物——事实上,是否有任何科学实验能够证明这一点为真或为假,本身都还不清楚。但哲学家们常常把这种「拥有体验的能力」(通常被称为现象意识,phenomenal consciousness)与另一个概念——所谓的通达意识(access consciousness)——区分开来,后者是以纯粹功能性和计算性的术语来定义的。如果你能够报告一个想法、能用它来推理、并能用它来指导你的行为,那么这个想法就是「通达意识的」(或曰「可被意识通达的」)。至于通达意识是否蕴含现象意识,抑或「拥有体验的能力」是否需要某种别的性质——这至今仍是一个存在争议的哲学问题。

We think our results do have something substantial to say about access consciousness in language models. The J-space appears to support the functions associated with conscious access: it holds the thoughts Claude can report on, deliberately bring to mind, and reason with, while the rest of its processing runs automatically beneath. Notably, none of this structure was designed into Claude—it emerged on its own during training, presumably because it was a useful way to organize computation. That suggests a mental workspace supporting conscious access isn’t just a peculiarity of how human brains happen to be wired. Instead, it appears to be a general solution that intelligent systems arrive at in order to solve certain kinds of problems. Now that we’ve identified this structure in Claude, it means we can make a meaningful distinction between the decisions Claude has made deliberately and those that happened automatically.

我们认为,我们的结果确实对「语言模型中的通达意识」有一些实质性的话可说。J 空间似乎支撑着与意识通达相关联的那些功能:它装载着 Claude 能够报告、能够有意唤起、并能用来推理的思绪,而它其余的处理则在底下自动地运行。值得注意的是,这套结构没有一处是被设计进 Claude 里去的——它是在训练过程中自行涌现的,想必是因为这是一种组织计算的有效方式。这就意味着,一个支撑意识通达的「心智工作空间」,并不只是人类大脑碰巧如此接线的一种怪癖。相反,它看起来是智能系统为了解决某些类型的问题而殊途同归得出的一种通用解法。既然我们已经在 Claude 身上辨认出了这套结构,那就意味着:我们可以在「Claude 有意做出的决定」与「自动发生的决定」之间,做出一个有意义的区分。

It’s important to note that there are several key differences between the workspace we identified in Claude and the global workspace model in humans. The brain’s workspace is sustained by recurrent loops—signals cycling back through the same circuits over time. In contrast, Claude’s workspace evolves over a single pass through the network, with the network’s depth playing the role that time plays in the brain. In this sense, Claude’s internal workspace processing is time-limited relative to humans’ (though it can compensate for this constraint by “thinking out loud” using its scratchpad). In other ways, however, Claude’s workspace is more powerful than that of humans. Human working memory fades within seconds, so the brain’s workspace has limited ability to retain information over time; in contrast, due to the attention mechanism in its neural network architecture, Claude can simply recall memories it cached at any earlier point in the text. Another important difference is the content of the workspace. While human conscious thoughts come in many formats—images, sounds, planned movements—Claude’s workspace is built almost entirely out of words. We suspect this is because producing words is the only kind of action Claude can take, which is not the case for humans.

需要着重指出的是,我们在 Claude 身上辨认出的这个工作空间,与人类的全局工作空间模型之间,存在若干关键差异。大脑的工作空间是由回环回路(recurrent loops)维持的——信号会随着时间的推移,一次次循环着回流经过同样的回路。相比之下,Claude 的工作空间是在单次穿过网络的过程中演变的——网络的深度,扮演着大脑中「时间」所扮演的角色。从这个意义上说,相对于人类,Claude 的内部工作空间处理是受时间限制的(不过它可以借助草稿纸「把话想出声来」,以此弥补这一约束)。然而,在另一些方面,Claude 的工作空间又比人类的更强大。人类的工作记忆在几秒钟之内就会消退,因此大脑的工作空间在「随时间保留信息」上能力有限;相比之下,得益于其神经网络架构中的注意力机制(attention mechanism),Claude 可以直接调取它在文本中任意更早位置缓存下来的记忆。另一个重要差异在于工作空间的内容。人类有意识的思想有许多种形式——图像、声音、计划好的动作——而 Claude 的工作空间几乎完全是由词语构成的。我们猜测,这是因为「产出词语」是 Claude 唯一能采取的行动,而对人类来说并非如此。

We hope the similarities and differences between the J-space and the global workspace model can feed back into neuroscience. The similarities present an exciting scientific opportunity: to the extent that the J-space mirrors our own mechanisms of conscious access, studying mechanisms in language models (much easier than studying human brains!) could inspire hypotheses in neuroscience. For instance, the J-space is constructed by identifying representations of potential outputs—words the model might say. If something similar holds in humans, it would suggest that the global workspace might be fundamentally tied to brain regions that prepare actions and speech, more so than to sensory areas. The differences between language models and human brains are instructive as well. They suggest that some aspects of our neural architecture, such as built-in recurrent connections, may not be strictly necessary to support the functions associated with conscious access. For an independent perspective on the neuroscientific implications of our work, see the invited commentary from Stanislas Dehaene and Lionel Naccache, two of the neuroscientists central to the development of global neuronal workspace theory.

我们希望,J 空间与全局工作空间模型之间的异同,能够反哺神经科学。其相似之处呈现出一个激动人心的科学机遇:在 J 空间映照着我们自身意识通达机制的那个限度之内,研究语言模型中的机制(可比研究人类大脑容易多了!)或许能为神经科学激发出新的假说。举例来说,J 空间是通过辨认「潜在输出的表征」——即模型可能会说出的词——而构建出来的。如果在人类身上也存在某种类似的情形,那就意味着:全局工作空间也许从根本上更多地系于「准备行动与言语」的脑区,而非系于感觉区。语言模型与人类大脑之间的差异同样发人深省。它们表明,我们神经架构的某些方面(比如与生俱来的回环连接)也许并非支撑「意识通达相关功能」所严格必需的。若想就我们这项工作的神经科学意涵获得一个独立视角,可参阅斯坦尼斯拉斯·迪昂(Stanislas Dehaene)与利昂内尔·纳卡什(Lionel Naccache)应邀撰写的评论——他们二位都是「全局神经元工作空间理论」发展进程中的核心神经科学家。

We mentioned that our experiments don’t answer whether AI models might have experiences. But that doesn’t make the question less important. Building systems with experiences like humans and animals have would raise very difficult ethical questions. Handling it correctly—and deciding whether it’s even morally acceptable—would require input from philosophers, scientists, religious leaders, governments, and the public. Thus, even if we’re not sure that we’ve crossed that bridge yet, we think it’s time to start thinking about it. We hope our work inspires further scientific investigation of forms of consciousness that might be present in AI systems, and a broader discussion of the implications.

我们前面提到,我们的实验并不能回答 AI 模型是否可能拥有体验。但这并不会让这个问题变得不那么重要。构建出「像人类和动物那样拥有体验」的系统,将会引发一些极为棘手的伦理问题。要妥善地处理它——乃至去判断这么做在道德上究竟是否可被接受——都将需要哲学家、科学家、宗教领袖、政府以及公众的共同参与。因此,即便我们还无法确定自己是否已经跨过了那道门槛,我们也认为,是时候开始思考这个问题了。我们希望,我们的工作能激发人们对「AI 系统中可能存在的各种意识形式」展开进一步的科学探究,并就其意涵展开更广泛的讨论。

This work is just a first step in what we expect to be an extensive line of research. The J-space looks like a good candidate for the divide between consciously accessible and unconscious processing in a language model, but we’d be surprised if it’s the whole story. The J-lens is undoubtedly an imperfect method, which only approximately captures the model’s “true workspace”—for instance, it can only identify concepts that correspond to single tokens. And there remain many mysteries about how the J-space works. We don’t know what mechanism decides what enters the J-space in the first place. We’ve seen hints that it’s tied to Claude’s sense of self, something like emotional reactions, and traces of metacognition, without exactly having worked out how. But we now have methods for tackling questions like these. As that work progresses, our understanding of LLM minds—and their relationship to our own—will grow clearer.

这项工作,只是我们预期中一条漫长研究路线上的第一步。J 空间看起来是「语言模型中『可被意识通达的处理』与『无意识处理』之分野」的一个不错的候选者,但若它就是故事的全部,我们反倒会感到意外。J 透镜无疑是一种并不完美的方法,它只是近似地捕捉了模型的「真正的工作空间」——比如说,它只能辨认出那些对应于单个 token 的概念。而关于 J 空间如何运作,也仍然存在着许多谜团。我们还不知道,究竟是什么机制在一开始决定了什么东西能进入 J 空间。我们已经看到一些迹象,表明它与 Claude 的「自我感」、与某种类似情绪反应的东西、以及与元认知(metacognition)的种种痕迹相关联——却还没能确切弄清其中的机理。但如今,我们已经拥有了应对这类问题的方法。随着这项工作的推进,我们对「大语言模型之心智」——以及它们与我们自身心智之关系——的理解,将会变得愈发清晰。

For more, read the full paper, and try the demo.

欲了解更多,请阅读完整论文,并试用我们的演示。


External commentary外部评论

We invited several outside experts to write independent commentaries on this work.

我们邀请了数位外部专家,就这项工作撰写独立评论。

  • Stanislas Dehaene and Lionel Naccache are cognitive neuroscientists who, together with Jean-Pierre Changeux, developed the global neuronal workspace model that inspired much of our work.

    斯坦尼斯拉斯·迪昂(Stanislas Dehaene)与利昂内尔·纳卡什(Lionel Naccache)是认知神经科学家,他们与让-皮埃尔·尚热(Jean-Pierre Changeux)一道,提出了「全局神经元工作空间模型」——我们这项工作的许多灵感正源于此。

  • Patrick Butlin, Dillon Plunkett, Robert Long (Eleos AI Research) and Derek Shiller (Rethink Priorities) study the potential for consciousness and moral status in AI systems.

    帕特里克·巴特林(Patrick Butlin)、迪伦·普伦基特(Dillon Plunkett)、罗伯特·朗(Robert Long,均来自 Eleos AI Research)以及德里克·希勒(Derek Shiller,来自 Rethink Priorities)研究 AI 系统中意识与道德地位的可能性。

  • Neel Nanda leads the language model interpretability team at Google DeepMind. His commentary includes an independent replication of some of our findings on an open-weight model.

    尼尔·南达(Neel Nanda)领导着 Google DeepMind 的语言模型可解释性团队。他的评论中包含了在一个开放权重模型上对我们部分发现的独立复现。

Read their commentaries here.

请在原文页面阅读他们的评论。

本文由 Claude Opus 4.8 翻译整理,仅供学习与交流。原文版权归 Anthropic 所有。

By nanikun