Anthropic's new research reports that the Claude language model exhibits a 'global workspace'-like structure similar to the conscious and unconscious partitioning observed in the brain, in which only a small fraction of internal states become explicit. The finding is important for model interpretability and AI safety because it may open new ways to understand and regulate internal processes.
Anthropic: discovery of global-workspace-like operation in the Claude language model
Anthropic's new research reports that the Claude language model exhibits a 'global workspace'-like structure similar to the conscious and unconscious partitioning observed in the brain, in which only a small fraction of internal states become explicit.



