Research

Anthropic development: Natural-language decoding of Claude's activations

Anthropic's new research called Natural Language Autoencoders translates the internal activations of the Claude language model into human-readable text.

Anthropic's new research called Natural Language Autoencoders translates the internal activations of the Claude language model into human-readable text. The work aims to translate the numerical activations that encode the model's "thoughts" into natural language, increasing interpretability and safety.