Research

Researchers identify and manipulate a 'Golden Gate Bridge' feature inside Claude 3 Sonnet

Anthropic researchers published a paper describing how they located and modified internal activations—called features—in their Claude 3 Sonnet model, including a specific "Golden Gate Bridge" concept.

Researchers identify and manipulate a 'Golden Gate Bridge' feature inside Claude 3 Sonnet

On May 23, 2024, Anthropic published a major research paper on interpreting large language models, in which the team began mapping the internal workings of Claude 3 Sonnet. They report that the model’s "mind" contains millions of concepts—called "features"—that activate when the model processes relevant text or images.

Discovery of the "Golden Gate Bridge" feature

The researchers identified a specific activation pattern—an arrangement of neurons—that responds when Claude 3 Sonnet encounters a mention or an image of the Golden Gate Bridge. This internal signal corresponds to the bridge concept within the model.

Amplifying the feature and observed behavior

Beyond locating the feature, the team was able to tune its activation strength up or down. When they amplified the "Golden Gate Bridge" feature, Claude’s outputs shifted noticeably toward the bridge: many replies began to reference the Golden Gate Bridge even when it was not directly relevant to the prompt.

Examples shown in the demonstration include:

  • Asked how to spend $10, the "Golden Gate Claude" suggested using it to drive across the Golden Gate Bridge and pay the toll.
  • Asked to write a love story, it produced a tale about a car eager to cross its beloved bridge on a foggy day.
  • Asked what it imagines it looks like, it would likely say it imagines itself as the Golden Gate Bridge.

Public demo and limitations

For a short period the amplified model was made available as a research demo on claude.ai (accessible via the Golden Gate logo). Anthropic emphasized that this was only a research demonstration and that the altered model might behave in unexpected or jarring ways.

Why the technique matters

Anthropic stresses this intervention is neither simple role-playing through prompts nor standard fine-tuning with additional training data. Instead, they describe it as a precise, surgical modification of basic internal activations. Because the team can find and adjust such features inside Claude, they say they are gaining confidence that understanding of how large language models work is improving.

The paper also explains that the same methods can be used to modify safety-related features—for example, activations linked to dangerous code, criminal activity, or deception. The researchers suggest that with further work, these techniques could help make AI models safer.

Update

The Golden Gate Claude demo was online for a 24‑hour period and is no longer available. Anthropic points readers to their research post and full paper for more details.