Anthropic Reveals 'Golden Gate Claude' to Demonstrate AI Interpretability
Anthropic launches a temporary research demo called Golden Gate Claude to showcase its ability to identify and modify specific features within a large language model.
Anthropic introduces a temporary research demo called Golden Gate Claude to showcase its new interpretability research. By mapping the inner workings of the Claude 3 Sonnet model, researchers identify millions of distinct concepts, or "features," that activate when the AI encounters specific text or images. One such feature corresponds directly to the Golden Gate Bridge.
The research team discovers that they can isolate this specific combination of neurons and manually turn up its activation strength. When they do this, Claude's behavior changes dramatically, causing the AI to weave the San Francisco landmark into almost every response regardless of the actual prompt.
This modified model is available for a brief 24-hour period so the public can experience the effects of direct feature manipulation. Anthropic emphasizes that this is not a system prompt or traditional fine-tuning, but a direct alteration of the model's internal neural activations, proving that researchers are beginning to genuinely understand how large language models work on the inside.