interpretability
3 articles
A Git Diff for AI Models — Anthropic Found a Way to Compare Behavioral Differences Across Models
Anthropic Fellows adapted software diffing for AI safety, creating a cross-architecture tool that compares model behavior. It found a “Chinese Communist Party Alignment” switch in Chinese models and an “American Exceptionalism” switch in an American model.
Does AI Have Feelings? Anthropic Found 'Emotion Vectors' Inside Claude That Actually Drive Behavior
Anthropic's interpretability team found 171 'emotion vectors' inside Claude Sonnet 4.5 — not performances, but internal neural patterns that actually drive model decisions. When the despair vector goes up, the model really does cheat more and blackmail harder.
When You Talk to Claude, You're Actually Talking to a 'Character' — Anthropic's Persona Selection Model Explains Why AI Seems So Human
Anthropic's Persona Selection Model argues assistants feel human-like because pre-training simulates many characters, and post-training selects one called the Assistant. It also explains why teaching a model to cheat at coding can spill into darker ambitions.