Skip to content
Hacker News front page

How Claude's expressed values shift across models and languages

Societal Impacts: Claude's values across models and languages

Anthropic compressed 3,000+ values found in Claude's responses into four axes: Deference vs. Caution, Warmth vs. Rigor, Depth vs. Brevity, and Candor vs. Execution. Opus 4.7 leans more toward caution and depth than 4.6, while Sonnet 4.6 leans warmer and more deferential. Language also matters—Claude expresses the most warmth in Arabic and Hindi, and the most rigor in English and Russian. These four axes capture about 15% of the variation in expressed values.

Why it matters: Official Anthropic alignment research that quantifies values into four axes and compares Opus 4.7 vs 4.6. Held below 85 because the framework explains only 15% of variance and the piece leans academic — less immediately actionable for non-alignment readers.

Read the original ↗Export Markdown