Study finds models take a moral route when judging malicious code
What happened
On October 7, Hacker News featured a study finding that when models are asked whether code is malicious, their routing looks closer to moral judgment than to plain code or vulnerability judgment. The study covered three open models, including OLMoE and DeepSeek-V2-Lite, and found malicious questions routed closest to moral questions, loading moral-path weights several times those of correctness questions. It also found that although routing resembled moral questions, concepts decoded from working memory included attacker and malware, showing a gap between routing and decoded concepts.
Written by AI from the coverage · updated 1 hour ago
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pageAsk a model if code is malicious and it reaches for its morals
研究发现,当被问及代码是否恶意时,模型走的是道德判断路径,而非单纯的代码或漏洞判断路径。在 OLMoE 和 DeepSeek-V2-Lite 等三个开源模型上,恶意问题的路由距离道德问题最近,加载道德路径的权重是正确性问题的数倍。模型路由像道德问题,但工作记忆解码出的却是攻击者、恶意软件等概念。
Heat over time
Not enough continuous observations to draw a trend yet.