Test scores how closely 12 AI models agree on recommendations
What happened
On October 8 the Hacker News front page covered a Magic Numbers test of recommendation consistency across 12 AI models. The test used hundreds of English questions and compared how close each model's recommendations were to the consensus of the others. The five Chinese models in the sample averaged a consistency score of 0.47, the six US models 0.38 and Mistral 0.36, so the Chinese models as a group sat closer to the other models' consensus. Among individual models DeepSeek scored highest and Claude Haiku lowest.
Written by AI from the coverage · updated 1 hour ago
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pageAI Model Groupthink
Magic Numbers 对12个 AI 模型进行推荐答案一致性测试,样本中的中国模型整体更接近其他模型的共识。基于数百个英语问题,5个中国模型的平均一致性得分为0.47,6个美国模型为0.38,Mistral 为0.36;DeepSeek 得分最高,Claude Haiku 最低。
Heat over time
Not enough continuous observations to draw a trend yet.