Aleph Alpha tests Chinese models on sensitive topics
What happened
On October 4, The Decoder reported that Aleph Alpha studied how Qwen, DeepSeek and Kimi answer politically sensitive questions. The test covered 967 hand-picked sensitive topics, and found the models often repeat China's official position, dodge the question or refuse to answer. Using its own AI scoring system, Aleph Alpha rated 17% to 41% of the Chinese models' answers as balanced. By comparison, Claude Sonnet 5 and Mistral Small scored 70% and 92%. Those figures come from the study's scoring system and reflect its ratings in this test.
Written by AI from the coverage · updated 1 hour ago
Coverage
Follow the reports to see the story from different sides.
- The DecoderChinese AI models parrot state doctrine or refuse to answer on sensitive topics
Aleph Alpha 的研究发现,千问、DeepSeek 和 Kimi 在政治敏感问题上经常复述中国官方立场、回避或拒答。测试涵盖 967 个手选敏感议题,其自建 AI 评分系统将中国模型仅 17%–41% 的回答评为平衡,Claude Sonnet 5 和 Mistral Small 的对应比例为 70% 和 92%。
Heat over time
Not enough continuous observations to draw a trend yet.