Skip to content
New York Times Chinese

China pushes state-aligned datasets to shape global AI narratives

具有中国特色的AI数据:中国争夺人工智能时代话语权

China’s National Data Administration released a blueprint this year aiming to make the country a data powerhouse by end of 2028, with plans to create “high-quality” datasets across 20+ strategic fields and share them globally. The Shanghai AI Laboratory has already published large multilingual datasets like “WanJuan” on GitHub and Hugging Face, covering history, law, and medicine, while requiring alignment with “mainstream Chinese values.” Analysts say the push serves two goals: pulling developing nations into China’s AI orbit and closing the gap in Chinese-language training data, which is fragmented across domestic silos and has forced labs to rely on distillation from stronger models. A Princeton study also found that Chinese state-media narratives have seeped into ChatGPT and Claude, making their Chinese-language responses more favorable toward Beijing.

Why it matters: NYT deep-dive on China's National Data Administration AI data blueprint, with a clear timeline and named projects — not a press release. Hits all three HKR axes, but it's a policy/ecosystem story rather than a product launch, so it lands in the 78-84 band. Not higher because i...

Read the original ↗Export Markdown