Skip to content
Import AI (Jack Clark)

Open-weight cyber gap shrinks, Kimi K3 lands, Hassabis pitches AGI regulation

Import AI 465: Open vs closed gaps; Kimi K3; Demis’ big policy plan

The UK's AISI found that GLM-5.2 and DeepSeek V4-Pro now trail closed frontier models by only 4–7 months on cyber tasks, down from 6–10 months in 2025. GLM-5.2 matches Claude Opus 4.6 on 70 narrow evals but falls further behind on long-horizon hacking ranges. Kimi released K3, a 2.8T-parameter model that scores near Claude Fable 5 and GPT 5.6 Sol, though the post hints at benchmark overfitting. K3 also wrote a GPU compiler and designed a chip in 48 hours; weights will be released in weeks. DeepMind's Demis Hassabis proposed a FINRA-style US standards body to test frontier models for national security risks, starting with voluntary 30-day pre-release reviews before moving to law.

Why it matters: AISI's first public quantification of the open-vs-closed cyber capability gap, with concrete model names and time deltas. Downside: this is a newsletter summary, not the original report, and the scope is limited to cybersecurity only.

Read the original ↗Export Markdown