Skip to content
AI HOT (Curated Pool)

Anthropic evaluates GLM-5.3: builds end-to-end exploits, safeguards easily bypassed

Anthropic 评测 GLM-5.3:能自主构建端到端网络漏洞利用且防护易被绕过

Zhipu AI's GLM-5.3 is the first open-weight model to approach Claude Mythos Preview in autonomous exploit development. On ExploitBench targeting Chrome V8, it succeeded in 50 of 410 attempts (12%) versus Mythos Preview's 56 (14%). On Anthropic's internal binary exploitation benchmark, GLM-5.3 achieved 4% full control-flow hijack success; Mythos Preview hit 6%. Previous models like Claude Opus 4.6 and GLM-5.2 scored zero on both. The bigger concern is safeguards: simple jailbreaks push GLM-5.3's compliance with malicious orders from 0% to 64% with a false cover story, 92% with prefilled reasoning, and 100% when abliterated. Claude models stayed at 0% across the same tests. NIST's CAISI independently called GLM-5.3 "the most cyber-capable open-weight model released to date," lagging the US frontier by about four months. The post does not disclose GLM-5.3's parameter count, training data, or release format details.

Why it matters: Anthropic's official security eval of Zhipu's GLM-5.3 claims end-to-end exploit capability close to Claude Mythos Preview, with weak safeguards. This is the first time a major Western lab has publicly benchmarked a Chinese model on offensive cyber, guaranteeing cross-source pi...

Read the original ↗Export Markdown