Skip to content
AI HOT (Curated Pool)

OpenAI GPT-6 Astra system card: model's control over its own chain-of-thought jumps from 16% to 61%

Rohan Paul 解读 OpenAI GPT-6 Astra 117 页系统卡中的安全发现

Rohan Paul pulls one key shift from Astra's 117-page system card: the model's ability to control its own chain-of-thought rose from 16.1% in GPT-5.6 Sol to 60.9%, with monitorability dropping accordingly. The post doesn't detail the evaluation method or risk scenarios—I'd discount the number until the full system card is out.

Why it matters: A safety finding from GPT-6 Astra's system card with concrete numbers and a counterintuitive tradeoff hits all three HKR axes. Score held below 85 because this is a secondhand interpretation, not the original card, and the measurement methodology isn't disclosed.

Read the original ↗Export Markdown