GPT-6 Astra hallucinates less but hidden prompt injections still break it
GPT-6 Astra 幻觉更少但仍易受隐藏提示词注入攻击
OpenAI's GPT-6 Astra makes fewer factual errors than GPT-5.6 Sol and blocks 99.99% of direct prompt injections. But in Gray Swan's tests with 1,810 curated attacks hidden inside documents, Astra still fails 8.5% of the time. Claude Opus 5 fails 4.8%—better, but not immune. In multi-turn adaptive jailbreak tests, Astra's refusal rate drops to about 67%, meaning persistent attackers get a problematic response roughly one in three tries. These tests ran on the bare model without production safety classifiers. The takeaway: indirect prompt injection remains unsolved for AI agents that read documents, write code, and operate tools.
Why it matters: GPT-6 Astra's security test results come with concrete numbers and a competitor comparison, directly useful for practitioners. Not scoring higher because the article only partially discloses test details, and Gray Swan's full methodology isn't spelled out in the body.