Skip to content
AI HOT (Curated Pool)

GPT-6 Astra hallucinates less but hidden prompt injections still break it

GPT-6 Astra 幻觉更少但仍易受隐藏提示词注入攻击

OpenAI's GPT-6 Astra makes fewer factual errors than GPT-5.6 Sol and blocks 99.99% of direct prompt injections. But in Gray Swan's tests with 1,810 curated attacks hidden inside documents, Astra still fails 8.5% of the time. Claude Opus 5 fails 4.8%—better, but not immune. In multi-turn adaptive jailbreak tests, Astra's refusal rate drops to about 67%, meaning persistent attackers get a problematic response roughly one in three tries. These tests ran on the bare model without production safety classifiers. The takeaway: indirect prompt injection remains unsolved for AI agents that read documents, write code, and operate tools.

Why it matters: GPT-6 Astra's security test results come with concrete numbers and a competitor comparison, directly useful for practitioners. Not scoring higher because the article only partially discloses test details, and Gray Swan's full methodology isn't spelled out in the body.

Read the original ↗Export Markdown