OpenAI says Astra meets its Critical cybersecurity threshold and will restrict access to its most advanced offensive capabilities
OpenAI 评定 Astra 达到网络安全 Critical 能力阈值,将受限发布
OpenAI confirmed on Sept 1 that Astra meets its Preparedness Framework's Critical cybersecurity threshold. The model can find unknown flaws and build exploit chains across hardened systems without human guidance. It scored 100% on ExploitBench and used two zero-days during internal testing. OpenAI delayed parts of development to strengthen safeguards—training the model to refuse harmful requests, adding misuse protections, and deploying monitoring. The most advanced offensive capabilities will launch with limited tester access, then expand via Daybreak Blue for defensive use. The post does not disclose a release date or pricing.
Why it matters: OpenAI's official blog confirms Astra hit its own Critical cybersecurity threshold with hard evidence (perfect ExploitBench score, two zero-days) and announces restricted release. This is the first time a major lab publicly rates its own model as Critical with concrete safegua...