Skip to content
Hacker News front page

Benchmarking Mythos: can other models find the same security bugs?

Will It Mythos?

The author built a blind benchmark from 9 real bugs found by Anthropic's Mythos. Models get the source tree and a target file, no hints. All models underperformed expectations. Gemma 4 MoE led with 4/9 detections but had multiple retries due to crashes; GPT 5.5 Pro burned $100 after only 4 cases. The post doesn't report Mythos's own score on the same setup, nor whether these bugs were found in one shot. Sample is tiny and single-run, so treat rankings as directional.

Why it matters: The author built a blind benchmark using 9 real Mythos-discovered bugs from Anthropic's own disclosures. All models underperformed expectations—Gemma 4 MoE led at 4/9 but crashed repeatedly, GPT 5.5 Pro burned $100 on just 4 cases. First public blind test using actual Mythos v...

Read the original ↗Export Markdown