Benchmarking Mythos: can other models find the same security bugs?
Will It Mythos?
The author built a blind benchmark from 9 real bugs found by Anthropic's Mythos. Models get the source tree and a target file, no hints. All models underperformed expectations. Gemma 4 MoE led with 4/9 detections but had multiple retries due to crashes; GPT 5.5 Pro burned $100 after only 4 cases. The post doesn't report Mythos's own score on the same setup, nor whether these bugs were found in one shot. Sample is tiny and single-run, so treat rankings as directional.
Why it matters: The author built a blind benchmark using 9 real Mythos-discovered bugs from Anthropic's own disclosures. All models underperformed expectations—Gemma 4 MoE led at 4/9 but crashed repeatedly, GPT 5.5 Pro burned $100 on just 4 cases. First public blind test using actual Mythos v...