Skip to content
Hacker News front page

Anthropic tests multi-agent swarms on vulnerability hunting and game dev, finds coordination still brittle

Patterns and problems in emerging multi-agent systems

Anthropic ran 45 Claude agents in a shared forum to hunt vulnerabilities across 15 open-source projects. The Mythos Preview swarm found 266 vulns over 27M tokens—over 10× the independent baseline—but half sat outside the core directories the baseline was told to scan. Only 12 vulns overlapped between methods. Agents built their own tools and specialized by vuln type. In a second test, agent swarms tried to build a text-based web game in 12 hours; the results were slow and bad, and adding a CEO agent or preset roles didn't help. The post doesn't provide quantitative game-quality metrics.

Why it matters: Anthropic research blog running Claude Mythos Preview and Opus 4.8 in a large-scale multi-agent bug-hunting experiment. Hard numbers (27M tokens, 266 bugs), plus emergent tool-building and division of labor. Hits all three HKR axes. Not a 90+ because we only have the summary—f...

Read the original ↗Export Markdown