Three AI agents tested on US and Iran procurement data
What happened
On October 3, a Hacker News front-page report covered a test of three AI agents on public procurement data. Researchers used the web versions of Meta Muse, Anthropic Claude Cowork (Opus 5.5 Medium) and OpenAI GPT 6.1 Sol (Medium) to fill in World Bank public procurement data for the US and Iran, comparing language, information access and authorization differences across the tasks. The report says results on the Iran task were clearly weaker than on the US task. It does not list each agent's specific scores or detail how those differences affected the results.
Written by AI from the coverage · updated 1 hour ago
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pageThree AI agents, two countries, and one uneven world wide web
研究者用 Meta Muse、Anthropic Claude Cowork(Opus 5.5 Medium)和 OpenAI GPT 6.1 Sol(Medium)的网页版补全世界银行美伊公共采购数据,发现伊朗任务结果明显弱于美国。
Heat over time
Not enough continuous observations to draw a trend yet.