Skip to content
Trending storyDeveloping

Three AI agents tested on US and Iran procurement data

1 report1 sourceupdated 1 hour ago

What happened

AI digest

On October 3, a Hacker News front-page report covered a test of three AI agents on public procurement data. Researchers used the web versions of Meta Muse, Anthropic Claude Cowork (Opus 5.5 Medium) and OpenAI GPT 6.1 Sol (Medium) to fill in World Bank public procurement data for the US and Iran, comparing language, information access and authorization differences across the tasks. The report says results on the Iran task were clearly weaker than on the US task. It does not list each agent's specific scores or detail how those differences affected the results.

Written by AI from the coverage · updated 1 hour ago

Coverage

Follow the reports to see the story from different sides.

Oct 3
  1. Hacker News front page
    Three AI agents, two countries, and one uneven world wide web

    研究者用 Meta Muse、Anthropic Claude Cowork(Opus 5.5 Medium)和 OpenAI GPT 6.1 Sol(Medium)的网页版补全世界银行美伊公共采购数据,发现伊朗任务结果明显弱于美国。

Heat over time

Not enough continuous observations to draw a trend yet.