BrowseComp: a benchmark for browsing agents
OpenAI open-sourced BrowseComp, a 1,266-question benchmark for measuring how well AI browsing agents find hard-to-locate information. Tasks require short, uniquely gradable answers; annotators checked that GPT-4o, o1, and an early deep research model failed, and that five searches did not reveal the answer on first-page results. The key signal is “hard to find, easy to verify,” which tests persistence, search strategy, and factual verification rather than basic retrieval.
Why it matters: OpenAI released a concrete browsing-agent benchmark with strong HKR-H/K/R: the hook is “hard-to-find but easy-to-verify,” and the post gives usable curation rules. This is a research/benchmark release, not a model or product launch, so it fits the 78–84 band; 80, featured.