OpenRouter launches live web search benchmarks to compare engines, depth, and models
OpenRouter 推出实时网页搜索基准测试:如何为智能体选择引擎、深度与模型
OpenRouter published live leaderboards testing web search combos across Exa, Parallel, Perplexity, and native lab engines. The biggest quality lever is search budget: on BrowseComp, Claude Opus 5 with Perplexity jumped from 35.8% at 1 turn to 89.0% at 25 turns, while cost rose only 2.5–7×. On easier tasks like HLE, extra turns barely helped—GPT-5.6 Sol scored similarly at 1 and 25 turns but cost 3× more. Models also burn through their full budget when they can't find an answer, driving up worst-case costs. The leaderboards update live; the post recommends testing against your own workload.
Why it matters: OpenRouter publishing its own web search benchmark with cross-engine comparisons is genuinely useful for agent builders. The headline finding—more search turns beats a model upgrade on cost—is actionable. Score isn't higher because this is a platform-run benchmark, not an inde...