OpenRouter launches Ori Eval: benchmark models against your own prompts to find the best fit
OpenRouter 推出 Ori Eval:用你的数据证明哪款模型最适合你的项目
OpenRouter released Ori Eval on July 31, a tool that benchmarks models directly inside your codebase. It scans every place your code calls a model, asks whether you care more about accuracy, latency, or cost, then auto-generates eval files and runs your real prompts against five recent models. The output is a table showing bug catch rate, p50 latency, and cost per PR — the post's example lists Claude Opus 5 at 94% catch, 38s p50, $0.041 per PR. The eval file is code you can run in CI to block regressions and re-run when new models drop. You start by telling your coding agent a single curl command; no eval-writing experience needed.
Why it matters: OpenRouter shipped a practical tool that lets devs benchmark models against their own codebase and real prompts, outputting bug catch rate, latency, and cost. The mechanism is concrete and the pain point is real — useful for anyone picking models day to day. Not scored higher ...