OpenRouter benchmarks Jev against LLMs for customer support: 100x cheaper but cannot generate text
What happened
OpenRouter 拿 TypeSafe 的 Jev 1.13 跟 GPT Luna、Claude Opus 跑了 60 张客服工单。Jev 做分类和判断要不要升级,每 1000 张票只要 0.025 美元,中位延迟 194 毫秒;Luna 要 0.09 美元,Opus 要 2.88 美元。Jev 直接返回带概率的类型化结果,不用再解析 JSON。但...
From AI HOT 精选
Coverage
Follow the reports to see the story from different sides.
- AI HOT (Curated Pool)OpenRouter benchmark: Jev 1.13 trails Claude Opus 5 by 3.3 points on Banking77 classification, but is 13x faster and 22x cheaper
OpenRouter tested Jev 1.13 and Claude Opus 5 on 3,080 Banking77 utterances across 77 intents. Jev hit 81.0% accuracy vs. Opus at 84.4%—a 3.3-point gap. Median latency: 175 ms for Jev, 2,266 ms for Opus. Cost per 1,000 requests: $0.11 vs. $2.42. Neither model produced malformed outputs. On compromised_card, Jev scored 95.0% while Opus got 70.0%. The post does not disclose Jev's parameter count or training details, and does not claim these results generalize to other classification tasks.
- AI HOT (Curated Pool)OpenRouter benchmarks Jev vs LLMs: 100x cheaper for decisions, but it can't write replies
OpenRouter benchmarked TypeSafe's Jev 1.13 against GPT Luna and Claude Opus on 60 support tickets. Jev classified and flagged escalation at $0.025 per 1,000 tickets with 194 ms median latency, versus $0.09 for Luna and $2.88 for Opus. It returns typed probabilities directly—no JSON parsing needed. Jev is text-only, weak at arithmetic, and can't generate prose. The post recommends routing with Jev first, then handing off to an LLM for replies.