Skip to content
Trending storyPast story

OpenRouter launches Batch API with 50% off for bundled inference

1 report1 sourceupdated 7 days ago

What happened

Summary

OpenRouter 新上线了 Batch API,你把一堆请求打包提交,模型供应商可以在 24 小时内挑空闲时间处理,作为交换,每 token 的费用直接砍半,部分模型甚至更低。两周测试期跑了超过 23 万个批次,中位数 7 分钟就出结果,90% 的批次在一小时内完成。提交时间比请求数量更影响速度:太平洋时间早上 5 点到中午 12 点最慢,最差的那...

Coverage

Follow the reports to see the story from different sides.

Sep 22
  1. AI HOT (Curated Pool)
    OpenRouter launches Batch API with 50% off for bundled inference

    OpenRouter's new Batch API lets you bundle requests so providers can process them within a 24-hour window, cutting per-token price by 50% or more. Across 230k+ batches during a two-week beta, the median finished in 7 minutes and 90% within an hour. Submission time matters more than batch size: batches sent 5am–noon Pacific are slowest, with the worst tenth taking 2–4.5 hours; after 6pm Pacific, 90% finish under 50 minutes. Over 70 models are supported for chat completions, messages, and embeddings—good for labeling, back-filling vectors, eval scoring, or summarizing ticket backlogs.