Skip to content
Computing Life · Share · Yage

GPT-4o mini hits 10M+ daily calls, not for chat or code

每天一千万次调用的 GPT-4o mini,不聊天也不写代码

On Aug 13, 2026, GPT-4o mini handled 17.61M requests on OpenRouter, averaging just 92 output tokens per call with an 18.6:1 input-to-output ratio. This read-heavy, write-light pattern maps to four pipeline roles: request routing, structured extraction, guard checks, and offline batch jobs—not chat or coding. Open-source small models like Qwen 27B barely appear on paid cloud routes because devs run them locally. The post doesn't disclose which specific customers or products drive those 17.61M calls.

Why it matters: A solid traffic analysis using public OpenRouter data, reframing GPT-4o mini from 'cheap substitute' to pipeline sorting station with real numbers and a four-category taxonomy. Downside: single-author analysis without cross-source verification, and the body excerpt cuts off be...

Read the original ↗Export Markdown