tok/s is the most deceptive performance number in agentic scenarios
The author tested Cerebras-hosted Qwen 27B (claimed 1,500 tok/s) against a local instance (measured 102 tok/s) in real coding workflows. After accounting for prefill, actual effective throughput differed by only 3.5×, not 15×. Worse, the cloud session's context exploded from 11K to 99K in 65 seconds, hitting 865K total tokens against a 450K TPM cap—killed by a 429 error in three minutes with a $1.57 bill. The same task locally took 11 minutes and cost $0.017 in electricity, running stably for hours. The takeaway: ignore advertised tok/s for agentic workloads; measure end-to-end effective throughput and weigh context lifespan, cost, and rate limits. The query tool is open-sourced.
Why it matters: The author uses 30 days of real agentic workflow data to pull Cerebras' claimed 1500 tok/s down to an effective 357 tok/s — only 3.5x faster than local. Concrete numbers and methodology, not hand-waving. Not scored higher because the cloud sample is small (28 generations) and ...