Most AI Work Can Wait
构建AI智能体应优先设计路由
Tomasz Tunguz argues teams should design the router before picking a model. A small routing layer decides which tier handles each request; done right, 70–80% of traffic runs on near-free local or async models, cutting AI spend by 90%+. He cites Coinbase halving AI costs while token usage grew, using better defaults, routing, and caching. The routing stack has three layers: a skill classifier for intent, a router that assigns tiers by complexity and context size, and a model selector that picks the cheapest option within a tier. Their agent runtime now adds synchronous failure-mode signals and nightly closed-loop feedback to keep improving the router.
Why it matters: Tunguz flips the 'model-first' inertia by arguing for routing-first architecture, anchored by the concrete claim that 70-80% of traffic can run on free models, cutting costs by 90%+. Coinbase's real case adds credibility. The deduction is that this reads more as an architectur...