Skip to content
Computing Life · Share · Yage

To cut AI token costs, don't start by swapping to a cheaper model

When AI bills spike, swapping to a cheaper model is the wrong first move. The post splits AI costs into internal efficiency (Copilot, Cursor) and customer delivery (support AI, Duolingo Max), and argues each needs a different knife. For internal costs, cut idle seats and cap agent loops first. For delivery, track cost per outcome so you don't trim gross margin along with token spend. China Merchants Bank burns 33B tokens daily, yet AI coding takes only ~5% of compute—a reminder that enterprise token volume often lives in support and ops. Doubao hit 180T daily calls, but analysts question how much is paid production vs. free trial quota. The real sequence: attribute costs to teams and outcomes first, negotiate model pricing last.

Why it matters: Splits AI costs into internal efficiency vs. customer delivery, uses CMB's real numbers to show coding tokens may be far lower than customer service and ops — directly useful for anyone managing AI budgets. Missing concrete how-to on cutting idle seats and agent loops; the pos...

Read the original ↗Export Markdown