RTK claims token savings, but our cost benchmarks disagree
Quesma spent over $1,500 running Terminal-Bench 2.1 with Claude Code + Fable 5.0 and OpenCode + DeepSeek V4 Pro 0813, with and without RTK. Fable's total cost dropped 5%, but nearly all savings came from one task finishing in half the turns. DeepSeek's cost rose 17% on average. RTK's built-in `rtk gain` metric is misleading: a single `head -1` call was credited as saving 120.5M tokens, though the actual bill didn't change. A bug in v0.45.0 caused 339 consecutive errors in one attempt; the post says v0.46.0 fixed it. Compressing terminal output does not equal cheaper coding, and can sometimes cost more.
Why it matters: Quesma spent $1,500 running Terminal-Bench 2.1 to benchmark RTK's real cost impact across two toolchains, with results contradicting RTK's claimed 60% token savings. Concrete numbers, clear methodology, and a direct conflict with the prevailing narrative — all three HKR axes h...