Fireworks AI launches FireRouter with Opus, cutting coding costs by 57%
What happened
Fireworks AI 把自家的缓存感知路由器 FireRouter 单独拿出来,做成了一个叫 FireRouter with Opus 的模型端点。它会在每次对话回合里,自动从 Claude Opus 5.5、GLM 5.3 和 GLM 5.3 Flash 三个模型里挑一个最划算的来用。内部 A/B 测试跑了一个多月,主要看写代码任务:跟纯用 Op...
Coverage
Follow the reports to see the story from different sides.
- AI HOT (Curated Pool)PickFireworks AI launches FireRouter with Opus, cutting coding costs by 57%
Fireworks AI packaged its router as a standalone model endpoint that picks between Claude Opus 5.5, GLM 5.3, and GLM 5.3 Flash per turn. In over a month of internal A/B testing on coding tasks, it retained 98.1% of Opus-only accuracy while dropping per-session cost from $15.36 to $6.63—a 57% cut. The cache-aware router deliberately trades roughly 3.6 percentage points of cache hit rate for lower spend, ending at a 94.2% hit rate. It works with Claude Code, Codex, Cursor IDE, and others via two CLI commands.
Heat over time
Not enough continuous observations to draw a trend yet.