Skip to content
AI HOT (Curated Pool)

Fireworks AI launches FireRouter with Opus, cutting coding costs by 57%

Fireworks AI 推出 FireRouter with Opus,编码任务成本降 57%

Fireworks AI packaged its router as a standalone model endpoint that picks between Claude Opus 5.5, GLM 5.3, and GLM 5.3 Flash per turn. In over a month of internal A/B testing on coding tasks, it retained 98.1% of Opus-only accuracy while dropping per-session cost from $15.36 to $6.63—a 57% cut. The cache-aware router deliberately trades roughly 3.6 percentage points of cache hit rate for lower spend, ending at a 94.2% hit rate. It works with Claude Code, Codex, Cursor IDE, and others via two CLI commands.

Why it matters: Fireworks cut Opus 5.5 routing cost by 57% with internal A/B data — real savings for devs coding with Claude. Not p1 because it's a routing-layer optimization, not a model capability leap, and only Fireworks' own numbers, no external validation.

Read the original ↗Export Markdown