Fireworks AI releases Ember-1, a post-trained Kimi K3 that uses ~40% fewer tokens
Fireworks AI 发布 Ember-1:基于 Kimi K3 的后训练模型,推理 Token 减少约 40%
Fireworks AI post-trained Kimi K3 into Ember-1, cutting reasoning tokens by ~40% without losing accuracy. K3 sometimes spends over 90% of tokens on internal reasoning, which compounds cost in multi-turn agent workloads. Ember-1 keeps useful self-correction but drops redundant loops. On Terminal Bench 2.1 it scores 82%, beating K3's max-effort setting by 1.1 points while costing 51.9% less. Only available via Fireworks serverless API—weights and training code are not released.
Why it matters: Fireworks post-trained Kimi K3 to cut ~40% reasoning tokens without accuracy loss, with concrete numbers and mechanism details—high practical value for agent builders. Capped below 85 because it's a third-party fine-tune, not a base model release, and the source is a MarktechP...