Skip to content
Trending storyWatching

Fireworks releases Ember-1 model with 40% fewer reasoning tokens at same quality

2 reports1 sourceupdated 2 days ago

What happened

From the coverage

Fireworks Research 基于 Kimi K3 训了个新模型 Ember-1,把推理时产生的内部思考 token 压掉了 35% 到 50%,但准确率没掉。他们跑了 50 多次训练实验,发现 K3 超过九成的 token 都花在内部推理上,其中很多是多余的。Ember-1 保留了有用的自我纠错,跳过了没产出的思考循环。在多轮 agent 任...

From AI HOT 精选

Coverage

Follow the reports to see the story from different sides.

Sep 28
  1. AI HOT (Curated Pool)Pick
    Fireworks AI releases Ember-1, a post-trained Kimi K3 that uses ~40% fewer tokens

    Fireworks AI post-trained Kimi K3 into Ember-1, cutting reasoning tokens by ~40% without losing accuracy. K3 sometimes spends over 90% of tokens on internal reasoning, which compounds cost in multi-turn agent workloads. Ember-1 keeps useful self-correction but drops redundant loops. On Terminal Bench 2.1 it scores 82%, beating K3's max-effort setting by 1.1 points while costing 51.9% less. Only available via Fireworks serverless API—weights and training code are not released.

Sep 24
  1. AI HOT (Curated Pool)Pick
    Fireworks launches Ember-1, matching Kimi K3 quality with 40% fewer tokens

    Fireworks Research released Ember-1, a model built on Kimi K3 that cuts reasoning tokens by 35–50% while keeping accuracy. Across 7 benchmarks and live A/B tests with two customers, quality held. The team ran 50+ training experiments and found K3 spends over 90% of tokens on internal reasoning, much of it unnecessary. Ember-1 preserves useful self-correction and skips unproductive loops. Savings compound in multi-turn agent tasks where prior reasoning is re-read each turn. The model is live on Fireworks' platform as the first in their own model series.

Heat over time

Not enough continuous observations to draw a trend yet.