Fireworks benchmarks DeepSeek-V4.1-Flash: matches GPT-6 Astra on DeepSWE at 1/15th the cost
Fireworks 评测 DeepSeek-V4.1-Flash:DeepSWE 达 GPT-6 Astra 水准、成本仅 1/15
Fireworks ran a full benchmark suite on DeepSeek-V4.1-Flash. On DeepSWE, it scores 74.34% pass@1, in the same band as GPT-6 Astra at 74.12%, but costs $0.43 per task—15x cheaper. The model uses a 552B MoE with a split activation design: 8B active for input, 16B for output, plus KV cache optimizations. On Terminal-Bench 2.1 it trails Astra by 1 point while costing 12x less. The post mentions an HLE and oracle router eval but does not disclose the actual scores.
Why it matters: DeepSeek V4.1-Flash matching GPT-6 Astra on DeepSWE at an order-of-magnitude lower cost is the strongest price-performance signal this week. Docked because the source is Fireworks' own benchmark, not an independent eval, and the body is truncated by a cookie wall with no full ...