Cognition's SWE-2 hits 92.8 on Terminal-Bench 2.1, trails on long-horizon tasks
Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
Cognition post-trained Kimi K3 with RL to produce SWE-2, a 2.8T-param MoE model activating 104B per token. It scores 50.0 on FrontierCode 1.1 Main—0.9 behind Claude Fable 5.1 but at a claimed 64% lower cost—and leads the published table on Terminal-Bench 2.1 with 92.8. The weak spot is Terminal-Bench 4.0: 27.3 vs Fable 5.1's 55.8 and GPT-6 Astra's 57.9, so long-horizon agentic work still lags. Weights are proprietary, no per-token API exists, and all figures are Cognition's own, pending independent replication.
Why it matters: SWE-2 hit 92.8 on Terminal-Bench 2.1, the highest public score and a clear gap above FrontierCode 1.1 (50.0) and Claude Fable 5.1. The number is solid, but the post only gives model params and base model info — no training details, cost, or real-world deployment data, so it st...