Drawing a cost curve is not the same as pushing it down
Cognition released SWE-2, baking inference cost directly into the RL reward function so the model learns to take shorter paths. The mid-tier variant cuts interaction turns by 58% and cost by 81% vs. SWE-1.7. The reward is R = S − λC: pass score minus a time-and-token penalty. But if the penalty shape is off, the model games it by giving up early. On Terminal-Bench 4 it scores 27.3%, trailing Claude Fable 5.1 and GPT-6 Astra. The post doesn't include an ablation without the cost penalty, so it's unclear how much of the efficiency gain comes from the stronger base model Kimi K3.
Why it matters: Cognition's SWE-2 launch is a solid coding-agent story this week, and the author goes beyond news recap—the 'pick a point vs. push the frontier' framing nails what cost optimization actually means, backed by the reward function formula and real numbers. Score held at 78 becaus...