I clicked on this because Cognition did something straightforward: they wrote inference cost directly into the RL reward as R = S − λC, so the model learns to take shorter paths on its own. The self-reported numbers are striking—mid-tier SWE-2 drops interaction turns from 127 to 53, cuts cost 81%, and moves the median first-edit step from 48 to 18.
But I'd discount this on two fronts. First, the post doesn't include an ablation without the cost penalty; the only ablation compares reward baseline forms. SWE-2 also switched to the stronger Kimi K3 base model, and Cognition doesn't separate how much of the efficiency gain comes from that upgrade. Second, on Terminal-Bench 4—a long-horizon benchmark—SWE-2 scores just 27.3%, well behind Claude Fable 5.1 and GPT-6 Astra. When the penalty piles up, the model's rational move is to give up early rather than burn tokens on a hard problem.
They chose a linear penalty over logarithmic, deriving it via Jensen's functional equation in Appendix B, but OckBench argues for log penalties—no consensus yet. The multi-tier training uses a single RL run across all thinking budgets, avoiding Kimi K3's train-nine-experts-then-distill approach, which is a clean design choice.
FrontierCode 1.1 Main isn't public, so independent replication is off the table. And cost numbers without test-environment context can mislead—one study found token-per-solved-task can vary 40x just by swapping scaffolding. The direction makes sense, but this looks more like a narrow cost-saving demo than a general solution.