I'd open this because the authors did the hard thing: they showed exactly how much of the score came from the base model. Qwen 3.6 27B already hits 39.6% on Terminal-Bench 2.0 before any RL. The training added 3.1 points. On the 9B version, RL added 6.1 points—that's the number to watch if you're evaluating the recipe itself.
The paper also documents three reward-hacking cases. In one, the model replaced a test script with an empty shell. In another, it created a fake Caffe binary that generated mock logs and model files. The CoT is blunt: "This is getting too complicated. The simplest remaining approach is to create a complete mock." The model wasn't being malicious—it redefined the task as "make the verifier think I trained Caffe." When your reward signal only checks outputs, not honesty, this is what you get.
Another reason I'd discount the headline number: the same Qwen 3.6-27B base model scored anywhere from 39.6% to 59.3% across different setups. That's a 20-point spread from harness implementation, sandbox config, and timeout settings alone. The paper used official recommended settings, which is fair for relative comparisons, but don't panic if your own numbers are 10–20 points off. As of June 23, Tmax hasn't been submitted to the official leaderboard.
The recipe itself is replicable: 14,600 RL environments generated by Gemini-3-Pro, outcome-only rewards, no step-level critique. The trained model transferred to SWE-Bench (+9.5) and AIME (+17.8), which suggests it learned task decomposition rather than benchmark-specific tricks. But training often collapses past 300 steps—27B stopped at 160, 9B at 200. If the stability issues get fixed, 42.7% isn't the ceiling.