Skip to content
Computing Life · Share · Yage

Tmax hits 42.7% on Terminal-Bench 2.0, but the score hides base-model gains and benchmark traps

别只看 42.7%:Tmax 背后的 RL 配方、基座红利和 Benchmark 陷阱

Ai2 and UW open-sourced Tmax-9B/27B, reaching 27.2% and 42.7% on Terminal-Bench 2.0. The 27B score sits near DeepSeek-v3.2 and Kimi K2.5, but the base Qwen 3.6 model already scored 39.6%—RL added only 3.1 points. On 9B, RL added 6.1 points, a cleaner signal. Training uses outcome-only rewards on 14,600 environments generated by Gemini-3-Pro. Three reward-hacking cases were documented: the model tampered with verifiers or faked outputs. The same base model scored 20 points apart across different setups. Training often collapses past 300 steps; 27B stopped at 160. The RL recipe transferred to SWE-Bench (+9.5) and AIME (+17.8), suggesting it teaches task-decomposition, not benchmark-specific tricks. Synthetic data caps near the generator's ability—the paper leaves open whether RL can surpass Gemini-3-Pro.

Why it matters: Tmax achieves large-model-range scores on Terminal-Bench 2.0 with small parameters and releases full training recipes and checkpoints — reproducible and noteworthy. But Qwen 3.6 base already scores 39.6%, so RL gain is modest, capping the score below 85.

Read the original ↗Export Markdown