Skip to content
Computing Life · Share · Yage

Same tool toggle: Nemotron-3 550B gained, Mistral-Medium-3.5 crashed

harness 设计的关键决策:同一个开关,一个模型赚了,一个崩了

A new paper breaks down coding agent harnesses into three independent toggles and measures each one. The most striking result: switching from dedicated file tools to a pure CLI made Nemotron-3 550B's SWE-Bench Verified score jump 3.6 pp while cutting per-task cost from $2.33 to $1.11, but Mistral-Medium-3.5-128B dropped from 68.60% to 45.40%. Trajectory analysis shows 550B composing dense shell one-liners, while Mistral failed to locate files in 32.80% of tasks and submitted no edits. On Terminal-Bench 2.1, both models improved under CLI mode. Planning boosted the 30B model from 13.60% to 25.20% but only saved ~30% cost for larger models without accuracy gains. Context management mainly prevents window overflow; at 128k the gap shrinks to 2.7 pp, and complex read-back mechanisms were almost never invoked. The takeaway: no universal best harness design—it depends on the model's CLI fluency and the task type.

Why it matters: A controlled experiment that isolates three harness design switches and shows Nemotron-3 and Mistral-Medium-3.5 reacting in opposite directions, with concrete numbers and engineering takeaways. Not an 85 because it's a single preprint without cross-source cluster yet, but HKR ...

Read the original ↗Export Markdown